[{"id":"doi:10.1038/s41598-025-27472-1","type":"article-journal","title":"Splitting smarter: Differential privacy for secure healthcare federated learning.","abstract":"Split Federated Learning (SplitFed) has emerged as a decentralized method of training ML models that enables multiple healthcare parties to collaboratively share models without sharing their raw data. This method, however, is vulnerable to label inference attacks, which can compromise patient privacy. Previous research efforts have attempted to address the question. However, these works do not conduct a detailed vulnerability analysis of SplitFed against label inference attacks. Additionally, some of these efforts propose differential privacy (DP) as a solution; the works focus on distributed learning paradigms where labels used for training the model are available to the clients, which is not a practical assumption. To address this, in this paper, we investigate the vulnerability of SplitFed models to label inference attacks in biomedical imaging. We propose a solution that incorporates DP into SplitFed to protect against label inference attacks. Additionally, we also provide a detailed vulnerability analysis of SplitFed against label inference attacks specific to healthcare applications. Finally, we propose a DP-based method for mitigating label inference attacks against Split-Fed models. Results indicate the efficacy of the SplitFed model under multiple conditions and found that the label inference accuracy changes from [Formula: see text] (No-DP) to [Formula: see text] (with DP). This indicates that integration of DP offers a robust mechanism for protecting patient privacy. Additionally, the usage of Cauchy noise in DP provides the best protection out of all noise categories, with a label inference accuracy of 0% while Exponential noise was the worst, resulting in a label inference accuracy of 68%.","author":[{"family":"Onireti","given":"Munirat"},{"family":"Shukla","given":"Raj"},{"family":"Das","given":"Tapadhir"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-27472-1","URL":"https://doi.org/10.1038/s41598-025-27472-1","source":"europepmc"},{"id":"doi:10.48550/arxiv.2608.28379","type":"manuscript","title":"Quantum Federated Learning Based on Bures--Uhlmann Geometry for Heterogeneous Noisy Clients","abstract":"Quantum federated learning enables collaborative model training across quantum devices without sharing raw data, and it faces the data and hardware heterogeneity inherent to noisy quantum devices. Utilizing the quantum geometric tensor is a natural remedy, yet pure-state approaches and diagonal approximations discard the correlations that encode parameter incompatibility. To address this, we extend the parameter-space geometry to the mixed states that noisy clients actually prepare. The real part of the resulting mixed-state geometric tensor is the Bures metric, which measures how fast the physical state changes under parameter variation, and the imaginary part is the mean Uhlmann curvature, which quantifies the incompatibility of estimating multiple parameters simultaneously. Accordingly, we employ the Bures metric as a local preconditioner and use the mean Uhlmann curvature to develop an achievable-precision aggregation rule that dynamically down-weights unreliable clients. Furthermore, we establish theoretical guarantees by proving a convergence theorem and a variance-dominance proposition. Empirical evaluations on a trapped-ion quantum emulator demonstrate that the proposed method maintains high accuracy across diverse device-heterogeneity conditions and outperforms standard federated averaging, whose accuracy degrades under strong noise.","author":[{"family":"Emori","given":"Haruki"},{"family":"Uchihara","given":"Masaki"},{"family":"Tokunaga","given":"Yuuki"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.28379","URL":"https://doi.org/10.48550/arxiv.2608.28379","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.27856","type":"manuscript","title":"FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling","abstract":"Recent advances in large language models are enabling autonomous clinical agents to perform increasingly complex electronic health record (EHR) modeling workflows. However, agents deployed at individual hospitals remain constrained by institution-specific data and modeling environments, while direct cross-hospital collaboration is restricted by the sensitivity of patient-level EHR data. Although federated learning (FL) provides a natural foundation for privacy-preserving collaboration, existing approaches remain predominantly model-centric, limiting federation to prediction models or their updates while overlooking the richer modeling experience accumulated by autonomous agents. To address this limitation, we propose FedEHR-Agents, an experience-centric federated agentic optimization framework for automated EHR modeling. Each hospital deploys an autonomous clinical EHR agent that performs data preprocessing and model development while refining local clinical modeling experience through historical memory, task-specific evaluation, and TextGrad-based prompt refinement. The federated server performs evidence-guided experience aggregation to integrate reliable and complementary modeling experience across heterogeneous hospitals and distills the aggregated experience into global meta-prompts for subsequent local refinement. Extensive experiments on real-world multi-hospital EHR benchmarks demonstrate that FedEHR-Agents consistently outperforms local and federated baselines across diverse clinical prediction tasks and remains robust across different federation scales and LLM backbones. These results establish clinical modeling experience as a promising collaborative object beyond conventional parameter-centric FL and point toward federated autonomous clinical intelligence.","author":[{"family":"Bai","given":"Jun"},{"family":"Wang","given":"Ruilin"},{"family":"Li","given":"Yue"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.27856","URL":"https://doi.org/10.48550/arxiv.2608.27856","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.27791","type":"manuscript","title":"Initialization Is Critical: Advancing Federated Short-Term Load Forecasting under Load Heterogeneity via Model Initialization","abstract":"Short-term load forecasting (STLF) provides essential information for numerous applications in modern power systems. However, accurate STLF often relies on fine-grained smart-meter data from distributed users, raising increasing concerns about data privacy. Federated learning (FL) has therefore emerged as a promising privacy-preserving paradigm for STLF. Nevertheless, this paper reveals structured heterogeneity in clients' load data. Specifically, clients exhibit different responses to exogenous factors and distinct temporal load profiles, which can degrade forecasting performance in FL. To mitigate these issues, this paper studies the role of model initialization in federated STLF, and proposes two initialization strategies from global and local perspectives. For global model initialization, when auxiliary public load data are available, a pretrained initialization strategy is developed to initialize the global model before federated training, thereby reducing client drift during the training process. For local model initialization, we propose SLIAvg, a sequential local initialization strategy that promotes a more consistent training process by allowing participating clients to start from progressively adapted models within each communication round. Since the proposed strategies only modify the initialization process, they are compatible with most existing FL frameworks and privacy-enhancing techniques. Experiments on real smart-meter data with two representative forecasting architectures demonstrate that the proposed strategies effectively improve forecasting performance, as evidenced by reduced client drift, improved convergence behavior, and lower forecasting errors.","author":[{"family":"Chen","given":"Jianing"},{"family":"Farhadi","given":"Vajiheh"},{"family":"Li","given":"Yan"},{"family":"La Porta","given":"Thomas"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.27791","URL":"https://doi.org/10.48550/arxiv.2608.27791","source":"datacite"},{"id":"doi:10.5281/zenodo.20697352","type":"article-journal","title":"AdaptFedAvg: Privacy-Preserving Topic Difficulty Modeling in Heterogeneous Smart Classroom Networks","abstract":"AdaptFedAvg is a privacy-preserving federated learning framework for topic difficulty modeling in heterogeneous smart classroom networks. Educational institutions generate large volumes of student interaction data, but privacy regulations such as GDPR, FERPA, and India’s DPDP Act restrict centralized data collection. This work proposes AdaptFedAvg, an enhanced Federated Averaging approach that improves performance under non-IID client distributions using three mechanisms: local proximal regularization, server-side momentum, and variance-based client quality weighting. The framework also introduces a seven-class student struggle taxonomy that maps struggling learners to targeted pedagogical interventions. Experimental evaluation on a simulated ten-school dataset under mild, medium, and severe non-IID settings shows that AdaptFedAvg converges within 30–40 communication rounds and achieves lower MAE than standard FedAvg, with especially strong robustness under severe heterogeneity. Results demonstrate the effectiveness of adaptive aggregation for privacy-preserving educational analytics.","author":[{"family":"Korra","given":"Kiran"},{"family":"Ch","given":"Sarayu"},{"family":"Bonthu","given":"Akshay"},{"family":"Thaduri","given":"Shiva"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20697352","URL":"https://doi.org/10.5281/zenodo.20697352","source":"datacite"},{"id":"doi:10.5281/zenodo.20697353","type":"article-journal","title":"AdaptFedAvg: Privacy-Preserving Topic Difficulty Modeling in Heterogeneous Smart Classroom Networks","abstract":"AdaptFedAvg is a privacy-preserving federated learning framework for topic difficulty modeling in heterogeneous smart classroom networks. Educational institutions generate large volumes of student interaction data, but privacy regulations such as GDPR, FERPA, and India’s DPDP Act restrict centralized data collection. This work proposes AdaptFedAvg, an enhanced Federated Averaging approach that improves performance under non-IID client distributions using three mechanisms: local proximal regularization, server-side momentum, and variance-based client quality weighting. The framework also introduces a seven-class student struggle taxonomy that maps struggling learners to targeted pedagogical interventions. Experimental evaluation on a simulated ten-school dataset under mild, medium, and severe non-IID settings shows that AdaptFedAvg converges within 30–40 communication rounds and achieves lower MAE than standard FedAvg, with especially strong robustness under severe heterogeneity. Results demonstrate the effectiveness of adaptive aggregation for privacy-preserving educational analytics.","author":[{"family":"Korra","given":"Kiran"},{"family":"Ch","given":"Sarayu"},{"family":"Bonthu","given":"Akshay"},{"family":"Thaduri","given":"Shiva"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20697353","URL":"https://doi.org/10.5281/zenodo.20697353","source":"datacite"},{"id":"doi:10.1038/s41598-026-39149-4","type":"article-journal","title":"MedLedgerFL: a hybrid blockchain-federated learning framework for secure remote healthcare services.","abstract":"The fast growth of telemedicine has made it clear that we need safe, cooperative, and privacy-protecting ways to handle sensitive medical information. Traditional centralised predictive model training methods have problems including data leaks, ownership disputes, and tight rules that make it hard for institutions to work together. To solve these problems, this paper suggests MedLedgerFL, a hybrid platform that combines blockchain with federated learning (FL) to make remote healthcare analytics safe and reliable. The blockchain layer makes sure that the model updates can be audited, can’t be changed, and can only be accessed by authorised healthcare institutions. The federated learning layer lets healthcare institutions work together to train the model without sharing patient data, which makes sure that they follow data protection laws like GDPR. Tests show that MedLedgerFL is better than other methods at making accurate predictions, communicating quickly, and keeping information private. Combining blockchain consensus with federated model aggregation makes things more open, lowers the danger of data exposure, and makes models more reliable. Future endeavours will concentrate on enhancing scalability among institutions, integrating sophisticated privacy-preserving techniques like differential privacy and homomorphic encryption, and assessing the framework’s efficacy in extensive, practical telemedicine applications.","author":[{"family":"Murala","given":"Dileep"},{"family":"Vemulapalli","given":"Lavanya"},{"family":"Balagoni","given":"Yadaiah"},{"family":"Patnala","given":"Eswar"},{"family":"Romeo","given":"BM"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-39149-4","URL":"https://doi.org/10.1038/s41598-026-39149-4","source":"europepmc"},{"id":"doi:10.3389/frai.2025.1644844","type":"article-journal","title":"Implementing federated learning for privacy-preserving emotion detection in educational environments.","abstract":"Emotion detection has become an essential tool in educational settings, where understanding and responding to students' emotions is crucial to improving their engagement, academic performance, and emotional well-being. However, traditional emotion detection systems, such as DeepFace, and hybrid transformer-based models face significant data privacy and scalability limitations. These models rely on transferring sensitive data to central servers, compromising student confidentiality and making deployment in large or diverse populations difficult. In this work, we propose a federated learning-based model designed to detect emotions in educational settings, preserving data privacy by processing them locally on students' devices (smartphones, tablets, and laptops). The model was integrated into the Moodle platform, allowing its evaluation in a conventional educational environment. Advanced anonymization and preprocessing techniques were implemented to ensure the security of emotional data and optimize its quality. The results demonstrate that the proposed model achieves a precision of 87%, a recall of 85%, and an F1-score of 86%, maintaining its performance under adverse conditions, such as low lighting and ambient noise. In addition, a 15% increase in academic participation and a 12% improvement in the average academic performance of students were observed, highlighting the system's positive impact on educational dynamics. This innovative method combines privacy, scalability, and performance, positioning itself as a viable and sustainable solution for emotion detection in contemporary educational environments.","author":[{"family":"Gutiérrez","given":"Rommel"},{"family":"Villegas-Ch","given":"William"},{"family":"Lujánmora","given":"Sergio"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3389/frai.2025.1644844","URL":"https://doi.org/10.3389/frai.2025.1644844","source":"europepmc"},{"id":"doi:10.1038/s41598-025-30150-x","type":"article-journal","title":"Edge-AI integrated secure wireless IoT architecture for real time healthcare monitoring and federated anomaly detection.","abstract":"The rapid digitalization of healthcare demands intelligent, low-latency, and privacy-preserving systems capable of operating at the network edge. This study introduces a unified Edge-AI framework that seamlessly combines dual wireless connectivity (LoRaWAN + 5G), federated learning (FL), Proof-of-Authority (PoA) block chain, and homomorphic encryption (HE) to achieve secure real-time anomaly detection in patient monitoring. A quantized CNN-LSTM model was deployed on NVIDIA Jetson Nano devices and trained using a synthetic dataset statistically modelled from the MIT-BIH Arrhythmia Database, capturing vital signals such as heart rate, temperature, and oxygen saturation. The integrated system attained 91.9% accuracy and 90.8% F1-score, with only an 8.7% latency overhead attributed to HE operations. A paired two-tailed t-test (p < 0.01) confirmed that these gains are both statistically and clinically significant, indicating reliable diagnostic performance under constrained conditions. Beyond performance, the proposed framework ensures end-to-end data confidentiality, tamper-proof auditability, and energy-efficient edge inference, offering a scalable pathway toward trustworthy, next-generation smart-healthcare ecosystems.","author":[{"family":"Prabha","given":"M"},{"family":"Nandhini","given":"S"},{"family":"Dayanidhy","given":"M"},{"family":"Pradeep","given":"R"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-30150-x","URL":"https://doi.org/10.1038/s41598-025-30150-x","source":"europepmc"},{"id":"doi:10.1038/s41598-025-26510-2","type":"article-journal","title":"A unified AI-driven framework for quantum-secured 6G THz networks with intelligent reflecting surfaces and federated edge learning.","abstract":"The main contribution of this manuscript is an innovative framework for integrating Artificial Intelligence (AI) in 6G wireless systems. With increased complexity, including bursty traffic, network complexity, and dynamic variability, there is a need for intelligence. This study develops and validates an AI-driven approach that enhances network performance through quantum communication decoding, beamforming, and decentralized edge processing. Kalman filtering predictive models are used to estimate variable channel conditions in a Terahertz (THz) network to support beamforming to optimize beamforming. Artificial Intelligence exploits smart reflective surfaces (IRS) strengthening signals and improving their coverage. Also, strong security of Quantum Key Distribution (QKD) protocols due to AI enhanced error correction technology, and rapid, yet privacy information conducting at edge nodes due to decentralised processing through federated learning are examples of enhanced capabilities. Extensive ns-3 simulations across 100 independent runs validate the framework's effectiveness and prove the system in practical 6G deployment scenarios including THz links, IRS component and edge nodes. The simulation results demonstrate that the proposed framework achieves superior performance compared to conventional approaches, with statistical validation across multiple deployment scenarios. The system decreases latency by 30%, and adds 25% to spectral efficiency. In bursty traffic, the energy efficiency is increased by 20% and packets delivery ratio (PDR) is boosted by 15%. The AI algorithms work effectively to regulate the channel estimation, beamforming, and resource allocation, and, as a result, showed an improvement in the order of magnitudes over previous studies. These results support the fact that AI demonstrates significant potential for transformative impact to a 6G network. The framework has been efficient in addressing problems of channel estimation, beamforming and distributed processing and novel calculations in quantum communication security protocols. Such findings can be used as the foundation of the further inclusion of AI-based technologies in 6G systems, which will help to deploy robust, resilient, and autonomous wireless networks to address the needs of a connective society.","author":[{"family":"Balaji","given":"CG"},{"family":"Menaka","given":"S"},{"family":"Rajeswari","given":"G"},{"family":"Ponnusamy","given":"Sivaram"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-26510-2","URL":"https://doi.org/10.1038/s41598-025-26510-2","source":"europepmc"},{"id":"doi:10.48550/arxiv.2606.23500","type":"manuscript","title":"Development and Design of FLKit: A Structured Onboarding Toolkit for Federated Learning in Health and Life Sciences","abstract":"Federated learning lets institutions train shared models without moving their data, which makes it a natural fit for health and life sciences research under strict privacy regulation. The methods are maturing fast, but the practical barrier now comes earlier: a team starting a federated project meets a scattered mix of frameworks, governance obligations, and unfamiliar roles, with no structured place to begin that fits its own background. FLKit closes that gap. It is an open, community-maintained onboarding toolkit that takes a multidisciplinary team through the full federated learning lifecycle and gives every contributor, clinical, legal, governance, or technical, a role-aware entry point instead of assuming fluency across all four. We modeled it on the ELIXIR Research Data Management Kit and built it with a multidisciplinary core team, a wider consortium supplying milestone reviews and roadmap direction, and external practitioners interviewed to keep the content grounded in real practice. FLKit sits on four lifecycle stages, Governance, Infrastructure, Wrangling, and Analysis, and connects them through 11 role-specific entry points, a cross-disciplinary glossary, a reusable FAIR-aligned FL Story template for planning and documenting projects, and a curated directory of tools, frameworks, and communities. Since the December 2024 demo it has grown to 39 pages across eight sections, with seven FL Stories documenting completed and ongoing projects in multiple sclerosis disability prediction, inflammatory bowel disease, genomics, and brain-computer interfaces. It is openly available at https://uhasselt-biomedicaldatasciences.github.io/federated-learning-toolkit/ and welcomes contributions from across the life sciences.","author":[{"family":"Pirmani","given":"Ashkan"},{"family":"Vermeulen","given":"Ilse"},{"family":"Vinterhalter","given":"Goran"},{"family":"Geys","given":"Lotte"},{"family":"Faes","given":"Axel"},{"family":"Ali","given":"Muhammad"},{"family":"Sattanathan","given":"Nishkala"},{"family":"Vandeweyer","given":"Geert"},{"family":"Moreau","given":"Yves"},{"family":"Peeters","given":"Liesbet"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.23500","URL":"https://doi.org/10.48550/arxiv.2606.23500","source":"datacite"},{"id":"doi:10.17632/g3tfmstg3v.2","type":"article-journal","title":"Dataset for a Systematic Mapping Study of Deep Learning-Based DDoS Detection in SDN and 5G/B5G Networks (2018–2024)","abstract":"This dataset contains the primary studies used in a systematic mapping study (SMS) on deep learning-based Distributed Denial of Service (DDoS) detection in Software-Defined Networking (SDN), 5G, and beyond-5G (B5G) network environments. The dataset includes 205 peer-reviewed publications published between 2018 and 2024, retrieved from Scopus and IEEE Xplore using a structured search query. Each entry corresponds to a single study and includes bibliographic information (title, authors, year, publication venue etc.), as well as manually curated and derived attributes. In addition to bibliographic metadata, the dataset provides structured annotations extracted from the title, abstract, and manual tags, including: learning_type (Supervised, Unsupervised, Reinforcement, Federated, Not specified; multi-label where applicable) dl_architecture (CNN, LSTM, BiLSTM, GRU, RNN, DNN/MLP, Autoencoder, GAN, Transformer, DBN, Echo State Network, Hybrid, Not specified; multi-label where applicable) network_context (SDN, 5G/6G, IoT/Edge, NFV, Not specified; multi-label where applicable) dataset_used (CICDDoS2019, InSDN, CIC-IDS-2017, CSE-CIC-IDS2018, NSL-KDD, UNSW-NB15, Edge-IIoTset, SDN-SlowRate-DDoS, ToN-IoT, ISCX, CICIoT2023, 5G-NIDD, Other/Custom, Not specified; multi-label where applicable) evaluation_setting (Offline, Testbed-Controller, Real-network, Not specified; multi-label where applicable) The annotations were generated through a semi-automated process combining script-based extraction and manual validation based on titles, abstracts, and manually assigned tags, without full-text analysis. In cases where specific information was not explicitly stated, the corresponding fields were marked as \"Not specified\". This dataset is intended to support reproducibility and transparency of the associated systematic mapping study, as well as to facilitate further research on deep learning-based DDoS detection in programmable networks.","author":[{"family":"Dikaros","given":"Nikolaos"},{"family":"Tzanakakis","given":"Ioannis"},{"family":"Siavvas","given":"Miltiadis"},{"family":"Karagiannidis","given":"George"},{"family":"Ioannou","given":"Konstantinos"}],"issued":{"date-parts":[[2026]]},"DOI":"10.17632/g3tfmstg3v.2","URL":"https://doi.org/10.17632/g3tfmstg3v.2","source":"datacite"},{"id":"doi:10.17632/g3tfmstg3v","type":"article-journal","title":"Dataset for a Systematic Mapping Study of Deep Learning-Based DDoS Detection in SDN and 5G/B5G Networks (2018–2024)","abstract":"This dataset contains the primary studies used in a systematic mapping study (SMS) on deep learning-based Distributed Denial of Service (DDoS) detection in Software-Defined Networking (SDN), 5G, and beyond-5G (B5G) network environments. The dataset includes 205 peer-reviewed publications published between 2018 and 2024, retrieved from Scopus and IEEE Xplore using a structured search query. Each entry corresponds to a single study and includes bibliographic information (title, authors, year, publication venue etc.), as well as manually curated and derived attributes. In addition to bibliographic metadata, the dataset provides structured annotations extracted from the title, abstract, and manual tags, including: learning_type (Supervised, Unsupervised, Reinforcement, Federated, Not specified; multi-label where applicable) dl_architecture (CNN, LSTM, BiLSTM, GRU, RNN, DNN/MLP, Autoencoder, GAN, Transformer, DBN, Echo State Network, Hybrid, Not specified; multi-label where applicable) network_context (SDN, 5G/6G, IoT/Edge, NFV, Not specified; multi-label where applicable) dataset_used (CICDDoS2019, InSDN, CIC-IDS-2017, CSE-CIC-IDS2018, NSL-KDD, UNSW-NB15, Edge-IIoTset, SDN-SlowRate-DDoS, ToN-IoT, ISCX, CICIoT2023, 5G-NIDD, Other/Custom, Not specified; multi-label where applicable) evaluation_setting (Offline, Testbed-Controller, Real-network, Not specified; multi-label where applicable) The annotations were generated through a semi-automated process combining script-based extraction and manual validation based on titles, abstracts, and manually assigned tags, without full-text analysis. In cases where specific information was not explicitly stated, the corresponding fields were marked as \"Not specified\". This dataset is intended to support reproducibility and transparency of the associated systematic mapping study, as well as to facilitate further research on deep learning-based DDoS detection in programmable networks.","author":[{"family":"Dikaros","given":"Nikolaos"},{"family":"Tzanakakis","given":"Ioannis"},{"family":"Siavvas","given":"Miltiadis"},{"family":"Karagiannidis","given":"George"},{"family":"Ioannou","given":"Konstantinos"}],"issued":{"date-parts":[[2026]]},"DOI":"10.17632/g3tfmstg3v","URL":"https://doi.org/10.17632/g3tfmstg3v","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.06332","type":"manuscript","title":"Aggregating Low Rank Adapters in Federated Fine-tuning","abstract":"Fine-tuning large language models requires high computational and memory resources, and is therefore associated with significant costs. When training on federated datasets, an increased communication effort is also needed. For this reason, parameter-efficient methods (PEFT) are becoming increasingly important. In this context, very good results have already been achieved by fine-tuning with low-rank adaptation methods (LoRA). The application of LoRA methods in Federated Learning, and especially the aggregation of adaptation matrices, is a current research field. In this article, we propose a novel aggregation method and compare it with different existing aggregation methods of low rank adapters trained in a federated fine-tuning of large machine learning models and evaluate their performance with respect to selected GLUE benchmark datasets.","author":[{"family":"Trautmann","given":"Evelyn"},{"family":"Hales","given":"Ian"},{"family":"Volk","given":"Martin"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.06332","URL":"https://doi.org/10.48550/arxiv.2501.06332","source":"datacite"},{"id":"doi:10.5281/zenodo.19852082","type":"article-journal","title":"The Federated Tumor Segmentation Challenge 2027","abstract":"International challenges have become the standard for validation of biomedical image analysis methods. We argue, though, that the actual performance even of the winning algorithms on ``real-world`` clinical data often remains unclear, as the data included in these challenges are usually acquired in very controlled settings at few institutions. The seemingly obvious solution of just collecting increasingly more data from more institutions in such challenges does not scale well due to privacy and ownership hurdles. We build upon the first-ever proposed federated learning challenge, Federated Tumor Segmentation (FeTS) 2021 and its follow-up in 2022 (published in Nature Communications), and in 2024 (published in MELBA), intending to address these hurdles, for the creation of tumor segmentation models. Specifically, the FeTS 2027 challenge will use clinically acquired, multi-institutional multi-parametric magnetic resonance imaging (mpMRI) scans from the BraTS 2025 Lighthouse challenge. The FeTS 2027 challenge focuses on innovating at the level of federated aggregation where locally trained models combine to form the consensus model for the segmentation of intrinsically heterogeneous (in appearance, shape, and histology) brain tumors, namely gliomas and meningiomas both in the pre-operative and post-operative setting. Compared to the BraTS 2025 Lighthouse challenge, the ultimate goal of the FeTS challenge is the creation of a consensus segmentation model that has gained knowledgefrom data of multiple institutions without pooling their data together (i.e., by retaining the data within each institution). Since the conception of the FeTS 2021 and its conduction of FeTS 2022 and 2024, Federated Learning has matured to a more active research field in biomedical AI. What separates FeTS 2027 challenge from those of previous years is that in 2027 we plan to further broaden the challenge in 3 ways: 1) We completely change the software infrastructure of the challenge and move from the previously custom code and OpenFL to NVIDIA FLARE, an enterprise-grade federated learning framework that streamlines development and supports an efficient transition from research prototypes to realworld deployment; 2) following the success of the BraTS 2025 lighthouse challenge, FeTS 2027 moves from purely preoperative MRI brain tumor scans to a combination of both pre-operative and post-operative settings, including resectioncavities; 3) the generalizability of the participants' aggregation methods will be evaluated beyond the challenge's segmentation task that the participants have access to, to a hidden (to the participants) task. This added evaluation will be a significant part of the final challenge paper which will provide detailed meta-analysis and provide furtherinsights about the developed aggregation methods. For fairness, since the participants will only have access to the segmentation data, this added evaluation on a new task will not be considered for the ranking. We will however announce performance on this hidden task during the challenge results presentation at MICCAI.","author":[{"family":"Elbatel","given":"Marawan"},{"family":"Yassin","given":"Aya"},{"family":"Li","given":"Xiaomeng"},{"family":"Mao","given":"Jiaji"},{"family":"Shen","given":"Jun"},{"family":"Ghonim","given":"Mohanad"},{"family":"Ghonim","given":"Mohamed"},{"family":"Tantawy","given":"Salma"},{"family":"Tamer","given":"Mariam"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19852082","URL":"https://doi.org/10.5281/zenodo.19852082","source":"datacite"},{"id":"doi:10.5281/zenodo.19852083","type":"article-journal","title":"The Federated Tumor Segmentation Challenge 2027","abstract":"International challenges have become the standard for validation of biomedical image analysis methods. We argue, though, that the actual performance even of the winning algorithms on ``real-world`` clinical data often remains unclear, as the data included in these challenges are usually acquired in very controlled settings at few institutions. The seemingly obvious solution of just collecting increasingly more data from more institutions in such challenges does not scale well due to privacy and ownership hurdles. We build upon the first-ever proposed federated learning challenge, Federated Tumor Segmentation (FeTS) 2021 and its follow-up in 2022 (published in Nature Communications), and in 2024 (published in MELBA), intending to address these hurdles, for the creation of tumor segmentation models. Specifically, the FeTS 2027 challenge will use clinically acquired, multi-institutional multi-parametric magnetic resonance imaging (mpMRI) scans from the BraTS 2025 Lighthouse challenge. The FeTS 2027 challenge focuses on innovating at the level of federated aggregation where locally trained models combine to form the consensus model for the segmentation of intrinsically heterogeneous (in appearance, shape, and histology) brain tumors, namely gliomas and meningiomas both in the pre-operative and post-operative setting. Compared to the BraTS 2025 Lighthouse challenge, the ultimate goal of the FeTS challenge is the creation of a consensus segmentation model that has gained knowledgefrom data of multiple institutions without pooling their data together (i.e., by retaining the data within each institution). Since the conception of the FeTS 2021 and its conduction of FeTS 2022 and 2024, Federated Learning has matured to a more active research field in biomedical AI. What separates FeTS 2027 challenge from those of previous years is that in 2027 we plan to further broaden the challenge in 3 ways: 1) We completely change the software infrastructure of the challenge and move from the previously custom code and OpenFL to NVIDIA FLARE, an enterprise-grade federated learning framework that streamlines development and supports an efficient transition from research prototypes to realworld deployment; 2) following the success of the BraTS 2025 lighthouse challenge, FeTS 2027 moves from purely preoperative MRI brain tumor scans to a combination of both pre-operative and post-operative settings, including resectioncavities; 3) the generalizability of the participants' aggregation methods will be evaluated beyond the challenge's segmentation task that the participants have access to, to a hidden (to the participants) task. This added evaluation will be a significant part of the final challenge paper which will provide detailed meta-analysis and provide furtherinsights about the developed aggregation methods. For fairness, since the participants will only have access to the segmentation data, this added evaluation on a new task will not be considered for the ranking. We will however announce performance on this hidden task during the challenge results presentation at MICCAI.","author":[{"family":"Elbatel","given":"Marawan"},{"family":"Yassin","given":"Aya"},{"family":"Li","given":"Xiaomeng"},{"family":"Mao","given":"Jiaji"},{"family":"Shen","given":"Jun"},{"family":"Ghonim","given":"Mohanad"},{"family":"Ghonim","given":"Mohamed"},{"family":"Tantawy","given":"Salma"},{"family":"Tamer","given":"Mariam"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19852083","URL":"https://doi.org/10.5281/zenodo.19852083","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.04772","type":"manuscript","title":"Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 Challenge","abstract":"Developing generalizable surgical AI requires multi-institutional data, yet patient privacy constraints preclude direct data sharing, making Federated Learning (FL) a natural candidate solution. The application of FL to complex, spatiotemporal surgical video data remains largely unbenchmarked. We present the FedSurg Challenge, the first international benchmarking initiative dedicated to FL in surgical vision, evaluated as a proof-of-concept on a multi-center laparoscopic appendectomy dataset (preliminary subset of Appendix300). Three submissions were evaluated on generalization to an unseen center and center-specific adaptation. Centralized and Swarm Learning baselines isolate the contributions of task difficulty and decentralization to observed performance. Even with all data pooled centrally, the task achieved only 26.31\\% F1-score on the unseen center, while decentralized training introduced an additional, separable performance penalty. Temporal modeling emerges as the dominant architectural factor: video-level spatiotemporal models consistently outperformed frame-level approaches regardless of aggregation strategy. Naive local fine-tuning leads to classifier collapse on imbalanced local data; structured personalized FL with parameter-efficient fine-tuning represents a more principled path toward center-specific adaptation. By characterizing current FL limitations through rigorous statistical analysis, this work establishes a methodological reference point for robust, privacy-preserving AI systems in surgical video analysis.","author":[{"family":"Kirchner","given":"Max"},{"family":"Hoffmann","given":"Hanna"},{"family":"Jenke","given":"Alexander"},{"family":"Saldanha","given":"Oliver"},{"family":"Pfeiffer","given":"Kevin"},{"family":"Kanjo","given":"Weam"},{"family":"Alekseenko","given":"Julia"},{"family":"De Boer","given":"Claas"},{"family":"Kolamuri","given":"Santhi"},{"family":"Mazza","given":"Lorenzo"},{"family":"Padoy","given":"Nicolas"},{"family":"Bano","given":"Sophia"},{"family":"Reinke","given":"Annika"},{"family":"Maier-Hein","given":"Lena"},{"family":"Stoyanov","given":"Danail"},{"family":"Kather","given":"Jakob"},{"family":"Kolbinger","given":"Fiona"},{"family":"Bodenstedt","given":"Sebastian"},{"family":"Speidel","given":"Stefanie"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.04772","URL":"https://doi.org/10.48550/arxiv.2510.04772","source":"datacite"},{"id":"doi:10.3390/cancers17213450","type":"article-journal","title":"Federated Learning Architecture for 3D Breast Cancer Image Classification.","abstract":"BACKGROUDS: Breast cancer remains a major global health challenge, with early diagnosis playing a crucial role in improving patient survival rates. Among the available diagnostic techniques, mammography is widely employed for early detection. However, its effectiveness is often constrained by the complexity of image interpretation, which makes automated detection methods increasingly vital. METHODS: In this study, we propose an advanced approach that leverages 3D mammographic imaging and integrates Federated Learning (FL) to enable decentralized, privacy-preserving model training across multiple institutions. To evaluate the effectiveness of this approach, we assess various machine learning models, including Convolutional Neural Networks (CNNs), Transfer Learning architectures (VGG16, VGG19, ResNet50), and AutoEncoders (AEs), using 3D mammographic data. RESULTS: Our results indicate that the CNN model achieves an accuracy of 97.30%, which improves slightly to 97.37% when the model is combined with Federated Learning, highlighting both the predictive performance and privacy-preserving advantages of our method. In contrast, Transfer Learning models and AutoEncoders exhibit lower accuracies that range from 48.83% to 89.24%, revealing their limitations in the context of this specific task. CONCLUSIONS: These findings underscore the effectiveness of the CNN-FL framework as a robust tool for breast cancer detection, showing that this approach offers a promising balance between diagnostic accuracy and data security-two critical factors in medical imaging.","author":[{"family":"Alhussan","given":"Amel"},{"family":"Nhidi","given":"Wiem"},{"family":"Filali","given":"Imen"},{"family":"Benhmida","given":"Faten"},{"family":"Ejbali","given":"Ridha"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/cancers17213450","URL":"https://doi.org/10.3390/cancers17213450","source":"europepmc"},{"id":"doi:10.5281/zenodo.21432910","type":"article-journal","title":"Extraction Workbook — Trustworthy Knowledge Distillation in Federated Learning: A Systematic Review of Proxy Data Strategies","abstract":"This dataset contains the complete data-extraction workbook of the systematic literature review \"Trustworthy Knowledge Distillation in Federated Learning: A Systematic Review of Proxy Data Strategies\". The review, conducted under the Kitchenham/Charters guidelines and the PRISMA 2020 reporting standard, covers 51 primary studies (2020–2025) on knowledge distillation in federated learning using synthetic or otherwise non-private proxy data, retrieved from five databases (IEEE Xplore, ACM Digital Library, Scopus, ScienceDirect, and Springer Link). The workbook (.xlsx) is organized as follows: Study_Register — bibliographic and contextual metadata for all 57 screened full-text studies (51 included, 6 excluded with reasons): identifier, title, year, venue, publication type, DOI, final eligibility, application domain, dataset summary, federated-learning setting, heterogeneity type, number of clients, and code/data availability. RQ1–RQ6 sheets — research-question-specific extraction fields for every study, covering proxy use and ownership (RQ1), generation mechanisms and curation policies (RQ2), effects on communication, privacy, and performance (RQ3), proxy-quality criteria and metrics (RQ4), threat models and defenses (RQ5), and methodological gaps and reproducibility (RQ6). Quality_Assessment — per-study scores for the eight-item quality checklist (Q1–Q8; Yes/Partial/No), with computed total scores (mean 6.70, median 7.00, range 4.0–8.0 across the included corpus). Controlled_Vocabulary — the harmonized terminology enforced across all categorical fields (proxy type, distillation flow, ownership scope, mechanism family, threat type, defense mechanism). Dashboard / README — corpus composition summaries and usage notes. Every extracted field is accompanied by a KeyEvidence entry pointing to the page, figure, or table of the source article that supports it, making each synthesis claim in the review traceable to primary evidence. The workbook is the primary resource for replicating, auditing, or extending the review. The deposit also includes the Python script used to generate all figures of the review directly from the workbook.","author":[{"family":"Freire","given":"Agostinho"},{"family":"De Andrade","given":"João"},{"family":"Silva","given":"Leandro"},{"family":"Lira","given":"Juan"},{"family":"Fisichella","given":"Marco"},{"family":"Fernandes","given":"Bruno"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21432910","URL":"https://doi.org/10.5281/zenodo.21432910","source":"datacite"},{"id":"doi:10.5281/zenodo.21432911","type":"article-journal","title":"Extraction Workbook — Trustworthy Knowledge Distillation in Federated Learning: A Systematic Review of Proxy Data Strategies","abstract":"This dataset contains the complete data-extraction workbook of the systematic literature review \"Trustworthy Knowledge Distillation in Federated Learning: A Systematic Review of Proxy Data Strategies\". The review, conducted under the Kitchenham/Charters guidelines and the PRISMA 2020 reporting standard, covers 51 primary studies (2020–2025) on knowledge distillation in federated learning using synthetic or otherwise non-private proxy data, retrieved from five databases (IEEE Xplore, ACM Digital Library, Scopus, ScienceDirect, and Springer Link). The workbook (.xlsx) is organized as follows: Study_Register — bibliographic and contextual metadata for all 57 screened full-text studies (51 included, 6 excluded with reasons): identifier, title, year, venue, publication type, DOI, final eligibility, application domain, dataset summary, federated-learning setting, heterogeneity type, number of clients, and code/data availability. RQ1–RQ6 sheets — research-question-specific extraction fields for every study, covering proxy use and ownership (RQ1), generation mechanisms and curation policies (RQ2), effects on communication, privacy, and performance (RQ3), proxy-quality criteria and metrics (RQ4), threat models and defenses (RQ5), and methodological gaps and reproducibility (RQ6). Quality_Assessment — per-study scores for the eight-item quality checklist (Q1–Q8; Yes/Partial/No), with computed total scores (mean 6.70, median 7.00, range 4.0–8.0 across the included corpus). Controlled_Vocabulary — the harmonized terminology enforced across all categorical fields (proxy type, distillation flow, ownership scope, mechanism family, threat type, defense mechanism). Dashboard / README — corpus composition summaries and usage notes. Every extracted field is accompanied by a KeyEvidence entry pointing to the page, figure, or table of the source article that supports it, making each synthesis claim in the review traceable to primary evidence. The workbook is the primary resource for replicating, auditing, or extending the review. The deposit also includes the Python script used to generate all figures of the review directly from the workbook.","author":[{"family":"Freire","given":"Agostinho"},{"family":"De Andrade","given":"João"},{"family":"Silva","given":"Leandro"},{"family":"Lira","given":"Juan"},{"family":"Fisichella","given":"Marco"},{"family":"Fernandes","given":"Bruno"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21432911","URL":"https://doi.org/10.5281/zenodo.21432911","source":"datacite"},{"id":"doi:10.5281/zenodo.21353400","type":"article-journal","title":"Agentic AI for Autonomous Arbitrage Trading in Equity and Cryptocurrency Markets: Concepts, Architecture, and Algorithmic Framework","abstract":"Arbitrage trading exploits temporary price discrepancies across financial markets to generate low-risk profit opportunities. However, conventional arbitrage systems primarily rely on predefined rules and lack the intelligence to adapt to rapidly changing market conditions. This chapter proposes an Agentic AI-based framework for autonomous arbitrage trading across equity and cryptocurrency markets. The framework integrates multiple intelligent agents responsible for market scanning, opportunity detection, risk assessment, decision-making, trade execution, monitoring, and continuous learning under the supervision of an Agentic AI Orchestrator. Real-time market indicators, including price spreads, liquidity, volatility, order book depth, momentum, and transaction costs, are analyzed to identify executable arbitrage opportunities while minimizing execution risk. An algorithmic workflow is presented to support autonomous, explainable, and adaptive trading across multiple exchanges. The chapter also discusses future research directions involving reinforcement learning, explainable AI, federated learning, blockchain-based settlement, and large language model-driven autonomous agents. The proposed framework provides a scalable foundation for next-generation intelligent financial trading systems.","author":[{"family":"Chakraborty","given":"Mohuya"},{"family":"Ghosh","given":"Ajanta"},{"family":"Palit","given":"Sudip"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21353400","URL":"https://doi.org/10.5281/zenodo.21353400","source":"datacite"},{"id":"doi:10.5281/zenodo.21353401","type":"article-journal","title":"Agentic AI for Autonomous Arbitrage Trading in Equity and Cryptocurrency Markets: Concepts, Architecture, and Algorithmic Framework","abstract":"Arbitrage trading exploits temporary price discrepancies across financial markets to generate low-risk profit opportunities. However, conventional arbitrage systems primarily rely on predefined rules and lack the intelligence to adapt to rapidly changing market conditions. This chapter proposes an Agentic AI-based framework for autonomous arbitrage trading across equity and cryptocurrency markets. The framework integrates multiple intelligent agents responsible for market scanning, opportunity detection, risk assessment, decision-making, trade execution, monitoring, and continuous learning under the supervision of an Agentic AI Orchestrator. Real-time market indicators, including price spreads, liquidity, volatility, order book depth, momentum, and transaction costs, are analyzed to identify executable arbitrage opportunities while minimizing execution risk. An algorithmic workflow is presented to support autonomous, explainable, and adaptive trading across multiple exchanges. The chapter also discusses future research directions involving reinforcement learning, explainable AI, federated learning, blockchain-based settlement, and large language model-driven autonomous agents. The proposed framework provides a scalable foundation for next-generation intelligent financial trading systems.","author":[{"family":"Chakraborty","given":"Mohuya"},{"family":"Ghosh","given":"Ajanta"},{"family":"Palit","given":"Sudip"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21353401","URL":"https://doi.org/10.5281/zenodo.21353401","source":"datacite"},{"id":"doi:10.5281/zenodo.15397540","type":"article-journal","title":"Federated Learning Applications in Medical Imaging: A Systematic Review Data Extraction Sheet","abstract":"Data Extraction Sheet for a Systematic Literature Review on Federated Learning in Medical Imaging (2025) Summary:This dataset presents a comprehensive extraction of peer-reviewed publications used in a Systematic Literature Review (SLR) focusing on the application of Federated Learning (FL) in medical imaging. The review was conducted in accordance with established SLR protocols, aiming to assess the current landscape, technical approaches, challenges, and future directions of FL in medical image analysis. The data extraction sheet includes 254 studies, each carefully analyzed across 25 core attributes. The sheet was designed to facilitate the structured evaluation of research output between approximately 2020 and 2025, capturing a wide range of information from bibliographic details to technical methodologies and study outcomes. 1. Purpose and Scope The primary aim of this extraction was to map and synthesize knowledge from existing literature to understand how Federated Learning is utilized in medical imaging, particularly focusing on: Disease diagnosis and classification Model architectures and aggregation strategies Privacy-preservation techniques Dataset usage and imaging modalities Study limitations, results, and recommendations The extracted data supports meta-analysis, gap identification, and evidence-based recommendations for future research in the field. 2. Structure and Fields in the Dataset The spreadsheet is structured with 26 columns (including one empty column) and 254 rows, each corresponding to an individual research paper. Below is a description of the most relevant fields: Paper ID: Unique numerical identifier for each entry. Citation: Full reference or title for traceability. Author(s) and Authors+Year: Provides author attribution with a standardized citation format (e.g., “Smith et al. (2023)”). Year of Publication: Indicates the temporal trend in research output. Type of Publication: Distinguishes between journal articles, conference proceedings, and others. Publisher: Captures the publishing entity (e.g., Elsevier, IEEE, Springer). Nature of Study: Describes the study type, including experimental studies (ES), theoretical/conceptual studies (TCS), or applied case studies (ACS). Abstract and Paper Title: Provide concise summaries and titles for reference. 3. Federated Learning-Specific Attributes Key technical attributes were extracted to capture the distinct components of FL applications: FL Problem Identified: Specifies the central problem the study addresses within FL (e.g., data heterogeneity, communication overhead, or privacy concerns). Research Approach/Methodology: Describes the design of the study, empirical evaluation, framework development, or simulations. ML Technique/Model/Architecture: Lists models used such as CNNs, ResNet, MobileNet, or novel FL-specific architectures. FL Aggregation Methods Used: Includes approaches like FedAvg, FedProx, or novel aggregation algorithms. (Data missing in 2 of 254 cases.) FL Framework/Tool Used: Indicates toolkits such as TensorFlow Federated, PySyft, Flower, or custom platforms (missing in 18 cases). Privacy-Preserving Technique: Captures strategies such as Differential Privacy (DP), Homomorphic Encryption (HE), or Secure Multi-Party Computation (SMPC). 4. Medical Imaging Domain-Specific Features Each study was analyzed in terms of its medical imaging context: Image Dataset Used: Specifies public or proprietary datasets (e.g., COVIDx, ChestX-ray14, TCGA). Imaging Modalities: Captures image types like X-rays, CT scans, MRIs, WSIs. Disease/Health Domain: Indicates general medical focus areas like pulmonology, oncology, neurology, dermatology. Specific Disease: Lists targeted conditions such as COVID-19, lung cancer, Alzheimer’s, diabetic retinopathy, etc. 5. Study Outcome Fields To synthesize conclusions and study significance, the following were extracted: Pros/Advantages/Strengths: Highlights benefits such as privacy preservation, improved generalizability","author":[{"family":"Abayomi-Alli","given":"Adebayo"},{"family":"Rocha","given":"Artur"},{"family":"Abayomi-Alli","given":"Olusola"},{"family":"Aguiar","given":"Ademar"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15397540","URL":"https://doi.org/10.5281/zenodo.15397540","source":"datacite"},{"id":"doi:10.5281/zenodo.15397541","type":"article-journal","title":"Federated Learning Applications in Medical Imaging: A Systematic Review Data Extraction Sheet","abstract":"Data Extraction Sheet for a Systematic Literature Review on Federated Learning in Medical Imaging (2025) Summary:This dataset presents a comprehensive extraction of peer-reviewed publications used in a Systematic Literature Review (SLR) focusing on the application of Federated Learning (FL) in medical imaging. The review was conducted in accordance with established SLR protocols, aiming to assess the current landscape, technical approaches, challenges, and future directions of FL in medical image analysis. The data extraction sheet includes 254 studies, each carefully analyzed across 25 core attributes. The sheet was designed to facilitate the structured evaluation of research output between approximately 2020 and 2025, capturing a wide range of information from bibliographic details to technical methodologies and study outcomes. 1. Purpose and Scope The primary aim of this extraction was to map and synthesize knowledge from existing literature to understand how Federated Learning is utilized in medical imaging, particularly focusing on: Disease diagnosis and classification Model architectures and aggregation strategies Privacy-preservation techniques Dataset usage and imaging modalities Study limitations, results, and recommendations The extracted data supports meta-analysis, gap identification, and evidence-based recommendations for future research in the field. 2. Structure and Fields in the Dataset The spreadsheet is structured with 26 columns (including one empty column) and 254 rows, each corresponding to an individual research paper. Below is a description of the most relevant fields: Paper ID: Unique numerical identifier for each entry. Citation: Full reference or title for traceability. Author(s) and Authors+Year: Provides author attribution with a standardized citation format (e.g., “Smith et al. (2023)”). Year of Publication: Indicates the temporal trend in research output. Type of Publication: Distinguishes between journal articles, conference proceedings, and others. Publisher: Captures the publishing entity (e.g., Elsevier, IEEE, Springer). Nature of Study: Describes the study type, including experimental studies (ES), theoretical/conceptual studies (TCS), or applied case studies (ACS). Abstract and Paper Title: Provide concise summaries and titles for reference. 3. Federated Learning-Specific Attributes Key technical attributes were extracted to capture the distinct components of FL applications: FL Problem Identified: Specifies the central problem the study addresses within FL (e.g., data heterogeneity, communication overhead, or privacy concerns). Research Approach/Methodology: Describes the design of the study, empirical evaluation, framework development, or simulations. ML Technique/Model/Architecture: Lists models used such as CNNs, ResNet, MobileNet, or novel FL-specific architectures. FL Aggregation Methods Used: Includes approaches like FedAvg, FedProx, or novel aggregation algorithms. (Data missing in 2 of 254 cases.) FL Framework/Tool Used: Indicates toolkits such as TensorFlow Federated, PySyft, Flower, or custom platforms (missing in 18 cases). Privacy-Preserving Technique: Captures strategies such as Differential Privacy (DP), Homomorphic Encryption (HE), or Secure Multi-Party Computation (SMPC). 4. Medical Imaging Domain-Specific Features Each study was analyzed in terms of its medical imaging context: Image Dataset Used: Specifies public or proprietary datasets (e.g., COVIDx, ChestX-ray14, TCGA). Imaging Modalities: Captures image types like X-rays, CT scans, MRIs, WSIs. Disease/Health Domain: Indicates general medical focus areas like pulmonology, oncology, neurology, dermatology. Specific Disease: Lists targeted conditions such as COVID-19, lung cancer, Alzheimer’s, diabetic retinopathy, etc. 5. Study Outcome Fields To synthesize conclusions and study significance, the following were extracted: Pros/Advantages/Strengths: Highlights benefits such as privacy preservation, improved generalizability","author":[{"family":"Abayomi-Alli","given":"Adebayo"},{"family":"Rocha","given":"Artur"},{"family":"Abayomi-Alli","given":"Olusola"},{"family":"Aguiar","given":"Ademar"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15397541","URL":"https://doi.org/10.5281/zenodo.15397541","source":"datacite"},{"id":"doi:10.5281/zenodo.20138488","type":"article-journal","title":"Comparative Review of Offline First AI Architectures for Smart Systems in Low Connectivity and Disaster Prone Environments","abstract":"Smart systems that rely on the cloud become unusable when the network goes down in rural and disaster vulnerable settings, thus interrupting the monitoring system as well as slowing the timely responding to emergencies. This work sought to find offline-first AI systems that can be reliable and power efficient (when the network is unavailable or unreliable). Technology or Method: A systematic review of peer-reviewed articles in the period 2015-2025 compared three big architectures, namely TinyML on microcontrollers, Edge AI on single board computers, and federated learning frameworks. The comparisons were on inference latency, energy use, computational ability, reliability and cost of deployment in the application of agriculture, disaster observation and rural health care. Conclusions: TinyML showed milliwatt operation which allowed constant battery, even at the lowest levels, means of operation, but due to the model complexity, there was a limitation. Edge AI platforms were offering less computation time inference through tasks that demanded more computation (real time object detection) but with high energy consumption. Federated learning was effective in enhancing privacy of data, and adversarial distributed model refinement without the transmission of raw data, but its performance was limited in situations of long time disconnection because of its reliance on model updates, which takes place on distributed nodes, and synchronizes only in the case of connectivity, which is intermittently available. This paper seeks to do a comparative review of these offline first AI-based systems, assessing how suitably each one can be completed to fit the smart systems themselves that are in low connectivity and disaster prone environments.","author":[{"family":"Sarait","given":"John"},{"family":"Mencias","given":"Leen"},{"family":"Espela","given":"Leander"},{"family":"Esparcia","given":"Jay"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20138488","URL":"https://doi.org/10.5281/zenodo.20138488","source":"datacite"},{"id":"doi:10.5281/zenodo.20138489","type":"article-journal","title":"Comparative Review of Offline First AI Architectures for Smart Systems in Low Connectivity and Disaster Prone Environments","abstract":"Smart systems that rely on the cloud become unusable when the network goes down in rural and disaster vulnerable settings, thus interrupting the monitoring system as well as slowing the timely responding to emergencies. This work sought to find offline-first AI systems that can be reliable and power efficient (when the network is unavailable or unreliable). Technology or Method: A systematic review of peer-reviewed articles in the period 2015-2025 compared three big architectures, namely TinyML on microcontrollers, Edge AI on single board computers, and federated learning frameworks. The comparisons were on inference latency, energy use, computational ability, reliability and cost of deployment in the application of agriculture, disaster observation and rural health care. Conclusions: TinyML showed milliwatt operation which allowed constant battery, even at the lowest levels, means of operation, but due to the model complexity, there was a limitation. Edge AI platforms were offering less computation time inference through tasks that demanded more computation (real time object detection) but with high energy consumption. Federated learning was effective in enhancing privacy of data, and adversarial distributed model refinement without the transmission of raw data, but its performance was limited in situations of long time disconnection because of its reliance on model updates, which takes place on distributed nodes, and synchronizes only in the case of connectivity, which is intermittently available. This paper seeks to do a comparative review of these offline first AI-based systems, assessing how suitably each one can be completed to fit the smart systems themselves that are in low connectivity and disaster prone environments.","author":[{"family":"Sarait","given":"John"},{"family":"Mencias","given":"Leen"},{"family":"Espela","given":"Leander"},{"family":"Esparcia","given":"Jay"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20138489","URL":"https://doi.org/10.5281/zenodo.20138489","source":"datacite"},{"id":"doi:10.5281/zenodo.20619417","type":"article-journal","title":"BioMutFed+: Reproducibility Artifacts (TMC-2025-09-2616)","abstract":"Reproducibility artifacts for \"BioMutFed+: Mutation-Driven Federated Learning for IIoT,\" IEEE Transactions on Mobile Computing (TMC-2025-09-2616). Two Jupyter notebooks reproduce the ten-seed Wisconsin Breast Cancer evaluation (Tables VI and VII) and the lambda sensitivity sweep within the full BioMutFed+ pipeline (Supplementary Section SIII-F, Table S2). All experiments use Dirichlet alpha=0.1 partitioning across 20 clients with 20% gradient-ascent adversarial clients. Each notebook is self-contained and produces the exact numerical values reported in the paper. See README.md for environment details and runtime.","author":[{"family":"Tallat","given":"Raiha"},{"family":"Wang","given":"Xingfu"},{"family":"Hawbani","given":"Ammar"},{"family":"Xiaohua","given":"Xu"},{"family":"Miao","given":"Fuyou"},{"family":"Zhao","given":"Liang"},{"family":"Wang","given":"Jiantao"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20619417","URL":"https://doi.org/10.5281/zenodo.20619417","source":"datacite"},{"id":"doi:10.5281/zenodo.20619418","type":"article-journal","title":"BioMutFed+: Reproducibility Artifacts (TMC-2025-09-2616)","abstract":"Reproducibility artifacts for \"BioMutFed+: Mutation-Driven Federated Learning for IIoT,\" IEEE Transactions on Mobile Computing (TMC-2025-09-2616). Two Jupyter notebooks reproduce the ten-seed Wisconsin Breast Cancer evaluation (Tables VI and VII) and the lambda sensitivity sweep within the full BioMutFed+ pipeline (Supplementary Section SIII-F, Table S2). All experiments use Dirichlet alpha=0.1 partitioning across 20 clients with 20% gradient-ascent adversarial clients. Each notebook is self-contained and produces the exact numerical values reported in the paper. See README.md for environment details and runtime.","author":[{"family":"Tallat","given":"Raiha"},{"family":"Wang","given":"Xingfu"},{"family":"Hawbani","given":"Ammar"},{"family":"Xiaohua","given":"Xu"},{"family":"Miao","given":"Fuyou"},{"family":"Zhao","given":"Liang"},{"family":"Wang","given":"Jiantao"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20619418","URL":"https://doi.org/10.5281/zenodo.20619418","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.00694","type":"manuscript","title":"Forecasting Energy Availability in Local Energy Communities via LSTM Federated Learning","abstract":"Local Energy Communities are emerging as crucial players in the landscape of sustainable development. A significant challenge for these communities is achieving self-sufficiency through effective management of the balance between energy production and consumption. To meet this challenge, it is essential to develop and implement forecasting models that deliver accurate predictions, which can then be utilized by optimization and planning algorithms. However, the application of forecasting solutions is often hindered by privacy constrains and regulations as the users participating in the Local Energy Community can be (rightfully) reluctant sharing their consumption patterns with others. In this context, the use of Federated Learning (FL) can be a viable solution as it allows to create a forecasting model without the need to share privacy sensitive information among the users. In this study, we demonstrate how FL and long short-term memory (LSTM) networks can be employed to achieve this objective, highlighting the trade-off between data sharing and forecasting accuracy.","author":[{"family":"Turazza","given":"Fabio"},{"family":"Pietri","given":"Marcello"},{"family":"Hadjidimitriou","given":"Natalia"},{"family":"Mamei","given":"Marco"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.00694","URL":"https://doi.org/10.48550/arxiv.2602.00694","source":"datacite"},{"id":"doi:10.5281/zenodo.18316378","type":"article-journal","title":"Scripts and mappings to harmonise dementia cohorts into the OMOP Common Data Model","abstract":"This resource contains the scripts, mappings, and other materials that were used to harmonise data from nine Dutch cohorts relevant to dementia research to the OMOP (Observational Medical Outcomes Partnership) Common Data Model (CDM). These materials were developed within the Netherlands Consortium of Dementia Cohorts (NCDC) project. NCDC included the following nine cohorts: Amsterdam Dementia Cohort, Doetinchem Cohort Study, EMIF-AD 90+ study, EMIF-AD PreclinAD study, Leiden Longevity Study, Longitudinal Aging Study Amsterdam, The Maastricht Study, Rotterdam Study, and SMART. Scope The materials in this publication were used to transform the data from the NCDC cohorts into the OMOP CDM. However, the overall approach and workflow are transferable to other cohort harmonisation efforts. As such, this work can serve as a reference implementation or starting point for the harmonisation of cohort data to any common data model. Harmonising data to a common data model is crucial in (dementia) research as it allows to combine studies, increase statistical power, and enhance the reliability of findings. Reuse of this resource is governed by the license specified in this Zenodo record and the associated GitHub repository. Who is it for? This work is primarily intended for researchers and (research) software engineers looking to harmonise multi-centre cohort data into a common data model. Effective reuse of this resource requires technical expertise in data transformation, familiarity with common data models, and in-depth domain knowledge of the underlying cohort data. It is also useful for researchers who want to work with data that are structured according to the OMOP CDM, or with data from (one of) the nine cohorts that were part of NCDC. What does it include? The resource includes materials covering all steps required to transform cohort data into the OMOP CDM, including the design of mappings and the execution of the ETL (Extract, Transform, Load) process. The first step in the harmonisation process is the collection of metadata. Variables from all cohort studies are identified, after which a set of harmonised variables is defined and mapped to OMOP concepts (the destination mapping). The variables of each individual cohort are then mapped to these harmonised variables (the source mapping). The folder examples/ contains an example destination mapping, source mapping, and dataset, and can be used to guide users in this initial harmonisation step The folder ncdc_mappings/ contains the destination mapping and source mappings for the nine NCDC cohorts In the subsequent step, the destination and source mappings are used to transform the cohort data and to set up and populate a database according to the common data model. The folder cdm_parser/ contains scripts that create and populate a PostgreSQL database based on the OMOP CDM. The scripts accept file-based datasets in CSV, SPSS, or SAS (Statistical Analysis Software) formats as input The folder scripts/ contains scripts that generate summary statistics to assess the correctness and completeness of the data transformation The subfolder examples/data-retrieval/ contains examples illustrating how to extract data from and query an OMOP CDM database. These examples can be used as an educational resource or as a starting point for researchers working with OMOP-formatted data. Further usage instructions can be found in the included README file, and software requirements are listed in the requirements file. Associated materials This repository was developed for the publication: Mateus P, Moonen J, Beran M, Jaarsma E, van der Landen SM, Heuvelink J, Birhanu M, Harms AGJ, Bron E, Wolters FJ, Cats D, Mei H, Oomens J, Jansen W, Schram MT, Dekker A, Bermejo I. Data harmonization and federated learning for multi-cohort dementia research using the OMOP common data model: A Netherlands consortium of dementia cohorts case study. J Biomed Inform. 2024 Jul;155:104661. doi: 10.1016/j.jbi.2024.104661. Epub","author":[{"family":"Da Costa Mateus","given":"Pedro"},{"family":"Moonen","given":"Justine"},{"family":"Beran","given":"Magdalena"},{"family":"Jaarsma","given":"Eva"},{"family":"Van Der Landen","given":"Sophie"},{"family":"Heuvelink","given":"Joost"},{"family":"Mahlet","given":"Birhanu"},{"family":"Harms","given":"Alexander"},{"family":"Bron","given":"Esther"},{"family":"Wolters","given":"Frank"},{"family":"Cats","given":"Davy"},{"family":"Mei","given":"Hailiang"},{"family":"Oomens","given":"Julie"},{"family":"Jansen","given":"Willemijn"},{"family":"Schram","given":"Miranda"},{"family":"Dekker","given":"Andre"},{"family":"Bermejo","given":"Iñigo"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18316378","URL":"https://doi.org/10.5281/zenodo.18316378","source":"datacite"},{"id":"doi:10.5281/zenodo.18316379","type":"article-journal","title":"Scripts and mappings to harmonise dementia cohorts into the OMOP Common Data Model","abstract":"This resource contains the scripts, mappings, and other materials that were used to harmonise data from nine Dutch cohorts relevant to dementia research to the OMOP (Observational Medical Outcomes Partnership) Common Data Model (CDM). These materials were developed within the Netherlands Consortium of Dementia Cohorts (NCDC) project. NCDC included the following nine cohorts: Amsterdam Dementia Cohort, Doetinchem Cohort Study, EMIF-AD 90+ study, EMIF-AD PreclinAD study, Leiden Longevity Study, Longitudinal Aging Study Amsterdam, The Maastricht Study, Rotterdam Study, and SMART. Scope The materials in this publication were used to transform the data from the NCDC cohorts into the OMOP CDM. However, the overall approach and workflow are transferable to other cohort harmonisation efforts. As such, this work can serve as a reference implementation or starting point for the harmonisation of cohort data to any common data model. Harmonising data to a common data model is crucial in (dementia) research as it allows to combine studies, increase statistical power, and enhance the reliability of findings. Reuse of this resource is governed by the license specified in this Zenodo record and the associated GitHub repository. Who is it for? This work is primarily intended for researchers and (research) software engineers looking to harmonise multi-centre cohort data into a common data model. Effective reuse of this resource requires technical expertise in data transformation, familiarity with common data models, and in-depth domain knowledge of the underlying cohort data. It is also useful for researchers who want to work with data that are structured according to the OMOP CDM, or with data from (one of) the nine cohorts that were part of NCDC. What does it include? The resource includes materials covering all steps required to transform cohort data into the OMOP CDM, including the design of mappings and the execution of the ETL (Extract, Transform, Load) process. The first step in the harmonisation process is the collection of metadata. Variables from all cohort studies are identified, after which a set of harmonised variables is defined and mapped to OMOP concepts (the destination mapping). The variables of each individual cohort are then mapped to these harmonised variables (the source mapping). The folder examples/ contains an example destination mapping, source mapping, and dataset, and can be used to guide users in this initial harmonisation step The folder ncdc_mappings/ contains the destination mapping and source mappings for the nine NCDC cohorts In the subsequent step, the destination and source mappings are used to transform the cohort data and to set up and populate a database according to the common data model. The folder cdm_parser/ contains scripts that create and populate a PostgreSQL database based on the OMOP CDM. The scripts accept file-based datasets in CSV, SPSS, or SAS (Statistical Analysis Software) formats as input The folder scripts/ contains scripts that generate summary statistics to assess the correctness and completeness of the data transformation The subfolder examples/data-retrieval/ contains examples illustrating how to extract data from and query an OMOP CDM database. These examples can be used as an educational resource or as a starting point for researchers working with OMOP-formatted data. Further usage instructions can be found in the included README file, and software requirements are listed in the requirements file. Associated materials This repository was developed for the publication: Mateus P, Moonen J, Beran M, Jaarsma E, van der Landen SM, Heuvelink J, Birhanu M, Harms AGJ, Bron E, Wolters FJ, Cats D, Mei H, Oomens J, Jansen W, Schram MT, Dekker A, Bermejo I. Data harmonization and federated learning for multi-cohort dementia research using the OMOP common data model: A Netherlands consortium of dementia cohorts case study. J Biomed Inform. 2024 Jul;155:104661. doi: 10.1016/j.jbi.2024.104661. Epub","author":[{"family":"Da Costa Mateus","given":"Pedro"},{"family":"Moonen","given":"Justine"},{"family":"Beran","given":"Magdalena"},{"family":"Jaarsma","given":"Eva"},{"family":"Van Der Landen","given":"Sophie"},{"family":"Heuvelink","given":"Joost"},{"family":"Mahlet","given":"Birhanu"},{"family":"Harms","given":"Alexander"},{"family":"Bron","given":"Esther"},{"family":"Wolters","given":"Frank"},{"family":"Cats","given":"Davy"},{"family":"Mei","given":"Hailiang"},{"family":"Oomens","given":"Julie"},{"family":"Jansen","given":"Willemijn"},{"family":"Schram","given":"Miranda"},{"family":"Dekker","given":"Andre"},{"family":"Bermejo","given":"Iñigo"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18316379","URL":"https://doi.org/10.5281/zenodo.18316379","source":"datacite"},{"id":"doi:10.5281/zenodo.18453954","type":"article-journal","title":"Federated learning for privacy-preserving, secure and scalable data intelligence in hybrid cloud systems","abstract":"The convergence of federated learning and hybrid cloud computing represents a transformative paradigm for privacy-preserving data intelligence. This review examines federated learning implementations in hybrid cloud environments, analyzing security mechanisms, privacy-preserving capabilities, and scalability challenges. We explore architectural frameworks and deployment strategies while analyzing security and privacy challenges from technical, organizational, and regulatory perspectives. The study highlights synergistic benefits of combining federated learning with hybrid cloud infrastructure and discusses emerging trends including homomorphic encryption, differential privacy, and blockchain integration. Through comprehensive literature analysis of publications from 2016 to 2024, key findings reveal that federated learning in hybrid clouds offers unprecedented opportunities for privacy-preserving analytics while introducing unique challenges in communication efficiency and cross-environment orchestration. Organizations can effectively leverage federated learning by implementing layered security architectures and maintaining continuous adaptation to evolving privacy regulations. This analysis provides valuable insights for practitioners and researchers navigating the intersection of federated learning and hybrid cloud computing.","author":[{"family":"Ezeakile","given":"Emmanuel"},{"family":"Disu","given":"Abdulateef"},{"family":"Alabi","given":"Cynthia"},{"family":"Mustapha","given":"Toyosi"},{"family":"Odewale","given":"Moses"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18453954","URL":"https://doi.org/10.5281/zenodo.18453954","source":"datacite"},{"id":"doi:10.5281/zenodo.18453953","type":"article-journal","title":"Federated learning for privacy-preserving, secure and scalable data intelligence in hybrid cloud systems","abstract":"The convergence of federated learning and hybrid cloud computing represents a transformative paradigm for privacy-preserving data intelligence. This review examines federated learning implementations in hybrid cloud environments, analyzing security mechanisms, privacy-preserving capabilities, and scalability challenges. We explore architectural frameworks and deployment strategies while analyzing security and privacy challenges from technical, organizational, and regulatory perspectives. The study highlights synergistic benefits of combining federated learning with hybrid cloud infrastructure and discusses emerging trends including homomorphic encryption, differential privacy, and blockchain integration. Through comprehensive literature analysis of publications from 2016 to 2024, key findings reveal that federated learning in hybrid clouds offers unprecedented opportunities for privacy-preserving analytics while introducing unique challenges in communication efficiency and cross-environment orchestration. Organizations can effectively leverage federated learning by implementing layered security architectures and maintaining continuous adaptation to evolving privacy regulations. This analysis provides valuable insights for practitioners and researchers navigating the intersection of federated learning and hybrid cloud computing.","author":[{"family":"Ezeakile","given":"Emmanuel"},{"family":"Disu","given":"Abdulateef"},{"family":"Alabi","given":"Cynthia"},{"family":"Mustapha","given":"Toyosi"},{"family":"Odewale","given":"Moses"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18453953","URL":"https://doi.org/10.5281/zenodo.18453953","source":"datacite"},{"id":"doi:10.48550/arxiv.2601.07901","type":"manuscript","title":"Decentralized Online Convex Optimization with Unknown Feedback Delays","abstract":"Decentralized online convex optimization (D-OCO), where multiple agents within a network collaboratively learn optimal decisions in real-time, arises naturally in applications such as federated learning, sensor networks, and multi-agent control. In this paper, we study D-OCO under unknown, time-and agent-varying feedback delays. While recent work has addressed this problem (Nguyen et al., 2024), existing algorithms assume prior knowledge of the total delay over agents and still suffer from suboptimal dependence on both the delay and network parameters. To overcome these limitations, we propose a novel algorithm that achieves an improved regret bound of O N $\\sqrt$ d tot + N $\\sqrt$ T (1-$σ$2) 1/4 , where T is the total horizon, d tot denotes the average total delay across agents, N is the number of agents, and 1 -$σ$ 2 is the spectral gap of the network. Our approach builds upon recent advances in D-OCO (Wan et al., 2024a), but crucially incorporates an adaptive learning rate mechanism via a decentralized communication protocol. This enables each agent to estimate delays locally using a gossip-based strategy without the prior knowledge of the total delay. We further extend our framework to the strongly convex setting and derive a sharper regret bound of O N $δ$max ln T $α$ , where $α$ is the strong convexity parameter and $δ$ max is the maximum number of missing observations averaged over agents. We also show that our upper bounds for both settings are tight up to logarithmic factors. Experimental results validate the effectiveness of our approach, showing improvements over existing benchmark algorithms.","author":[{"family":"Qiu","given":"Hao"},{"family":"Zhang","given":"Mengxiao"},{"family":"Achddou","given":"Juliette"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2601.07901","URL":"https://doi.org/10.48550/arxiv.2601.07901","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.20610","type":"manuscript","title":"FedPOD: the deployable units of training for federated learning","abstract":"This paper proposes FedPOD, which ranked first in the 2024 Federated Tumor Segmentation (FeTS) Challenge, for optimizing learning efficiency and communication cost in federated learning among multiple clients. Inspired by FedPIDAvg, we define a round-wise task for FedPOD to enhance training efficiency. FedPIDAvg achieved performance improvement by incorporating the training loss reduction for prediction entropy as weights using differential terms. Furthermore, by modeling data distribution with a Poisson distribution and using a PID controller, it reduced communication costs even in skewed data distribution. However, excluding participants classified as outliers based on the Poisson distribution can limit data utilization. Additionally, PID controller requires the same participants to be maintained throughout the federated learning process as it uses previous rounds' learning information in the current round. In our approach, FedPOD addresses these issues by including participants excluded as outliers, eliminating dependency on previous rounds' learning information, and applying a method for calculating validation loss at each round. In this challenge, FedPOD presents comparable performance to FedPIDAvg in metrics of Dice score, 0.78, 0.71 and 0.72 for WT, ET and TC in average, and projected convergence score, 0.74 in average. Furthermore, the concept of FedPOD draws inspiration from Kubernetes' smallest computing unit, POD, designed to be compatible with Kubernetes auto-scaling. Extending round-wise tasks of FedPOD to POD units allows flexible design by applying scale-out similar to Kubernetes' auto-scaling. This work demonstrated the potentials of FedPOD to enhance federated learning by improving efficiency, flexibility, and performance in metrics.","author":[{"family":"Kim","given":"Daewoon"},{"family":"Yie","given":"Si"},{"family":"Lee","given":"Jae"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.20610","URL":"https://doi.org/10.48550/arxiv.2512.20610","source":"datacite"},{"id":"doi:10.3217/m750p-pqw83","type":"article-journal","title":"Accessibility for academic teaching staff (Report 05/2025)","abstract":"The report describes a pilot “Video course on accessibility” created by Wrocław University of Science and Technology (Wrocław Tech) for academic teaching staff, including laboratory instructors. The course links the university’s local Moodle‑based LMS (ePortal PWr) with the Unite! Metacampus platform via Learning Tools Interoperability (LTI). The self‑study course comprises seven short videos covering mental‑health crises, autism, visual and hearing impairments, physical disability, and invisible disability. Each video is followed by a mandatory knowledge‑check question; a final five‑question test awards a certificate of completion. Total duration is about 50 minutes. Technical integration uses eduGAIN single‑sign‑on, with ePortal acting as the LTI tool provider and Metacampus as the LTI consumer. Secure credential exchange enables automatic grade pass‑back and synchronized progress, while configuration responsibilities were delegated to ePortal administrators to simplify the instructor’s role. The pilot enrolled the target of ≥ 50 academic staff. Survey feedback shows unanimous agreement that the course enhanced understanding of disability needs, awareness of support mechanisms, and confidence in implementing inclusive teaching practices. Participants valued the practical, student‑driven examples and intuitive navigation, though they suggested making the final test visible earlier, improving promotion, and optimizing the interface for small screens. Key lessons highlight the benefit of delegating LTI setup, the need for clear communication of course milestones, and the importance of responsive design. Recommendations include scheduling future trainings during high‑availability periods, offering introductory webinars, expanding interactive video tools, and providing richer LTI‑configuration guides for external admins. The series “Unite! Digital Teaching and Learning Success Story Report” was initiated by Cm.2 Digital Campus, led by TU Graz (Martin Ebner) as part of the idea to collect, spread and enhance the possibilities, usages and lessons learned of teaching and learning offers using the federated learning management system Metacampus in 11/2024 till the end of the current Erasmus+ funding period 10/2026. All reports are available under open license and originally published at the TU Graz repository. Copyright holders are the authors.","author":[{"family":"Forycka","given":"Renata"},{"family":"Herczak-Ciara","given":"Agnieszka"},{"family":"Jach","given":"Katarzyna"},{"family":"Krysiak","given":"Jarosław"},{"family":"Kosela","given":"Anna"},{"family":"Zawada","given":"Justyna"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3217/m750p-pqw83","URL":"https://doi.org/10.3217/m750p-pqw83","source":"datacite"},{"id":"doi:10.3217/163xg-5h074","type":"article-journal","title":"Accessibility for academic teaching staff (Report 05/2025)","abstract":"The report describes a pilot “Video course on accessibility” created by Wrocław University of Science and Technology (Wrocław Tech) for academic teaching staff, including laboratory instructors. The course links the university’s local Moodle‑based LMS (ePortal PWr) with the Unite! Metacampus platform via Learning Tools Interoperability (LTI). The self‑study course comprises seven short videos covering mental‑health crises, autism, visual and hearing impairments, physical disability, and invisible disability. Each video is followed by a mandatory knowledge‑check question; a final five‑question test awards a certificate of completion. Total duration is about 50 minutes. Technical integration uses eduGAIN single‑sign‑on, with ePortal acting as the LTI tool provider and Metacampus as the LTI consumer. Secure credential exchange enables automatic grade pass‑back and synchronized progress, while configuration responsibilities were delegated to ePortal administrators to simplify the instructor’s role. The pilot enrolled the target of ≥ 50 academic staff. Survey feedback shows unanimous agreement that the course enhanced understanding of disability needs, awareness of support mechanisms, and confidence in implementing inclusive teaching practices. Participants valued the practical, student‑driven examples and intuitive navigation, though they suggested making the final test visible earlier, improving promotion, and optimizing the interface for small screens. Key lessons highlight the benefit of delegating LTI setup, the need for clear communication of course milestones, and the importance of responsive design. Recommendations include scheduling future trainings during high‑availability periods, offering introductory webinars, expanding interactive video tools, and providing richer LTI‑configuration guides for external admins. The series “Unite! Digital Teaching and Learning Success Story Report” was initiated by Cm.2 Digital Campus, led by TU Graz (Martin Ebner) as part of the idea to collect, spread and enhance the possibilities, usages and lessons learned of teaching and learning offers using the federated learning management system Metacampus in 11/2024 till the end of the current Erasmus+ funding period 10/2026. All reports are available under open license and originally published at the TU Graz repository. Copyright holders are the authors.","author":[{"family":"Forycka","given":"Renata"},{"family":"Herczak-Ciara","given":"Agnieszka"},{"family":"Jach","given":"Katarzyna"},{"family":"Krysiak","given":"Jarosław"},{"family":"Kosela","given":"Anna"},{"family":"Zawada","given":"Justyna"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3217/163xg-5h074","URL":"https://doi.org/10.3217/163xg-5h074","source":"datacite"},{"id":"doi:10.6084/m9.figshare.28776466","type":"article-journal","title":"Artificial intelligence and its application in clinical microbiology","abstract":"Traditional microbiological diagnostics face challenges in pathogen identification speed and antimicrobial resistance (AMR) evaluation. Artificial intelligence (AI) offers transformative solutions, necessitating a comprehensive review of its applications, advancements, and integration challenges in clinical microbiology. This review examines AI-driven methodologies, including machine learning (ML), deep learning (DL), and convolutional neural networks (CNNs), for enhancing pathogen detection, AMR prediction, and diagnostic imaging. Applications in virology (e.g. COVID-19 RT-PCR optimization), parasitology (e.g. malaria detection), and bacteriology (e.g. automated colony counting) are analyzed. A literature search was conducted using PubMed, Scopus, and Web of Science (2018–2024), prioritizing peer-reviewed studies on AI’s diagnostic accuracy, workflow efficiency, and clinical validation. AI significantly improves diagnostic precision and operational efficiency but requires robust validation to address data heterogeneity, model interpretability, and ethical concerns. Future success hinges on interdisciplinary collaboration to develop standardized, equitable AI tools tailored for global healthcare settings. Advancing explainable AI and federated learning frameworks will be critical for bridging current implementation gaps and maximizing AI’s potential in combating infectious diseases.","author":[{"family":"Mairi","given":"Assia"},{"family":"Hamza","given":"Lamia"},{"family":"Touati","given":"Abdelaziz"}],"issued":{"date-parts":[[2025]]},"DOI":"10.6084/m9.figshare.28776466","URL":"https://doi.org/10.6084/m9.figshare.28776466","source":"datacite"},{"id":"doi:10.6084/m9.figshare.28776466.v1","type":"article-journal","title":"Artificial intelligence and its application in clinical microbiology","abstract":"Traditional microbiological diagnostics face challenges in pathogen identification speed and antimicrobial resistance (AMR) evaluation. Artificial intelligence (AI) offers transformative solutions, necessitating a comprehensive review of its applications, advancements, and integration challenges in clinical microbiology. This review examines AI-driven methodologies, including machine learning (ML), deep learning (DL), and convolutional neural networks (CNNs), for enhancing pathogen detection, AMR prediction, and diagnostic imaging. Applications in virology (e.g. COVID-19 RT-PCR optimization), parasitology (e.g. malaria detection), and bacteriology (e.g. automated colony counting) are analyzed. A literature search was conducted using PubMed, Scopus, and Web of Science (2018–2024), prioritizing peer-reviewed studies on AI’s diagnostic accuracy, workflow efficiency, and clinical validation. AI significantly improves diagnostic precision and operational efficiency but requires robust validation to address data heterogeneity, model interpretability, and ethical concerns. Future success hinges on interdisciplinary collaboration to develop standardized, equitable AI tools tailored for global healthcare settings. Advancing explainable AI and federated learning frameworks will be critical for bridging current implementation gaps and maximizing AI’s potential in combating infectious diseases.","author":[{"family":"Mairi","given":"Assia"},{"family":"Hamza","given":"Lamia"},{"family":"Touati","given":"Abdelaziz"}],"issued":{"date-parts":[[2025]]},"DOI":"10.6084/m9.figshare.28776466.v1","URL":"https://doi.org/10.6084/m9.figshare.28776466.v1","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.06206","type":"manuscript","title":"The MICCAI Federated Tumor Segmentation (FeTS) Challenge 2024: Efficient and Robust Aggregation Methods for Federated Learning","abstract":"We present the design and results of the MICCAI Federated Tumor Segmentation (FeTS) Challenge 2024, which focuses on federated learning (FL) for glioma sub-region segmentation in multi-parametric MRI and evaluates new weight aggregation methods aimed at improving robustness and efficiency. Six participating teams were evaluated using a standardized FL setup and a multi-institutional dataset derived from the BraTS glioma benchmark, consisting of 1,251 training cases, 219 validation cases, and 570 hidden test cases with segmentations for enhancing tumor (ET), tumor core (TC), and whole tumor (WT). Teams were ranked using a cumulative scoring system that considered both segmentation performance, measured by Dice Similarity Coefficient (DSC) and the 95th percentile Hausdorff Distance (HD95), and communication efficiency assessed through the convergence score. A PID-controller-based method achieved the top overall ranking, obtaining mean DSC values of 0.733, 0.761, and 0.751 for ET, TC, and WT, respectively, with corresponding HD95 values of 33.922 mm, 33.623 mm, and 32.309 mm, while also demonstrating the highest communication efficiency with a convergence score of 0.764. These findings advance the state of federated learning for medical imaging, surpassing top-performing methods from previous challenge iterations and highlighting PID controllers as effective mechanisms for stabilizing and optimizing weight aggregation in FL. The challenge code is available at https://github.com/FeTS-AI/Challenge.","author":[{"family":"Linardos","given":"Akis"},{"family":"Pati","given":"Sarthak"},{"family":"Baid","given":"Ujjwal"},{"family":"Edwards","given":"Brandon"},{"family":"Foley","given":"Patrick"},{"family":"Ta","given":"Kevin"},{"family":"Chung","given":"Verena"},{"family":"Sheller","given":"Micah"},{"family":"Khan","given":"Muhammad"},{"family":"Jafaritadi","given":"Mojtaba"},{"family":"Kontio","given":"Elina"},{"family":"Khan","given":"Suleiman"},{"family":"Mächler","given":"Leon"},{"family":"Ezhov","given":"Ivan"},{"family":"Shit","given":"Suprosanna"},{"family":"Paetzold","given":"Johannes"},{"family":"Grimberg","given":"Gustav"},{"family":"Nickel","given":"Manuel"},{"family":"Naccache","given":"David"},{"family":"Siomos","given":"Vasilis"},{"family":"Passerat-Palmbach","given":"Jonathan"},{"family":"Tarroni","given":"Giacomo"},{"family":"Kim","given":"Daewoon"},{"family":"Klausmann","given":"Leonard"},{"family":"Shah","given":"Prashant"},{"family":"Menze","given":"Bjoern"},{"family":"Makris","given":"Dimitrios"},{"family":"Bakas","given":"Spyridon"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.06206","URL":"https://doi.org/10.48550/arxiv.2512.06206","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.03287","type":"manuscript","title":"Multi-Frequency Federated Learning for Human Activity Recognition Using Head-Worn Sensors","abstract":"Human Activity Recognition (HAR) benefits various application domains, including health and elderly care. Traditional HAR involves constructing pipelines reliant on centralized user data, which can pose privacy concerns as they necessitate the uploading of user data to a centralized server. This work proposes multi-frequency Federated Learning (FL) to enable: (1) privacy-aware ML; (2) joint ML model learning across devices with varying sampling frequency. We focus on head-worn devices (e.g., earbuds and smart glasses), a relatively unexplored domain compared to traditional smartwatch- or smartphone-based HAR. Results have shown improvements on two datasets against frequency-specific approaches, indicating a promising future in the multi-frequency FL-HAR task. The proposed network's implementation is publicly available for further research and development.","author":[{"family":"Fenoglio","given":"Dario"},{"family":"Li","given":"Mohan"},{"family":"Casnici","given":"Davide"},{"family":"Laporte","given":"Matias"},{"family":"Gashi","given":"Shkurta"},{"family":"Santini","given":"Silvia"},{"family":"Gjoreski","given":"Martin"},{"family":"Langheinrich","given":"Marc"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.03287","URL":"https://doi.org/10.48550/arxiv.2512.03287","source":"datacite"},{"id":"doi:10.48550/arxiv.2511.22616","type":"manuscript","title":"Federated Learning Survey: A Multi-Level Taxonomy of Aggregation Techniques, Experimental Insights, and Future Frontiers","abstract":"The integration of IoT and AI has unlocked innovation across industries, but growing privacy concerns and data isolation hinder progress. Traditional centralized ML struggles to overcome these challenges, which has led to the rise of Federated Learning (FL), a decentralized paradigm that enables collaborative model training without sharing local raw data. FL ensures data privacy, reduces communication overhead, and supports scalability, yet its heterogeneity adds complexity compared to centralized approaches. This survey focuses on three main FL research directions: personalization, optimization, and robustness, offering a structured classification through a hybrid methodology that combines bibliometric analysis with systematic review to identify the most influential works. We examine challenges and techniques related to heterogeneity, efficiency, security, and privacy, and provide a comprehensive overview of aggregation strategies, including architectures, synchronization methods, and diverse federation objectives. To complement this, we discuss practical evaluation approaches and present experiments comparing aggregation methods under IID and non-IID data distributions. Finally, we outline promising research directions to advance FL, aiming to guide future innovation in this rapidly evolving field.","author":[{"family":"Arbaoui","given":"Meriem"},{"family":"Brahmia","given":"Mohamed"},{"family":"Rahmoun","given":"Abdellatif"},{"family":"Zghal","given":"Mourad"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2511.22616","URL":"https://doi.org/10.48550/arxiv.2511.22616","source":"datacite"},{"id":"doi:10.48550/arxiv.2511.16822","type":"manuscript","title":"A Robust Federated Learning Approach for Combating Attacks Against IoT Systems Under non-IID Challenges","abstract":"In the context of the growing proliferation of user devices and the concurrent surge in data volumes, the complexities arising from the substantial increase in data have posed formidable challenges to conventional machine learning model training. Particularly, this is evident within resource-constrained and security-sensitive environments such as those encountered in networks associated with the Internet of Things (IoT). Federated Learning has emerged as a promising remedy to these challenges by decentralizing model training to edge devices or parties, effectively addressing privacy concerns and resource limitations. Nevertheless, the presence of statistical heterogeneity in non-Independently and Identically Distributed (non-IID) data across different parties poses a significant hurdle to the effectiveness of FL. Many FL approaches have been proposed to enhance learning effectiveness under statistical heterogeneity. However, prior studies have uncovered a gap in the existing research landscape, particularly in the absence of a comprehensive comparison between federated methods addressing statistical heterogeneity in detecting IoT attacks. In this research endeavor, we delve into the exploration of FL algorithms, specifically FedAvg, FedProx, and Scaffold, under different data distributions. Our focus is on achieving a comprehensive understanding of and addressing the challenges posed by statistical heterogeneity. In this study, We classify large-scale IoT attacks by utilizing the CICIoT2023 dataset. Through meticulous analysis and experimentation, our objective is to illuminate the performance nuances of these FL methods, providing valuable insights for researchers and practitioners in the domain.","author":[{"family":"Gad","given":"Eyad"},{"family":"Fadlullah","given":"Zubair"},{"family":"Fouda","given":"Mostafa"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2511.16822","URL":"https://doi.org/10.48550/arxiv.2511.16822","source":"datacite"},{"id":"doi:10.5281/zenodo.17587337","type":"article-journal","title":"A Comprehensive Review of Data Privacy Challenges in Social Media Platforms","abstract":"Social media privacy has become one of the most urgent 21st-century digital concerns, as sites frame communication, identity, and public discourse while also facilitating surveillance, profiling, and exploitation. This review examines twenty peer-reviewed articles from 2003 to 2024, drawn from IEEE, Springer, ACM, and Scopus. The research was grouped into four broad categories: user behavior and awareness, legal and regulatory environment, risks and threats, and privacy enhancing technologies. The findings suggest that while technical solutions (e.g., encryption, differential privacy, federated learning) and policy tools (e.g., GDPR, CCPA, DPDP) are changing, they are not aligned with user understanding, cultural environments, and platform incentives. User literacy deficits, ineffective regulation, and data monetization-based business models remain eroding privacy protections. The review underscores the imperative for interdisciplinary approaches that merge legal, technical, and social insights, ensuring privacy-by-design systems that are easy to use and culturally sensitive.","author":[{"family":"Manikantan","given":"R"},{"family":"Meghana","given":"J"},{"family":"Padmavathi","given":"C"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17587337","URL":"https://doi.org/10.5281/zenodo.17587337","source":"datacite"},{"id":"doi:10.5281/zenodo.17587338","type":"article-journal","title":"A Comprehensive Review of Data Privacy Challenges in Social Media Platforms","abstract":"Social media privacy has become one of the most urgent 21st-century digital concerns, as sites frame communication, identity, and public discourse while also facilitating surveillance, profiling, and exploitation. This review examines twenty peer-reviewed articles from 2003 to 2024, drawn from IEEE, Springer, ACM, and Scopus. The research was grouped into four broad categories: user behavior and awareness, legal and regulatory environment, risks and threats, and privacy enhancing technologies. The findings suggest that while technical solutions (e.g., encryption, differential privacy, federated learning) and policy tools (e.g., GDPR, CCPA, DPDP) are changing, they are not aligned with user understanding, cultural environments, and platform incentives. User literacy deficits, ineffective regulation, and data monetization-based business models remain eroding privacy protections. The review underscores the imperative for interdisciplinary approaches that merge legal, technical, and social insights, ensuring privacy-by-design systems that are easy to use and culturally sensitive.","author":[{"family":"Manikantan","given":"R"},{"family":"Meghana","given":"J"},{"family":"Padmavathi","given":"C"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17587338","URL":"https://doi.org/10.5281/zenodo.17587338","source":"datacite"},{"id":"doi:10.48550/arxiv.2511.01800","type":"manuscript","title":"Bayesian Coreset Optimization for Personalized Federated Learning","abstract":"In a distributed machine learning setting like Federated Learning where there are multiple clients involved which update their individual weights to a single central server, often training on the entire individual client's dataset for each client becomes cumbersome. To address this issue we propose $\\methodprop$: a personalized coreset weighted federated learning setup where the training updates for each individual clients are forwarded to the central server based on only individual client coreset based representative data points instead of the entire client data. Through theoretical analysis we present how the average generalization error is minimax optimal up to logarithm bounds (upper bounded by $\\mathcal{O}(n_k^{-\\frac{2 β}{2 β+\\boldsymbolΛ}} \\log ^{2 δ^{\\prime}}(n_k))$) and lower bounds of $\\mathcal{O}(n_k^{-\\frac{2 β}{2 β+\\boldsymbolΛ}})$, and how the overall generalization error on the data likelihood differs from a vanilla Federated Learning setup as a closed form function ${\\boldsymbol{\\Im}}(\\boldsymbol{w}, n_k)$ of the coreset weights $\\boldsymbol{w}$ and coreset sample size $n_k$. Our experiments on different benchmark datasets based on a variety of recent personalized federated learning architectures show significant gains as compared to random sampling on the training data followed by federated learning, thereby indicating how intelligently selecting such training samples can help in performance. Additionally, through experiments on medical datasets our proposed method showcases some gains as compared to other submodular optimization based approaches used for subset selection on client's data.","author":[{"family":"Chanda","given":"Prateek"},{"family":"Modi","given":"Shrey"},{"family":"Ramakrishnan","given":"Ganesh"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2511.01800","URL":"https://doi.org/10.48550/arxiv.2511.01800","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.25233","type":"manuscript","title":"FedCLF -- Towards Efficient Participant Selection for Federated Learning in Heterogeneous IoV Networks","abstract":"Federated Learning (FL) is a distributed machine learning technique that preserves data privacy by sharing only the trained parameters instead of the client data. This makes FL ideal for highly dynamic, heterogeneous, and time-critical applications, in particular, the Internet of Vehicles (IoV) networks. However, FL encounters considerable challenges in such networks owing to the high data and device heterogeneity. To address these challenges, we propose FedCLF, i.e., FL with Calibrated Loss and Feedback control, which introduces calibrated loss as a utility in the participant selection process and a feedback control mechanism to dynamically adjust the sampling frequency of the clients. The envisaged approach (a) enhances the overall model accuracy in case of highly heterogeneous data and (b) optimizes the resource utilization for resource constrained IoV networks, thereby leading to increased efficiency in the FL process. We evaluated FedCLF vis-à-vis baseline models, i.e., FedAvg, Newt, and Oort, using CIFAR-10 dataset with varying data heterogeneity. Our results depict that FedCLF significantly outperforms the baseline models by up to a 16% improvement in high data heterogeneity-related scenarios with improved efficiency via reduced sampling frequency.","author":[{"family":"Wijethilake","given":"Kasun"},{"family":"Mahmood","given":"Adnan"},{"family":"Sheng","given":"Quan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.25233","URL":"https://doi.org/10.48550/arxiv.2509.25233","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.20193","type":"manuscript","title":"FairEquityFL -- A Fair and Equitable Client Selection in Federated Learning for Heterogeneous IoV Networks","abstract":"Federated Learning (FL) has been extensively employed for a number of applications in machine learning, i.e., primarily owing to its privacy preserving nature and efficiency in mitigating the communication overhead. Internet of Vehicles (IoV) is one of the promising applications, wherein FL can be utilized to train a model more efficiently. Since only a subset of the clients can participate in each FL training round, challenges arise pertinent to fairness in the client selection process. Over the years, a number of researchers from both academia and industry have proposed numerous FL frameworks. However, to the best of our knowledge, none of them have employed fairness for FL-based client selection in a dynamic and heterogeneous IoV environment. Accordingly, in this paper, we envisage a FairEquityFL framework to ensure an equitable opportunity for all the clients to participate in the FL training process. In particular, we have introduced a sampling equalizer module within the selector component for ensuring fairness in terms of fair collaboration opportunity for all the clients in the client selection process. The selector is additionally responsible for both monitoring and controlling the clients' participation in each FL training round. Moreover, an outlier detection mechanism is enforced for identifying malicious clients based on the model performance in terms of considerable fluctuation in either accuracy or loss minimization. The selector flags suspicious clients and temporarily suspend such clients from participating in the FL training process. We further evaluate the performance of FairEquityFL on a publicly available dataset, FEMNIST. Our simulation results depict that FairEquityFL outperforms baseline models to a considerable extent.","author":[{"family":"Islam","given":"Fahmida"},{"family":"Mahmood","given":"Adnan"},{"family":"Mukhtiar","given":"Noorain"},{"family":"Wijethilake","given":"Kasun"},{"family":"Sheng","given":"Quan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.20193","URL":"https://doi.org/10.48550/arxiv.2509.20193","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.19403","type":"manuscript","title":"Recovering Clinical Utility Under Differential Privacy: Empirical Validation of Adaptive Federated Aggregation on Heterogeneous Cardiovascular Datasets","abstract":"Validating federated learning frameworks on real clinical data is an essential step between proof-of-concept demonstrations in controlled synthetic environments and deployment in real multicenter healthcare settings. A prior architectural study by the same authors (Tertulino and Alencar, 2026) demonstrated, on a synthetic six-feature benchmark, that server-side adaptive optimization acts as a temporal denoiser for Differential Privacy noise, answering an open challenge identified in the original pipeline work (Tertulino, 2025). That study used synthetically generated data and explicitly identified real-world validation as a priority future direction. The present work addresses this gap by validating the FedCVR framework on five publicly available real cardiovascular datasets (Framingham, Cleveland, Hungarian, Switzerland, and Long Beach VA), harmonized to the 13-attribute UCI Heart Disease schema and configured as a heterogeneous federated scenario with leave-one-institution-out cross-validation. Results demonstrate that FedCVR preserves its adaptive advantage on real data, achieving an F1-Score of 79.2% and AUC of 0.96 under the operational privacy budget (noise multiplier = 0.8, privacy budget epsilon approximately 4.2), while statistically outperforming standard FedAvg on all evaluated metrics (paired t-tests, all p &lt;= 0.003, significant under the Bonferroni-corrected threshold). The measured privacy cost on real data confirms the graceful degradation pattern observed in the synthetic experiments, providing empirical evidence of the framework's clinical viability in genuine multicenter contexts.","author":[{"family":"Tertulino","given":"Rodrigo"},{"family":"Alencar","given":"Laercio"},{"family":"Almeida","given":"Ricardo"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.19403","URL":"https://doi.org/10.48550/arxiv.2607.19403","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.14416","type":"manuscript","title":"Federated Learning for Feature Generalization with Convex Constraints","abstract":"Federated learning (FL) often struggles with generalization due to heterogeneous client data. Local models are prone to overfitting their local data distributions, and even transferable features can be distorted during aggregation. To address these challenges, we propose FedCONST, an approach that adaptively modulates update magnitudes based on the parameter strength of the global model. This prevents over-emphasizing well-learned parameters while reinforcing underdeveloped ones. Specifically, FedCONST employs linear convex constraints to ensure training stability and preserve locally learned generalization capabilities during aggregation. A Gradient Signal to Noise Ratio (GSNR) analysis further validates the effectiveness of FedCONST in enhancing feature transferability and robustness. As a result, FedCONST effectively aligns local and global objectives, mitigating overfitting and promoting stronger generalization across diverse FL environments, achieving state-of-the-art performance.","author":[{"family":"Kim","given":"Dongwon"},{"family":"Kim","given":"Donghee"},{"family":"Shyn","given":"Sung"},{"family":"Kim","given":"Kwangsu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.14416","URL":"https://doi.org/10.48550/arxiv.2606.14416","source":"datacite"},{"id":"doi:10.48550/arxiv.2505.02540","type":"manuscript","title":"Lazy But Effective: Collaborative Personalized Federated Learning with Heterogeneous Data","abstract":"In Federated Learning, heterogeneity in client data distributions often means that a single global model does not have the best performance for individual clients. Consider for example training a next-word prediction model for keyboards: user-specific language patterns due to demographics (dialect, age, etc.), language proficiency, and writing style result in a highly non-IID dataset across clients. Other examples are medical images taken with different machines, or driving data from different vehicle types. To address this, we propose a simple yet effective personalized federated learning framework (pFedLIA) that utilizes a computationally efficient influence approximation, called `Lazy Influence', to cluster clients in a distributed manner before model aggregation. Within each cluster, data owners collaborate to jointly train a model that captures the specific data patterns of the clients. Our method has been shown to successfully recover the global model's performance drop due to the non-IID-ness in various synthetic and real-world settings, specifically a next-word prediction task on the Nordic languages as well as several benchmark tasks. It matches the performance of a hypothetical Oracle clustering, and significantly improves on existing baselines, e.g., an improvement of 17% on CIFAR100.","author":[{"family":"Rokvic","given":"Ljubomir"},{"family":"Danassis","given":"Panayiotis"},{"family":"Faltings","given":"Boi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2505.02540","URL":"https://doi.org/10.48550/arxiv.2505.02540","source":"datacite"},{"id":"doi:10.48550/arxiv.2605.17219","type":"manuscript","title":"Integration of AI in Cybersecurity: Current Trends with a Focused Look at Intrusion Detection Applications","abstract":"Artificial Intelligence (AI) is widely adopted today for its ability to detect patterns, automate tasks, and reduce time and cost across various applications. Its integration into Cybersecurity has garnered significant attention, particularly in areas such as intrusion detection, malware analysis, and phishing or spam detection. As AI and cybersecurity evolve, new methods and approaches emerge regularly. Current trends include the use of Generative AI, Natural Language Processing, Federated Learning for privacy-preserving collaborative training, and eXplainable AI to ensure interpretability and trust, which are vital in cybersecurity. This paper presents an interesting review of current AI-based cybersecurity trends, focusing on intrusion detection approaches and aiming to uncover meaningful insights through comparative analysis based on the employed AI techniques and reported performance.","author":[{"family":"Tazili","given":"S"},{"family":"Mansour","given":"A"},{"family":"Chkouri","given":"MY"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.17219","URL":"https://doi.org/10.48550/arxiv.2605.17219","source":"datacite"},{"id":"doi:10.48550/arxiv.2605.27385","type":"manuscript","title":"Personalized Observation Normalization for Federated Reinforcement Learning in Simulation Environments with Heterogeneity","abstract":"Federated reinforcement learning (FedRL) enables multiple agents to collaboratively train a global policy without sharing raw data, making it ideal for privacy-sensitive applications. However, FedRL faces challenges in heterogeneous environments where differing state-transition dynamics lead to non-identical input distributions and imbalanced parameter updates during aggregation. Therefore, this paper develops a personalized observation normalization (PON) method, allowing each agent to locally normalize raw state inputs using a continuously updated running mean and variance. This design ensures consistent scaling of local feature without overshadowing across agents during aggregation. Furthermore, we demonstrate that sharing normalization parameters across agents is ineffective due to the diverse local input distributions, which highlights the necessity of personalized statistics. Experiments on heterogeneous MuJoCo tasks show that our developed PON accelerates training and achieves superior performance compared to baseline methods.","author":[{"family":"Pang","given":"Yiran"},{"family":"Ni","given":"Zhen"},{"family":"Zhong","given":"Xiangnan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.27385","URL":"https://doi.org/10.48550/arxiv.2605.27385","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.14769","type":"manuscript","title":"Federated Distillation on Edge Devices: Efficient Client-Side Filtering for Non-IID Data","abstract":"Federated distillation has emerged as a promising collaborative machine learning approach, offering enhanced privacy protection and reduced communication compared to traditional federated learning by exchanging model outputs (soft logits) rather than full model parameters. However, existing methods employ complex selective knowledge-sharing strategies that require clients to identify in-distribution proxy data through computationally expensive statistical density ratio estimators. Additionally, server-side filtering of ambiguous knowledge introduces latency to the process. To address these challenges, we propose a robust, resource-efficient EdgeFD method that reduces the complexity of the client-side density ratio estimation and removes the need for server-side filtering. EdgeFD introduces an efficient KMeans-based density ratio estimator for effectively filtering both in-distribution and out-of-distribution proxy data on clients, significantly improving the quality of knowledge sharing. We evaluate EdgeFD across diverse practical scenarios, including strong non-IID, weak non-IID, and IID data distributions on clients, without requiring a pre-trained teacher model on the server for knowledge distillation. Experimental results demonstrate that EdgeFD outperforms state-of-the-art methods, consistently achieving accuracy levels close to IID scenarios even under heterogeneous and challenging conditions. The significantly reduced computational overhead of the KMeans-based estimator is suitable for deployment on resource-constrained edge devices, thereby enhancing the scalability and real-world applicability of federated distillation. The code is available online for reproducibility.","author":[{"family":"Mujtaba","given":"Ahmed"},{"family":"Radchenko","given":"Gleb"},{"family":"Prodan","given":"Radu"},{"family":"Masana","given":"Marc"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.14769","URL":"https://doi.org/10.48550/arxiv.2508.14769","source":"datacite"},{"id":"doi:10.48550/arxiv.2605.10515","type":"manuscript","title":"SoK: A Systematic Bidirectional Literature Review of AI &amp; DLT Convergence","abstract":"The integration of Artificial Intelligence (AI) with Distributed Ledger Technology (DLT) has become a growing research area, yet contributions tend to cluster around specific application domains or examine only one direction of the integration, leaving the broader architectural interplay between the two technologies poorly understood. This work addresses that gap through a structured, bidirectional review of peer-reviewed studies published between 2020 and 2025. We classify contributions along two directions: AI-enhanced DLT, and DLT-enhanced AI. In the first case, we examine how AI techniques improve DLT systems across five layers: data, network, consensus, execution, and application layers. In the second case, we analyse how DLT supports AI systems across five layers: infrastructure, data, model, inference, and application layers, with particular attention to federated learning, model evaluation, and multi-agent coordination. The analysis reveals that most works concentrate on a small subset of layers: execution and consensus for AI-enhanced DLT, data and model for DLT-enhanced AI. Other layers remain comparatively neglected. Despite reported improvements in controlled settings, no study demonstrates deployment at production scale, and the field has not yet offered satisfying answers to fundamental questions around scalability, interoperability, and verifiable execution. We argue that progress will require cross-layer co-design and empirical validation in real-world settings.","author":[{"family":"Kathia","given":"Ali"},{"family":"Erinle","given":"Yimika"},{"family":"Satybaldy","given":"Abylay"},{"family":"Tasca","given":"Paolo"},{"family":"Vadgama","given":"Nikhil"},{"family":"Javarone","given":"Marco"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.10515","URL":"https://doi.org/10.48550/arxiv.2605.10515","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.12672","type":"manuscript","title":"Robust Federated Learning under Adversarial Attacks via Loss-Based Client Clustering","abstract":"Federated Learning (FL) enables collaborative model training across multiple clients without sharing private data. We consider FL scenarios wherein FL clients are subject to adversarial (Byzantine) attacks, while the FL server is trusted (honest) and has a trustworthy side dataset. This may correspond to, e.g., cases where the server possesses trusted data prior to federation, or to the presence of a trusted client that temporarily assumes the server role. Our approach requires only two honest participants, i.e., the server and one client, to function effectively, without prior knowledge of the number of malicious clients. Theoretical analysis demonstrates bounded optimality gaps even under strong Byzantine attacks. Experimental results show that our algorithm significantly outperforms standard and robust FL baselines such as Mean, Trimmed Mean, Median, Krum, and Multi-Krum under various attack strategies including label flipping, sign flipping, and Gaussian noise addition across MNIST, FMNIST, and CIFAR-10 benchmarks using the Flower framework.","author":[{"family":"Kritharakis","given":"Emmanouil"},{"family":"Jakovetic","given":"Dusan"},{"family":"Makris","given":"Antonios"},{"family":"Tserpes","given":"Konstantinos"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.12672","URL":"https://doi.org/10.48550/arxiv.2508.12672","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.20016","type":"manuscript","title":"FedSWA: Improving Generalization in Federated Learning with Highly Heterogeneous Data via Momentum-Based Stochastic Controlled Weight Averaging","abstract":"For federated learning (FL) algorithms such as FedSAM, their generalization capability is crucial for real-word applications. In this paper, we revisit the generalization problem in FL and investigate the impact of data heterogeneity on FL generalization. We find that FedSAM usually performs worse than FedAvg in the case of highly heterogeneous data, and thus propose a novel and effective federated learning algorithm with Stochastic Weight Averaging (called \\texttt{FedSWA}), which aims to find flatter minima in the setting of highly heterogeneous data. Moreover, we introduce a new momentum-based stochastic controlled weight averaging FL algorithm (\\texttt{FedMoSWA}), which is designed to better align local and global models. Theoretically, we provide both convergence analysis and generalization bounds for \\texttt{FedSWA} and \\texttt{FedMoSWA}. We also prove that the optimization and generalization errors of \\texttt{FedMoSWA} are smaller than those of their counterparts, including FedSAM and its variants. Empirically, experimental results on CIFAR10/100 and Tiny ImageNet demonstrate the superiority of the proposed algorithms compared to their counterparts. Open source code at: https://github.com/junkangLiu0/FedSWA.","author":[{"family":"Junkang","given":"Liu"},{"family":"Liu","given":"Yuanyuan"},{"family":"Shang","given":"Fanhua"},{"family":"Liu","given":"Hongying"},{"family":"Liu","given":"Jin"},{"family":"Feng","given":"Wei"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.20016","URL":"https://doi.org/10.48550/arxiv.2507.20016","source":"datacite"},{"id":"doi:10.13025/29816","type":"article-journal","title":"An insight on the timely diagnosis of diabetic retinopathy using traditional and ai-driven approaches","abstract":"Diabetic Retinopathy is a progressive microvascular complication of diabetes that requires early detection to improve patient outcomes. Traditional screening techniques, including fundus photography and optical coherence tomography, provide valuable diagnostic insights but have some limitations, like cost and technical complexity. Artificial intelligence is transforming the detection of diabetic retinopathy, moving away from traditional machine learning models that rely on manually created features to deep learning methods that allow for automatic feature extraction from retinal images. This systematic review investigates the evolution and diagnostic performance of AI-based techniques for DR detection. Studies were included if they applied machine learning or deep learning methods to retinal fundus or OCT images for DR classification. A total of 116 studies were included following comprehensive searches in databases such as PubMed, ScienceDirect, and IEEE Xplore, covering publications up to February 2024. Risk of bias was assessed in a representative sample of six studies, indicating a significant overall risk due to inadequate reporting on blinding and selective outcome reporting. Federated learning emerged as a promising alternative, enabling decentralized collaboration without compromising data privacy. Additionally, the growing focus on Explainable AI helps address the \"black-box\" nature of deep learning models by providing visual and textual explanations for predictions, thereby enhancing clinician trust and facilitating informed decision-making. By incorporating artificial intelligence with standard diagnostic frameworks, this research highlights the possibility for more accurate, scalable, and reachable diabetic retinopathy detection, paving the way for considerable advancements in ophthalmic problem management and enhancing patient care.","author":[{"family":"Asif","given":"Malaika"},{"family":"Ur Rehman","given":"Fasih"},{"family":"Rashid","given":"Zoya"},{"family":"Hussain","given":"Altaf"},{"family":"Mirza","given":"Alina"},{"family":"Qureshi","given":"Waqar"}],"issued":{"date-parts":[[2025]]},"DOI":"10.13025/29816","URL":"https://doi.org/10.13025/29816","source":"datacite"},{"id":"doi:10.48550/arxiv.2504.16438","type":"manuscript","title":"POPri: Private Federated Learning using Preference-Optimized Synthetic Data","abstract":"In practical settings, differentially private Federated learning (DP-FL) is the dominant method for training models from private, on-device client data. Recent work has suggested that DP-FL may be enhanced or outperformed by methods that use DP synthetic data (Wu et al., 2024; Hou et al., 2024). The primary algorithms for generating DP synthetic data for FL applications require careful prompt engineering based on public information and/or iterative private client feedback. Our key insight is that the private client feedback collected by prior DP synthetic data methods (Hou et al., 2024; Xie et al., 2024) can be viewed as an RL (reinforcement learning) reward. Our algorithm, Policy Optimization for Private Data (POPri) harnesses client feedback using policy optimization algorithms such as Direct Preference Optimization (DPO) to fine-tune LLMs to generate high-quality DP synthetic data. To evaluate POPri, we release LargeFedBench, a new federated text benchmark for uncontaminated LLM evaluations on federated client data. POPri substantially improves the utility of DP synthetic data relative to prior work on LargeFedBench datasets and an existing benchmark from Xie et al. (2024). POPri closes the gap between next-token prediction accuracy in the fully-private and non-private settings by up to 58%, compared to 28% for prior synthetic data methods, and 3% for state-of-the-art DP federated learning methods. The code and data are available at https://github.com/meiyuw/POPri.","author":[{"family":"Hou","given":"Charlie"},{"family":"Wang","given":"Mei"},{"family":"Zhu","given":"Yige"},{"family":"Lazar","given":"Daniel"},{"family":"Fanti","given":"Giulia"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2504.16438","URL":"https://doi.org/10.48550/arxiv.2504.16438","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.21198","type":"manuscript","title":"Uncovering Gradient Inversion Risks in Practical Language Model Training","abstract":"The gradient inversion attack has been demonstrated as a significant privacy threat to federated learning (FL), particularly in continuous domains such as vision models. In contrast, it is often considered less effective or highly dependent on impractical training settings when applied to language models, due to the challenges posed by the discrete nature of tokens in text data. As a result, its potential privacy threats remain largely underestimated, despite FL being an emerging training method for language models. In this work, we propose a domain-specific gradient inversion attack named Grab (gradient inversion with hybrid optimization). Grab features two alternating optimization processes to address the challenges caused by practical training settings, including a simultaneous optimization on dropout masks between layers for improved token recovery and a discrete optimization for effective token sequencing. Grab can recover a significant portion (up to 92.9% recovery rate) of the private training data, outperforming the attack strategy of utilizing discrete optimization with an auxiliary model by notable improvements of up to 28.9% recovery rate in benchmark settings and 48.5% recovery rate in practical settings. Grab provides a valuable step forward in understanding this privacy threat in the emerging FL training mode of language models.","author":[{"family":"Feng","given":"Xinguo"},{"family":"Ma","given":"Zhongkui"},{"family":"Wang","given":"Zihan"},{"family":"Chegne","given":"Eu"},{"family":"Ma","given":"Mengyao"},{"family":"Abuadbba","given":"Alsharif"},{"family":"Bai","given":"Guangdong"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.21198","URL":"https://doi.org/10.48550/arxiv.2507.21198","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.03004","type":"manuscript","title":"CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics","abstract":"Recent research has highlighted the importance of data quality in scaling large language models (LLMs). However, automated data quality control faces unique challenges in collaborative settings where sharing is not allowed directly between data silos. To tackle this issue, this paper proposes a novel data quality control technique based on the notion of data influence on the training dynamics of LLMs, that high quality data are more likely to have similar training dynamics to the anchor dataset. We then leverage the influence of the training dynamics to select high-quality data from different private domains, with centralized model updates on the server side in a collaborative training fashion by either model merging or federated learning. As for the data quality indicator, we compute the per-sample gradients with respect to the private data and the anchor dataset, and use the trace of the accumulated inner products as a measurement of data quality. In addition, we develop a quality control evaluation tailored for collaborative settings with heterogeneous domain data. Experiments show that training on the high-quality data selected by our method can often outperform other data selection methods for collaborative fine-tuning of LLMs, across diverse private domain datasets, in medical, multilingual and financial settings. Our code is released at github.com/Ryan0v0/CLUES.","author":[{"family":"Zhao","given":"Wanru"},{"family":"Fan","given":"Hongxiang"},{"family":"Hu","given":"Shell"},{"family":"Zhou","given":"Wangchunshu"},{"family":"Chen","given":"Bofan"},{"family":"Lane","given":"Nicholas"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.03004","URL":"https://doi.org/10.48550/arxiv.2507.03004","source":"datacite"},{"id":"doi:10.48550/arxiv.2502.21266","type":"manuscript","title":"Supporting the development of Machine Learning for fundamental science in a federated Cloud with the AI_INFN platform","abstract":"Machine Learning (ML) is driving a revolution in the way scientists design, develop, and deploy data-intensive software. However, the adoption of ML presents new challenges for the computing infrastructure, particularly in terms of provisioning and orchestrating access to hardware accelerators for development, testing, and production. The INFN-funded project AI_INFN (\"Artificial Intelligence at INFN\") aims at fostering the adoption of ML techniques within INFN use cases by providing support on multiple aspects, including the provision of AI-tailored computing resources. It leverages cloud-native solutions in the context of INFN Cloud, to share hardware accelerators as effectively as possible, ensuring the diversity of the Institute's research activities is not compromised. In this contribution, we provide an update on the commissioning of a Kubernetes platform designed to ease the development of GPU-powered data analysis workflows and their scalability on heterogeneous, distributed computing resources, possibly federated as Virtual Kubelets with the interLink provider.","author":[{"family":"Anderlini","given":"Lucio"},{"family":"Barbetti","given":"Matteo"},{"family":"Bianchini","given":"Giulio"},{"family":"Ciangottini","given":"Diego"},{"family":"Pra","given":"Stefano"},{"family":"Michelotto","given":"Diego"},{"family":"Pellegrino","given":"Carmelo"},{"family":"Petrini","given":"Rosa"},{"family":"Pascolini","given":"Alessandro"},{"family":"Spiga","given":"Daniele"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2502.21266","URL":"https://doi.org/10.48550/arxiv.2502.21266","source":"datacite"},{"id":"doi:10.48550/arxiv.2504.02142","type":"manuscript","title":"Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds","abstract":"Group robustness has become a major concern in machine learning (ML) as conventional training paradigms were found to produce high error on minority groups. Without explicit group annotations, proposed solutions rely on heuristics that aim to identify and then amplify the minority samples during training. In our work, we first uncover a critical shortcoming of these methods: an inability to distinguish legitimate minority samples from poison samples in the training set. By amplifying poison samples as well, group robustness methods inadvertently boost the success rate of an adversary -- e.g., from $0\\%$ without amplification to over $97\\%$ with it. Notably, we supplement our empirical evidence with an impossibility result proving this inability of a standard heuristic under some assumptions. Moreover, scrutinizing recent poisoning defenses both in centralized and federated learning, we observe that they rely on similar heuristics to identify which samples should be eliminated as poisons. In consequence, minority samples are eliminated along with poisons, which damages group robustness -- e.g., from $55\\%$ without the removal of the minority samples to $41\\%$ with it. Finally, as they pursue opposing goals using similar heuristics, our attempt to alleviate the trade-off by combining group robustness methods and poisoning defenses falls short. By exposing this tension, we also hope to highlight how benchmark-driven ML scholarship can obscure the trade-offs among different metrics with potentially detrimental consequences.","author":[{"family":"Panaitescu-Liess","given":"Michael"},{"family":"Kaya","given":"Yigitcan"},{"family":"Zhu","given":"Sicheng"},{"family":"Huang","given":"Furong"},{"family":"Dumitras","given":"Tudor"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2504.02142","URL":"https://doi.org/10.48550/arxiv.2504.02142","source":"datacite"},{"id":"doi:10.48550/arxiv.2503.04054","type":"manuscript","title":"Controlled privacy leakage propagation throughout overlapping grouped learning","abstract":"Federated Learning (FL) is the standard protocol for collaborative learning. In FL, multiple workers jointly train a shared model. They exchange model updates calculated on their data, while keeping the raw data itself local. Since workers naturally form groups based on common interests and privacy policies, we are motivated to extend standard FL to reflect a setting with multiple, potentially overlapping groups. In this setup where workers can belong and contribute to more than one group at a time, complexities arise in understanding privacy leakage and in adhering to privacy policies. To address the challenges, we propose differential private overlapping grouped learning (DPOGL), a novel method to implement privacy guarantees within overlapping groups. Under the honest-but-curious threat model, we derive novel privacy guarantees between arbitrary pairs of workers. These privacy guarantees describe and quantify two key effects of privacy leakage in DP-OGL: propagation delay, i.e., the fact that information from one group will leak to other groups only with temporal offset through the common workers and information degradation, i.e., the fact that noise addition over model updates limits information leakage between workers. Our experiments show that applying DP-OGL enhances utility while maintaining strong privacy compared to standard FL setups.","author":[{"family":"Kiani","given":"Shahrzad"},{"family":"Boenisch","given":"Franziska"},{"family":"Draper","given":"Stark"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2503.04054","URL":"https://doi.org/10.48550/arxiv.2503.04054","source":"datacite"},{"id":"doi:10.48550/arxiv.2503.02017","type":"manuscript","title":"A Lightweight and Secure Deep Learning Model for Privacy-Preserving Federated Learning in Intelligent Enterprises","abstract":"The ever growing Internet of Things (IoT) connections drive a new type of organization, the Intelligent Enterprise. In intelligent enterprises, machine learning based models are adopted to extract insights from data. Due to the efficiency and privacy challenges of these traditional models, a new federated learning (FL) paradigm has emerged. In FL, multiple enterprises can jointly train a model to update a final model. However, firstly, FL trained models usually perform worse than centralized models, especially when enterprises training data is non-IID (Independent and Identically Distributed). Second, due to the centrality of FL and the untrustworthiness of local enterprises, traditional FL solutions are vulnerable to poisoning and inference attacks and violate privacy. Thirdly, the continuous transfer of parameters between enterprises and servers increases communication costs. To this end, the FedAnil+ model is proposed, a novel, lightweight, and secure Federated Deep Learning Model that includes three main phases. In the first phase, the goal is to solve the data type distribution skew challenge. Addressing privacy concerns against poisoning and inference attacks is covered in the second phase. Finally, to alleviate the communication overhead, a novel compression approach is proposed that significantly reduces the size of the updates. The experiment results validate that FedAnil+ is secure against inference and poisoning attacks with better accuracy. In addition, it shows improvements over existing approaches in terms of model accuracy (13%, 16%, and 26%), communication cost (17%, 21%, and 25%), and computation cost (7%, 9%, and 11%).","author":[{"family":"Fotohi","given":"Reza"},{"family":"Aliee","given":"Fereidoon"},{"family":"Farahani","given":"Bahar"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2503.02017","URL":"https://doi.org/10.48550/arxiv.2503.02017","source":"datacite"},{"id":"doi:10.48550/arxiv.2502.14205","type":"manuscript","title":"Accurate Forgetting for Heterogeneous Federated Continual Learning","abstract":"Recent years have witnessed a burgeoning interest in federated learning (FL). However, the contexts in which clients engage in sequential learning remain under-explored. Bridging FL and continual learning (CL) gives rise to a challenging practical problem: federated continual learning (FCL). Existing research in FCL primarily focuses on mitigating the catastrophic forgetting issue of continual learning while collaborating with other clients. We argue that the forgetting phenomena are not invariably detrimental. In this paper, we consider a more practical and challenging FCL setting characterized by potentially unrelated or even antagonistic data/tasks across different clients. In the FL scenario, statistical heterogeneity and data noise among clients may exhibit spurious correlations which result in biased feature learning. While existing CL strategies focus on a complete utilization of previous knowledge, we found that forgetting biased information is beneficial in our study. Therefore, we propose a new concept accurate forgetting (AF) and develop a novel generative-replay method~\\method~which selectively utilizes previous knowledge in federated networks. We employ a probabilistic framework based on a normalizing flow model to quantify the credibility of previous knowledge. Comprehensive experiments affirm the superiority of our method over baselines.","author":[{"family":"Wuerkaixi","given":"Abudukelimu"},{"family":"Cui","given":"Sen"},{"family":"Zhang","given":"Jingfeng"},{"family":"Yan","given":"Kunda"},{"family":"Han","given":"Bo"},{"family":"Niu","given":"Gang"},{"family":"Fang","given":"Lei"},{"family":"Zhang","given":"Changshui"},{"family":"Sugiyama","given":"Masashi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2502.14205","URL":"https://doi.org/10.48550/arxiv.2502.14205","source":"datacite"},{"id":"doi:10.48550/arxiv.2502.07951","type":"manuscript","title":"Federated Self-supervised Domain Generalization for Label-efficient Polyp Segmentation","abstract":"Employing self-supervised learning (SSL) methodologies assumes par-amount significance in handling unlabeled polyp datasets when building deep learning-based automatic polyp segmentation models. However, the intricate privacy dynamics surrounding medical data often preclude seamless data sharing among disparate medical centers. Federated learning (FL) emerges as a formidable solution to this privacy conundrum, yet within the realm of FL, optimizing model generalization stands as a pressing imperative. Robust generalization capabilities are imperative to ensure the model's efficacy across diverse geographical domains post-training on localized client datasets. In this paper, a Federated self-supervised Domain Generalization method is proposed to enhance the generalization capacity of federated and Label-efficient intestinal polyp segmentation, named LFDG. Based on a classical SSL method, DropPos, LFDG proposes an adversarial learning-based data augmentation method (SSADA) to enhance the data diversity. LFDG further proposes a relaxation module based on Source-reconstruction and Augmentation-masking (SRAM) to maintain stability in feature learning. We have validated LFDG on polyp images from six medical centers. The performance of our method achieves 3.80% and 3.92% better than the baseline and other recent FL methods and SSL methods, respectively.","author":[{"family":"Tan","given":"Xinyi"},{"family":"Wang","given":"Jiacheng"},{"family":"Wang","given":"Liansheng"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2502.07951","URL":"https://doi.org/10.48550/arxiv.2502.07951","source":"datacite"},{"id":"doi:10.3390/en18040936","type":"article-journal","title":"Federated Learning and Neural Circuit Policies: A Novel Framework for Anomaly Detection in Energy-Intensive Machinery","abstract":"In the realm of predictive maintenance for energy-intensive machinery, effective anomaly detection is crucial for minimizing downtime and optimizing operational efficiency. This paper introduces a novel approach that integrates federated learning (FL) with Neural Circuit Policies (NCPs) to enhance anomaly detection in compressors utilized in leather tanning operations. Unlike traditional Long Short-Term Memory (LSTM) networks, which rely heavily on historical data patterns and often struggle with generalization, NCPs incorporate physical constraints and system dynamics, resulting in superior performance. Our comparative analysis reveals that NCPs significantly outperform LSTMs in accuracy and interpretability within a federated learning framework. This innovative combination not only addresses pressing data privacy concerns but also facilitates collaborative learning across decentralized data sources. By showcasing the effectiveness of FL and NCPs, this research paves the way for advanced predictive maintenance strategies that prioritize both performance and data integrity in energy-intensive industries.","author":[{"family":"Palma","given":"Giulia"},{"family":"Geraci","given":"Giovanni"},{"family":"Rizzo","given":"Antonio"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/en18040936","URL":"https://doi.org/10.3390/en18040936","source":"openalex"},{"id":"doi:10.48550/arxiv.2603.30004","type":"manuscript","title":"From Patterns to Policy: A Scoping Review Based on Bibliometric Analysis (ScoRBA) of Intelligent and Secure Smart Hospital Ecosystems","abstract":"This study examines the evolution of Intelligent and Secure Smart Hospital Ecosystems using a Scoping Review with Bibliometric Analysis (ScoRBA) to map research patterns, identify gaps, and derive policy implications. Analyzing 891 journal articles from Scopus (2006-2025) through co-occurrence analysis, network visualization, overlay analysis, and the Enhanced Strategic Diagram (ESD), the study applies the PAGER framework to link Patterns, Advances, Gaps, Research directions, and Evidence-based policy implications. Findings reveal three interrelated clusters: AI-driven intelligent healthcare systems, decentralized privacy-preserving digital health ecosystems, and scalable cloud-edge infrastructures, showing a convergence toward integrated ecosystem architectures where intelligence, trust, and infrastructure reinforce each other. Despite progress in AI, blockchain, and cloud computing, gaps remain in interoperability, real-world implementation, governance, and cross-layer integration. Emerging themes such as explainable AI, federated learning, and privacy mechanisms highlight areas needing further research. Policy-relevant recommendations focus on coordinated governance, scalable infrastructure, and secure data ecosystems, particularly for developing country contexts. The study bridges bibliometric evidence with actionable policies, supporting informed decision-making in smart hospital development.","author":[{"family":"Wijaya","given":"Adi"},{"family":"Hermawan","given":"Budi"},{"family":"Baihaqi","given":"Wiga"},{"family":"Supriyanto","given":"Catur"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.30004","URL":"https://doi.org/10.48550/arxiv.2603.30004","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.24601","type":"manuscript","title":"FED-HARGPT: A Hybrid Centralized-Federated Approach of a Transformer-based Architecture for Human Context Recognition","abstract":"The study explores a hybrid centralized-federated approach for Human Activity Recognition (HAR) using a Transformer-based architecture. With the increasing ubiquity of edge devices, such as smartphones and wearables, a significant amount of private data from wearable and inertial sensors is generated, facilitating discreet monitoring of human activities, including resting, sleeping, and walking. This research focuses on deploying HAR technologies using mobile sensor data and leveraging Federated Learning within the Flower framework to evaluate the training of a federated model derived from a centralized baseline. The experimental results demonstrate the effectiveness of the proposed hybrid approach in improving the accuracy and robustness of HAR models while preserving data privacy in a non-IID data scenario. The federated learning setup demonstrated comparable performance to centralized models, highlighting the potential of federated learning to strike a balance between data privacy and model performance in real-world applications.","author":[{"family":"Gibaut","given":"Wandemberg"},{"family":"Osorio","given":"Alexandre"},{"family":"Munoz","given":"Amparo"},{"family":"Neto","given":"Sildolfo"},{"family":"Grassiotto","given":"Fabio"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.24601","URL":"https://doi.org/10.48550/arxiv.2603.24601","source":"datacite"},{"id":"doi:10.17605/osf.io/jxbms","type":"article-journal","title":"Artificial Intelligence and Radiomics for Differentiating Pseudoprogression from True Progression in High-Grade Gliomas: A Meta-Analysis","abstract":"Differentiating true progression (TP) from pseudoprogression (PsP) in high-grade gliomas (HGGs) on MRI is a critical clinical challenge. This meta-analysis evaluates the overall diagnostic accuracy of radiomics and artificial intelligence (AI) models to identify algorithmic determinants of optimal performance. A PRISMA-compliant search of PubMed, Ovid MEDLINE, and EMBASE (up to 2025) was conducted. Quality was assessed via QUADAS-2. Diagnostic metrics were pooled utilizing bivariate random-effects models and Summary Receiver Operating Characteristic (SROC) curves. Across 34 included studies, the overall pooled sensitivity and specificity were 82% and 79%, respectively. Shape-based and Deep Learning (DL) features achieved the highest Area Under the Curve (up to 0.96 and 0.95). First-order statistics and Gray-Level Co-occurrence Matrix (GLCM) yielded the highest consistency and accuracy peaks (up to 98%). Multiparametric MRI was the most robust and widely used imaging modality. Radiomics demonstrates high, reliable diagnostic potential for discriminating HGG PsP from TP. While advanced DL models and shape features maximize discriminative performance, classical handcrafted features remain the most extensively validated. Future studies must prioritize independent external validation, federated learning (FL), and Explainable AI (XAI) for clinical translation.","author":[{"family":"Facchinetti","given":"Giovanni"},{"family":"De Maria","given":"Lucio"},{"family":"Pagani","given":"Nicola"},{"family":"Ponzio","given":"Francesco"}],"issued":{"date-parts":[[2026]]},"DOI":"10.17605/osf.io/jxbms","URL":"https://doi.org/10.17605/osf.io/jxbms","source":"datacite"},{"id":"doi:10.48550/arxiv.2505.11191","type":"manuscript","title":"Multi-Modal Multi-Task (M3T) Federated Foundation Models for Embodied AI: Potentials and Challenges for Edge Integration","abstract":"As embodied AI systems become increasingly multi-modal, personalized, and interactive, they must learn effectively from diverse sensory inputs, adapt continually to user preferences, and operate safely under resource and privacy constraints. These challenges expose a pressing need for machine learning models capable of swift, context-aware adaptation while balancing model generalization and personalization. Here, two methods emerge as suitable candidates, each offering parts of these capabilities: multi-modal multi-task foundation models (M3T-FMs) provide a pathway toward generalization across tasks and modalities, whereas federated learning (FL) offers the infrastructure for distributed, privacy-preserving model updates and user-level model personalization. However, when used in isolation, each of these approaches falls short of meeting the complex and diverse capability requirements of real-world embodied AI environments. In this vision paper, we introduce multi-modal multi-task federated foundation models (M3T-FFMs) for embodied AI, a new paradigm that unifies the strengths of M3T-FMs with the privacy-preserving distributed training nature of FL, enabling intelligent systems at the wireless edge. We collect critical deployment dimensions of M3T-FFMs in embodied AI ecosystems under a unified framework, which we name \"EMBODY\": Embodiment heterogeneity, Modality richness and imbalance, Bandwidth and compute constraints, On-device continual learning, Distributed control and autonomy, and Yielding safety, privacy, and personalization. For each, we identify concrete challenges and envision actionable research directions. We also present an evaluation framework for deploying M3T-FFMs in embodied AI systems, along with the associated trade-offs. Finally, we present a prototype implementation of M3T-FFMs and evaluate their energy and latency performance.","author":[{"family":"Borazjani","given":"Kasra"},{"family":"Abdisarabshali","given":"Payam"},{"family":"Nadimi","given":"Fardis"},{"family":"Khosravan","given":"Naji"},{"family":"Liwang","given":"Minghui"},{"family":"Wang","given":"Xianbin"},{"family":"Hong","given":"Yiguang"},{"family":"Hosseinalipour","given":"Seyyedali"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2505.11191","URL":"https://doi.org/10.48550/arxiv.2505.11191","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.19674","type":"manuscript","title":"C^2Prompt: Class-aware Client Knowledge Interaction for Federated Continual Learning","abstract":"Federated continual learning (FCL) tackles scenarios of learning from continuously emerging task data across distributed clients, where the key challenge lies in addressing both temporal forgetting over time and spatial forgetting simultaneously. Recently, prompt-based FCL methods have shown advanced performance through task-wise prompt communication.In this study, we underscore that the existing prompt-based FCL methods are prone to class-wise knowledge coherence between prompts across clients. The class-wise knowledge coherence includes two aspects: (1) intra-class distribution gap across clients, which degrades the learned semantics across prompts, (2) inter-prompt class-wise relevance, which highlights cross-class knowledge confusion. During prompt communication, insufficient class-wise coherence exacerbates knowledge conflicts among new prompts and induces interference with old prompts, intensifying both spatial and temporal forgetting. To address these issues, we propose a novel Class-aware Client Knowledge Interaction (C${}^2$Prompt) method that explicitly enhances class-wise knowledge coherence during prompt communication. Specifically, a local class distribution compensation mechanism (LCDC) is introduced to reduce intra-class distribution disparities across clients, thereby reinforcing intra-class knowledge consistency. Additionally, a class-aware prompt aggregation scheme (CPA) is designed to alleviate inter-class knowledge confusion by selectively strengthening class-relevant knowledge aggregation. Extensive experiments on multiple FCL benchmarks demonstrate that C${}^2$Prompt achieves state-of-the-art performance. Our source code is available at https://github.com/zhoujiahuan1991/NeurIPS2025-C2Prompt","author":[{"family":"Xu","given":"Kunlun"},{"family":"Feng","given":"Yibo"},{"family":"Li","given":"Jiangmeng"},{"family":"Qi","given":"Yongsheng"},{"family":"Zhou","given":"Jiahuan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.19674","URL":"https://doi.org/10.48550/arxiv.2509.19674","source":"datacite"},{"id":"doi:10.21268/20250808-2","type":"article-journal","title":"Towards a generic and resource-efficient testbed for federated learning in wireless sensor networks","abstract":"This paper presents a lightweight, modular and portable testbed for evaluating Federated Learning (FL) on embedded wireless sensor network (WSN) nodes. The testbed is designed for Tiny Edge and Little Edge level devices and supports structured data flow, multi-threaded communication and flexible algorithm integration. It enables simulation of sensor data, local training of models using various artificial intelligence algorithms, and communication between clients and the server. It gives precise control over training conditions, node behavior during model training and the timing of individual operations - facilitating research into realistic FL scenarios in constrained environments. All with the strict rules of embedded programming optimization.","author":[{"family":"Turchan","given":"Krzysztof"},{"family":"Wołoszyn","given":"Kamil"},{"family":"Piotrowski","given":"Krzysztof"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21268/20250808-2","URL":"https://doi.org/10.21268/20250808-2","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.06060","type":"manuscript","title":"Simple Yet Effective: Extracting Private Data Across Clients in Federated Fine-Tuning of Large Language Models","abstract":"Federated large language models (FedLLMs) enable cross-silo collaborative training among institutions while preserving data locality, making them appealing for privacy-sensitive domains such as law, finance, and healthcare. However, the memorization behavior of LLMs can lead to privacy risks that may cause cross-client data leakage. In this work, we study the threat of cross-client data extraction, where a semi-honest participant attempts to recover personally identifiable information (PII) memorized from other clients' data. We propose three simple yet effective extraction strategies that leverage contextual prefixes from the attacker's local data, including frequency-based prefix sampling and local fine-tuning to amplify memorization. To evaluate these attacks, we construct a Chinese legal-domain dataset with fine-grained PII annotations consistent with CPIS, GDPR, and CCPA standards, and assess extraction performance using two metrics: coverage and efficiency. Experimental results show that our methods can recover up to 56.6% of victim-exclusive PII, where names, addresses, and birthdays are particularly vulnerable. These findings highlight concrete privacy risks in FedLLMs and establish a benchmark and evaluation framework for future research on privacy-preserving federated learning. Code and data are available at https://github.com/SMILELab-FL/FedPII.","author":[{"family":"Hu","given":"Yingqi"},{"family":"Zhang","given":"Zhuo"},{"family":"Zhang","given":"Jingyuan"},{"family":"Wang","given":"Jinghua"},{"family":"Wang","given":"Qifan"},{"family":"Qu","given":"Lizhen"},{"family":"Xu","given":"Zenglin"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.06060","URL":"https://doi.org/10.48550/arxiv.2506.06060","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.05137","type":"manuscript","title":"FedGIN: Federated Learning with Dynamic Global Intensity Non-linear Augmentation for Organ Segmentation using Multi-modal Images","abstract":"Medical image segmentation plays a crucial role in AI-assisted diagnostics, surgical planning, and treatment monitoring. Accurate and robust segmentation models are essential for enabling reliable, data-driven clinical decision making across diverse imaging modalities. Given the inherent variability in image characteristics across modalities, developing a unified model capable of generalizing effectively to multiple modalities would be highly beneficial. This model could streamline clinical workflows and reduce the need for modality-specific training. However, real-world deployment faces major challenges, including data scarcity, domain shift between modalities (e.g., CT vs. MRI), and privacy restrictions that prevent data sharing. To address these issues, we propose FedGIN, a Federated Learning (FL) framework that enables multimodal organ segmentation without sharing raw patient data. Our method integrates a lightweight Global Intensity Non-linear (GIN) augmentation module that harmonizes modality-specific intensity distributions during local training. We evaluated FedGIN using two types of datasets: an imputed dataset and a complete dataset. In the limited dataset scenario, the model was initially trained using only MRI data, and CT data was added to assess its performance improvements. In the complete dataset scenario, both MRI and CT data were fully utilized for training on all clients. In the limited-data scenario, FedGIN achieved a 12 to 18% improvement in 3D Dice scores on MRI test cases compared to FL without GIN and consistently outperformed local baselines. In the complete dataset scenario, FedGIN demonstrated near-centralized performance, with a 30% Dice score improvement over the MRI-only baseline and a 10% improvement over the CT-only baseline, highlighting its strong cross-modality generalization under privacy constraints.","author":[{"family":"Nagaraju","given":"Sachin"},{"family":"Moradi","given":"Ashkan"},{"family":"Abrahamsen","given":"Bendik"},{"family":"Elschot","given":"Mattijs"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.05137","URL":"https://doi.org/10.48550/arxiv.2508.05137","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.17625","type":"manuscript","title":"Catastrophic Forgetting Resilient One-Shot Incremental Federated Learning","abstract":"Modern big-data systems generate massive, heterogeneous, and geographically dispersed streams that are large-scale and privacy-sensitive, making centralization challenging. While federated learning (FL) provides a privacy-enhancing training mechanism, it assumes a static data flow and learns a collaborative model over multiple rounds, making learning with \\textit{incremental} data challenging in limited-communication scenarios. This paper presents One-Shot Incremental Federated Learning (OSI-FL), the first FL framework that addresses the dual challenges of communication overhead and catastrophic forgetting. OSI-FL communicates category-specific embeddings, devised by a frozen vision-language model (VLM) from each client in a single communication round, which a pre-trained diffusion model at the server uses to synthesize new data similar to the client's data distribution. The synthesized samples are used on the server for training. However, two challenges still persist: i) tasks arriving incrementally need to retrain the global model, and ii) as future tasks arrive, retraining the model introduces catastrophic forgetting. To this end, we augment training with Selective Sample Retention (SSR), which identifies and retains the top-p most informative samples per category and task pair based on sample loss. SSR bounds forgetting by ensuring that representative retained samples are incorporated into training in further iterations. The experimental results indicate that OSI-FL outperforms baselines, including traditional and one-shot FL approaches, in both class-incremental and domain-incremental scenarios across three benchmark datasets.","author":[{"family":"Zaland","given":"Obaidullah"},{"family":"Khan","given":"Zulfiqar"},{"family":"Bhuyan","given":"Monowar"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.17625","URL":"https://doi.org/10.48550/arxiv.2602.17625","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.17614","type":"manuscript","title":"Guarding the Middle: Protecting Intermediate Representations in Federated Split Learning","abstract":"Big data scenarios, where massive, heterogeneous datasets are distributed across clients, demand scalable, privacy-preserving learning methods. Federated learning (FL) enables decentralized training of machine learning (ML) models across clients without data centralization. Decentralized training, however, introduces a computational burden on client devices. U-shaped federated split learning (UFSL) offloads a fraction of the client computation to the server while keeping both data and labels on the clients' side. However, the intermediate representations (i.e., smashed data) shared by clients with the server are prone to exposing clients' private data. To reduce exposure of client data through intermediate data representations, this work proposes k-anonymous differentially private UFSL (KD-UFSL), which leverages privacy-enhancing techniques such as microaggregation and differential privacy to minimize data leakage from the smashed data transferred to the server. We first demonstrate that an adversary can access private client data from intermediate representations via a data-reconstruction attack, and then present a privacy-enhancing solution, KD-UFSL, to mitigate this risk. Our experiments indicate that, alongside increasing the mean squared error between the actual and reconstructed images by up to 50% in some cases, KD-UFSL also decreases the structural similarity between them by up to 40% on four benchmarking datasets. More importantly, KD-UFSL improves privacy while preserving the utility of the global model. This highlights its suitability for large-scale big data applications where privacy and utility must be balanced.","author":[{"family":"Zaland","given":"Obaidullah"},{"family":"Mistry","given":"Sajib"},{"family":"Bhuyan","given":"Monowar"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.17614","URL":"https://doi.org/10.48550/arxiv.2602.17614","source":"datacite"},{"id":"doi:10.48550/arxiv.2511.14406","type":"manuscript","title":"Watch Out for the Lifespan: Evaluating Backdoor Attacks Against Federated Model Adaptation","abstract":"Large models adaptation through Federated Learning (FL) addresses a wide range of use cases and is enabled by Parameter-Efficient Fine-Tuning techniques such as Low-Rank Adaptation (LoRA). However, this distributed learning paradigm faces several security threats, particularly to its integrity, such as backdoor attacks that aim to inject malicious behavior during the local training steps of certain clients. We present the first analysis of the influence of LoRA on state-of-the-art backdoor attacks targeting model adaptation in FL. Specifically, we focus on backdoor lifespan, a critical characteristic in FL, that can vary depending on the attack scenario and the attacker's ability to effectively inject the backdoor. A key finding in our experiments is that for an optimally injected backdoor, the backdoor persistence after the attack is longer when the LoRA's rank is lower. Importantly, our work highlights evaluation issues of backdoor attacks against FL and contributes to the development of more robust and fair evaluations of backdoor attacks, enhancing the reliability of risk assessments for critical FL systems. Our code is publicly available.","author":[{"family":"Vuillod","given":"Bastien"},{"family":"Moellic","given":"Pierre"},{"family":"Dutertre","given":"Jean"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2511.14406","URL":"https://doi.org/10.48550/arxiv.2511.14406","source":"datacite"},{"id":"doi:10.48550/arxiv.2511.11696","type":"manuscript","title":"Toward Dignity-Aware AI: Next-Generation Elderly Monitoring from Fall Detection to ADL","abstract":"This position paper envisions a next-generation elderly monitoring system that moves beyond fall detection toward the broader goal of Activities of Daily Living (ADL) recognition. Our ultimate aim is to design privacy-preserving, edge-deployed, and federated AI systems that can robustly detect and understand daily routines, supporting independence and dignity in aging societies. At present, ADL-specific datasets are still under collection. As a preliminary step, we demonstrate feasibility through experiments using the SISFall dataset and its GAN-augmented variants, treating fall detection as a proxy task. We report initial results on federated learning with non-IID conditions, and embedded deployment on Jetson Orin Nano devices. We then outline open challenges such as domain shift, data scarcity, and privacy risks, and propose directions toward full ADL monitoring in smart-room environments. This work highlights the transition from single-task detection to comprehensive daily activity recognition, providing both early evidence and a roadmap for sustainable and human-centered elderly care AI.","author":[{"family":"Shao","given":"Xun"},{"family":"Otani","given":"Aoba"},{"family":"Hirasuka","given":"Yuto"},{"family":"Cai","given":"Runji"},{"family":"Loke","given":"Seng"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2511.11696","URL":"https://doi.org/10.48550/arxiv.2511.11696","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.04384","type":"manuscript","title":"Blockchain Federated Learning for Sustainable Retail: Reducing Waste through Collaborative Demand Forecasting","abstract":"Effective demand forecasting is crucial for reducing food waste. However, data privacy concerns often hinder collaboration among retailers, limiting the potential for improved predictive accuracy. In this study, we explore the application of Federated Learning (FL) in Sustainable Supply Chain Management (SSCM), with a focus on the grocery retail sector dealing with perishable goods. We develop a baseline predictive model for demand forecasting and waste assessment in an isolated retailer scenario. Subsequently, we introduce a Blockchain-based FL model, trained collaboratively across multiple retailers without direct data sharing. Our preliminary results show that FL models have performance almost equivalent to the ideal setting in which parties share data with each other, and are notably superior to models built by individual parties without sharing data, cutting waste and boosting efficiency.","author":[{"family":"Turazza","given":"Fabio"},{"family":"Neri","given":"Alessandro"},{"family":"Pietri","given":"Marcello"},{"family":"Butturi","given":"Maria"},{"family":"Picone","given":"Marco"},{"family":"Mamei","given":"Marco"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.04384","URL":"https://doi.org/10.48550/arxiv.2602.04384","source":"datacite"},{"id":"doi:10.5281/zenodo.18482551","type":"article-journal","title":"Artificial Intelligence in Clinically Validated Medical Imaging: Transforming Radiological Practice and Diagnostic Accuracy","abstract":"Artificial intelligence (AI) is transforming diagnostic radiology by enhancing image interpretation,pattern recognition, and clinical decision-making across multiple imaging modalities. This systematicliterature review critically evaluates the diagnostic performance, workflow efficiency, andimplementation challenges of AI-based imaging systems. A comprehensive search of PubMed,Scopus, Web of Science, and IEEE Xplore identified 11 high- and moderate-quality studiespublished between 2015 and 2025, following Preferred Reporting Items for Systematic Reviews andMeta-Analyses (PRISMA) 2020 guidelines. The analysis revealed that deep learning algorithms,particularly convolutional neural networks, reported diagnostic accuracies ranging from 85% to 98%across individual studies, with sensitivities in several applications exceeding 90%; these valuesrepresent descriptive ranges extracted from the included studies rather than pooled or weightedsummary estimates. AI applications demonstrated superior reproducibility and clinical reliability inmammography, chest radiography, and neuroimaging, where multi-institutional validation supportedconsistent outcomes. Despite these advancements, limitations persist due to inconsistent externalvalidation, dataset imbalance, and inadequate methodological transparency. Emerging frameworkssuch as explainable AI (XAI) and federated learning show potential to enhance interpretability, datasecurity, and equity in clinical deployment. Furthermore, AI integration was associated with reducedinterpretation time and improved workflow efficiency without compromising diagnostic accuracy.Overall, this review underscores AI’s transition from experimental innovation to a clinicallyindispensable tool. By synthesizing evidence across technical, clinical, and ethical dimensions, itprovides a comprehensive foundation for developing standardized, transparent, and equitable AImodels in diagnostic imaging practice.","author":[{"family":"Alkahtani","given":"Shatha"},{"family":"Alkahtani","given":"Shatha"},{"family":"Alessa","given":"Raghad"},{"family":"Alnumani","given":"Nouf"},{"family":"Gorban","given":"Albatole"},{"family":"Alturaikhem","given":"Nour"},{"family":"Aljandan","given":"Asma"},{"family":"Alshamrani","given":"Mona"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18482551","URL":"https://doi.org/10.5281/zenodo.18482551","source":"datacite"},{"id":"doi:10.5281/zenodo.18441822","type":"article-journal","title":"Artificial Intelligence in Clinically Validated Medical Imaging: Transforming Radiological Practice and Diagnostic Accuracy","abstract":"Artificial intelligence (AI) is transforming diagnostic radiology by enhancing image interpretation,pattern recognition, and clinical decision-making across multiple imaging modalities. This systematicliterature review critically evaluates the diagnostic performance, workflow efficiency, andimplementation challenges of AI-based imaging systems. A comprehensive search of PubMed,Scopus, Web of Science, and IEEE Xplore identified 11 high- and moderate-quality studiespublished between 2015 and 2025, following Preferred Reporting Items for Systematic Reviews andMeta-Analyses (PRISMA) 2020 guidelines. The analysis revealed that deep learning algorithms,particularly convolutional neural networks, reported diagnostic accuracies ranging from 85% to 98%across individual studies, with sensitivities in several applications exceeding 90%; these valuesrepresent descriptive ranges extracted from the included studies rather than pooled or weightedsummary estimates. AI applications demonstrated superior reproducibility and clinical reliability inmammography, chest radiography, and neuroimaging, where multi-institutional validation supportedconsistent outcomes. Despite these advancements, limitations persist due to inconsistent externalvalidation, dataset imbalance, and inadequate methodological transparency. Emerging frameworkssuch as explainable AI (XAI) and federated learning show potential to enhance interpretability, datasecurity, and equity in clinical deployment. Furthermore, AI integration was associated with reducedinterpretation time and improved workflow efficiency without compromising diagnostic accuracy.Overall, this review underscores AI’s transition from experimental innovation to a clinicallyindispensable tool. By synthesizing evidence across technical, clinical, and ethical dimensions, itprovides a comprehensive foundation for developing standardized, transparent, and equitable AImodels in diagnostic imaging practice.","author":[{"family":"Alkahtani","given":"Shatha"},{"family":"Alkahtani","given":"Shatha"},{"family":"Alessa","given":"Raghad"},{"family":"Alnumani","given":"Nouf"},{"family":"Gorban","given":"Albatole"},{"family":"Alturaikhem","given":"Nour"},{"family":"Aljandan","given":"Asma"},{"family":"Alshamrani","given":"Mona"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18441822","URL":"https://doi.org/10.5281/zenodo.18441822","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.17634","type":"manuscript","title":"Federated Learning With Individualized Privacy Through Client Sampling","abstract":"With growing concerns about user data collection, individualized privacy has emerged as a promising solution to balance protection and utility by accounting for diverse user privacy preferences. Instead of enforcing a uniform level of anonymization for all users, this approach allows individuals to choose privacy settings that align with their comfort levels. Building on this idea, we propose an adapted method for enabling Individualized Differential Privacy (IDP) in Federated Learning (FL) by handling clients according to their personal privacy preferences. By extending the SAMPLE algorithm from centralized settings to FL, we calculate client-specific sampling rates based on their heterogeneous privacy budgets and integrate them into a modified IDP-FedAvg algorithm. We test this method under realistic privacy distributions and multiple datasets. The experimental results demonstrate that our approach achieves clear improvements over uniform DP baselines, reducing the trade-off between privacy and utility. Compared to the alternative SCALE method in related work, which assigns differing noise scales to clients, our method performs notably better. However, challenges remain for complex tasks with non-i.i.d. data, primarily stemming from the constraints of the decentralized setting.","author":[{"family":"Lange","given":"Lucas"},{"family":"Borchardt","given":"Ole"},{"family":"Rahm","given":"Erhard"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.17634","URL":"https://doi.org/10.48550/arxiv.2501.17634","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.01185","type":"manuscript","title":"FedBGS: A Blockchain Approach to Segment Gossip Learning in Decentralized Systems","abstract":"Privacy-Preserving Federated Learning (PPFL) is a Decentralized machine learning paradigm that enables multiple participants to collaboratively train a global model without sharing their data with the integration of cryptographic and privacy-based techniques to enhance the security of the global system. This privacy-oriented approach makes PPFL a highly suitable solution for training shared models in sectors where data privacy is a critical concern. In traditional FL, local models are trained on edge devices, and only model updates are shared with a central server, which aggregates them to improve the global model. However, despite the presence of the aforementioned privacy techniques, in the classical Federated structure, the issue of the server as a single-point-of-failure remains, leading to limitations both in terms of security and scalability. This paper introduces FedBGS, a fully Decentralized Blockchain-based framework that leverages Segmented Gossip Learning through Federated Analytics. The proposed system aims to optimize blockchain usage while providing comprehensive protection against all types of attacks, ensuring both privacy, security and non-IID data handling in Federated environments.","author":[{"family":"Turazza","given":"Fabio"},{"family":"Pietri","given":"Marcello"},{"family":"Picone","given":"Marco"},{"family":"Mamei","given":"Marco"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.01185","URL":"https://doi.org/10.48550/arxiv.2602.01185","source":"datacite"},{"id":"doi:10.5281/zenodo.18441823","type":"article-journal","title":"Artificial Intelligence in Clinically Validated Medical Imaging: Transforming Radiological Practice and Diagnostic Accuracy","abstract":"Artificial intelligence (AI) is transforming diagnostic radiology by enhancing image interpretation,pattern recognition, and clinical decision-making across multiple imaging modalities. This systematicliterature review critically evaluates the diagnostic performance, workflow efficiency, andimplementation challenges of AI-based imaging systems. A comprehensive search of PubMed,Scopus, Web of Science, and IEEE Xplore identified 11 high- and moderate-quality studiespublished between 2015 and 2025, following Preferred Reporting Items for Systematic Reviews andMeta-Analyses (PRISMA) 2020 guidelines. The analysis revealed that deep learning algorithms,particularly convolutional neural networks, reported diagnostic accuracies ranging from 85% to 98%across individual studies, with sensitivities in several applications exceeding 90%; these valuesrepresent descriptive ranges extracted from the included studies rather than pooled or weightedsummary estimates. AI applications demonstrated superior reproducibility and clinical reliability inmammography, chest radiography, and neuroimaging, where multi-institutional validation supportedconsistent outcomes. Despite these advancements, limitations persist due to inconsistent externalvalidation, dataset imbalance, and inadequate methodological transparency. Emerging frameworkssuch as explainable AI (XAI) and federated learning show potential to enhance interpretability, datasecurity, and equity in clinical deployment. Furthermore, AI integration was associated with reducedinterpretation time and improved workflow efficiency without compromising diagnostic accuracy.Overall, this review underscores AI’s transition from experimental innovation to a clinicallyindispensable tool. By synthesizing evidence across technical, clinical, and ethical dimensions, itprovides a comprehensive foundation for developing standardized, transparent, and equitable AImodels in diagnostic imaging practice.","author":[{"family":"Alkahtani","given":"Shatha"},{"family":"Alkahtani","given":"Shatha"},{"family":"Alessa","given":"Raghad"},{"family":"Alnumani","given":"Nouf"},{"family":"Gorban","given":"Albatole"},{"family":"Alturaikhem","given":"Nour"},{"family":"Aljandan","given":"Asma"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18441823","URL":"https://doi.org/10.5281/zenodo.18441823","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.00660","type":"manuscript","title":"Differential privacy for medical deep learning: methods, tradeoffs, and deployment implications","abstract":"Differential privacy (DP) is a key technique for protecting sensitive patient data in medical deep learning (DL). As clinical models grow more data-dependent, balancing privacy with utility and fairness has become a critical challenge. This scoping review synthesizes recent developments in applying DP to medical DL, with a particular focus on DP-SGD and alternative mechanisms across centralized and federated settings. Using a structured search strategy, we identified 74 studies published up to March 2025. Our analysis spans diverse data modalities, training setups, and downstream tasks, and highlights the tradeoffs between privacy guarantees, model accuracy, and subgroup fairness. We find that while DP-especially at strong privacy budgets-can preserve performance in well-structured imaging tasks, severe degradation often occurs under strict privacy, particularly in underrepresented or complex modalities. Furthermore, privacy-induced performance gaps disproportionately affect demographic subgroups, with fairness impacts varying by data type and task. A small subset of studies explicitly addresses these tradeoffs through subgroup analysis or fairness metrics, but most omit them entirely. Beyond DP-SGD, emerging approaches leverage alternative mechanisms, generative models, and hybrid federated designs, though reporting remains inconsistent. We conclude by outlining key gaps in fairness auditing, standardization, and evaluation protocols, offering guidance for future work toward equitable and clinically robust privacy-preserving DL systems in medicine.","author":[{"family":"Mohammadi","given":"Marziyeh"},{"family":"Vejdanihemmat","given":"Mohsen"},{"family":"Lotfinia","given":"Mahshad"},{"family":"Rusu","given":"Mirabela"},{"family":"Truhn","given":"Daniel"},{"family":"Maier","given":"Andreas"},{"family":"Arasteh","given":"Soroosh"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.00660","URL":"https://doi.org/10.48550/arxiv.2506.00660","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.09313","type":"manuscript","title":"Hetero-SplitEE: Split Learning of Neural Networks with Early Exits for Heterogeneous IoT Devices","abstract":"The continuous scaling of deep neural networks has fundamentally transformed machine learning, with larger models demonstrating improved performance across diverse tasks. This growth in model size has dramatically increased the computational resources required for the training process. Consequently, distributed approaches, such as Federated Learning and Split Learning, have become essential paradigms for scalable deployment. However, existing Split Learning approaches assume client homogeneity and uniform split points across all participants. This critically limits their applicability to real-world IoT systems where devices exhibit heterogeneity in computational resources. To address this limitation, this paper proposes Hetero-SplitEE, a novel method that enables heterogeneous IoT devices to train a shared deep neural network in parallel collaboratively. By integrating heterogeneous early exits into hierarchical training, our approach allows each client to select distinct split points (cut layers) tailored to its computational capacity. In addition, we propose two cooperative training strategies, the Sequential strategy and the Averaging strategy, to facilitate this collaboration among clients with different split points. The Sequential strategy trains clients sequentially with a shared server model to reduce computational overhead. The Averaging strategy enables parallel client training with periodic cross-layer aggregation. Extensive experiments on CIFAR-10, CIFAR-100, and STL-10 datasets using ResNet-18 demonstrate that our method maintains competitive accuracy while efficiently supporting diverse computational constraints, enabling practical deployment of collaborative deep learning in heterogeneous IoT ecosystems.","author":[{"family":"Oda","given":"Yuki"},{"family":"Ono","given":"Yuta"},{"family":"Nakamura","given":"Hiroshi"},{"family":"Takase","given":"Hideki"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.09313","URL":"https://doi.org/10.48550/arxiv.2512.09313","source":"datacite"},{"id":"doi:10.48550/arxiv.2601.17713","type":"manuscript","title":"FedCCA: Client-Centric Adaptation against Data Heterogeneity in Federated Learning on IoT Devices","abstract":"With the rapid development of the Internet of Things (IoT), AI model training on private data such as human sensing data is highly desired. Federated learning (FL) has emerged as a privacy-preserving distributed training framework for this purpuse. However, the data heterogeneity issue among IoT devices can significantly degrade the model performance and convergence speed in FL. Existing approaches limit in fixed client selection and aggregation on cloud server, making the privacy-preserving extraction of client-specific information during local training challenging. To this end, we propose Client-Centric Adaptation federated learning (FedCCA), an algorithm that optimally utilizes client-specific knowledge to learn a unique model for each client through selective adaptation, aiming to alleviate the influence of data heterogeneity. Specifically, FedCCA employs dynamic client selection and adaptive aggregation based on the additional client-specific encoder. To enhance multi-source knowledge transfer, we adopt an attention-based global aggregation strategy. We conducted extensive experiments on diverse datasets to assess the efficacy of FedCCA. The experimental results demonstrate that our approach exhibits a substantial performance advantage over competing baselines in addressing this specific problem.","author":[{"family":"Wang","given":"Kaile"},{"family":"Cao","given":"Jiannong"},{"family":"Yang","given":"Yu"},{"family":"Li","given":"Xiaoyin"},{"family":"Cao","given":"Yinfeng"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2601.17713","URL":"https://doi.org/10.48550/arxiv.2601.17713","source":"datacite"},{"id":"doi:10.48550/arxiv.2505.13643","type":"manuscript","title":"FedCTTA: A Collaborative Approach to Continual Test-Time Adaptation in Federated Learning","abstract":"Federated Learning (FL) enables collaborative model training across distributed clients without sharing raw data, making it ideal for privacy-sensitive applications. However, FL models often suffer performance degradation due to distribution shifts between training and deployment. Test-Time Adaptation (TTA) offers a promising solution by allowing models to adapt using only test samples. However, existing TTA methods in FL face challenges such as computational overhead, privacy risks from feature sharing, and scalability concerns due to memory constraints. To address these limitations, we propose Federated Continual Test-Time Adaptation (FedCTTA), a privacy-preserving and computationally efficient framework for federated adaptation. Unlike prior methods that rely on sharing local feature statistics, FedCTTA avoids direct feature exchange by leveraging similarity-aware aggregation based on model output distributions over randomly generated noise samples. This approach ensures adaptive knowledge sharing while preserving data privacy. Furthermore, FedCTTA minimizes the entropy at each client for continual adaptation, enhancing the model's confidence in evolving target distributions. Our method eliminates the need for server-side training during adaptation and maintains a constant memory footprint, making it scalable even as the number of clients or training rounds increases. Extensive experiments show that FedCTTA surpasses existing methods across diverse temporal and spatial heterogeneity scenarios.","author":[{"family":"Rajib","given":"Rakibul"},{"family":"Iftee","given":"Md"},{"family":"Hossain","given":"Mir"},{"family":"Rahman","given":"AKMM"},{"family":"Mistry","given":"Sajib"},{"family":"Amin","given":"MA"},{"family":"Ali","given":"Amin"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2505.13643","URL":"https://doi.org/10.48550/arxiv.2505.13643","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.12978","type":"manuscript","title":"Beyond Trade-offs: A Unified Framework for Privacy, Robustness, and Communication Efficiency in Federated Learning","abstract":"We propose Fed-DPRoC, a novel federated learning framework designed to jointly provide differential privacy (DP), Byzantine robustness, and communication efficiency. Central to our approach is the concept of robust-compatible compression, which allows reducing the bi-directional communication overhead without undermining the robustness of the aggregation. We instantiate our framework as RobAJoL, which integrates the Johnson-Lindenstrauss (JL)-based compression mechanism with robust averaging for robustness. Our theoretical analysis establishes the compatibility of JL transform with robust averaging, ensuring that RobAJoL maintains robustness guarantees, satisfies DP, and substantially reduces communication overhead. We further present simulation results on CIFAR-10, Fashion MNIST, and FEMNIST, validating our theoretical claims. We compare RobAJoL with a state-of-the-art communication-efficient and robust FL scheme augmented with DP for a fair comparison, demonstrating that RobAJoL outperforms existing methods in terms of robustness and utility under different Byzantine attacks.","author":[{"family":"Xia","given":"Yue"},{"family":"Jahani-Nezhad","given":"Tayyebeh"},{"family":"Bitar","given":"Rawad"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.12978","URL":"https://doi.org/10.48550/arxiv.2508.12978","source":"datacite"},{"id":"doi:10.5281/zenodo.18883693","type":"article-journal","title":"Dual asymmetric momentum improves federated class unlearning in edge systems (Code and Reproducibility Scripts)","abstract":"This Zenodo record provides the reference implementation of FedDAM, a parameter-efficient post-hoc federated class unlearning method designed for resource-constrained edge deployments. The release includes training and unlearning pipelines for CIFAR-10 and CIFAR-100 under federated non-IID Dirichlet partitions, matched-budget comparisons, and evaluation scripts to regenerate the key tables/figures reported in the associated Scientific Reports submission. Key features:(i) Auxiliary-head-only unlearning with frozen backbone and main classifier to reduce communication and client compute.(ii) Dual-asymmetric momentum buffers to decouple retain vs forget optimization dynamics.(iii) Reproducibility controls: fixed seeds, paired partitions and participation schedules, and configuration files for sweeps. New analyses included for revision:(A) Gradient cosine similarity analysis comparing FedAU vs FedDAM (retain vs forget gradient conflict and update alignment).(C) Forget-set heterogeneity statistics: distribution of ∣Du,k∣|D_{u,k}|∣Du,k∣ and fraction of clients with ∣Du,k∣=0|D_{u,k}|=0∣Du,k∣=0 across αD\\alpha_DαD. Dataset: CIFAR-10 and CIFAR-100 (public). No new datasets were generated.How to reproduce: See README.md in the archive.","author":[{"family":"Mayaluri","given":"Zefree"},{"family":"Patra","given":"Achirangshu"},{"family":"Sahoo","given":"Prabodh"},{"family":"Kumawat","given":"Gaurav"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18883693","URL":"https://doi.org/10.5281/zenodo.18883693","source":"datacite"},{"id":"doi:10.5281/zenodo.18883694","type":"article-journal","title":"Dual asymmetric momentum improves federated class unlearning in edge systems (Code and Reproducibility Scripts)","abstract":"This Zenodo record provides the reference implementation of FedDAM, a parameter-efficient post-hoc federated class unlearning method designed for resource-constrained edge deployments. The release includes training and unlearning pipelines for CIFAR-10 and CIFAR-100 under federated non-IID Dirichlet partitions, matched-budget comparisons, and evaluation scripts to regenerate the key tables/figures reported in the associated Scientific Reports submission. Key features:(i) Auxiliary-head-only unlearning with frozen backbone and main classifier to reduce communication and client compute.(ii) Dual-asymmetric momentum buffers to decouple retain vs forget optimization dynamics.(iii) Reproducibility controls: fixed seeds, paired partitions and participation schedules, and configuration files for sweeps. New analyses included for revision:(A) Gradient cosine similarity analysis comparing FedAU vs FedDAM (retain vs forget gradient conflict and update alignment).(C) Forget-set heterogeneity statistics: distribution of ∣Du,k∣|D_{u,k}|∣Du,k∣ and fraction of clients with ∣Du,k∣=0|D_{u,k}|=0∣Du,k∣=0 across αD\\alpha_DαD. Dataset: CIFAR-10 and CIFAR-100 (public). No new datasets were generated.How to reproduce: See README.md in the archive.","author":[{"family":"Mayaluri","given":"Zefree"},{"family":"Patra","given":"Achirangshu"},{"family":"Sahoo","given":"Prabodh"},{"family":"Kumawat","given":"Gaurav"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18883694","URL":"https://doi.org/10.5281/zenodo.18883694","source":"datacite"},{"id":"doi:10.5281/zenodo.21097801","type":"article-journal","title":"Evaluación de Mecanismos Semiasíncronos Deterministas y Adaptativos en Aprendizaje Federado","abstract":"El aprendizaje automático ha evolucionado significativamente en los últimos años, expandiendo su aplicación a entornos de Internet de las Cosas (IoT), donde estas técnicas poseen un gran potencial en sectores como la mejora de diagnósticos médicos, el reconocimiento facial o la clasificación de imágenes. En este contexto, el Aprendizaje Federado Semiasíncrono (SAFL) surge como una extensión del Aprendizaje Federado (FL), diseñada para entornos con clientes heterogéneos y recursos limitados, lo que permite que los nodos participen en el entrenamiento sin necesidad de sincronizarse completamente en cada ronda, lo que mejora la eficiencia y la flexibilidad del sistema. Sin embargo, esta semiasincronía introduce nuevos desafíos, como la variabilidad en la contribución de los clientes y la inestabilidad en la convergencia, que requieren estrategias específicas de gestión y mitigación. Este trabajo evalúa diversas estrategias semiasíncronas, contrastando enfoques deterministas, basados en un grado de semiasincronía M estático, frente a una propuesta dinámica que se autoajusta según la disponibilidad y el comportamiento de los clientes. La contribución principal consiste en analizar la viabilidad de una estrategia basada en una heurística de ajuste dinámico, que permite al sistema adaptarse a un entorno heterogéneo de recursos limitados. Este estudio preliminar sienta las bases para futuras estrategias federadas adaptativas que incluyan posibles mecanismos de mitigación de la obsolescencia y otras ineficiencias derivadas de la semiasincronía, orientadas a infraestructuras distribuidas de bajo coste y escenarios IoT.","author":[{"family":"Hidalgo Izquierdo","given":"Víctor"},{"family":"Caminero","given":"Maria"},{"family":"Carrión Espinosa","given":"María"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21097801","URL":"https://doi.org/10.5281/zenodo.21097801","source":"datacite"},{"id":"doi:10.5281/zenodo.21097802","type":"article-journal","title":"Evaluación de Mecanismos Semiasíncronos Deterministas y Adaptativos en Aprendizaje Federado","abstract":"El aprendizaje automático ha evolucionado significativamente en los últimos años, expandiendo su aplicación a entornos de Internet de las Cosas (IoT), donde estas técnicas poseen un gran potencial en sectores como la mejora de diagnósticos médicos, el reconocimiento facial o la clasificación de imágenes. En este contexto, el Aprendizaje Federado Semiasíncrono (SAFL) surge como una extensión del Aprendizaje Federado (FL), diseñada para entornos con clientes heterogéneos y recursos limitados, lo que permite que los nodos participen en el entrenamiento sin necesidad de sincronizarse completamente en cada ronda, lo que mejora la eficiencia y la flexibilidad del sistema. Sin embargo, esta semiasincronía introduce nuevos desafíos, como la variabilidad en la contribución de los clientes y la inestabilidad en la convergencia, que requieren estrategias específicas de gestión y mitigación. Este trabajo evalúa diversas estrategias semiasíncronas, contrastando enfoques deterministas, basados en un grado de semiasincronía M estático, frente a una propuesta dinámica que se autoajusta según la disponibilidad y el comportamiento de los clientes. La contribución principal consiste en analizar la viabilidad de una estrategia basada en una heurística de ajuste dinámico, que permite al sistema adaptarse a un entorno heterogéneo de recursos limitados. Este estudio preliminar sienta las bases para futuras estrategias federadas adaptativas que incluyan posibles mecanismos de mitigación de la obsolescencia y otras ineficiencias derivadas de la semiasincronía, orientadas a infraestructuras distribuidas de bajo coste y escenarios IoT.","author":[{"family":"Hidalgo Izquierdo","given":"Víctor"},{"family":"Caminero","given":"Maria"},{"family":"Carrión Espinosa","given":"María"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21097802","URL":"https://doi.org/10.5281/zenodo.21097802","source":"datacite"},{"id":"doi:10.24406/publica-5187","type":"article-journal","title":"Scaling Smart Cities with Federated Learning: Balancing Accuracy and Privacy for Building Energy Performance Prediction","abstract":"The building energy sector is a significant contributor to carbon emissions, thereby playing a crucial role in driving global sustainability efforts to achieve the net-zero targets outlined in the Paris Climate Agreement. Precise predictions of building energy performance are imperative for effective planning and investment decisions aimed at enhancing energy efficiency. While data-driven methods, primarily leveraging machine learning techniques, offer promising predictive capabilities, they heavily rely on large datasets for accurate assessments. However, a prevalent challenge arises as energy consultants and agencies often lack expansive datasets, and if they do, they are reluctant to share their data. To overcome these hurdles, the study implements a decentralized, privacy-preserving machine learning approach known as federated learning. This approach was applied to a dataset encompassing over 25,000 residential buildings featuring diverse construction attributes and energy sources. The simulation involved mimicking different energy agencies by segmenting geographic regions. The study compared the prediction performance of federated learning with that of a model accessing the entire dataset and a fully isolated local model. The findings demonstrate that federated learning achieves a 12% improvement in prediction performance compared to the isolated model. This outcome underscores federated learning’s capacity to leverage the full potential of scaling data-driven methodologies, providing a pathway to unlock new business models in both research and practice, while aligning with net-zero aspirations.","author":[{"family":"Delgado Fernandez","given":"Joaquin"},{"family":"Willburger","given":"Lukas"},{"family":"Wiethe","given":"Christian"},{"family":"Wenninger","given":"Simon"},{"family":"Fridgen","given":"Gilbert"},{"family":"Unav"}],"issued":{"date-parts":[[2026]]},"DOI":"10.24406/publica-5187","URL":"https://doi.org/10.24406/publica-5187","source":"datacite"},{"id":"doi:10.5281/zenodo.22148845","type":"article-journal","title":"The Connect.CI Portal: Building and Sustaining a Cyberinfrastructure Community Platform","abstract":"ConnectCI is a web-based platform that facilitates collaboration across the national research computing ecosystem. Since its introduction at PEARC21, the platform has evolved from a project management tool for regional cyberteams to a comprehensive community infrastructure hosting thirteen community sites, with four major programs under active development. This paper presents five years of growth in new features and new communities, and through the addition of AI-powered support tools. These capabilities reflect three core functions: connecting professionals across institutional boundaries, enabling learning through shared knowledge and events, and advancing careers through recognition, mentorship, and workforce development pathways. Technical advances include a multi-tenant Drupal architecture enabling per-community customization within a single codebase, CILogon-based federated authentication, and a retrieval-augmented generation chatbot and Model Context Protocol servers for programmatic access to community data. We report on adoption metrics, lessons learned from operating a multi-stakeholder platform, and ConnectCI’s role in strengthening the research computing workforce.","author":[{"family":"Ma","given":"Julie"},{"family":"Fein","given":"Lissie"},{"family":"Pasquale","given":"Andrew"},{"family":"Gazula","given":"Vikram"},{"family":"Brandt","given":"Kevin"},{"family":"Chakravorty","given":"Dhruva"},{"family":"Chalker","given":"Alan"},{"family":"Figurelle","given":"Wayne"},{"family":"Sherman","given":"Andrew"},{"family":"Ghahramani","given":"Forough"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22148845","URL":"https://doi.org/10.5281/zenodo.22148845","source":"datacite"},{"id":"doi:10.5281/zenodo.22148846","type":"article-journal","title":"The Connect.CI Portal: Building and Sustaining a Cyberinfrastructure Community Platform","abstract":"ConnectCI is a web-based platform that facilitates collaboration across the national research computing ecosystem. Since its introduction at PEARC21, the platform has evolved from a project management tool for regional cyberteams to a comprehensive community infrastructure hosting thirteen community sites, with four major programs under active development. This paper presents five years of growth in new features and new communities, and through the addition of AI-powered support tools. These capabilities reflect three core functions: connecting professionals across institutional boundaries, enabling learning through shared knowledge and events, and advancing careers through recognition, mentorship, and workforce development pathways. Technical advances include a multi-tenant Drupal architecture enabling per-community customization within a single codebase, CILogon-based federated authentication, and a retrieval-augmented generation chatbot and Model Context Protocol servers for programmatic access to community data. We report on adoption metrics, lessons learned from operating a multi-stakeholder platform, and ConnectCI’s role in strengthening the research computing workforce.","author":[{"family":"Ma","given":"Julie"},{"family":"Fein","given":"Lissie"},{"family":"Pasquale","given":"Andrew"},{"family":"Gazula","given":"Vikram"},{"family":"Brandt","given":"Kevin"},{"family":"Chakravorty","given":"Dhruva"},{"family":"Chalker","given":"Alan"},{"family":"Figurelle","given":"Wayne"},{"family":"Sherman","given":"Andrew"},{"family":"Ghahramani","given":"Forough"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22148846","URL":"https://doi.org/10.5281/zenodo.22148846","source":"datacite"},{"id":"doi:10.24433/co.2155937.v1","type":"article-journal","title":"FusionNet Lite: Lightweight Sequence Fusion for Predictive Maintenance in Industrial IoT","abstract":"Reproducibility capsule for FusionNet Lite, a lightweight sequence-fusion deep-learning pipeline for predictive maintenance in industrial IoT systems. The campaign evaluates FusionNet Lite alongside CNN, BiLSTM, MLP, and CNN-LSTM baselines using centralized and FedAvg federated training across five seeds. The capsule includes leakage-safe preprocessing, stratified AI4I splitting, event-aware chronological MetroPT3 splitting with all rows retained, Dirichlet label non-IID client allocation, integrated-gradients analysis, statistical comparisons, publication tables and figures, and a consistency audit. Dataset files are placeholders in this capsule. Full reproducible execution requires attaching the authorised AI4I 2020 and MetroPT3 datasets under /data before selecting full mode.","author":[{"family":"Sharma","given":"Aman"},{"family":"Sim","given":"Kwan"},{"family":"Chandrasekaran","given":"Siva"}],"issued":{"date-parts":[[2026]]},"DOI":"10.24433/co.2155937.v1","URL":"https://doi.org/10.24433/co.2155937.v1","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.27063","type":"manuscript","title":"An Accurate and Single-Communication Federated Inference Algorithm","abstract":"Joint analyses across multiple institutions are increasingly important in biomedical and epidemiological research, particularly for rare diseases where datasets are typical small. However, privacy regulations and institutional policies often prevent the sharing of individual-level patient data. In this paper we present an accurate and single-communication federated inference algorithm. Single-communication federated inference enables statistical analyses through a single exchange of summary statistics between participating centers and a coordinating server, preserving privacy while reducing communication and computational costs compared with iterative federated learning. We extend a recently proposed single-communication federated inference strategy that is based on second-order Taylor expansions by using third-order expansions to better approximate local log-likelihood functions. The proposed method is evaluated through simulation studies based on real data and compared with existing federated inference strategies. The simulation studies assess the performance of the proposed method, with a particular focus on scenarios involving small local sample sizes, where quadratic approximations may fail to capture skewness and other higher-order characteristics of the log-likelihood function. They demonstrate that incorporating higher-order information of the log-likelihood function improves the accuracy while preserving the privacy, communication efficiency, and scalability required for collaborative biomedical and epidemiological research.","author":[{"family":"Montagnani","given":"Laura"},{"family":"Coolen","given":"Anthony"},{"family":"Jonker","given":"Marianne"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.27063","URL":"https://doi.org/10.48550/arxiv.2608.27063","source":"datacite"},{"id":"doi:10.5281/zenodo.20625872","type":"article-journal","title":"An Overview of Extreme Learning Machine-Based Intelligent Fault Localization Approaches","abstract":"Abstract The reliable operation of modern power systems depends heavily on the rapid detection and accurate localization of faults occurring in transmission and distribution networks. With the increasing integration of renewable energy resources, smart grid technologies, distributed generation, and advanced communication infrastructures, conventional fault localization methods face significant challenges related to network complexity, dynamic operating conditions, and measurement uncertainties. In recent years, Artificial Intelligence (AI)-based techniques have emerged as promising alternatives for enhancing the accuracy and efficiency of fault localization processes. Among these techniques, the Extreme Learning Machine (ELM) has gained considerable attention due to its fast learning capability, low computational complexity, excellent generalization performance, and suitability for real-time applications. This paper presents a comprehensive overview of ELM-based intelligent fault localization approaches developed for power system protection and monitoring. The review discusses the fundamental principles of fault localization and the theoretical foundations of ELM, including its architecture, learning mechanism, and major variants such as Online Sequential ELM, Kernel ELM, Weighted ELM, and Deep ELM. Furthermore, existing research contributions employing ELM for transmission lines, distribution networks, microgrids, renewable energy-integrated systems, and smart grid environments are systematically analyzed and compared. The paper also examines various signal processing and feature extraction techniques used in conjunction with ELM, including wavelet transforms, empirical mode decomposition, and phasor measurement unit-based approaches. Performance metrics, implementation challenges, and comparative advantages of ELM over traditional machine learning methods are critically evaluated. Finally, current research gaps and future directions are identified, highlighting opportunities in explainable artificial intelligence, federated learning, digital twins, edge computing, and hybrid intelligent fault localization frameworks. The findings indicate that ELM-based approaches offer a promising and computationally efficient solution for next-generation intelligent fault localization systems, contributing significantly to the development of reliable, adaptive, and resilient smart power grids. Keywords: Fault Localization, Extreme Learning Machine, Artificial Intelligence, Smart Grid, Distribution Networks, Power System Protection, Machine Learning, Intelligent Fault Diagnosis","author":[{"family":"Raut","given":"Priyanka"},{"family":"Dongre","given":"Kiran"},{"family":"Bhagat","given":"Amol"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20625872","URL":"https://doi.org/10.5281/zenodo.20625872","source":"datacite"},{"id":"doi:10.5281/zenodo.20625873","type":"article-journal","title":"An Overview of Extreme Learning Machine-Based Intelligent Fault Localization Approaches","abstract":"Abstract The reliable operation of modern power systems depends heavily on the rapid detection and accurate localization of faults occurring in transmission and distribution networks. With the increasing integration of renewable energy resources, smart grid technologies, distributed generation, and advanced communication infrastructures, conventional fault localization methods face significant challenges related to network complexity, dynamic operating conditions, and measurement uncertainties. In recent years, Artificial Intelligence (AI)-based techniques have emerged as promising alternatives for enhancing the accuracy and efficiency of fault localization processes. Among these techniques, the Extreme Learning Machine (ELM) has gained considerable attention due to its fast learning capability, low computational complexity, excellent generalization performance, and suitability for real-time applications. This paper presents a comprehensive overview of ELM-based intelligent fault localization approaches developed for power system protection and monitoring. The review discusses the fundamental principles of fault localization and the theoretical foundations of ELM, including its architecture, learning mechanism, and major variants such as Online Sequential ELM, Kernel ELM, Weighted ELM, and Deep ELM. Furthermore, existing research contributions employing ELM for transmission lines, distribution networks, microgrids, renewable energy-integrated systems, and smart grid environments are systematically analyzed and compared. The paper also examines various signal processing and feature extraction techniques used in conjunction with ELM, including wavelet transforms, empirical mode decomposition, and phasor measurement unit-based approaches. Performance metrics, implementation challenges, and comparative advantages of ELM over traditional machine learning methods are critically evaluated. Finally, current research gaps and future directions are identified, highlighting opportunities in explainable artificial intelligence, federated learning, digital twins, edge computing, and hybrid intelligent fault localization frameworks. The findings indicate that ELM-based approaches offer a promising and computationally efficient solution for next-generation intelligent fault localization systems, contributing significantly to the development of reliable, adaptive, and resilient smart power grids. Keywords: Fault Localization, Extreme Learning Machine, Artificial Intelligence, Smart Grid, Distribution Networks, Power System Protection, Machine Learning, Intelligent Fault Diagnosis","author":[{"family":"Raut","given":"Priyanka"},{"family":"Dongre","given":"Kiran"},{"family":"Bhagat","given":"Amol"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20625873","URL":"https://doi.org/10.5281/zenodo.20625873","source":"datacite"},{"id":"doi:10.5281/zenodo.22119317","type":"article-journal","title":"Coded evidence base for a PRISMA-guided systematic review of reporting fragmentation in federated-learning-based intrusion detection systems","abstract":"The complete coded evidence base underlying a PRISMA-guided systematic review of reporting fragmentation in federated-learning-based network intrusion detection systems (FL-NIDS): the per-study extraction table for all 105 included primary studies, the coding schema and codebook, the database search strings, the PRISMA flow and provenance note, and the supplementary records for the adversarial-evaluation and explainability dimensions. What this deposit supports. Computational reproducibility: the released table is sufficient to reproduce every aggregate numerical result reported in the manuscript. The included script verify_headline_results.py computes and reports the check count at runtime; in the validated run it reported 23 checks, 23 matched, 0 mismatched. Document identity: through title, authors, year, venue and DOI for all 105 studies. Independent re-coding: a reader with lawful access to the original publications can re-code selected studies using the published schema. What it does not support. Source-quote auditability: source excerpts and exact source passages are withheld from the public deposit (see WITHHELD_FIELDS.md). Auditing an individual coding decision against its source passage still requires access to the original publication. Coding was performed by a single coder; no inter-rater reliability statistic exists. Licence scope. CC BY 4.0 applies to the depositors' original material only: the coded values, extraction schema, codebook, search strings, schema-trigger metadata, analysis scripts, manifests and documentation. The record also includes bibliographic metadata identifying the reviewed publications; rights in those underlying publications remain with their respective authors and rights holders. No publisher PDFs, article full texts, copyrighted tables or figures, or verbatim excerpts from the reviewed publications are included. No licence is granted by the depositors over any underlying third-party work.","author":[{"family":"Alzoubi","given":"Mohamad"},{"family":"Serrão","given":"Carlos"},{"family":"Pavia","given":"João"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22119317","URL":"https://doi.org/10.5281/zenodo.22119317","source":"datacite"},{"id":"doi:10.5281/zenodo.22119316","type":"article-journal","title":"Coded evidence base for a PRISMA-guided systematic review of reporting fragmentation in federated-learning-based intrusion detection systems","abstract":"The complete coded evidence base underlying a PRISMA-guided systematic review of reporting fragmentation in federated-learning-based network intrusion detection systems (FL-NIDS): the per-study extraction table for all 105 included primary studies, the coding schema and codebook, the database search strings, the PRISMA flow and provenance note, and the supplementary records for the adversarial-evaluation and explainability dimensions. What this deposit supports. Computational reproducibility: the released table is sufficient to reproduce every aggregate numerical result reported in the manuscript. The included script verify_headline_results.py computes and reports the check count at runtime; in the validated run it reported 23 checks, 23 matched, 0 mismatched. Document identity: through title, authors, year, venue and DOI for all 105 studies. Independent re-coding: a reader with lawful access to the original publications can re-code selected studies using the published schema. What it does not support. Source-quote auditability: source excerpts and exact source passages are withheld from the public deposit (see WITHHELD_FIELDS.md). Auditing an individual coding decision against its source passage still requires access to the original publication. Coding was performed by a single coder; no inter-rater reliability statistic exists. Licence scope. CC BY 4.0 applies to the depositors' original material only: the coded values, extraction schema, codebook, search strings, schema-trigger metadata, analysis scripts, manifests and documentation. The record also includes bibliographic metadata identifying the reviewed publications; rights in those underlying publications remain with their respective authors and rights holders. No publisher PDFs, article full texts, copyrighted tables or figures, or verbatim excerpts from the reviewed publications are included. No licence is granted by the depositors over any underlying third-party work.","author":[{"family":"Alzoubi","given":"Mohamad"},{"family":"Serrão","given":"Carlos"},{"family":"Pavia","given":"João"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22119316","URL":"https://doi.org/10.5281/zenodo.22119316","source":"datacite"},{"id":"doi:10.5281/zenodo.20067237","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20067237","URL":"https://doi.org/10.5281/zenodo.20067237","source":"datacite"},{"id":"doi:10.5281/zenodo.20119054","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20119054","URL":"https://doi.org/10.5281/zenodo.20119054","source":"datacite"},{"id":"doi:10.5281/zenodo.20134134","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20134134","URL":"https://doi.org/10.5281/zenodo.20134134","source":"datacite"},{"id":"doi:10.5281/zenodo.20558450","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20558450","URL":"https://doi.org/10.5281/zenodo.20558450","source":"datacite"},{"id":"doi:10.5281/zenodo.20557954","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20557954","URL":"https://doi.org/10.5281/zenodo.20557954","source":"datacite"},{"id":"doi:10.5281/zenodo.7221216","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.7221216","URL":"https://doi.org/10.5281/zenodo.7221216","source":"datacite"},{"id":"doi:10.5281/zenodo.20021611","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20021611","URL":"https://doi.org/10.5281/zenodo.20021611","source":"datacite"},{"id":"doi:10.4230/lipics.itc.2026.8","type":"article-journal","title":"Fast Bounded-Independence Functions and Their Duals","abstract":"We continue the study of fast functions, computable by linear-size circuits, that share useful properties of random functions. Motivated by cryptographic applications, we generalize and improve on previous results in this area, obtaining the following results: - For any constant t, we construct a fast t-wise independent hash function with algebraic degree log₂ t (over F₂), simultaneously optimizing both asymptotic circuit size and degree. - We simplify and improve a recent construction (ITCS 2026) of a family of fast codes with fast duals, both meeting the Gilbert-Varshamov bound. Unlike the previous construction, our construction has negligible failure probability, can accommodate general fields and rates, supports a systematic encoding, and admits fast universal encoders. - We strengthen the above to support stronger random-like properties, such as optimal combinatorial list-decoding. This is achieved by constructing, for any constant t, a family of fast linear functions that map any t linearly independent inputs to uniform and statistically independent outputs. Prior to our work, this was only known for t = 1. We demonstrate the usefulness of the above results to cryptography. This includes the first nontrivial protocols for perfectly secure multiparty computation whose circuit complexity scales linearly with the number of parties, as well as protocols for computing encrypted matrix-vector products with optimal asymptotic circuit complexity.","author":[{"family":"Brehm","given":"Martijn"},{"family":"Ishai","given":"Yuval"},{"family":"Resch","given":"Nicolas"}],"issued":{"date-parts":[[2026]]},"DOI":"10.4230/lipics.itc.2026.8","URL":"https://doi.org/10.4230/lipics.itc.2026.8","source":"datacite"},{"id":"doi:10.4230/lipics.itc.2026.7","type":"article-journal","title":"Compressing Correlations via Secret Replication: PCFs from Symmetric Cryptography","abstract":"We revisit the question of securely compressing multiparty correlations using only symmetric cryptography. A linear correlation C, defined by a linear subspace C ⊆ 𝔽ⁿ, samples a secret random 𝐜 ∈ C and assigns to each party a fixed subset of the entries of 𝐜. Gilboa and Ishai (Crypto 1999) and Cramer, Damgård and Ishai (TCC 2005) provide a general technique for securely compressing many independent samples from C by replicating independent keys of a pseudorandom function (PRF) among the parties. This implies a pseudorandom correlation function (PCF) for C from any PRF, where the PCF key size scales with the number of minimal-support codewords in C. We observe that the above generalizes to other types of useful target correlations C_T by using a secret replication pattern obtained via a random secret assignment of parties in C to parties in C_T. We present several corollaries of this general blueprint. These include a re-derivation of two-party PCF constructions for VOLE and subfield-VOLE over small domains (Roy, Crypto 2022) as well as new multiparty PCFs for small-domain VOLE-style correlations, including scalar-vector multiplication triples and their authenticated variants. Finally, we discuss applications to secure computation.","author":[{"family":"Ishai","given":"Yuval"},{"family":"Krawczyk","given":"Hugo"},{"family":"Rabin","given":"Tal"}],"issued":{"date-parts":[[2026]]},"DOI":"10.4230/lipics.itc.2026.7","URL":"https://doi.org/10.4230/lipics.itc.2026.7","source":"datacite"},{"id":"doi:10.5281/zenodo.20810243","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20810243","URL":"https://doi.org/10.5281/zenodo.20810243","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.22115","type":"manuscript","title":"Game-Theoretic Framework for Private Data Sharing in Vehicular Networks","abstract":"We present a novel game-theoretic framework designed to enhance privacy and scalability in decentralized vehicular data collection systems. The proposed hybrid architecture comprises vehicles that supply sensor data, independent servers that process data via secure multiparty computation, a coordinator node that manages data flow, and data consumers that set economic incentives. Crucially, our framework ensures that only the data consumer can access the fully aggregated data, preventing individual raw data exposure and significantly reducing privacy risks. By integrating principles of the Stackelberg competition from game theory, our approach dynamically balances privacy and economic incentives, enabling vehicles to make participation decisions based on perceived privacy risks and incentives. We empirically validate our framework using real-world vehicular location data, quantifying privacy risks by evaluating the accuracy with which a potential adversary can reconstruct a vehicle's path using only a subset of the shared data. This paper details the development and deployment of a data-trading platform within this framework, introducing a practical and privacy-preserving marketplace for profitable vehicle data sharing. Through experiments and simulations, we evaluate the effectiveness of the system in preserving privacy and explore the dynamics that influence vehicle participation. Our findings highlight the robustness of the proposed framework in preserving privacy while supporting an active data market.","author":[{"family":"Alsaqabi","given":"Yousef"},{"family":"Zhou","given":"Yinan"},{"family":"Nawab","given":"Faisal"},{"family":"Krishnamachari","given":"Bhaskar"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.22115","URL":"https://doi.org/10.48550/arxiv.2606.22115","source":"datacite"},{"id":"doi:10.5281/zenodo.20798902","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20798902","URL":"https://doi.org/10.5281/zenodo.20798902","source":"datacite"},{"id":"doi:10.5281/zenodo.20556781","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20556781","URL":"https://doi.org/10.5281/zenodo.20556781","source":"datacite"},{"id":"doi:10.5281/zenodo.20540410","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20540410","URL":"https://doi.org/10.5281/zenodo.20540410","source":"datacite"},{"id":"doi:10.5281/zenodo.20508037","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20508037","URL":"https://doi.org/10.5281/zenodo.20508037","source":"datacite"},{"id":"doi:10.5281/zenodo.20492877","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20492877","URL":"https://doi.org/10.5281/zenodo.20492877","source":"datacite"},{"id":"doi:10.5281/zenodo.20289411","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20289411","URL":"https://doi.org/10.5281/zenodo.20289411","source":"datacite"},{"id":"doi:10.5281/zenodo.20283494","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20283494","URL":"https://doi.org/10.5281/zenodo.20283494","source":"datacite"},{"id":"doi:10.5281/zenodo.20135122","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20135122","URL":"https://doi.org/10.5281/zenodo.20135122","source":"datacite"},{"id":"doi:10.5281/zenodo.20123414","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20123414","URL":"https://doi.org/10.5281/zenodo.20123414","source":"datacite"},{"id":"doi:10.5281/zenodo.20119622","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20119622","URL":"https://doi.org/10.5281/zenodo.20119622","source":"datacite"},{"id":"doi:10.5281/zenodo.20117892","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20117892","URL":"https://doi.org/10.5281/zenodo.20117892","source":"datacite"},{"id":"doi:10.5281/zenodo.20086001","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20086001","URL":"https://doi.org/10.5281/zenodo.20086001","source":"datacite"},{"id":"doi:10.5281/zenodo.20021628","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20021628","URL":"https://doi.org/10.5281/zenodo.20021628","source":"datacite"},{"id":"doi:10.5281/zenodo.19913411","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19913411","URL":"https://doi.org/10.5281/zenodo.19913411","source":"datacite"},{"id":"doi:10.5281/zenodo.19734601","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19734601","URL":"https://doi.org/10.5281/zenodo.19734601","source":"datacite"},{"id":"doi:10.5281/zenodo.19730079","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19730079","URL":"https://doi.org/10.5281/zenodo.19730079","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.13474","type":"manuscript","title":"Secure and Privacy-Preserving Vertical Federated Learning","abstract":"We propose a novel end-to-end privacy-preserving framework, instantiated by three efficient protocols for different deployment scenarios, covering both input and output privacy, for the vertically split scenario in federated learning (FL), where features are split across clients and labels are not shared by all parties. We do so by distributing the role of the aggregator in FL into multiple servers and having them run secure multiparty computation (MPC) protocols to perform model and feature aggregation and apply differential privacy (DP) to the final released model. While a naive solution would have the clients delegating the entirety of training to run in MPC between the servers, our optimized solution, which supports purely global and also global-local models updates with privacy-preserving, drastically reduces the amount of computation and communication performed using multiparty computation. The experimental results also show the effectiveness of our protocols.","author":[{"family":"Jin","given":"Shan"},{"family":"Rachuri","given":"Sai"},{"family":"Wang","given":"Yizhen"},{"family":"Nascimento","given":"Anderson"},{"family":"Cai","given":"Yiwei"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.13474","URL":"https://doi.org/10.48550/arxiv.2604.13474","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.09975","type":"manuscript","title":"EncFormer: Secure and Efficient Transformer Inference over Encrypted Data","abstract":"Transformer inference in machine-learning-as-a-service (MLaaS) raises privacy concerns for sensitive user inputs. Prior secure solutions that combine fully homomorphic encryption (FHE) and secure multiparty computation (MPC) are bottlenecked by inefficient FHE kernels, communication-heavy MPC protocols, and expensive FHE-MPC conversions. We present EncFormer, a two-party private Transformer inference framework that introduces Stage Compatible Patterns so that FHE kernels compose efficiently, reducing repacking and conversions. EncFormer also provides a cost analysis model built around a minimal-conversion baseline, enabling principled selection of FHE-MPC boundaries. To further reduce communication, EncFormer proposes a secure complex CKKS-MPC conversion protocol and designs communication-efficient MPC protocols for nonlinearities. With GPU optimizations, evaluations on GPT- and BERT-style models show that EncFormer achieves 1.4x-30.4x lower online MPC communication and 1.3x-9.8x lower end-to-end latency against prior hybrid FHE-MPC systems, and 1.9x-3.5x lower end-to-end latency on BERT-base than FHE-only pipelines under a matched backend, while maintaining near-plaintext accuracy on selected GLUE tasks.","author":[{"family":"Zhu","given":"Yufan"},{"family":"Jin","given":"Chao"},{"family":"Aung","given":"Khin"},{"family":"Xiao","given":"Xiaokui"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.09975","URL":"https://doi.org/10.48550/arxiv.2604.09975","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.13570","type":"manuscript","title":"Privacy-Preserving Machine Learning for IoT: A Cross-Paradigm Survey and Future Roadmap","abstract":"The rapid proliferation of the Internet of Things has intensified demand for robust privacy-preserving machine learning mechanisms to safeguard sensitive data generated by large-scale, heterogeneous, and resource-constrained devices. Unlike centralized environments, IoT ecosystems are inherently decentralized, bandwidth-limited, and latency-sensitive, exposing privacy risks across sensing, communication, and distributed training pipelines. These characteristics render conventional anonymization and centralized protection strategies insufficient for practical deployments. This survey presents a comprehensive IoT-centric, cross-paradigm analysis of privacy-preserving machine learning. We introduce a structured taxonomy spanning perturbation-based mechanisms such as differential privacy, distributed paradigms such as federated learning, cryptographic approaches including homomorphic encryption and secure multiparty computation, and generative synthesis techniques based on generative adversarial networks. For each paradigm, we examine formal privacy guarantees, computational and communication complexity, scalability under heterogeneous device participation, and resilience against threats including membership inference, model inversion, gradient leakage, and adversarial manipulation. We further analyze deployment constraints in wireless IoT environments, highlighting trade-offs between privacy, communication overhead, model convergence, and system efficiency within next-generation mobile architectures. We also consolidate evaluation methodologies, summarize representative datasets and open-source frameworks, and identify open challenges including hybrid privacy integration, energy-aware learning, privacy-preserving large language models, and quantum-resilient machine learning.","author":[{"family":"Zaman","given":"Zakia"},{"family":"Gauravaram","given":"Praveen"},{"family":"Hassan","given":"Mahbub"},{"family":"Jha","given":"Sanjay"},{"family":"Hu","given":"Wen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.13570","URL":"https://doi.org/10.48550/arxiv.2603.13570","source":"datacite"},{"id":"doi:10.5281/zenodo.19328321","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19328321","URL":"https://doi.org/10.5281/zenodo.19328321","source":"datacite"},{"id":"doi:10.5281/zenodo.19203631","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19203631","URL":"https://doi.org/10.5281/zenodo.19203631","source":"datacite"},{"id":"doi:10.5281/zenodo.19202486","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19202486","URL":"https://doi.org/10.5281/zenodo.19202486","source":"datacite"},{"id":"doi:10.5281/zenodo.19200883","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19200883","URL":"https://doi.org/10.5281/zenodo.19200883","source":"datacite"},{"id":"doi:10.5281/zenodo.18921535","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18921535","URL":"https://doi.org/10.5281/zenodo.18921535","source":"datacite"},{"id":"doi:10.5281/zenodo.18921288","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18921288","URL":"https://doi.org/10.5281/zenodo.18921288","source":"datacite"},{"id":"doi:10.5281/zenodo.20492206","type":"article-journal","title":"Autonomous Procurement Systems for Resilient Global Supply Networks","abstract":"Abstract Global supply chains are increasingly exposed to disruptions arising from geopolitical conflicts, trade sanctions, climate-related events, logistics bottlenecks, supplier financial instability, and evolving ESG requirements. Traditional procurement systems remain largely reactive, relying on periodic supplier assessments, static risk scoring mechanisms, and fragmented decision-making processes that are insufficient for managing modern multi-tier supply networks. This whitepaper introduces Autonomous Procurement Systems (APS), a conceptual research framework that integrates Multi-Source Risk Intelligence, Graph-Based Supply Network Analysis, Digital Twin Simulation, Multi-Objective Optimization, and Agentic AI into a unified architecture for resilient procurement decision-making. The framework proposes a five-layer model capable of continuously monitoring heterogeneous risk signals, modeling risk propagation across supplier networks, simulating disruption scenarios, optimizing sourcing strategies under uncertainty, and supporting governed autonomous procurement actions. A key contribution of the framework is its explicit consideration of non-traditional procurement risks, including climate emergencies, warfare, sanctions, geopolitical instability, infrastructure disruptions, and systemic supply chain shocks. The paper introduces several novel concepts, including a Conflict Impact Propagation Model (CIPM), a Procurement Disruption Scenario Ontology (PDSO), a Graduated Autonomy Framework for Procurement AI, and a Synthetic-to-Real Data Strategy for future machine learning research in procurement. Rather than advocating a specific algorithmic approach, the framework adopts an algorithm-agnostic research philosophy, identifying and comparing candidate techniques from machine learning, graph neural networks, operations research, reinforcement learning, digital twins, federated learning, and quantum-inspired optimization. The objective is to provide a rigorous foundation for future empirical research, enterprise innovation initiatives, and the development of next-generation procurement intelligence platforms. This publication is intended as a foundational research artifact for academics, operations research practitioners, supply chain professionals, innovation teams, and enterprise architects exploring the future of resilient and intelligent procurement systems. Keywords Procurement AI; Supply Chain Resilience; Supplier Risk Management; Graph Neural Networks; Digital Twin; Operations Research; Reinforcement Learning; Agentic AI; Strategic Sourcing; Supply Chain Risk Intelligence; Geopolitical Risk; Climate Risk; ESG Procurement; Multi-Objective Optimization; Federated Learning; Autonomous Procurement Systems. Authors Somnath Banerjee, Subhamoy Bhaduri, and Binayak Mukherjee. Version Version 1.0 (Foundation Concept Paper), June 2026.","author":[{"family":"Banerjee","given":"Somnath"},{"family":"Bhaduri","given":"Subhamoy"},{"family":"Mukherjee","given":"Binayak"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20492206","URL":"https://doi.org/10.5281/zenodo.20492206","source":"datacite"},{"id":"doi:10.5281/zenodo.20492207","type":"article-journal","title":"Autonomous Procurement Systems for Resilient Global Supply Networks","abstract":"Abstract Global supply chains are increasingly exposed to disruptions arising from geopolitical conflicts, trade sanctions, climate-related events, logistics bottlenecks, supplier financial instability, and evolving ESG requirements. Traditional procurement systems remain largely reactive, relying on periodic supplier assessments, static risk scoring mechanisms, and fragmented decision-making processes that are insufficient for managing modern multi-tier supply networks. This whitepaper introduces Autonomous Procurement Systems (APS), a conceptual research framework that integrates Multi-Source Risk Intelligence, Graph-Based Supply Network Analysis, Digital Twin Simulation, Multi-Objective Optimization, and Agentic AI into a unified architecture for resilient procurement decision-making. The framework proposes a five-layer model capable of continuously monitoring heterogeneous risk signals, modeling risk propagation across supplier networks, simulating disruption scenarios, optimizing sourcing strategies under uncertainty, and supporting governed autonomous procurement actions. A key contribution of the framework is its explicit consideration of non-traditional procurement risks, including climate emergencies, warfare, sanctions, geopolitical instability, infrastructure disruptions, and systemic supply chain shocks. The paper introduces several novel concepts, including a Conflict Impact Propagation Model (CIPM), a Procurement Disruption Scenario Ontology (PDSO), a Graduated Autonomy Framework for Procurement AI, and a Synthetic-to-Real Data Strategy for future machine learning research in procurement. Rather than advocating a specific algorithmic approach, the framework adopts an algorithm-agnostic research philosophy, identifying and comparing candidate techniques from machine learning, graph neural networks, operations research, reinforcement learning, digital twins, federated learning, and quantum-inspired optimization. The objective is to provide a rigorous foundation for future empirical research, enterprise innovation initiatives, and the development of next-generation procurement intelligence platforms. This publication is intended as a foundational research artifact for academics, operations research practitioners, supply chain professionals, innovation teams, and enterprise architects exploring the future of resilient and intelligent procurement systems. Keywords Procurement AI; Supply Chain Resilience; Supplier Risk Management; Graph Neural Networks; Digital Twin; Operations Research; Reinforcement Learning; Agentic AI; Strategic Sourcing; Supply Chain Risk Intelligence; Geopolitical Risk; Climate Risk; ESG Procurement; Multi-Objective Optimization; Federated Learning; Autonomous Procurement Systems. Authors Somnath Banerjee, Subhamoy Bhaduri, and Binayak Mukherjee. Version Version 1.0 (Foundation Concept Paper), June 2026.","author":[{"family":"Banerjee","given":"Somnath"},{"family":"Bhaduri","given":"Subhamoy"},{"family":"Mukherjee","given":"Binayak"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20492207","URL":"https://doi.org/10.5281/zenodo.20492207","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.25794","type":"manuscript","title":"Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data","abstract":"Federated Learning (FL) enables distributed training of machine learning models while preserving data privacy. However, FL struggles with heterogeneous, non-IID client data distributions, resulting in sub-optimal and biased global models. In this paper, we propose pFedMARL, a novel approach leveraging Multi-Agent Reinforcement Learning (MARL) with Twin Delayed Deep Deterministic Policy Gradient (TD3) to dynamically adapt aggregation strategies in FL settings. Our method employs a server-side agent adjusting client contributions to optimize global model robustness and client-side agents balancing global and local updates to personalize models effectively without pre-training. We demonstrate superior performance of pFedMARL for training a semi-supervised audio spectrogram transformer, matching or outperforming FedAvg, Ditto, and local training approaches across multiple non-IID scenarios and in the presence of adversarial clients. Our results indicate that pFedMARL actively improves accuracy, robustness, and fairness, making it suitable for real-world deployments.","author":[{"family":"Glitza","given":"Rene"},{"family":"Becker","given":"Luca"},{"family":"Martin","given":"Rainer"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.25794","URL":"https://doi.org/10.48550/arxiv.2608.25794","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.25514","type":"manuscript","title":"Joint Beamforming Design and Port Selection in Fluid Antenna-Assisted Multi-Cell Networks: A Personalized Federated Learning Approach","abstract":"This paper investigates joint beamforming and port selection in multi-cell fluid antenna-assisted (FAS) networks. In such networks, active beamforming and discrete FA port selection are coupled through intra-cell and inter-cell interference and are jointly optimized to maximize the weighted sum-rate (WSR). We develop a federated representation learning (FedRep) framework with a position-aware dual-branch deep neural network (PA-DNN). The PA-DNN uses channel state information and port positional encoding as inputs, and jointly outputs beamforming vectors and port selections through two task-specific branches. To support decentralized training across heterogeneous cells, the FedRep framework shares global beamforming-related parameters among base stations while keeping port-selection parameters local for cell-specific adaptation. Simulation results show that the proposed scheme achieves a higher weighted sum-rate than conventional FL and port-selection benchmark schemes.","author":[{"family":"Gao","given":"Liwen"},{"family":"Zheng","given":"Li"},{"family":"Hao","given":"Xing"},{"family":"Chen","given":"Ziru"},{"family":"Cai","given":"Lin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.25514","URL":"https://doi.org/10.48550/arxiv.2608.25514","source":"datacite"},{"id":"doi:10.5281/zenodo.20426044","type":"article-journal","title":"HemaFAIR & HELIOS Federated Analysis Hackathon (Cyprus, May 2026): Training Resources and Datasets","abstract":"This Zenodo record contains the training and educational materials developed and used during the HemaFAIR and HELIOS Federated Data Analysis Training School and Hackathon, held at Alion Beach Hotel in Ayia Napa, Cyprus, on 11–14 May 2026. The event brought together researchers, clinicians, data scientists, and FAIR data experts to provide hands-on training on federated data analysis, FAIR data principles, semantic interoperability, common data models, and privacy-preserving approaches for biomedical and rare disease research, with a particular focus on hemoglobinopathies. The uploaded materials include: Training presentations and workshop materials delivered by the trainers Hands-on exercises and demonstration resources Example datasets used during the practical sessions Dataset dictionaries and metadata documentation Supporting materials related to FAIRification, semantic web technologies, OMOP common data models, and federated learning approaches The materials are shared to support capacity building, reproducibility, FAIR data practices, and collaborative research in rare diseases and hemoglobinopathies. These resources may be useful for researchers, clinicians, students, and institutions interested in federated analysis infrastructures, data interoperability, and responsible data sharing approaches. We would like to thank all trainers, organisers, and participants who contributed to the successful implementation of the training school and hackathon.","author":[{"family":"Waagmeester","given":"Andra"},{"family":"Kersloot","given":"Martijn"},{"family":"Cremonesi","given":"Francesco"},{"family":"Wijnbergen","given":"Daphne"},{"family":"Tamana","given":"Stella"},{"family":"Orphanou","given":"Kalia"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20426044","URL":"https://doi.org/10.5281/zenodo.20426044","source":"datacite"},{"id":"doi:10.5281/zenodo.20426045","type":"article-journal","title":"HemaFAIR & HELIOS Federated Analysis Hackathon (Cyprus, May 2026): Training Resources and Datasets","abstract":"This Zenodo record contains the training and educational materials developed and used during the HemaFAIR and HELIOS Federated Data Analysis Training School and Hackathon, held at Alion Beach Hotel in Ayia Napa, Cyprus, on 11–14 May 2026. The event brought together researchers, clinicians, data scientists, and FAIR data experts to provide hands-on training on federated data analysis, FAIR data principles, semantic interoperability, common data models, and privacy-preserving approaches for biomedical and rare disease research, with a particular focus on hemoglobinopathies. The uploaded materials include: Training presentations and workshop materials delivered by the trainers Hands-on exercises and demonstration resources Example datasets used during the practical sessions Dataset dictionaries and metadata documentation Supporting materials related to FAIRification, semantic web technologies, OMOP common data models, and federated learning approaches The materials are shared to support capacity building, reproducibility, FAIR data practices, and collaborative research in rare diseases and hemoglobinopathies. These resources may be useful for researchers, clinicians, students, and institutions interested in federated analysis infrastructures, data interoperability, and responsible data sharing approaches. We would like to thank all trainers, organisers, and participants who contributed to the successful implementation of the training school and hackathon.","author":[{"family":"Waagmeester","given":"Andra"},{"family":"Kersloot","given":"Martijn"},{"family":"Cremonesi","given":"Francesco"},{"family":"Wijnbergen","given":"Daphne"},{"family":"Tamana","given":"Stella"},{"family":"Orphanou","given":"Kalia"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20426045","URL":"https://doi.org/10.5281/zenodo.20426045","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.24073","type":"manuscript","title":"ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud Removal","abstract":"Low-earth-orbit (LEO) satellites enable high-resolution, large-scale Earth observation for applications such as disaster monitoring and environmental surveillance. However, cloud coverage often obscures the Earth's surface, and conventional cloud-removal pipelines that download cloudy images to ground stations for processing suffer from limited contact windows, constrained satellite-to-ground bandwidth, and high latency. In this work, we propose a novel satellite federated learning framework for cloud removal across LEO constellations, named orbital attention leaky integrate-and-fire (OrbitALIF). OrbitALIF performs both onboard training and inference using a compact 2.30,M-parameter spiking neural network (SNN) backbone with an adaptive gated fusion module (AGFM) and a spectral-spatial hybrid attention module (SHAM), combined with a decentralized federated learning strategy that shares model weights via inter-satellite links. Our experiments show that OrbitALIF achieves competitive cloud removal quality while consuming only 0.287,mJ per inference on neuromorphic hardware, a 72.3 times (98.6%) energy reduction versus an equivalent artificial neural network (ANN).","author":[{"family":"Zhang","given":"Bohan"},{"family":"Xu","given":"Chenyu"},{"family":"Mao","given":"Yijie"},{"family":"Shi","given":"Yuanming"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.24073","URL":"https://doi.org/10.48550/arxiv.2608.24073","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.22552","type":"manuscript","title":"Model-Consistent Byzantine-Resilient Decentralized Federated Learning for Collaborative Missions","abstract":"Decentralized federated learning (DFL) is a promising paradigm for autonomous nodes to collaboratively train AI models without relying on a central server. However, existing DFL solutions do not guarantee global model consistency, a critical requirement for collaborative mission-critical scenarios where model divergence undermines decision uniformity and safety. This lack of consistency also amplifies vulnerability to Byzantine adversaries, who exploit the decentralized network topology and weak synchrony to perform equivocation and model poisoning attacks against individual victims. This paper introduces DFL-C, a novel Byzantine-resilient DFL architecture that enables decentralized nodes to perform collaborative training with global model consistency. At its core, DFL-C integrates an asynchronous common subset (ACS) consensus protocol into the DFL workflow to ensure all nodes aggregate a uniform set of model updates to establish global model consistency, despite individual Byzantine equivocation. DFL-C further implements a dual-domain trust scoring mechanism to provide resilience against data-domain Byzantine manipulations including model poisoning attacks. This mechanism complements the consensus protocol, significantly reducing the latter's runtime. Our experimental results demonstrate that DFL-C maintains model accuracy while achieving global model consistency under Byzantine behaviors with moderate consensus overhead. Notably, when compared with the state-of-the-art DFL solution BALANCE (Fang et al.) that does not provide model consistency, DFL-C achieves better model accuracy against untargeted model poisoning attacks and comparable resilience against backdoor attacks, with the advantage widened under non-IID scenarios.","author":[{"family":"Li","given":"Yue"},{"family":"Bhujel","given":"Sudip"},{"family":"Lira","given":"Cameron"},{"family":"Wang","given":"Ning"},{"family":"Xiao","given":"Yang"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.22552","URL":"https://doi.org/10.48550/arxiv.2608.22552","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.20038","type":"manuscript","title":"An Inclusive and Lightweight Approach to Federated Continual Learning for Cultural Heritage","abstract":"Artificial intelligence can support cultural heritage and digital humanities through large-scale retrieval and analysis of digitized collections. However, cultural heritage data are often distributed across institutions, constrained by ownership and access restrictions, and continuously evolving over time. Federated Continual Learning (FCL) is well suited to this setting, as it enables models to learn from distributed and sequential data without sharing raw collections. In this paper, we propose FedCurv-DR, a lightweight, regularisation-based FCL strategy. The method accumulates parameter-importance estimates across clients and experiences to protect learned knowledge, while updating them only at fixed intervals to minimize communication and computation overhead. We evaluate FedCurv-DR in a continual learning scenario using the WikiArt image dataset for genre classification with evolving styles, reporting performance, energy, and fairness metrics. Our results show that FedCurv- DR reduces forgetting and balances performance, fairness, and energy efficiency for sustainable AI in cultural heritage.","author":[{"family":"Theologitis","given":"Ioannis"},{"family":"Meng","given":"Debin"},{"family":"Eleftheriadis","given":"Stylianos"},{"family":"Lolis","given":"Vasileios"},{"family":"Votis","given":"Konstantinos"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.20038","URL":"https://doi.org/10.48550/arxiv.2608.20038","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.18311","type":"manuscript","title":"FedCoRe: Target-Adaptive Completion for Missing Modalities in Healthcare Federated Learning","abstract":"Federated multimodal models often assume every site has every modality, although hospitals differ in access to EHRs, chest radiographs, and ECGs. We study this setting on a MIMIC-derived respiratory deterioration task with simulated FL clients and introduce FedCoRe (Federated Cross-Modal Representation Completion). FedCoRe learns representation- or logit-space corrections rather than generating synthetic ECGs or CXR images. When a client observes a modality that may be missing at deployment, it evaluates the same example with and without that modality to obtain paired supervision. Only clients with such pairs update the completion module, and validation may retain the unchanged prediction. We freeze the trained multimodal predictor during evaluation so that measured differences come only from completion. Hiding ECG reduced AUROC by about 0.085; paired-example FedAvg restored 0.0415 AUROC, or 49.0% of the lost performance. We therefore report two distinct effects: paired-example FedAvg partially recovers the missing-ECG gap, while validation-selected completion is a task-specific classifier-logit correction rather than literal ECG recovery. For CXR, effect-aware completion recovers 52.8% of the loss in a controlled test where CXR is hidden. Paired-example FedAvg transfers part of this effect, but validation keeps the no-completion baseline for deployment cases whose inputs lack CXR. Thus, FedCoRe should be read as a validation-gated completion/correction framework: it can recover missing-modality signal in supported settings, but it should be deployed only when paired examples and validation evidence support that modality.","author":[{"family":"Roth","given":"Holger"},{"family":"Xu","given":"Ziyue"},{"family":"Cnudde","given":"Peter"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.18311","URL":"https://doi.org/10.48550/arxiv.2608.18311","source":"datacite"},{"id":"doi:10.5281/zenodo.18662863","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18662863","URL":"https://doi.org/10.5281/zenodo.18662863","source":"datacite"},{"id":"doi:10.5281/zenodo.18662165","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18662165","URL":"https://doi.org/10.5281/zenodo.18662165","source":"datacite"},{"id":"doi:10.5281/zenodo.18378157","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18378157","URL":"https://doi.org/10.5281/zenodo.18378157","source":"datacite"},{"id":"doi:10.5281/zenodo.18377441","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18377441","URL":"https://doi.org/10.5281/zenodo.18377441","source":"datacite"},{"id":"doi:10.5281/zenodo.18374973","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18374973","URL":"https://doi.org/10.5281/zenodo.18374973","source":"datacite"},{"id":"doi:10.48550/arxiv.2601.09916","type":"manuscript","title":"Learning-Augmented Perfectly Secure Collaborative Matrix Multiplication","abstract":"This paper presents a perfectly secure matrix multiplication (PSMM) protocol for multiparty computation (MPC) of $\\mathrm{A}^{\\top}\\mathrm{B}$ over finite fields. The proposed scheme guarantees correctness and information-theoretic privacy against threshold-bounded, semi-honest colluding agents, under explicit local storage constraints. Our scheme encodes submatrices as evaluations of sparse masking polynomials and combines coefficient alignment with Beaver-style randomness to ensure perfect secrecy. We demonstrate that any colluding set of parties below the security threshold observes uniformly random shares, and that the recovery threshold is optimal, matching existing information-theoretic limits. Building on this framework, we introduce a learning-augmented extension that integrates tensor-decomposition-based local block multiplication, capturing both classical and learned low-rank methods. We demonstrate that the proposed learning-based PSMM preserves privacy and recovery guarantees for MPC, while providing scalable computational efficiency gains (up to $80\\%$) as the matrix dimensions grow.","author":[{"family":"He","given":"Zixuan"},{"family":"Salehi","given":"Mohammad"},{"family":"Malak","given":"Derya"},{"family":"Stavrou","given":"Photios"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2601.09916","URL":"https://doi.org/10.48550/arxiv.2601.09916","source":"datacite"},{"id":"doi:10.5281/zenodo.18017581","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.18017581","URL":"https://doi.org/10.5281/zenodo.18017581","source":"datacite"},{"id":"doi:10.5281/zenodo.17988137","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17988137","URL":"https://doi.org/10.5281/zenodo.17988137","source":"datacite"},{"id":"doi:10.5445/ir/1000188270","type":"article-journal","title":"PETs and AI: Privacy Washing and the Need for a PETs Evaluation Framework (Dagstuhl Seminar 25112)","abstract":"As public awareness of data collection practices and regulatory frameworks grows, privacy-enhancing technologies (PETs) have emerged as a promising approach to reconciling data utility with individual privacy rights. PETs underpin privacy-preserving machine learning (PPML), integrating tools like differential privacy, homomorphic encryption, and secure multiparty computation to safeguard data throughout the AI lifecycle. However, despite significant technical progress, PETs face critical policy and governance challenges. Recent works have raised concerns about efficacy and deployment of PETs, observing that fundamental rights of people are continually being harmed, including, paradoxically, privacy. PETs have been used in surveillance applications and as a privacy washing tool. Current approaches often fail to address broader harms beyond data protection, highlighting the need for a more comprehensive privacy evaluation framework. This Dagstuhl Seminar brought together scholars in computer science and law, along with policymakers, regulators, and industry leaders, to discuss privacy washing and the challenges of detecting privacy washing through PETs and explored pathways toward a framework to address these challenges.","author":[{"family":"Cristofaro","given":"Emiliano"},{"family":"Shrishak","given":"Kris"},{"family":"Strufe","given":"Thorsten"},{"family":"Troncoso","given":"Carmela"},{"family":"Morsbach","given":"Felix"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5445/ir/1000188270","URL":"https://doi.org/10.5445/ir/1000188270","source":"datacite"},{"id":"doi:10.5281/zenodo.17895475","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17895475","URL":"https://doi.org/10.5281/zenodo.17895475","source":"datacite"},{"id":"doi:10.5281/zenodo.17882317","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17882317","URL":"https://doi.org/10.5281/zenodo.17882317","source":"datacite"},{"id":"doi:10.5281/zenodo.17801567","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17801567","URL":"https://doi.org/10.5281/zenodo.17801567","source":"datacite"},{"id":"doi:10.5281/zenodo.17801124","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17801124","URL":"https://doi.org/10.5281/zenodo.17801124","source":"datacite"},{"id":"doi:10.5281/zenodo.17800374","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17800374","URL":"https://doi.org/10.5281/zenodo.17800374","source":"datacite"},{"id":"doi:10.5281/zenodo.17800098","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17800098","URL":"https://doi.org/10.5281/zenodo.17800098","source":"datacite"},{"id":"doi:10.5281/zenodo.17799841","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17799841","URL":"https://doi.org/10.5281/zenodo.17799841","source":"datacite"},{"id":"doi:10.5281/zenodo.17799229","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17799229","URL":"https://doi.org/10.5281/zenodo.17799229","source":"datacite"},{"id":"doi:10.5281/zenodo.17792555","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17792555","URL":"https://doi.org/10.5281/zenodo.17792555","source":"datacite"},{"id":"doi:10.5281/zenodo.17792444","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17792444","URL":"https://doi.org/10.5281/zenodo.17792444","source":"datacite"},{"id":"doi:10.5281/zenodo.17709279","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17709279","URL":"https://doi.org/10.5281/zenodo.17709279","source":"datacite"},{"id":"doi:10.5281/zenodo.17590348","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17590348","URL":"https://doi.org/10.5281/zenodo.17590348","source":"datacite"},{"id":"doi:10.5281/zenodo.17588916","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17588916","URL":"https://doi.org/10.5281/zenodo.17588916","source":"datacite"},{"id":"doi:10.5281/zenodo.17482705","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17482705","URL":"https://doi.org/10.5281/zenodo.17482705","source":"datacite"},{"id":"doi:10.5281/zenodo.17464296","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17464296","URL":"https://doi.org/10.5281/zenodo.17464296","source":"datacite"},{"id":"doi:10.5281/zenodo.17456557","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17456557","URL":"https://doi.org/10.5281/zenodo.17456557","source":"datacite"},{"id":"doi:10.5281/zenodo.17425251","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17425251","URL":"https://doi.org/10.5281/zenodo.17425251","source":"datacite"},{"id":"doi:10.5281/zenodo.17425207","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17425207","URL":"https://doi.org/10.5281/zenodo.17425207","source":"datacite"},{"id":"doi:10.5281/zenodo.17380264","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17380264","URL":"https://doi.org/10.5281/zenodo.17380264","source":"datacite"},{"id":"doi:10.5281/zenodo.17350374","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17350374","URL":"https://doi.org/10.5281/zenodo.17350374","source":"datacite"},{"id":"doi:10.4230/dagrep.15.3.77","type":"article-journal","title":"PETs and AI: Privacy Washing and the Need for a PETs Evaluation Framework (Dagstuhl Seminar 25112)","abstract":"As public awareness of data collection practices and regulatory frameworks grows, privacy-enhancing technologies (PETs) have emerged as a promising approach to reconciling data utility with individual privacy rights. PETs underpin privacy-preserving machine learning (PPML), integrating tools like differential privacy, homomorphic encryption, and secure multiparty computation to safeguard data throughout the AI lifecycle. However, despite significant technical progress, PETs face critical policy and governance challenges. Recent works have raised concerns about efficacy and deployment of PETs, observing that fundamental rights of people are continually being harmed, including, paradoxically, privacy. PETs have been used in surveillance applications and as a privacy washing tool. Current approaches often fail to address broader harms beyond data protection, highlighting the need for a more comprehensive privacy evaluation framework. This Dagstuhl Seminar brought together scholars in computer science and law, along with policymakers, regulators, and industry leaders, to discuss privacy washing and the challenges of detecting privacy washing through PETs and explored pathways toward a framework to address these challenges.","author":[{"family":"De Cristofaro","given":"Emiliano"},{"family":"Shrishak","given":"Kris"},{"family":"Strufe","given":"Thorsten"},{"family":"Troncoso","given":"Carmela"},{"family":"Morsbach","given":"Felix"}],"issued":{"date-parts":[[2025]]},"DOI":"10.4230/dagrep.15.3.77","URL":"https://doi.org/10.4230/dagrep.15.3.77","source":"datacite"},{"id":"doi:10.5281/zenodo.17289282","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17289282","URL":"https://doi.org/10.5281/zenodo.17289282","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.03218","type":"manuscript","title":"Cheat-Penalised Quantum Weak Coin-Flipping","abstract":"Coin-flipping is a fundamental task in two-party cryptography where two remote mistrustful parties wish to generate a shared uniformly random bit. While quantum protocols promising near-perfect security exist for weak coin-flipping -- when the parties want opposing outcomes -- it has been shown that they must be inefficient in terms of their round complexity, and it is an open question of how space efficient they can be. In this work, we consider a variant called cheat-penalised weak coin-flipping in which if a party gets caught cheating, they lose $Λ$ points (compared to $0$ in the standard definition). We find that already for a small cheating penalty, the landscape of coin-flipping changes dramatically. For example, with $Λ=0.01$, we exhibit a protocol where neither Alice nor Bob can bias the result in their favour beyond $1/2 + 10^{-8}$, which uses $24$ qubits and $10^{16}$ rounds of communication (provably $10^{7}$ times better than any weak coin-flipping protocol with matching security). For the same space requirements, we demonstrate how one can choose between lowering how much a malicious party can bias the result (down to $1/2 + 10^{-10}$) and reducing the rounds of communication (down to $25,180$), depending on what is preferred. To find these protocols, we make two technical contributions. First, we extend the point game-protocol correspondence introduced by Kitaev and Mochon, to incorporate: (i) approximate point games, (ii) the cheat-penalised setting, and (iii) round and space complexity. Second, we give the first (to the best of our knowledge) numerical algorithm for constructing (approximate) point games that correspond to high security and low complexity. Our results open up the possibility of having secure and practical quantum protocols for multiparty computation.","author":[{"family":"Arora","given":"Atul"},{"family":"Miller","given":"Carl"},{"family":"Morales","given":"Mauro"},{"family":"Sikora","given":"Jamie"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.03218","URL":"https://doi.org/10.48550/arxiv.2510.03218","source":"datacite"},{"id":"doi:10.5281/zenodo.17233165","type":"article-journal","title":"Implementing federated learning with privacy-preserving encryption to secure patient-derived imaging and sequencing data from cyber intrusions","abstract":"The growing adoption of artificial intelligence (AI) and data-driven analytics in healthcare has accelerated the integration of large-scale patient-derived imaging and genomic sequencing data into clinical workflows. However, this surge in biomedical data sharing has intensified cybersecurity challenges, particularly in protecting sensitive patient information from unauthorized access and cyber intrusions. Traditional centralized machine learning models, which aggregate data into a single repository, pose significant privacy risks and increase the attack surface for malicious actors. To address these challenges, federated learning (FL) has emerged as a transformative paradigm, enabling collaborative model training across decentralized nodes without transferring raw data. Yet, while FL mitigates some privacy concerns, it remains vulnerable to inference attacks, gradient leakage, and model inversion tactics. This paper explores the implementation of federated learning frameworks integrated with privacy-preserving encryption techniques, such as homomorphic encryption, differential privacy, and secure multiparty computation, specifically for safeguarding patient-derived medical imaging and sequencing datasets. These technologies ensure that sensitive genetic markers, radiographic scans, and multi-omic features remain encrypted throughout model training and aggregation processes. We examine recent advances in privacy-enhancing technologies, discuss system architectures suited for cross-institutional healthcare collaboration, and evaluate their performance trade-offs in terms of computational cost, model accuracy, and security guarantees. Furthermore, we propose a hybrid encryption-aware federated learning workflow tailored to radiogenomic applications, highlighting its resilience against adversarial threats while maintaining diagnostic precision. By narrowing focus to clinical implementations, this work provides a scalable and secure foundation for AI-driven biomedical research, enhancing trust and compliance in digital health ecosystems.","author":[{"family":"Kalejaiye","given":"Adebayo"},{"family":"Shallom","given":"Kigbu"},{"family":"Chukwuani","given":"Elvis"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17233165","URL":"https://doi.org/10.5281/zenodo.17233165","source":"datacite"},{"id":"doi:10.5281/zenodo.17233164","type":"article-journal","title":"Implementing federated learning with privacy-preserving encryption to secure patient-derived imaging and sequencing data from cyber intrusions","abstract":"The growing adoption of artificial intelligence (AI) and data-driven analytics in healthcare has accelerated the integration of large-scale patient-derived imaging and genomic sequencing data into clinical workflows. However, this surge in biomedical data sharing has intensified cybersecurity challenges, particularly in protecting sensitive patient information from unauthorized access and cyber intrusions. Traditional centralized machine learning models, which aggregate data into a single repository, pose significant privacy risks and increase the attack surface for malicious actors. To address these challenges, federated learning (FL) has emerged as a transformative paradigm, enabling collaborative model training across decentralized nodes without transferring raw data. Yet, while FL mitigates some privacy concerns, it remains vulnerable to inference attacks, gradient leakage, and model inversion tactics. This paper explores the implementation of federated learning frameworks integrated with privacy-preserving encryption techniques, such as homomorphic encryption, differential privacy, and secure multiparty computation, specifically for safeguarding patient-derived medical imaging and sequencing datasets. These technologies ensure that sensitive genetic markers, radiographic scans, and multi-omic features remain encrypted throughout model training and aggregation processes. We examine recent advances in privacy-enhancing technologies, discuss system architectures suited for cross-institutional healthcare collaboration, and evaluate their performance trade-offs in terms of computational cost, model accuracy, and security guarantees. Furthermore, we propose a hybrid encryption-aware federated learning workflow tailored to radiogenomic applications, highlighting its resilience against adversarial threats while maintaining diagnostic precision. By narrowing focus to clinical implementations, this work provides a scalable and secure foundation for AI-driven biomedical research, enhancing trust and compliance in digital health ecosystems.","author":[{"family":"Kalejaiye","given":"Adebayo"},{"family":"Shallom","given":"Kigbu"},{"family":"Chukwuani","given":"Elvis"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17233164","URL":"https://doi.org/10.5281/zenodo.17233164","source":"datacite"},{"id":"doi:10.5281/zenodo.17225939","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17225939","URL":"https://doi.org/10.5281/zenodo.17225939","source":"datacite"},{"id":"doi:10.5281/zenodo.17208446","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17208446","URL":"https://doi.org/10.5281/zenodo.17208446","source":"datacite"},{"id":"doi:10.5281/zenodo.17201334","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17201334","URL":"https://doi.org/10.5281/zenodo.17201334","source":"datacite"},{"id":"doi:10.5281/zenodo.17190502","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17190502","URL":"https://doi.org/10.5281/zenodo.17190502","source":"datacite"},{"id":"doi:10.5281/zenodo.17190331","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17190331","URL":"https://doi.org/10.5281/zenodo.17190331","source":"datacite"},{"id":"doi:10.5281/zenodo.21815685","type":"article-journal","title":"Supporting Data for State-Aware Upload Gating for Personalized Federated Learning in Mobile Edge Environments","abstract":"This dataset contains the client-partition files, per-seed raw experimental logs, processed summary data, and analysis scripts supporting the results reported in “State-Aware Upload Gating for Personalized Federated Learning in Mobile Edge Environments.” The package covers the CIFAR-10 and CIFAR-100 main experiments, Random-Gate comparisons, threshold experiments, and upload-budget-matched state-factor ablations. The corresponding source code, environment specifications, and verified execution commands are archived separately in SAUG-PFLlib, Version 1.0.0(https://doi.org/10.5281/zenodo.21807089). The files are restricted during peer review. Access is provided to editors and reviewers through a private link. The files will be made publicly available upon publication of the associated article.","author":[{"family":"Long","given":"Bangfa"},{"family":"Wang","given":"Xiaohan"},{"family":"Wu","given":"Nianhua"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21815685","URL":"https://doi.org/10.5281/zenodo.21815685","source":"datacite"},{"id":"doi:10.5281/zenodo.21815686","type":"article-journal","title":"Supporting Data for State-Aware Upload Gating for Personalized Federated Learning in Mobile Edge Environments","abstract":"This dataset contains the client-partition files, per-seed raw experimental logs, processed summary data, and analysis scripts supporting the results reported in “State-Aware Upload Gating for Personalized Federated Learning in Mobile Edge Environments.” The package covers the CIFAR-10 and CIFAR-100 main experiments, Random-Gate comparisons, threshold experiments, and upload-budget-matched state-factor ablations. The corresponding source code, environment specifications, and verified execution commands are archived separately in SAUG-PFLlib, Version 1.0.0(https://doi.org/10.5281/zenodo.21807089). The files are restricted during peer review. Access is provided to editors and reviewers through a private link. The files will be made publicly available upon publication of the associated article.","author":[{"family":"Long","given":"Bangfa"},{"family":"Wang","given":"Xiaohan"},{"family":"Wu","given":"Nianhua"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21815686","URL":"https://doi.org/10.5281/zenodo.21815686","source":"datacite"},{"id":"doi:10.5281/zenodo.20161043","type":"article-journal","title":"A Hybrid Machine Learning Framework for Anomaly Detection  and Electricity Theft Identification in Smart Distribution  Networks: A Comprehensive Review","abstract":"This paper presents a review of hybrid machine learning approaches for electricity theft detection in smart grids using Advanced Metering Infrastructure (AMI) data. The study analyzes supervised, unsupervised, and hybrid anomaly detection techniques including Isolation Forest and Histogram-based Gradient Boosting. Experimental evaluation on a synthetic smart meter dataset demonstrates improved detection accuracy and reduced false positives using hybrid models. The paper also discusses future directions such as Explainable AI (XAI), Federated Learning, and Graph Neural Networks (GNNs) for enhancing smart grid security.","author":[{"family":"Lakdeswar","given":"Anjali"},{"family":"Shamkule","given":"Devashree"},{"family":"Awachat","given":"Mansi"},{"family":"Urade","given":"Bobby"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20161043","URL":"https://doi.org/10.5281/zenodo.20161043","source":"datacite"},{"id":"doi:10.5281/zenodo.20161044","type":"article-journal","title":"A Hybrid Machine Learning Framework for Anomaly Detection  and Electricity Theft Identification in Smart Distribution  Networks: A Comprehensive Review","abstract":"This paper presents a review of hybrid machine learning approaches for electricity theft detection in smart grids using Advanced Metering Infrastructure (AMI) data. The study analyzes supervised, unsupervised, and hybrid anomaly detection techniques including Isolation Forest and Histogram-based Gradient Boosting. Experimental evaluation on a synthetic smart meter dataset demonstrates improved detection accuracy and reduced false positives using hybrid models. The paper also discusses future directions such as Explainable AI (XAI), Federated Learning, and Graph Neural Networks (GNNs) for enhancing smart grid security.","author":[{"family":"Lakdeswar","given":"Anjali"},{"family":"Shamkule","given":"Devashree"},{"family":"Awachat","given":"Mansi"},{"family":"Urade","given":"Bobby"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20161044","URL":"https://doi.org/10.5281/zenodo.20161044","source":"datacite"},{"id":"doi:10.5281/zenodo.17184062","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17184062","URL":"https://doi.org/10.5281/zenodo.17184062","source":"datacite"},{"id":"doi:10.5281/zenodo.17157969","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17157969","URL":"https://doi.org/10.5281/zenodo.17157969","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.08683","type":"manuscript","title":"Perfectly-Private Analog Secure Aggregation in Federated Learning","abstract":"In federated learning, multiple parties train models locally and share their parameters with a central server, which aggregates them to update a global model. To address the risk of exposing sensitive data through local models, secure aggregation via secure multiparty computation has been proposed to enhance privacy. At the same time, perfect privacy can only be achieved by a uniform distribution of the masked local models to be aggregated. This raises a problem when working with real valued data, as there is no measure on the reals that is invariant under the masking operation, and hence information leakage is bound to occur. Shifting the data to a finite field circumvents this problem, but as a downside runs into an inherent accuracy complexity tradeoff issue due to fixed point modular arithmetic as opposed to floating point numbers that can simultaneously handle numbers of varying magnitudes. In this paper, a novel secure parameter aggregation method is proposed that employs the torus rather than a finite field. This approach guarantees perfect privacy for each party's data by utilizing the uniform distribution on the torus, while avoiding accuracy losses. Experimental results show that the new protocol performs similarly to the model without secure aggregation while maintaining perfect privacy. Compared to the finite field secure aggregation, the torus-based protocol can in some cases significantly outperform it in terms of model accuracy and cosine similarity, hence making it a safer choice.","author":[{"family":"Jaramillo-Velez","given":"Delio"},{"family":"Rajput","given":"Charul"},{"family":"Freij-Hollanti","given":"Ragnar"},{"family":"Hollanti","given":"Camilla"},{"family":"Amat","given":"Alexandre"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.08683","URL":"https://doi.org/10.48550/arxiv.2509.08683","source":"datacite"},{"id":"doi:10.5281/zenodo.17141429","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17141429","URL":"https://doi.org/10.5281/zenodo.17141429","source":"datacite"},{"id":"doi:10.5281/zenodo.16984013","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16984013","URL":"https://doi.org/10.5281/zenodo.16984013","source":"datacite"},{"id":"doi:10.5281/zenodo.16902938","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16902938","URL":"https://doi.org/10.5281/zenodo.16902938","source":"datacite"},{"id":"doi:10.5281/zenodo.16902478","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16902478","URL":"https://doi.org/10.5281/zenodo.16902478","source":"datacite"},{"id":"doi:10.5281/zenodo.16778947","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16778947","URL":"https://doi.org/10.5281/zenodo.16778947","source":"datacite"},{"id":"doi:10.5281/zenodo.16778588","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16778588","URL":"https://doi.org/10.5281/zenodo.16778588","source":"datacite"},{"id":"doi:10.5281/zenodo.16778091","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16778091","URL":"https://doi.org/10.5281/zenodo.16778091","source":"datacite"},{"id":"doi:10.5281/zenodo.16777843","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16777843","URL":"https://doi.org/10.5281/zenodo.16777843","source":"datacite"},{"id":"doi:10.5281/zenodo.16637557","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16637557","URL":"https://doi.org/10.5281/zenodo.16637557","source":"datacite"},{"id":"doi:10.5281/zenodo.16634806","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16634806","URL":"https://doi.org/10.5281/zenodo.16634806","source":"datacite"},{"id":"doi:10.5281/zenodo.16568888","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16568888","URL":"https://doi.org/10.5281/zenodo.16568888","source":"datacite"},{"id":"doi:10.5281/zenodo.16367523","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16367523","URL":"https://doi.org/10.5281/zenodo.16367523","source":"datacite"},{"id":"doi:10.5281/zenodo.16364043","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16364043","URL":"https://doi.org/10.5281/zenodo.16364043","source":"datacite"},{"id":"doi:10.5281/zenodo.16363992","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16363992","URL":"https://doi.org/10.5281/zenodo.16363992","source":"datacite"},{"id":"doi:10.5281/zenodo.16363494","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16363494","URL":"https://doi.org/10.5281/zenodo.16363494","source":"datacite"},{"id":"doi:10.5281/zenodo.16312200","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16312200","URL":"https://doi.org/10.5281/zenodo.16312200","source":"datacite"},{"id":"doi:10.5281/zenodo.16307283","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16307283","URL":"https://doi.org/10.5281/zenodo.16307283","source":"datacite"},{"id":"doi:10.5281/zenodo.16304714","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16304714","URL":"https://doi.org/10.5281/zenodo.16304714","source":"datacite"},{"id":"doi:10.5281/zenodo.16304294","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16304294","URL":"https://doi.org/10.5281/zenodo.16304294","source":"datacite"},{"id":"doi:10.5281/zenodo.15834616","type":"article-journal","title":"Recent Advances in Privacy-Preserving Query Processing Techniques for Encrypted Relational Databases in Cloud Infrastructure","abstract":"Abstract: The growing reliance on cloud infrastructure for storing and managing relational databases has introduced critical challenges in preserving data confidentiality while enabling efficient query execution. Traditional encryption schemes offer strong data protection but often impede the ability to perform complex SQL operations without decryption, creating a trade-off between security and functionality. Recent advances in privacy-preserving query processing techniques—including homomorphic encryption, searchable encryption, oblivious RAM, and secure multiparty computation—have revolutionized the field by enabling secure computation over encrypted data with minimal performance penalties. This review paper systematically analyzes the state-of-the-art mechanisms that support encrypted SQL query processing in cloud-hosted relational databases. It evaluates the theoretical foundations, computational overheads, query expressiveness, and practical deployment scenarios of these privacy-preserving methods. Furthermore, it explores hybrid approaches that combine cryptographic and hardware-based techniques to balance performance with security guarantees. The paper also highlights emerging trends such as federated SQL processing, data provenance tracking, and secure hardware enclaves that are reshaping privacy-preserving architectures. The study aims to provide a comprehensive understanding of the current capabilities, limitations, and future research directions in secure query execution for encrypted relational databases deployed in cloud environments. Keywords: Encrypted SQL Query Processing, Homomorphic Encryption, Secure Cloud Databases, Privacy-Preserving Computation, Searchable Encryption, Cloud Data Confidentiality. Title: Recent Advances in Privacy-Preserving Query Processing Techniques for Encrypted Relational Databases in Cloud Infrastructure Author: Onuh Matthew Ijiga, Nonso Okika, Semirat Abidemi Balogun,. Ogboji James Agbo, Lawrence Anebi Enyejo International Journal of Computer Science and Information Technology Research ISSN 2348-1196 (print), ISSN 2348-120X (online) Vol. 13, Issue 3, July 2025 - September 2025 Page No: 62-80 Research Publish Journals Website: www.researchpublish.com Published Date: 08-July-2025 DOI: https://doi.org/10.5281/zenodo.15834617 Paper Download Link (Source) https://www.researchpublish.com/papers/recent-advances-in-privacy-preserving-query-processing-techniques-for-encrypted-relational-databases-in-cloud-infrastructure","author":[{"family":"Ijiga","given":"Onuh"},{"family":"Okika","given":"Nonso"},{"family":"Balogun","given":"Semirat"},{"family":"Agbo","given":"Ogboji"},{"family":"Enyejo","given":"Lawrence"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15834616","URL":"https://doi.org/10.5281/zenodo.15834616","source":"datacite"},{"id":"doi:10.5281/zenodo.15834617","type":"article-journal","title":"Recent Advances in Privacy-Preserving Query Processing Techniques for Encrypted Relational Databases in Cloud Infrastructure","abstract":"Abstract: The growing reliance on cloud infrastructure for storing and managing relational databases has introduced critical challenges in preserving data confidentiality while enabling efficient query execution. Traditional encryption schemes offer strong data protection but often impede the ability to perform complex SQL operations without decryption, creating a trade-off between security and functionality. Recent advances in privacy-preserving query processing techniques—including homomorphic encryption, searchable encryption, oblivious RAM, and secure multiparty computation—have revolutionized the field by enabling secure computation over encrypted data with minimal performance penalties. This review paper systematically analyzes the state-of-the-art mechanisms that support encrypted SQL query processing in cloud-hosted relational databases. It evaluates the theoretical foundations, computational overheads, query expressiveness, and practical deployment scenarios of these privacy-preserving methods. Furthermore, it explores hybrid approaches that combine cryptographic and hardware-based techniques to balance performance with security guarantees. The paper also highlights emerging trends such as federated SQL processing, data provenance tracking, and secure hardware enclaves that are reshaping privacy-preserving architectures. The study aims to provide a comprehensive understanding of the current capabilities, limitations, and future research directions in secure query execution for encrypted relational databases deployed in cloud environments. Keywords: Encrypted SQL Query Processing, Homomorphic Encryption, Secure Cloud Databases, Privacy-Preserving Computation, Searchable Encryption, Cloud Data Confidentiality. Title: Recent Advances in Privacy-Preserving Query Processing Techniques for Encrypted Relational Databases in Cloud Infrastructure Author: Onuh Matthew Ijiga, Nonso Okika, Semirat Abidemi Balogun,. Ogboji James Agbo, Lawrence Anebi Enyejo International Journal of Computer Science and Information Technology Research ISSN 2348-1196 (print), ISSN 2348-120X (online) Vol. 13, Issue 3, July 2025 - September 2025 Page No: 62-80 Research Publish Journals Website: www.researchpublish.com Published Date: 08-July-2025 DOI: https://doi.org/10.5281/zenodo.15834617 Paper Download Link (Source) https://www.researchpublish.com/papers/recent-advances-in-privacy-preserving-query-processing-techniques-for-encrypted-relational-databases-in-cloud-infrastructure","author":[{"family":"Ijiga","given":"Onuh"},{"family":"Okika","given":"Nonso"},{"family":"Balogun","given":"Semirat"},{"family":"Agbo","given":"Ogboji"},{"family":"Enyejo","given":"Lawrence"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15834617","URL":"https://doi.org/10.5281/zenodo.15834617","source":"datacite"},{"id":"doi:10.5281/zenodo.15775004","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15775004","URL":"https://doi.org/10.5281/zenodo.15775004","source":"datacite"},{"id":"doi:10.5281/zenodo.15755596","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15755596","URL":"https://doi.org/10.5281/zenodo.15755596","source":"datacite"},{"id":"doi:10.5281/zenodo.15735468","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15735468","URL":"https://doi.org/10.5281/zenodo.15735468","source":"datacite"},{"id":"doi:10.5281/zenodo.15705598","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15705598","URL":"https://doi.org/10.5281/zenodo.15705598","source":"datacite"},{"id":"doi:10.5281/zenodo.15680899","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15680899","URL":"https://doi.org/10.5281/zenodo.15680899","source":"datacite"},{"id":"doi:10.5281/zenodo.15645951","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15645951","URL":"https://doi.org/10.5281/zenodo.15645951","source":"datacite"},{"id":"doi:10.5281/zenodo.15641796","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15641796","URL":"https://doi.org/10.5281/zenodo.15641796","source":"datacite"},{"id":"doi:10.5281/zenodo.15640102","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15640102","URL":"https://doi.org/10.5281/zenodo.15640102","source":"datacite"},{"id":"doi:10.5281/zenodo.14930826","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.14930826","URL":"https://doi.org/10.5281/zenodo.14930826","source":"datacite"},{"id":"doi:10.5281/zenodo.14645969","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.14645969","URL":"https://doi.org/10.5281/zenodo.14645969","source":"datacite"},{"id":"doi:10.5281/zenodo.14830617","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.14830617","URL":"https://doi.org/10.5281/zenodo.14830617","source":"datacite"},{"id":"doi:10.5281/zenodo.15479320","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15479320","URL":"https://doi.org/10.5281/zenodo.15479320","source":"datacite"},{"id":"doi:10.5281/zenodo.15235159","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15235159","URL":"https://doi.org/10.5281/zenodo.15235159","source":"datacite"},{"id":"doi:10.5281/zenodo.15235328","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15235328","URL":"https://doi.org/10.5281/zenodo.15235328","source":"datacite"},{"id":"doi:10.5281/zenodo.14960634","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.14960634","URL":"https://doi.org/10.5281/zenodo.14960634","source":"datacite"},{"id":"doi:10.5281/zenodo.15056091","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15056091","URL":"https://doi.org/10.5281/zenodo.15056091","source":"datacite"},{"id":"doi:10.5281/zenodo.15235171","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15235171","URL":"https://doi.org/10.5281/zenodo.15235171","source":"datacite"},{"id":"doi:10.5281/zenodo.15524567","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15524567","URL":"https://doi.org/10.5281/zenodo.15524567","source":"datacite"},{"id":"doi:10.5281/zenodo.15637804","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15637804","URL":"https://doi.org/10.5281/zenodo.15637804","source":"datacite"},{"id":"doi:10.5281/zenodo.15043904","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15043904","URL":"https://doi.org/10.5281/zenodo.15043904","source":"datacite"},{"id":"doi:10.5281/zenodo.14724343","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.14724343","URL":"https://doi.org/10.5281/zenodo.14724343","source":"datacite"},{"id":"doi:10.5281/zenodo.15044005","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15044005","URL":"https://doi.org/10.5281/zenodo.15044005","source":"datacite"},{"id":"doi:10.5281/zenodo.14988748","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.14988748","URL":"https://doi.org/10.5281/zenodo.14988748","source":"datacite"},{"id":"doi:10.5281/zenodo.15236473","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15236473","URL":"https://doi.org/10.5281/zenodo.15236473","source":"datacite"},{"id":"doi:10.5281/zenodo.15043885","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15043885","URL":"https://doi.org/10.5281/zenodo.15043885","source":"datacite"},{"id":"doi:10.5281/zenodo.15016534","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15016534","URL":"https://doi.org/10.5281/zenodo.15016534","source":"datacite"},{"id":"doi:10.5281/zenodo.15043808","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15043808","URL":"https://doi.org/10.5281/zenodo.15043808","source":"datacite"},{"id":"doi:10.5281/zenodo.15638289","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15638289","URL":"https://doi.org/10.5281/zenodo.15638289","source":"datacite"},{"id":"doi:10.5281/zenodo.15013534","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15013534","URL":"https://doi.org/10.5281/zenodo.15013534","source":"datacite"},{"id":"doi:10.5281/zenodo.15082284","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15082284","URL":"https://doi.org/10.5281/zenodo.15082284","source":"datacite"},{"id":"doi:10.5281/zenodo.14652557","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.14652557","URL":"https://doi.org/10.5281/zenodo.14652557","source":"datacite"},{"id":"doi:10.5281/zenodo.19417528","type":"article-journal","title":"Edge AI: A Survey of Next-Generation Waste Classification and Routing Architectures","abstract":"The recent advances in edge computing and computer vision are completely transforming the way cities handle solid waste. In this review, close attention will be paid to the way smart city systems have changed to separate various types of waste and direct users to the appropriate bins in real-time. Our particular focus is the fact that the trend has shifted to less heavy and bulky smart bins and to the touchless mobile-edge systems. We go through the decisions of hardware, the move away towards large neural networks (as in VGG16) to thinner ones (such as MobileNetV2 and YOLOv8n) and the usage of location-finding algorithms such as Haversine formulas. The traditional smart-city systems are usually based on expensive embedded sensors, thus becoming difficult to implement and expand to larger cities. In the present-day, however, devices such as WebGL, Tensorflow.js, and an ordinary smartphone camera are utilized to perform the heavy lifting on the device. This decentralized design is faster, less expensive, more secret and more effortless to expand. Reviewing the studies of 2015 to 2026, we mention the constant problem of obtaining correct classification in poor lighting and solid communication systems are necessary. Finally, we outline the current state of edgedriven waste management, indicate the weaknesses of cloudheavy systems and consider the future of this technology, such as federated learning, blockchain rewards, and enhanced IoT synchronization.","author":[{"family":"Saini","given":"Yash"},{"family":"Tyagi","given":"Anmol"},{"family":"Kumar","given":"Priynshu"},{"family":"Sagar"},{"family":"Priyanka"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19417528","URL":"https://doi.org/10.5281/zenodo.19417528","source":"datacite"},{"id":"doi:10.5281/zenodo.19425963","type":"article-journal","title":"Edge AI: A Survey of Next-Generation Waste Classification and Routing Architectures","abstract":"The recent advances in edge computing and computer vision are completely transforming the way cities handle solid waste. In this review, close attention will be paid to the way smart city systems have changed to separate various types of waste and direct users to the appropriate bins in real-time. Our particular focus is the fact that the trend has shifted to less heavy and bulky smart bins and to the touchless mobile-edge systems. We go through the decisions of hardware, the move away towards large neural networks (as in VGG16) to thinner ones (such as MobileNetV2 and YOLOv8n) and the usage of location-finding algorithms such as Haversine formulas. The traditional smart-city systems are usually based on expensive embedded sensors, thus becoming difficult to implement and expand to larger cities. In the present-day, however, devices such as WebGL, Tensorflow.js, and an ordinary smartphone camera are utilized to perform the heavy lifting on the device. This decentralized design is faster, less expensive, more secret and more effortless to expand. Reviewing the studies of 2015 to 2026, we mention the constant problem of obtaining correct classification in poor lighting and solid communication systems are necessary. Finally, we outline the current state of edgedriven waste management, indicate the weaknesses of cloudheavy systems and consider the future of this technology, such as federated learning, blockchain rewards, and enhanced IoT synchronization.","author":[{"family":"Saini","given":"Yash"},{"family":"Tyagi","given":"Anmol"},{"family":"Kumar","given":"Priyanshu"},{"family":"Sagar"},{"family":"Priyanka"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19425963","URL":"https://doi.org/10.5281/zenodo.19425963","source":"datacite"},{"id":"doi:10.5281/zenodo.19416593","type":"article-journal","title":"Edge AI: A Survey of Next-Generation Waste Classification and Routing Architectures","abstract":"The recent advances in edge computing and computer vision are completely transforming the way cities handle solid waste. In this review, close attention will be paid to the way smart city systems have changed to separate various types of waste and direct users to the appropriate bins in real-time. Our particular focus is the fact that the trend has shifted to less heavy and bulky smart bins and to the touchless mobile-edge systems. We go through the decisions of hardware, the move away towards large neural networks (as in VGG16) to thinner ones (such as MobileNetV2 and YOLOv8n) and the usage of location-finding algorithms such as Haversine formulas. The traditional smart-city systems are usually based on expensive embedded sensors, thus becoming difficult to implement and expand to larger cities. In the present-day, however, devices such as WebGL, Tensorflow.js, and an ordinary smartphone camera are utilized to perform the heavy lifting on the device. This decentralized design is faster, less expensive, more secret and more effortless to expand. Reviewing the studies of 2015 to 2026, we mention the constant problem of obtaining correct classification in poor lighting and solid communication systems are necessary. Finally, we outline the current state of edgedriven waste management, indicate the weaknesses of cloudheavy systems and consider the future of this technology, such as federated learning, blockchain rewards, and enhanced IoT synchronization.","author":[{"family":"Saini","given":"Yash"},{"family":"Tyagi","given":"Anmol"},{"family":"Kumar","given":"Priyanshu"},{"family":"Sagar"},{"family":"Priyanka"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19416593","URL":"https://doi.org/10.5281/zenodo.19416593","source":"datacite"},{"id":"doi:10.5281/zenodo.19416594","type":"article-journal","title":"Edge AI: A Survey of Next-Generation Waste Classification a nd R outing Architectures","abstract":"The recent advances in edge computing and computer vision are completely transforming the way cities handle solid waste. In this review, close attention will be paid to the way smart city systems have changed to separate various types of waste and direct users to the appropriate bins in real-time. Our particular focus is the fact that the trend has shifted to less heavy and bulky smart bins and to the touchless mobile-edge systems. We go through the decisions of hardware, the move away towards large neural networks (as in VGG16) to thinner ones (such as MobileNetV2 and YOLOv8n) and the usage of location-finding algorithms such as Haversine formulas. The traditional smart-city systems are usually based on expensive embedded sensors, thus becoming difficult to implement and expand to larger cities. In the present-day, however, devices such as WebGL, Tensorflow.js, and an ordinary smartphone camera are utilized to perform the heavy lifting on the device. This decentralized design is faster, less expensive, more secret and more effortless to expand. Reviewing the studies of 2015 to 2026, we mention the constant problem of obtaining correct classification in poor lighting and solid communication systems are necessary. Finally, we outline the current state of edgedriven waste management, indicate the weaknesses of cloudheavy systems and consider the future of this technology, such as federated learning, blockchain rewards, and enhanced IoT synchronization.","author":[{"family":"Saini","given":"Yash"},{"family":"Tyagi","given":"Anmol"},{"family":"Kumar","given":"Priynshu"},{"family":"Sagar"},{"family":"Priyanka"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19416594","URL":"https://doi.org/10.5281/zenodo.19416594","source":"datacite"},{"id":"doi:10.5281/zenodo.21724220","type":"article-journal","title":"Proceedings of the 7th Summer School on Cyber-Physical Systems and Internet-of-Things Including AIoT Academy 2026","abstract":"Natalie Simson and Johannes EckerPart 1: A Systematic of Digital Design................................................. 1Part 2: Interface-Based Design Flow................................................. 41Part 3: Handshake-Based Design.................................................... 56Part 4: CPU Examples................................................................ 77 Muhammad Shafique and Alberto MarchisioFrom Energy-Efficient and Secure AI to Emerging Trends in Quantum Machine Learning: Challenges, Methods, Applications, and Systems.................................... 93 Abdelhakim Baouya, Brahim Hamid, Otmane Ait Mohamed and Saddek BensalemSecure and Resilient Engineering of Cyber-Physical Systems (CPS) via Stochastic Games................................................................................... 216 Nabil AbdennadherIntroduction to Quantum Computing from a Software Engineering Perspective .... 263 Dražen JurišićFractional-Order Systems and Applications......................................... 310 Andrej ŠkrabaESP32 Hands-on: From Edge Device Control to Cloud and AI in CPS/IoT Systems... 426 Matija StojanovićLaw and AI - An Overview of Some Dilemmas and Cases........................... 471 Vladimir MladenovićCyber Security and Federated Learning for Autonomous Vehicle Networks......... 490 Radovan Stojanović and Jovan ĐurkovićAIoT in Medical Wearables......................................................... 533 Vesna Maraš, Mitar Otašević, Dejan Zejak and Jovan ĐurkovićAIoT in Precise Agriculture......................................................... 567 Summer School on CPS&IoT’2026 Schedule............................................... 583 Summer School on CPS&IoT’2026 7th Generation (Students and Teachers)................ 587 Certificate of Attendance.................................................................. 588 Author Index.............................................................................. 589 Photo Gallery 590","author":[{"family":"Stojanovic","given":"Radovan"},{"family":"Škraba","given":"Andrej"},{"family":"Đurković","given":"Jovan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21724220","URL":"https://doi.org/10.5281/zenodo.21724220","source":"datacite"},{"id":"doi:10.5281/zenodo.21724219","type":"article-journal","title":"Proceedings of the 7th Summer School on Cyber-Physical Systems and Internet-of-Things Including AIoT Academy 2026","abstract":"Natalie Simson and Johannes EckerPart 1: A Systematic of Digital Design................................................. 1Part 2: Interface-Based Design Flow................................................. 41Part 3: Handshake-Based Design.................................................... 56Part 4: CPU Examples................................................................ 77 Muhammad Shafique and Alberto MarchisioFrom Energy-Efficient and Secure AI to Emerging Trends in Quantum Machine Learning: Challenges, Methods, Applications, and Systems.................................... 93 Abdelhakim Baouya, Brahim Hamid, Otmane Ait Mohamed and Saddek BensalemSecure and Resilient Engineering of Cyber-Physical Systems (CPS) via Stochastic Games................................................................................... 216 Nabil AbdennadherIntroduction to Quantum Computing from a Software Engineering Perspective .... 263 Dražen JurišićFractional-Order Systems and Applications......................................... 310 Andrej ŠkrabaESP32 Hands-on: From Edge Device Control to Cloud and AI in CPS/IoT Systems... 426 Matija StojanovićLaw and AI - An Overview of Some Dilemmas and Cases........................... 471 Vladimir MladenovićCyber Security and Federated Learning for Autonomous Vehicle Networks......... 490 Radovan Stojanović and Jovan ĐurkovićAIoT in Medical Wearables......................................................... 533 Vesna Maraš, Mitar Otašević, Dejan Zejak and Jovan ĐurkovićAIoT in Precise Agriculture......................................................... 567 Summer School on CPS&IoT’2026 Schedule............................................... 583 Summer School on CPS&IoT’2026 7th Generation (Students and Teachers)................ 587 Certificate of Attendance.................................................................. 588 Author Index.............................................................................. 589 Photo Gallery 590","author":[{"family":"Stojanovic","given":"Radovan"},{"family":"Škraba","given":"Andrej"},{"family":"Đurković","given":"Jovan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21724219","URL":"https://doi.org/10.5281/zenodo.21724219","source":"datacite"},{"id":"doi:10.5281/zenodo.21671708","type":"article-journal","title":"Accurately Inferring Missing Factor VIII Concentrate Dosing Information from Health-Record Data in Patients with Haemophilia A","abstract":"This poster, presented at the ISTH 2026 Congress (International Society on Thrombosis and Haemostasis), showcases research conducted within the PHEMS project to address missing treatment information in electronic health records (EHRs) for patients with haemophilia A. The study addresses the challenge of missing treatment information by developing a methodology to infer factor VIII dosing from routinely collected clinical data. Using data from children with haemophilia A treated at Erasmus MC Sophia Children's Hospital, the approach was evaluated for its ability to reconstruct treatment histories and generate more complete datasets for research. By improving the quality and usability of EHR data, this work supports the development of privacy-preserving machine learning models for rare diseases and contributes to the broader PHEMS mission of advancing federated paediatric health data research across Europe.","author":[{"family":"Janssen","given":"Alexander"},{"family":"Mathot","given":"Ron"},{"family":"Cnossen","given":"Marjon"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21671708","URL":"https://doi.org/10.5281/zenodo.21671708","source":"datacite"},{"id":"doi:10.5281/zenodo.21671709","type":"article-journal","title":"Accurately Inferring Missing Factor VIII Concentrate Dosing Information from Health-Record Data in Patients with Haemophilia A","abstract":"This poster, presented at the ISTH 2026 Congress (International Society on Thrombosis and Haemostasis), showcases research conducted within the PHEMS project to address missing treatment information in electronic health records (EHRs) for patients with haemophilia A. The study addresses the challenge of missing treatment information by developing a methodology to infer factor VIII dosing from routinely collected clinical data. Using data from children with haemophilia A treated at Erasmus MC Sophia Children's Hospital, the approach was evaluated for its ability to reconstruct treatment histories and generate more complete datasets for research. By improving the quality and usability of EHR data, this work supports the development of privacy-preserving machine learning models for rare diseases and contributes to the broader PHEMS mission of advancing federated paediatric health data research across Europe.","author":[{"family":"Janssen","given":"Alexander"},{"family":"Mathot","given":"Ron"},{"family":"Cnossen","given":"Marjon"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21671709","URL":"https://doi.org/10.5281/zenodo.21671709","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.25107","type":"manuscript","title":"MOSAIC-FL, a micro-service based privacy-preserving framework with application to genomics","abstract":"Security and privacy are primordial requirements for Federated Learning (FL), especially in fields such as healthcare and genomics where sensitive information has to be analyzed. Our FL framework is designed to address these challenges while proposing a modular, flexible and micro-service architecture. More precisely, it integrates an efficient gRPC communication layer and a Finite State Machine to ensure robust component synchronization and threat detection, while relying on a fault-tolerant secure aggregation protocol using a Threshold variant of the CKKS homomorphic cryptosystem. This allows blind model aggregation by an orchestration server, requiring a minimum of $t$-out-of-$N$ active clients for decryption while minimizing communication overhead thanks to both cryptographic and network protocols. We ensure IND-CPA-D security through noise flooding and mitigate the recent key-recovery attack on synchronized decryptors by renewing the collective key material at every round. We demonstrate the framework's effectiveness through diverse use cases, ranging from standard image recognition (EMNIST) to complex genomic classification including breast cancer subtyping on TCGA, evaluating system performance across different threshold values and model scales.","author":[{"family":"Largillier","given":"Paul"},{"family":"Paygambar","given":"Karl"},{"family":"Gouy-Pailler","given":"Cédric"},{"family":"Meyer","given":"Vincent"},{"family":"Mziou","given":"Mallek"},{"family":"Stan","given":"Oana"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.25107","URL":"https://doi.org/10.48550/arxiv.2607.25107","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.11974","type":"manuscript","title":"Poison to Detect: Detection of Targeted Overfitting in Federated Learning","abstract":"Federated Learning (FL) enables collaborative model training among clients without centralising data, making it a widely adopted privacy-enhancing technology (PET). Despite its privacy benefits, FL remains vulnerable to orchestrator-driven privacy attacks. In this paper, we study an underexplored threat in which a dishonest orchestrator intentionally manipulates the aggregation process to induce targeted overfitting in local models of specific clients. Although prior work focuses on reducing information leakage during training, we emphasise early client-side detection of targeted overfitting, allowing clients to disengage before significant harm occurs. To this end, we propose three detection techniques -- label flipping, backdoor trigger injection, and model fingerprinting -- which enable clients to verify the integrity of the global aggregation. We evaluated our methods across multiple datasets and attack scenarios. In single-client attacks, all three methods detect orchestrator-induced overfitting within 1-2 training rounds with F1 scores up to 0.7. Scalability experiments further show that detection effectiveness is influenced by cohort composition and method parameters. These results demonstrate that client-side integrity testing can provide early, effective, and scalable detection, supporting safer deployment of FL systems.","author":[{"family":"Mestari","given":"Soumia"},{"family":"Zuziak","given":"Maciej"},{"family":"Lenzini","given":"Gabriele"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.11974","URL":"https://doi.org/10.48550/arxiv.2509.11974","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.28282","type":"manuscript","title":"Pre-Deployment Complexity Estimation for Federated Perception Systems","abstract":"Edge AI systems increasingly rely on federated learning to train perception models in distributed, privacy-preserving, and resource-constrained environments. Before training, however, practitioners often lack practical tools for estimating task difficulty in terms of expected accuracy and communication effort. We present a classifier-agnostic, pre-deployment framework that combines intrinsic data properties such as dimensionality, sparsity, and heterogeneity, with client-distribution composition to estimate learning complexity in federated perception systems. Using federated learning as a representative distributed training setting, we examine how learning difficulty varies across different federated configurations. Experiments on three MNIST variants show strong negative correlations between the combined complexity metric and maximum and average federated accuracy, while the intrinsic and distributed components exhibit consistent relationships with communication effort. These findings suggest that complexity estimation can serve as a practical diagnostic tool for resource planning, dataset assessment, and feasibility evaluation in edge-deployed perception systems.","author":[{"family":"Solaiman","given":"Kma"},{"family":"Islam","given":"Shafkat"},{"family":"De Oliveira","given":"Ruy"},{"family":"Bhargava","given":"Bharat"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.28282","URL":"https://doi.org/10.48550/arxiv.2603.28282","source":"datacite"},{"id":"doi:10.48448/wkv8-xd76","type":"article-journal","title":"CO-EVO: Co-evolving Semantic Anchoring and Style Diversification for Federated DG-ReID","abstract":"Federated domain generalization for person re-identification (FedDG-ReID) aims to collaboratively train a pedestrian retrieval model across multiple decentralized source domains such that it can generalize to unseen target environments without compromising raw data privacy. However, this task is significantly challenged by the inherent stylistic gaps across decentralized clients. Without global supervision, models easily succumb to shortcut learning where representations overfit to domain specific camera biases rather than universal identity features. We propose CO-EVO, a novel federated framework that resolves this semantic-style conflict through a co-evolutionary mechanism. On the semantic side, Camera-Invariant Semantic Anchoring (CSA) learns identity prompts with cross-camera consistency to establish purified and domain-agnostic anchors that filter out local imaging noise. On the visual side, Global Style Diversification (GSD), powered by a Global Camera-Style Bank (GCSB), synthesizes realistic perturbations to expand the visual boundaries of training data. The core of CO-EVO is its co-evolutionary loop where purified anchors act as gravitational centers to guide the image encoder toward robust anatomical attributes amidst diverse style variations. Extensive experiments demonstrate that CO-EVO achieves state-of-the-art (SOTA) performance, proving that the synergy between semantic purification and style expansion is essential for robust cross-domain generalization. Our code is available at: \\url{https://github.com/NanYiyuzurn/ACL-LGPS-2026}.","author":[{"family":"Hu","given":"Jianwei"},{"family":"Huang","given":"Tingxuan"},{"family":"Lai","given":"Jinshan"},{"family":"Ma","given":"Qiang"},{"family":"Xiang","given":"Liuyu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48448/wkv8-xd76","URL":"https://doi.org/10.48448/wkv8-xd76","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.14877","type":"manuscript","title":"Deep Reinforcement Learning for 6G AI-RAN: A Comprehensive Survey","abstract":"The evolution toward sixth-generation (6G) networks is transforming the radio access network (RAN) into a programmable and intelligent control platform that must continuously adapt to heterogeneous services, dynamic environments, and competing performance objectives. Open Radio Access Network (O-RAN) provides the open interfaces, disaggregated architecture, and multi-timescale control loops needed to support this transformation, while deep reinforcement learning (DRL) offers a natural framework for optimizing sequential decisions under uncertainty. However, existing surveys either address artificial intelligence (AI) and machine learning (ML) in O-RAN broadly or focus on isolated DRL use cases, leaving a gap in the systematic connection between DRL methodology, O-RAN architecture, and operational deployment. To the best of our knowledge, this article presents the first dedicated and comprehensive survey of DRL for Open AI-RAN. We review the foundations of model-free, model-based, offline, safe, multi-agent, federated, and transfer learning, and provide an O-RAN-aware framework for formulating RAN control problems through states, observations, actions, rewards, constraints, and temporal structure. We classify DRL applications across radio resource management, mobility management, interference control, traffic steering, energy efficiency, network slicing, integrated sensing and communication, security, and massive MIMO. We further examine multi-agent and federated coordination, foundation models and agentic AI, trustworthy DRL, sim-to-real transfer, continual adaptation, resource-efficient inference, and reinforcement learning operations. Finally, we review experimental platforms, benchmarks, standards, and industry activities, and identify research directions toward sample-efficient, safe, scalable, interoperable, and deployable DRL control for 6G Open AI-RAN.","author":[{"family":"Lu","given":"Jie"},{"family":"Yan","given":"Peihao"},{"family":"Wang","given":"Qijun"},{"family":"Lin","given":"Ruxin"},{"family":"Zeng","given":"Huacheng"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.14877","URL":"https://doi.org/10.48550/arxiv.2608.14877","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.13844","type":"manuscript","title":"Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future Directions","abstract":"Large language models (LLMs) have become core components of cloud-based intelligent services in academia and industry, yet their training and deployment are hindered by high computational costs, data centralization, and privacy concerns. Federated learning (FL) offers a decentralized training paradigm that enables clients to collaboratively train a learning model without sharing raw data, making it a promising solution for privacy-preserving LLM training and reasoning. This paper presents a comprehensive survey of federated prompt learning (FPL) to review recent advances in integrating the federated learning paradigm and large language models, answering the following research questions: RQ1: The fundamental motivations, characteristics, and enabling technologies of FPL, and how it differs from conventional FL and full-model federated fine-tuning; RQ2: The trade-offs FPL approaches exhibit in performance, communication efficiency, computational overhead, scalability, personalization, and heterogeneity handling; RQ3: The remaining security, privacy, robustness, and system challenges, along with key future research directions. To this end, we systematically examine existing FPL methods across the full model lifecycle: pre-training, fine-tuning, and practical applications, while discussing security, privacy, and robustness issues and summarizing existing defense mechanisms. Finally, we highlight open challenges and future directions, aiming to help readers understand how the insights drive research in FPL.","author":[{"family":"Yang","given":"Qinglin"},{"family":"Qiu","given":"Chen"},{"family":"Zhang","given":"Hongyuan"},{"family":"Li","given":"Pengdeng"},{"family":"Liu","given":"Yuan"},{"family":"Tian","given":"Zhihong"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.13844","URL":"https://doi.org/10.48550/arxiv.2608.13844","source":"datacite"},{"id":"doi:10.5281/zenodo.21801833","type":"article-journal","title":"Landslide Early Warning System using SAR Data and Machine Learning","abstract":"Abstract - Landslides continue to pose a significant threat to human lives, infrastructure, transportation networks, and ecological systems, particularly in mountainous and high-rainfall regions. Recent advances in artificial intelligence, remote sensing, and geospatial analytics have transformed conventional landslide monitoring into intelligent early warning systems capable of continuous environmental assessment and rapid decision-making. This survey presents a comprehensive review of machine learning-based landslide early warning systems with particular emphasis on the integration of Sentinel-1 Interferometric Synthetic Aperture Radar (InSAR), geo-fencing, meteorological information, ensemble learning, and real-time alert dissemination. The survey consolidates the methodologies, datasets, feature engineering strategies, prediction algorithms, deployment architectures, and evaluation metrics reported in recent literature while analyzing their strengths and limitations. Furthermore, the survey presents the LandSense framework as an integrated case study demonstrating how machine learning, satellite-derived deformation monitoring, rainfall analysis, secure backend services, geospatial risk mapping, evacuation route planning, and multi-channel alert mechanisms can be combined into a unified disaster management platform. Comparative analysis of Random Forest, XGBoost, ensemble learning, deep learning, and time-series forecasting techniques highlights current research trends and identifies remaining challenges related to data availability, computational complexity, model generalization, explainability, and operational deployment. The survey concludes by outlining future research directions involving explainable artificial intelligence, transformer architectures, graph neural networks, digital twins, federated learning, and edge-based disaster intelligence for next-generation landslide early warning systems. This survey reviews recent developments in AI-driven landslide prediction techniques and analyzes the integration of machine learning algorithms, Sentinel-1 InSAR, geo-fencing, rainfall monitoring, and emergency alert systems for disaster management. It also presents the LandSense framework as a comprehensive case study that combines ensemble learning, geospatial analysis, real-time monitoring, safe route planning, and multi-channel alert dissemination into a unified early warning platform. By comparing existing approaches and identifying current research gaps, this survey highlights future directions for developing scalable, explainable, and intelligent landslide early warning systems.","author":[{"family":"Fernandes","given":"Rohan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21801833","URL":"https://doi.org/10.5281/zenodo.21801833","source":"datacite"},{"id":"doi:10.5281/zenodo.21801834","type":"article-journal","title":"Landslide Early Warning System using SAR Data and Machine Learning","abstract":"Abstract - Landslides continue to pose a significant threat to human lives, infrastructure, transportation networks, and ecological systems, particularly in mountainous and high-rainfall regions. Recent advances in artificial intelligence, remote sensing, and geospatial analytics have transformed conventional landslide monitoring into intelligent early warning systems capable of continuous environmental assessment and rapid decision-making. This survey presents a comprehensive review of machine learning-based landslide early warning systems with particular emphasis on the integration of Sentinel-1 Interferometric Synthetic Aperture Radar (InSAR), geo-fencing, meteorological information, ensemble learning, and real-time alert dissemination. The survey consolidates the methodologies, datasets, feature engineering strategies, prediction algorithms, deployment architectures, and evaluation metrics reported in recent literature while analyzing their strengths and limitations. Furthermore, the survey presents the LandSense framework as an integrated case study demonstrating how machine learning, satellite-derived deformation monitoring, rainfall analysis, secure backend services, geospatial risk mapping, evacuation route planning, and multi-channel alert mechanisms can be combined into a unified disaster management platform. Comparative analysis of Random Forest, XGBoost, ensemble learning, deep learning, and time-series forecasting techniques highlights current research trends and identifies remaining challenges related to data availability, computational complexity, model generalization, explainability, and operational deployment. The survey concludes by outlining future research directions involving explainable artificial intelligence, transformer architectures, graph neural networks, digital twins, federated learning, and edge-based disaster intelligence for next-generation landslide early warning systems. This survey reviews recent developments in AI-driven landslide prediction techniques and analyzes the integration of machine learning algorithms, Sentinel-1 InSAR, geo-fencing, rainfall monitoring, and emergency alert systems for disaster management. It also presents the LandSense framework as a comprehensive case study that combines ensemble learning, geospatial analysis, real-time monitoring, safe route planning, and multi-channel alert dissemination into a unified early warning platform. By comparing existing approaches and identifying current research gaps, this survey highlights future directions for developing scalable, explainable, and intelligent landslide early warning systems.","author":[{"family":"Fernandes","given":"Rohan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21801834","URL":"https://doi.org/10.5281/zenodo.21801834","source":"datacite"},{"id":"doi:10.5281/zenodo.21409456","type":"article-journal","title":"Designing the Future of Secure Intelligence: A Vision for AI-Enabled Cyber-Resilient Ecosystems","abstract":"This vision paper presents a consolidated perspective developed through six European initiatives—VIGILANCE, CyberAId, ENFORCE, ACT4FOOD, CTIS4NIS, and 5G-TACTIC—working together through the Cyber-Resilience Synergy Network (CRSN). It examines how Artificial Intelligence, cybersecurity-by-design, autonomous decision-making, proactive threat detection, Zero Trust approaches, secure data sharing, and collaborative intelligence can support the development of trustworthy and cyber-resilient next-generation digital ecosystems. Drawing on shared methodologies, technological prototypes, pilot deployments, and validation activities, the paper proposes a common roadmap for integrating AI-driven automation, proactive cyber defence, federated and collaborative AI, secure and interoperable data spaces, and application-driven intelligence into future large-scale digital infrastructures. The discussion spans critical infrastructures, telecommunications, financial services, industrial environments, and food supply chains. The paper also considers trustworthy AI, human oversight, regulatory alignment, scalability, distributed learning, digital sovereignty, capacity building, and emerging research challenges. It concludes by presenting a common architectural vision and identifying research and policy directions for strengthening European digital autonomy and societal resilience. Version note: This Zenodo record contains the preprint version of the publication. Published version: The final publication appears in Artificial Intelligence Applications and Innovations: AIAI 2026 International Workshops, IFIP Advances in Information and Communication Technology, Volume 796, pp. 158–178, Springer Nature Switzerland AG, 2027.DOI: 10.1007/978-3-032-30507-7_10","author":[{"family":"Gizelis","given":"Christos"},{"family":"Marinakis","given":"Achilleas"},{"family":"Papamokos","given":"Stavros"},{"family":"Kefalogiannis","given":"Michalis"},{"family":"Mesogiti","given":"Ioanna"},{"family":"Liberopoulos","given":"Giorgos"},{"family":"Theodoropoulou","given":"Elina"},{"family":"Poulimenou","given":"Maria"},{"family":"Kachrimani","given":"Louiza"},{"family":"Tzanakaki","given":"Anna"},{"family":"Anastassopoulos","given":"Markos"},{"family":"Azrak","given":"Thomas"},{"family":"Dimitrakopoulou","given":"Maria"},{"family":"Kokkinos","given":"Odysseas"},{"family":"Bianconi","given":"Luca"},{"family":"Clerici","given":"Giulia"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21409456","URL":"https://doi.org/10.5281/zenodo.21409456","source":"datacite"},{"id":"doi:10.5281/zenodo.21409455","type":"article-journal","title":"Designing the Future of Secure Intelligence: A Vision for AI-Enabled Cyber-Resilient Ecosystems","abstract":"This vision paper presents a consolidated perspective developed through six European initiatives—VIGILANCE, CyberAId, ENFORCE, ACT4FOOD, CTIS4NIS, and 5G-TACTIC—working together through the Cyber-Resilience Synergy Network (CRSN). It examines how Artificial Intelligence, cybersecurity-by-design, autonomous decision-making, proactive threat detection, Zero Trust approaches, secure data sharing, and collaborative intelligence can support the development of trustworthy and cyber-resilient next-generation digital ecosystems. Drawing on shared methodologies, technological prototypes, pilot deployments, and validation activities, the paper proposes a common roadmap for integrating AI-driven automation, proactive cyber defence, federated and collaborative AI, secure and interoperable data spaces, and application-driven intelligence into future large-scale digital infrastructures. The discussion spans critical infrastructures, telecommunications, financial services, industrial environments, and food supply chains. The paper also considers trustworthy AI, human oversight, regulatory alignment, scalability, distributed learning, digital sovereignty, capacity building, and emerging research challenges. It concludes by presenting a common architectural vision and identifying research and policy directions for strengthening European digital autonomy and societal resilience. Version note: This Zenodo record contains the preprint version of the publication. Published version: The final publication appears in Artificial Intelligence Applications and Innovations: AIAI 2026 International Workshops, IFIP Advances in Information and Communication Technology, Volume 796, pp. 158–178, Springer Nature Switzerland AG, 2027.DOI: 10.1007/978-3-032-30507-7_10","author":[{"family":"Gizelis","given":"Christos"},{"family":"Marinakis","given":"Achilleas"},{"family":"Papamokos","given":"Stavros"},{"family":"Kefalogiannis","given":"Michalis"},{"family":"Mesogiti","given":"Ioanna"},{"family":"Liberopoulos","given":"Giorgos"},{"family":"Theodoropoulou","given":"Elina"},{"family":"Poulimenou","given":"Maria"},{"family":"Kachrimani","given":"Louiza"},{"family":"Tzanakaki","given":"Anna"},{"family":"Anastassopoulos","given":"Markos"},{"family":"Azrak","given":"Thomas"},{"family":"Dimitrakopoulou","given":"Maria"},{"family":"Kokkinos","given":"Odysseas"},{"family":"Bianconi","given":"Luca"},{"family":"Clerici","given":"Giulia"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21409455","URL":"https://doi.org/10.5281/zenodo.21409455","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.05553","type":"manuscript","title":"Federated Physics-Grounded Reinforcement Learning for Distributed Stability Control in Smart Grids","abstract":"Transient stability control in smart grids requires rapid post-fault damping of generator frequency and rotor angle deviations to prevent cascading failures. This paper proposes FedPPO-PG, a Federated Multi-Agent Proximal Policy Optimization framework with Physics-Grounded neighborhoods, which reformulates transient stability control as a cooperative multi-agent reinforcement learning problem optimized directly against closed-loop stability objectives. Each generator hosts an independent local actor augmented with the frequency deviations of its two most strongly coupled electrical neighbors, identified from the post-fault Kron-reduced susceptance matrix. A guided policy initialization phase warm-starts all actors from the classical decentralized controller, while a centralized critic guides advantage estimation under the centralized training--decentralized execution (CTDE) paradigm. Evaluated on a simulation of the IEEE 39-bus benchmark system across five training and three unseen fault contingencies, FedPPO-PG achieves 100% stabilization in all 24 trials, reduces mean stability time by 72.4%, and cuts the control power by 7-14 times compared to the centralized baseline. Each actor executes independently with no central coordinator at deployment, and the per-actor inference latency satisfies the IEEE/IEC 60255-118-1-2018 real-time reporting requirements.","author":[{"family":"Al-Refai","given":"Omar"},{"family":"Shahbaz","given":"Ibrahim"},{"family":"Husseinat","given":"Adam"},{"family":"Hammad","given":"Eman"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.05553","URL":"https://doi.org/10.48550/arxiv.2607.05553","source":"datacite"},{"id":"doi:10.5281/zenodo.19357843","type":"article-journal","title":"Ep. 128: AI's Dial-Up Era: Looking Back from 2036","abstract":"Episode summary: In this forward-thinking episode of My Weird Prompts, hosts Herman Poppleberry and Corn kick off the year 2026 by traveling a decade into the future. They imagine a world in 2036 where the \"cutting-edge\" AI of today is viewed as an adorable, clunky relic of the past—much like we view the screeching sounds of dial-up internet today. From the death of prompt engineering to the rise of zero-latency, embodied intelligence, the duo breaks down why our current obsession with context windows and text boxes is just a passing phase. They dive deep into the transition from \"command-based\" to \"intent-based\" computing, where AI understands your needs without the need for complex instructions. Herman explains the shift from monolithic models to federated swarms of specialized agents, and how the \"hallucination\" bug of the 2020s will eventually be seen as a primitive technical limitation. Whether you're curious about the future of robotics or the evolution of persistent holographic memory, this episode provides a fascinating roadmap for the next decade of innovation. Tune in to find out why your current smartphone might soon feel like a rotary phone. Show Notes As the calendar turned to January 1, 2026, *My Weird Prompts* hosts Herman Poppleberry and Corn took a moment to look not just at the year ahead, but a full decade into the future. Prompted by a thought experiment from their housemate Daniel, the duo spent the episode \"time traveling\" to 2036 to look back at the current state of artificial intelligence. Their conclusion? The sophisticated tools we use today—the LLMs, the image generators, and the coding assistants—are destined to become the \"dial-up modems\" of the future. ### The Death of the Prompt One of the most striking insights from the discussion was the predicted obsolescence of \"prompt engineering.\" In 2026, users pride themselves on their ability to craft complex instructions, using delimiters and \"chain-of-thought\" techniques to coax the best results out of a model. Herman argues that by 2036, this will seem as primitive as using a rotary phone. We are currently in a \"lossy\" phase of technology, where we must translate human intent into rigid strings of text. Herman suggests that the future lies in \"intent-based computing.\" In this future, AI will possess such deep context regarding a user's life, professional history, and personal preferences that it will no longer require a three-paragraph explanation. A simple glance or a vague suggestion will suffice, as the machine will already understand the nuances of what \"professional\" or \"creative\" means to that specific individual. ### From Context Windows to Holographic Memory The hosts also tackled the technical limitations of modern AI memory. Today, developers and users celebrate when a model's \"context window\" expands to a million tokens. However, Herman describes the current state of AI as a \"brilliant assistant who gets hit with an amnesia ray every time you walk out of the room.\" By 2036, the concept of a \"window\" will likely be replaced by what Herman calls \"persistent, holographic memory.\" Instead of a blank slate at the start of every chat, a personal AI will have a continuous, decade-long relationship with its user. It will remember a casual comment about architectural styles from years prior and seamlessly apply that knowledge to a current project. The manual management of AI memory will become a relic of a more cumbersome era. ### Zero Latency and the End of the \"Thinking\" Pause One of the most relatable points of the episode was the \"dial-up screech\" of 2026: latency. Even the fastest models today have a slight delay as they generate tokens. Herman predicts that 2036 will be the era of \"zero-latency intelligence.\" Powered by specialized hardware—potentially optical or neuromorphic chips—AI responses will be instantaneous or even predictive. The duo joked about how future generations will find it hilarious that we used to sit and watch text scroll a","author":[{"family":"Rosehill","given":"Daniel"},{"family":"Tts","given":"Chatterbox"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19357843","URL":"https://doi.org/10.5281/zenodo.19357843","source":"datacite"},{"id":"doi:10.48550/arxiv.2605.08121","type":"manuscript","title":"Performance and Energy Trade-Off Analysis of Hierarchical Federated Learning for Plant Disease Classification","abstract":"Early detection of plant diseases is critical for improving crop productivity, while it also facilitates the foundations of precision agriculture. Recent advances in distributed deep learning have enabled plant disease classification models to be trained across geographically distributed agricultural sensing infrastructures. However, deploying such systems in large-scale Internet of Things (IoT) environments, introduces significant challenges related to computational cost, energy consumption, and system efficiency. In this paper, we present a design-space exploration of hierarchical federated learning architectures for plant disease classification, with a particular focus on the trade-offs between predictive performance and energy efficiency. We further introduce a power- and energy-aware optimization framework that enables the systematic evaluation and selection of model-aggregator configurations under varying deployment constraints. The hierarchical federated architecture organizes distributed clients through intermediate aggregation layers, reducing communication and computational overhead. We evaluate multiple convolutional neural network architectures, including EfficientNet-B0, ResNet-50, and MobileNetV3-Large, in combination with different federated aggregation strategies such as FedAvg, FedProx, and FedAvgM. Experimental results demonstrate that different model-aggregator combinations exhibit distinct performance-energy trade-offs. Consequently, we highlight configurations that achieve competitive diagnostic accuracy and significantly reduce system resource requirements.","author":[{"family":"Papanikolaou","given":"Athanasios"},{"family":"Tziouvaras","given":"Athanasios"},{"family":"Stoikos","given":"Pavlos"},{"family":"Xenakis","given":"Apostolos"},{"family":"Parambath","given":"Shameem"},{"family":"Floros","given":"George"},{"family":"Zereik","given":"Enrica"},{"family":"Petrovic","given":"Ivan"},{"family":"Bonsignorio","given":"Fabio"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.08121","URL":"https://doi.org/10.48550/arxiv.2605.08121","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.25858","type":"manuscript","title":"Color Matters: Trigger Color Affects Success in Federated Backdoor Attacks","abstract":"Federated learning is vulnerable to backdoor attacks in which malicious clients inject poisoned updates while preserving benign-task performance. In this paper, we study a semantics-driven backdoor mechanism in which attackers use natural visual accessories as triggers and manipulate only the trigger color while keeping the attack pipeline fixed. Our framework considers semantic trigger objects such as masks and sunglasses, instantiated in black and white variants, and evaluates their effect in a controlled federated learning setting. Malicious clients construct poisoned samples by applying a trigger to source-class images and relabeling them to an attacker-chosen target class, while benign clients train only on clean data. We analyze this mechanism under both a standard poisoning objective and a stronger SABLE-based objective that combines clean classification loss, triggered target loss, feature-separation loss in the penultimate representation space, and regularization to keep malicious updates close to the global model. This design enables the attack to remain effective while reducing excessive update drift. Experiments on a four-class CelebA hair-color task show that trigger color significantly changes attack success rate even when trigger semantics, placement, and poisoning budget are unchanged. White triggers are more effective for attacks targeting the blond class, whereas black triggers perform better for attacks targeting the black class. The same trend persists under robust aggregation, showing that trigger color is a meaningful factor in the operation, persistence, and evaluation of semantic backdoor mechanisms in federated learning.","author":[{"family":"Herath","given":"Kavindu"},{"family":"Zhao","given":"Joshua"},{"family":"Bagchi","given":"Saurabh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.25858","URL":"https://doi.org/10.48550/arxiv.2606.25858","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.10124","type":"manuscript","title":"FedSteer: Taming Extreme Gradient Staleness in Federated Learning with Corrective Projections and Caching","abstract":"Federated learning (FL) is often subject to aggregation variance if clients do not consistently participate in training rounds. While reusing stale model updates from inactive clients is a common technique to reduce this variance, we find that with skewed client participation, the resulting update staleness can become severe enough to destabilize training. To remedy this, we propose FedSteer, a novel method that constructs a gradient subspace from a cache of recent client gradients to serve as a low-dimensional representation of the current optimization landscape. FedSteer projects an active client's true gradient onto this subspace to find a set of optimal coordinates. For an inactive client, FedSteer reuses these coordinates with the now-evolved subspace drifted by other active clients. This process effectively \"steers\" outdated gradients toward the current global objective. This is complemented by a selective caching strategy that identifies a representative client subset to form the subspace, reducing server memory. Experiments demonstrate that FedSteer significantly outperforms baselines, preventing performance collapse in challenging scenarios while delivering accuracy gains of over 7% in others.","author":[{"family":"Zhang","given":"Haoran"},{"family":"Pereira","given":"Cainã"},{"family":"Siew","given":"Marie"},{"family":"Liu","given":"Xutong"},{"family":"Joe-Wong","given":"Carlee"},{"family":"El-Azouzi","given":"Rachid"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.10124","URL":"https://doi.org/10.48550/arxiv.2606.10124","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.12970","type":"manuscript","title":"Probabilistic Feature Imputation and Uncertainty-Aware Multimodal Federated Aggregation","abstract":"Multimodal federated learning enables privacy-preserving collaborative model training across healthcare institutions. However, a fundamental challenge arises from modality heterogeneity: many clinical sites possess only a subset of modalities due to resource constraints or workflow variations. Existing approaches address this through feature imputation networks that synthesize missing modality representations, yet these methods produce point estimates without reliability measures, forcing downstream classifiers to treat all imputed features as equally trustworthy. In safety-critical medical applications, this limitation poses significant risks. We propose the Probabilistic Feature Imputation Network (P-FIN), which outputs calibrated uncertainty estimates alongside imputed features. This uncertainty is leveraged at two levels: (1) locally, through sigmoid gating that attenuates unreliable feature dimensions before classification, and (2) globally, through Fed-UQ-Avg, an aggregation strategy that prioritizes updates from clients with reliable imputation. Experiments on federated chest X-ray classification using CheXpert, NIH Open-I, and PadChest demonstrate consistent improvements over deterministic baselines, with +5.36% AUC gain in the most challenging configuration.","author":[{"family":"Shahid","given":"Nafis"},{"family":"Ahmed","given":"Maroof"},{"family":"Haider","given":"Md"},{"family":"Sagor","given":"Saidur"},{"family":"Rahman","given":"Aashnan"},{"family":"Hossain","given":"Md"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.12970","URL":"https://doi.org/10.48550/arxiv.2604.12970","source":"datacite"},{"id":"doi:10.48550/arxiv.2502.17748","type":"manuscript","title":"FinP: Fairness-in-Privacy in Federated Learning by Addressing Disparities in Privacy Risk","abstract":"Federated Learning (FL) inherently mitigates mass data centralization risks; however, its privacy protections are not equally distributed - leaving vulnerable individuals disproportionately exposed to sophisticated privacy attacks. Crucially, statistical heterogeneity in human-centric FL environments often results in an inequitable distribution of privacy risks, particularly affecting those whose sensitive attributes or behaviors make them outliers. To address this critical gap, we introduce FinP, a novel framework designed to formalize and enforce fairness-in-privacy by mitigating disproportionate client vulnerability to Source Inference Attacks (SIA). FinP operationalizes a two-pronged defense strategy that tackles both the symptoms and root causes of privacy disparity, ensuring that no group of clients bears an excessive privacy burden. It combines a server-side adaptive aggregation mechanism, which dynamically weights client contributions based on their estimated privacy risk, with a client-side regularization technique to curb localized overfitting that drives unique data memorization. Extensive empirical evaluations on FEMNIST, Human Activity Recognition (HAR), and CIFAR-10 datasets demonstrate that FinP effectively aligns privacy fairness with primary task utility. Notably, FinP successfully mitigates SIA risks and reduces disparities in privacy exposure, establishing that strong fairness-in-privacy guarantees need not compromise model utility. Ultimately, FinP establishes equitable privacy protections by reducing vulnerability disparities by up to 57.14%, while preserving global model utility within a marginal +/- 1.75% of standard federated baselines.","author":[{"family":"Zhao","given":"Tianyu"},{"family":"Srewa","given":"Mahmoud"},{"family":"Elmalaki","given":"Salma"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2502.17748","URL":"https://doi.org/10.48550/arxiv.2502.17748","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.22427","type":"manuscript","title":"CLoVE: Personalized Federated Learning through Clustering of Loss Vector Embeddings","abstract":"We propose CLoVE (Clustering of Loss Vector Embeddings), a novel algorithm for Clustered Federated Learning (CFL). In CFL, clients are naturally grouped into clusters based on their data distribution. However, identifying these clusters is challenging, as client assignments are unknown. CLoVE utilizes client embeddings derived from model losses on client data, and leverages the insight that clients in the same cluster share similar loss values, while those in different clusters exhibit distinct loss patterns. Based on these embeddings, CLoVE is able to iteratively identify and separate clients from different clusters and optimize cluster-specific models through federated aggregation. Key advantages of CLoVE over existing CFL algorithms are (1) its simplicity, (2) its applicability to both supervised and unsupervised settings, and (3) the fact that it eliminates the need for near-optimal model initialization, which makes it more robust and better suited for real-world applications. We establish theoretical convergence bounds, showing that CLoVE can recover clusters accurately with high probability in a single round and converges exponentially fast to optimal models in a linear setting. Our comprehensive experiments comparing with a variety of both CFL and generic Personalized Federated Learning (PFL) algorithms on different types of datasets and an extensive array of non-IID settings demonstrate that CLoVE achieves highly accurate cluster recovery in just a few rounds of training, along with state-of-the-art model accuracy, across a variety of both supervised and unsupervised PFL tasks.","author":[{"family":"Bhatia","given":"Randeep"},{"family":"Papadis","given":"Nikos"},{"family":"Kodialam","given":"Murali"},{"family":"Lakshman","given":"Tv"},{"family":"Chakrabarty","given":"Sayak"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.22427","URL":"https://doi.org/10.48550/arxiv.2506.22427","source":"datacite"},{"id":"doi:10.48550/arxiv.2505.19699","type":"manuscript","title":"Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments","abstract":"Federated Learning (FL) is a decentralized machine learning paradigm that enables clients to collaboratively train models while preserving data privacy. However, the coexistence of model and data heterogeneity gives rise to inconsistent representations and divergent optimization dynamics across clients, ultimately hindering robust global performance. To transcend these challenges, we propose Mosaic, a novel data-free knowledge distillation framework tailored for heterogeneous distributed environments. Mosaic first trains local generative models to approximate each client's personalized distribution, enabling synthetic data generation that safeguards privacy through strict separation from real data. Subsequently, Mosaic forms a Mixture-of-Experts (MoE) from client models based on their specialized knowledge, and distills it into a global model using the generated data. To further enhance the MoE architecture, Mosaic integrates expert predictions via a lightweight meta model trained on a few representative prototypes. Extensive experiments on standard image and multimodal benchmarks demonstrate that Mosaic consistently outperforms state-of-the-art approaches under both model and data heterogeneity. The source code has been published at https://github.com/Wings-Of-Disaster/Mosaic.","author":[{"family":"Liu","given":"Junming"},{"family":"Gao","given":"Yanting"},{"family":"Li","given":"Yuqi"},{"family":"Meng","given":"Siyuan"},{"family":"Sun","given":"Yifei"},{"family":"Wu","given":"Aoqi"},{"family":"Chen","given":"Yirong"},{"family":"Wang","given":"Ding"},{"family":"Wen","given":"Shiping"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2505.19699","URL":"https://doi.org/10.48550/arxiv.2505.19699","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.12845","type":"manuscript","title":"A Privacy-Preserving Framework Using Remote Data Science for Inter-Institutional Student Retention Prediction","abstract":"This study explores privacy-preserving machine learning (PPML) techniques using the PySyft platform to enable collaborative prediction of student retention between institutions. We developed a remote data science (RDS) framework with a semi-air-gapped architecture consisting of high-side and low-side servers, allowing researchers from three universities to build predictive models on sensitive student data without direct data access. Using historical data from a small private university (N=720), we evaluated three synthetic data generation approaches and validated the framework through inter-institutional collaboration. The results demonstrate consistent classification performance across institutions (Macro F1: 0.690--0.695) while maintaining strict Family Educational Rights and Privacy Act (FERPA) compliance. We also propose Data-Type-Aware Templates, a novel synthetic data method that prioritizes privacy over distributional fidelity. Our findings confirm that RDS-based PPML is technically feasible for educational settings and offers a practical alternative to federated learning for small-scale inter-institutional collaborations. The code is available at https://github.com/jtfields/NAIRR240195-Privacy-Preserving-Machine-Learning.","author":[{"family":"Fields","given":"John"},{"family":"Islam","given":"KMS"},{"family":"Thota","given":"Ruchitha"},{"family":"Chen","given":"Victor"},{"family":"Madiraju","given":"Praveen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.12845","URL":"https://doi.org/10.48550/arxiv.2606.12845","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.11556","type":"manuscript","title":"Privacy-Preserving Federated Autoencoder for ECG Anomaly Detection on Edge Devices","abstract":"Continuous electrocardiography (ECG) monitoring could surface rhythm abnormalities before they escalate into cardiovascular events. However, a deployable system must satisfy three requirements simultaneously: legal-grade privacy (GDPR, HIPAA), real-time inference on constrained edge hardware, and detection quality under non-IID cross-hospital data. We design and evaluate an end-to-end federated system addressing all three for unsupervised 12-lead ECG anomaly detection on PTB-XL dataset, combining three autoencoder families (VanillaAE, ConvAE, VAE), Flower-based federated averaging (FedAvg) across ten simulated hospitals, client-side differentially private SGD (DP-SGD) with a Rényi-DP accountant, and 8-bit integer (INT8) post-training quantization with Raspberry Pi 4 benchmarking. Our main contributions are: an empirical characterization of how these mechanisms compose, practical DP-specific recommendations, and technical and security insights for a clinically sensitive setting. Federated learning matches or exceeds the centralized baseline across all architectures (ConvAE federated area under the ROC curve, AUROC, $0.782$), and an $\\varepsilon$ sweep identifies $\\varepsilon=4$ as the recommended clinical operating point. INT8 quantization roughly halves model size and cuts Pi 4 latency by up to $44%$ with $&lt;0.12%$ AUROC loss. Crucially, DP and quantization penalties are empirically independent, so practitioners need not trade a strong privacy guarantee for a compact edge footprint. To our knowledge, this is the first system combining federated learning, formal $(\\varepsilon,δ)$-DP, unsupervised reconstruction-based detection, and quantized AArch64 deployment.","author":[{"family":"Akyol","given":"Kaan"},{"family":"Szeląg","given":"Jakub"},{"family":"Abadi","given":"Aydin"},{"family":"Alghamdi","given":"Maha"},{"family":"Albalawi","given":"Ghadah"},{"family":"Kaleelullah","given":"Ghouse"},{"family":"Tutus","given":"Hilal"},{"family":"Subaiei","given":"Sarah"},{"family":"Kapse","given":"Shardul"},{"family":"Raheeb","given":"Syed"},{"family":"Ahmed","given":"Mujeeb"},{"family":"Ullah","given":"Rehmat"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.11556","URL":"https://doi.org/10.48550/arxiv.2606.11556","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.10780","type":"manuscript","title":"Secure Aggregation with Top-K Sparsification in Decentralized Federated Learning","abstract":"Secure aggregation is a vital component for mitigating gradient leakage in federated learning, but its communication cost conventionally scales with the gradient dimension. This becomes prohibitive for large models and even more pronounced in decentralized federated learning with limited bandwidth and unreliable nodes. Top-K gradient sparsification is an effective approach to reduce communication by transmitting only a few entries of the full gradient, while maintaining competitive model accuracy. Nevertheless, the top-K entries selected by each user are unpredictable and vary across users, which poses a challenge for efficient sparse secure aggregation. This paper studies information-theoretic secure aggregation with top-K sparsification in decentralized federated learning under user dropouts and user collusion. We propose a communication-efficient sparse secure aggregation scheme that offloads dimension-dependent overhead to an offline phase and protects private gradients using random masks and permutations. Experimental results demonstrate that our scheme preserves accuracy comparable to full-gradient aggregation even with only 1% gradient sparsification, while substantially reducing the communication cost.","author":[{"family":"Tang","given":"Hengxuan"},{"family":"Zhu","given":"Jinbao"},{"family":"Tang","given":"Xiaohu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.10780","URL":"https://doi.org/10.48550/arxiv.2606.10780","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.09869","type":"manuscript","title":"QSplitFL: Capability Aware Deep Q-Learning for Optimal Split Point Selection in Split Federated Learning","abstract":"Federated Learning (FL) combined with Split Learning (SL) is a privacy preserving paradigm that enables training deep neural networks (DNNs) on resource constrained devices while reducing overall training cost. However, determining the optimal split point, meaning the layer where the model is divided still remains a critical challenge, especially when clients have heterogeneous hardware capabilities. Fixed split points can overload weak devices and increase the communication and server load, which slows convergence and reduces stability. This paper introduces QSplitFL, a novel capability-aware Deep Q-Network (DQN) framework for optimal split point selection in Split learning based Federated Learning (SFL) environments. Unlike existing approaches that rely on high-dimensional model weight representations, QSplitFL employs a lightweight state representation derived directly from client hardware metrics, including CPU utilization, memory, battery level, and network latency. The proposed framework incorporates a decayed loss-drop reward function that prioritizes early convergence, and a committee-based DQN architecture with majority voting to mitigate reward hacking. Extensive experiments on MNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100 datasets using CNN, ResNet50, MobileNetV4, and ConvNeXt architectures demonstrate that our approach achieves better convergence and higher accuracy compared to existing methods, while effectively adapting to heterogeneous device resources. The source code is publicly available at https://github.com/AIPO-Lab/QSplitFL.","author":[{"family":"Shadin","given":"Nazmus"},{"family":"Zhang","given":"Xinyue"},{"family":"Wang","given":"Jingyi"},{"family":"Pan","given":"Miao"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.09869","URL":"https://doi.org/10.48550/arxiv.2606.09869","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.09548","type":"manuscript","title":"Model Poisoning Against Federated Model Adaptation with Chain of Bit-Flips","abstract":"Federated Learning (FL) allows a set of clients to collectively train a global model without sharing local training data. Giving the responsibility of the training to decentralized actors may lead to poisoning attacks: clients controlled by malicious third party potentially poison the training dataset to install a backdoor in neural networks. In FL, these backdoor attacks rely solely on algorithmic approach, however, recent advances in hardware faults threats (e.g, Rowhammer) have widen the overall attack surface. In the context of federated model adaptation, we introduce a novel category of backdoor attack against FL systems that relies on model poisoning based on hardware-fault attacks. More precisely, we propose a task-agnostic backdoor attack that is implanted during the FL training time by inducing hardware faults (bit-flips) in parameters of a single local model. The backdoor is crafted during a previous offline phase from the pretrained model initially used by the FL system. Our results show that a backdoor can be successfully applied on different type of models and datasets. Typically, with up to 10 faults per malicious client occurrence and 19 total occurrences on a ResNet-18 are enough to reach 94% of attack success rate. Finally, we discuss the practicality and the robustness of the attack potential defenses, while putting into perspective the practical constraints of Rowhammer, which is the preferred attack vector for this type of threats.","author":[{"family":"Vuillod","given":"Bastien"},{"family":"Hector","given":"Kevin"},{"family":"Moellic","given":"Pierre"},{"family":"Dutertre","given":"Jean"},{"family":"Potin","given":"Olivier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.09548","URL":"https://doi.org/10.48550/arxiv.2606.09548","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.09227","type":"manuscript","title":"Trustworthy Smart Fabs via Professional Proxies: Scaling Safe and Sustainable by Design (SSbD) through Industrial Data Spaces","abstract":"The convergence of the 2026 European Union Safe and Sustainable by Design (SSbD) framework, Corporate Sustainability Due Diligence Directive (CSDDD), and Carbon Border Adjustment Mechanism (CBAM) introduce a severe governance bottleneck for advanced semiconductor manufacturing facilities (\"Smart Fabs\"). Regulatory compliance demands have surpassed the capacity of manual corporate reporting, creating a direct conflict between multi-stakeholder transparency and corporate data privacy. This paper addresses this challenge by introducing a zero-trust socio-technical orchestration framework that operationalizes a six-layer SSbD reference architecture within trustworthy industrial data spaces. We propose a shift from reactive automation to autonomous governance through \"Professional Proxies\"-role-based agentic workflows executing within hardware-isolated trust zones. Structured as an interoperable network protocol stack, the framework coordinates an automated, five-step \"relay race\" between Facility, Process Engineering, and Finance proxy teams to align factory-floor yield models with macro-level sustainability mandates. By executing Virtual Metrology (VM) predictions and Federated Machine Learning (FML) inside hardware-rooted Trusted Execution Environments (TEEs), this architecture resolves the Data Sovereignty Paradox, demonstrating how fabs can export cryptographically signed compliance tokens via International Data Spaces (IDS) connectors without exposing proprietary process recipes. Ultimately, this framework provides technology managers with a verifiable, evidence-based pathway toward resilient, net-zero Industry 5.0 ecosystems.","author":[{"family":"Liao","given":"Han"},{"family":"Kao","given":"Chang"},{"family":"Ang","given":"Karen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.09227","URL":"https://doi.org/10.48550/arxiv.2606.09227","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.08173","type":"manuscript","title":"AI-Native Closed-Loop Security for 6G-Enabled Cyber-Physical Systems: From Edge Detection to Network-Wide Mitigation","abstract":"In sixth-generation (6G) networks, billions of cyber-physical systems (CPSs) - autonomous vehicles, smart grids, industrial robots, and remote-surgical equipment - will run over ultra-reliable low-latency slices, collapsing the gap between a remote breach and physical harm to milliseconds, a budget perimeter firewalls and centralised security operations centres cannot meet. This survey reframes 6G CPS security as a closed-loop, AI-native pipeline that senses at the multi-access edge computing (MEC) tier, using minute-scale call-detail records (CDRs) for baseline learning and sub-millisecond RAN/Open-RAN (O-RAN) telemetry for the latency-critical path. It decides locally with compressed deep models, mitigates network-wide via SDN, NFV, and O-RAN controllers, and retrains through federated learning (FL) and digital-twin (DT) replay. We formalise a per-slice, tail-bounded latency contract on the sense, detect, and mitigate stages, enforced at a slice-dependent tail percentile (p99 for safety-critical URLLC slices). Organising 128 peer-reviewed studies (2017-2026) under a PRISMA 2020 protocol, we (i) map the 6G/CPS threat surface to MITRE ATT&amp;CK and a CDR-observable feature space; (ii) unify edge anomaly detection and DDoS classification across twelve datasets and statistical, graph, and transformer models; (iii) synthesise SDN/NFV/O-RAN primitives into one closed-loop reference architecture; (iv) treat FL, large language models (LLMs), DT, post-quantum cryptography (PQC), zero-trust architecture (ZTA), and explainable AI as cross-cutting enablers, not parallel pillars; and (v) consolidate open problems into five directions spanning data, latency, trust, standardisation, and evaluation.","author":[{"family":"Hussain","given":"Bilal"},{"family":"Bilal","given":"Muhammad"},{"family":"Li","given":"Tan"},{"family":"Pervaiz","given":"Haris"},{"family":"Tang","given":"Xiao"},{"family":"Du","given":"Qinghe"},{"family":"Ahmad","given":"Fawad"},{"family":"Azhar","given":"Muhammad"},{"family":"Zhang","given":"Jun"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.08173","URL":"https://doi.org/10.48550/arxiv.2606.08173","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.06154","type":"manuscript","title":"Amortizing Federated Adaptation: Hypernetwork Driven LoRA for Personalized Foundation Models","abstract":"Federated fine-tuning of foundation models using Low-Rank Adaptation (LoRA) offers a communication efficient solution for distributed learning. However, existing federated LoRA methods suffer from two fundamental limitations: (1) structural aggregation bias, where independently averaging low rank factors fails to approximate the true combined update, and (2) client side initialization lag, as clients repeatedly reinitialize LoRA parameters across communication rounds, slowing convergence. We propose HyperLoRA, a unified framework that addresses both issues through amortized federated adaptation through hypernetwork-driven LoRA generation and product space aggregation. Instead of iterative per-client optimization, HyperLoRA employs a learned generator that maps client distribution signatures to LoRA initializations, effectively amortizing per client adaptation. On the server side, we introduce a learned aggregation module that directly synthesizes updates in the low-rank product space, eliminating the inconsistencies of factor-wise averaging. A lightweight residual correction module further improves stability under heterogenous (non-IID) client distributions.By replacing iterative optimization and heuristic averaging with learned operators, HyperLoRA jointly enables efficient personalization, unbiased aggregation, and faster convergence. Experiments on federated vision and vision-language benchmarks show that HyperLoRA achieves improved convergence speed, greater robustness to distribution shift, and stronger personalization performance compared to prior federated LoRA methods.","author":[{"family":"Gupta","given":"Sunny"},{"family":"Shanker","given":"Shambhavi"},{"family":"Sethi","given":"Amit"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.06154","URL":"https://doi.org/10.48550/arxiv.2606.06154","source":"datacite"},{"id":"doi:10.48550/arxiv.2605.18936","type":"manuscript","title":"FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data","abstract":"Social media text data are often used to train Machine Learning (ML) models to identify users exhibiting high-risk mental health behaviors. However, sharing this sensitive data poses privacy risks and limits the growth of benchmark datasets. We comprehensively evaluate whether privacy-preserving ML techniques can enable safer data sharing while preserving performance. Specifically, we apply federated learning (FL) and Differentially Private FL for two widely-studied mental health prediction tasks: depression detection on X (Twitter) and suicide crisis detection on Reddit. We simulate realistic data-sharing scenarios by treating each user as a client in a non-IID setting, evaluating across different client fractions, aggregation strategies, and privacy budgets. While FL achieves comparable performance to centralized training (centralized F1 = 85.63; best FL model F1 = 83.16) on depression identification, we find that Differentially Private FL has a large performance-privacy trade-off (up to F1 = 27.01 drop) even with low levels of noise (epsilon = 50). This is due to the distortion of highly informative yet sparse mental health linguistic markers related to mental health, like health topics and emotion words. This research empirically demonstrates the potential and limitations of current privacy preservation techniques for mental health inference tasks.","author":[{"family":"Abdelkadir","given":"Nuredin"},{"family":"Ratnam","given":"Anjali"},{"family":"Talat","given":"Zeerak"},{"family":"Chancellor","given":"Stevie"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.18936","URL":"https://doi.org/10.48550/arxiv.2605.18936","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.01607","type":"manuscript","title":"FedMTFI: Feature Importance Based Optimized Multi Teacher Knowledge Distillation in Heterogeneous Federated Learning Environment","abstract":"Federated learning (FL) is a decentralized approach that enables collaborative model training without exposing raw data. Instead of transferring sensitive data, it allows devices to share only model weights, keeping personal data locally and secure. However, in real world settings, the data held by devices is often not evenly distributed and devices mostly differ in computing power and memory capacity. These differences make FL harder to maintain consistent performance across the system. To address these issues, we propose FedMTFI, a novel architecture that combines multi-teacher knowledge distillation (MTKD) with feature importance to improve the FL process in heterogeneous environments. In FedMTFI, clients are clustered based on similar hardware and model types. Each cluster trains a specific model on not independently and identically distributed (non-IID) data. Within a cluster, every client updates that model using only its own local private data. The server then aggregates the locally trained models in each cluster using FedAvg to form multiple prototype models. Then these prototypes serve as teacher models to train a global generalized student model using MTKD. What makes FedMTFI more unique is the integration of Shapley values (SHAP) to emphasize important features during distillation, which enhances both accuracy and interpretability. Experimental results show that FedMTFI achieves higher accuracy than traditional FL algorithms and performs more effectively under non-IID data conditions.","author":[{"family":"Shadin","given":"Nazmus"},{"family":"Cummings","given":"Aaron"},{"family":"Zhang","given":"Xinyue"},{"family":"Deng","given":"Bobin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.01607","URL":"https://doi.org/10.48550/arxiv.2606.01607","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.01186","type":"manuscript","title":"The Gaussian-Head OFL Family: One-Shot Federated Learning from Client Global Statistics","abstract":"Classical Federated Learning relies on a multi-round iterative process of model exchange and aggregation between server and clients, with high communication costs and privacy risks from repeated model transmissions. In contrast, one-shot federated learning (OFL) alleviates these limitations by reducing communication to a single round, thereby lowering overhead and enhancing practical deployability. Nevertheless, most existing one-shot approaches remain either impractical or constrained, for example, they often depend on the availability of a public dataset, assume homogeneous client models, or require uploading additional data or model information. To overcome these issues, we introduce the Gaussian-Head OFL (GH-OFL) family, a suite of one-shot federated methods that assume class-conditional Gaussianity of pretrained embeddings. Clients transmit only sufficient statistics (per-class counts and first/second-order moments) and the server builds heads via three components: (i) Closed-form Gaussian heads (NB/LDA/QDA) computed directly from the received statistics; (ii) FisherMix, a linear head with cosine margin trained on synthetic samples drawn in an estimated Fisher subspace; and (iii) Proto-Hyper, a lightweight low-rank residual head that refines Gaussian logits via knowledge distillation on those synthetic samples. In our experiments, GH-OFL methods deliver state-of-the-art robustness and accuracy under strong non-IID skew while remaining strictly data-free.","author":[{"family":"Turazza","given":"Fabio"},{"family":"Picone","given":"Marco"},{"family":"Mamei","given":"Marco"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.01186","URL":"https://doi.org/10.48550/arxiv.2602.01186","source":"datacite"},{"id":"doi:10.48550/arxiv.2505.18332","type":"manuscript","title":"An Attack to Break Permutation-Based Private Third-Party Inference Schemes for LLMs","abstract":"Recent advances in Large Language Models (LLMs) have led to the widespread adoption of third-party inference services, raising critical privacy concerns. Existing methods of performing private third-party inference, such as Secure Multiparty Computation (SMPC), often rely on cryptographic methods. However, these methods are thousands of times slower than standard unencrypted inference, and fail to scale to large modern LLMs. Therefore, recent lines of work have explored the replacement of expensive encrypted nonlinear computations in SMPC with statistical obfuscation methods - in particular, revealing permuted hidden states to the third parties, with accompanying strong claims of the difficulty of reversal into the unpermuted states. In this work, we begin by introducing a novel reconstruction technique that can recover original prompts from hidden states with nearly perfect accuracy across multiple state-of-the-art LLMs. We then show that extensions of our attack are nearly perfectly effective in reversing permuted hidden states of LLMs, demonstrating the insecurity of three recently proposed privacy schemes. We further dissect the shortcomings of prior theoretical `proofs' of permuation security which allow our attack to succeed. Our findings highlight the importance of rigorous security analysis in privacy-preserving LLM inference.","author":[{"family":"Thomas","given":"Rahul"},{"family":"Zahran","given":"Louai"},{"family":"Choi","given":"Erica"},{"family":"Potti","given":"Akilesh"},{"family":"Goldblum","given":"Micah"},{"family":"Pal","given":"Arka"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2505.18332","URL":"https://doi.org/10.48550/arxiv.2505.18332","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.23478","type":"manuscript","title":"ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour","abstract":"Fully homomorphic encryption (FHE) lets a server run inference on encrypted data with strong privacy guarantees, but running a Transformer under FHE is expensive. Its non-linear operations, such as softmax, normalization, and activation, must be replaced with polynomial approximations that the CKKS scheme supports, and the depth of these approximations dominates inference cost. Existing FHE Transformers use hand-tuned approximation settings, such as iteration count and polynomial degree, applied uniformly across layers, models, and tasks. Hand-tuning is slow and error-prone. Even a single uniform setting has about $10^7$ choices, and manual search cannot exploit layer-wise variation. AutoFHE, the only automated method with multi-objective search, targets ReLU-only CNNs and needs full fine-tuning per candidate, which is too costly for Transformers. Per-layer settings also push the search space to about $10^{85}$ for BERT and ViT and $10^{228}$ for LLaMA3, beyond both manual and fine-tuning-based search. We present ATLAS, a training-free framework that automates this search by treating each layer's approximation setting as a multi-objective optimization over latency and accuracy. The problem is hard: the decision space is large (96 or 256 variables), each configuration takes 70 to 1,000 seconds to evaluate even in cleartext, and 85 to 90 percent of configurations are invalid. ATLAS handles this with a two-stage optimization strategy and a surrogate model, completing the search in about one hour. Compared to an iterative softmax baseline, ATLAS cuts multiplicative depth and end-to-end latency by about 35 percent with little accuracy loss, and works across encoder-only, decoder-only, and vision Transformers, complementing parallel work on packing and matrix multiplication.","author":[{"family":"Xie","given":"Jianhang"},{"family":"Tan","given":"Sicheng"},{"family":"Boddeti","given":"Vishnu"},{"family":"Lu","given":"Zhichao"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.23478","URL":"https://doi.org/10.48550/arxiv.2607.23478","source":"datacite"},{"id":"doi:10.48441/4427.3459","type":"article-journal","title":"Deterministic homomorphic identity tokens for privacy-preserving attribute-based authentication","abstract":"Traditional multi-factor authentication (MFA) schemes, despite their layered defences, have been compromised in several common vulnerabilities and exposures (CVEs) due to their sequential verification process, which allows adversaries to target individual identity factors in isolation. In this work, we propose a cryptographic framework that enhances the IdentiToken model by integrating deterministic homomorphic encryption inspired by the Paillier scheme. Our approach generates a structured token composed of encrypted sub-tokens, each representing either a stable or volatile element of a device’s identity. These tokens can be recomputed in real-time and compared for authentication without exposing raw attribute values. By leveraging the additive homomorphism, we quantify the degree of change in core attributes across sessions, enabling similarity-based identity verification. Beyond authentication, this token architecture supports a wide range of security applications, including real-time fingerprinting, zero-trust access control, anomaly detection and behavioural malware analysis.","author":[{"family":"Tripathi","given":"Shashank"},{"family":"Wöhnert","given":"Kai"},{"family":"Skwarek","given":"Volker"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48441/4427.3459","URL":"https://doi.org/10.48441/4427.3459","source":"datacite"},{"id":"doi:10.26187/deakin.33346416","type":"article-journal","title":"VPFL: A verifiable privacy-preserving federated learning scheme for edge computing systems","abstract":"Federated learning for edge computing is a promising solution in the data booming era, which leverages the computation ability of each edge device to train local models and only shares the model gradients to the central server. However, the frequently transmitted local gradients could also leak the participants’ private data. To protect the privacy of local training data, lots of cryptographic-based Privacy-Preserving Federated Learning (PPFL) schemes have been proposed. However, due to the constrained resource nature of mobile devices and complex cryptographic operations, traditional PPFL schemes fail to provide efficient data confidentiality and lightweight integrity verification simultaneously. To tackle this problem, we propose a Verifiable Privacy-preserving Federated Learning scheme (VPFL) for edge computing systems to prevent local gradients from leaking over the transmission stage. Firstly, we combine the Distributed Selective Stochastic Gradient Descent (DSSGD) method with Paillier homomorphic cryptosystem to achieve the distributed encryption functionality, so as to reduce the computation cost of the complex cryptosystem. Secondly, we further present an online/offline signature method to realize the lightweight gradients integrity verification, where the offline part can be securely outsourced to the edge server. Comprehensive security analysis demonstrates the proposed VPFL can achieve data confidentiality, authentication, and integrity. At last, we evaluate both communication overhead and computation cost of the proposed VPFL scheme, the experimental results have shown VPFL has low computation costs and communication overheads while maintaining high training accuracy.","author":[{"family":"Zhang","given":"J"},{"family":"Liu","given":"Y"},{"family":"Wu","given":"D"},{"family":"Lou","given":"S"},{"family":"Chen","given":"B"},{"family":"Yu","given":"S"}],"issued":{"date-parts":[[2026]]},"DOI":"10.26187/deakin.33346416","URL":"https://doi.org/10.26187/deakin.33346416","source":"datacite"},{"id":"doi:10.24406/publica-9087","type":"article-journal","title":"Flexible Parallel Radix-4 MDC NTT for FHE","abstract":"Fully Homomorphic Encryption (FHE) is a promising technology that allows calculations to be performed on encrypted data. However, its widespread adoption is hindered by high computational requirements. In this paper, we present a design-time flexible FPGA accelerator that aims to speed up the Number Theoretic Transform (NTT), the most crucial bottleneck in FHE schemes. Our accelerator is throughput-oriented and implements a Multipath Delay Commutator (MDC) approach. While previous works focus on radix-2 implementations, we implement a parallel radix-4 NTT architecture. We exploit constants within radix-4 structures and reduce DSP slice utilization by more than 25%. We achieve that by using a recent constant modular multiplication technique by Bertels et al. and leverage special properties of goldilock primes. The polynomial degree, modulus and number of parallel compute cores are design-time configurable. Due to its configurability, our design allows for the evaluation of different area and speed trade-offs and targets various parameter sets and requirements. To demonstrate the applicability of our approach, we conduct a case study and evaluate different configurations and parameters that are applicable to FHE schemes such as CGGI or CKKS. All implementations are done on the AMD Alveo U55C. Compared to our reference software implementation based on OpenFHE, we achieve a speed-up of up to 23.75×. Compared to related works, with the same bandwidth requirements, we achieve up to 3.81× lower ATP and up to 4.33× higher TPS.","author":[{"family":"Stelzer","given":"Tobias"},{"family":"Karl","given":"Patrick"},{"family":"Seelos-Zankl","given":"Andreas"},{"family":"Unav"}],"issued":{"date-parts":[[2026]]},"DOI":"10.24406/publica-9087","URL":"https://doi.org/10.24406/publica-9087","source":"datacite"},{"id":"doi:10.5281/zenodo.20084740","type":"article-journal","title":"The Role of Cryptography in Network Security: A Systematic Review and Emerging Trends","abstract":"Cryptography is the backbone of modern network security, providing confidentiality, integrity, authentication, and non-repudiation for digital communication. However, the rapid evolution of cyber threats, particularly the looming arrival of large-scale quantum computers, poses serious challenges to the cryptographic algorithms that protect today's networks. This paper presents a systematic review of cryptography in network security, following the PRISMA 2020 guidelines. A total of 68 studies published between 2016 and 2025 were selected from five major academic databases: IEEE Xplore, ACM Digital Library, Scopus, Web of Science, and ScienceDirect. The review covers classical symmetric and asymmetric algorithms, widely deployed cryptographic protocols such as TLS 1.3, IPsec, and SSH, and the growing body of work on post-quantum cryptography (PQC). Key findings include the following: NIST finalized three post-quantum cryptographic standards (FIPS 203, 204, and 205) in August 2024; lightweight cryptography standards for IoT devices were published in 2025 with the selection of ASCON; and real-world deployment of hybrid classical/post-quantum schemes has already begun in major web browsers and messaging applications. This paper also examines emerging trends in homomorphic encryption, zero-knowledge proofs, and AI-driven cryptanalysis. Based on the findings, this review identifies critical gaps in PQC migration strategies, IoT security, and the integration of cryptography with artificial intelligence, and proposes directions for future research.","author":[{"family":"Makolo","given":"Daniel"},{"family":"Desmond","given":"Obafemi"},{"family":"Anibe","given":"Dauda"},{"family":"Ikoojo","given":"Ejiga"},{"family":"Adinoyi","given":"Lawal"},{"family":"Okoli","given":"Patience"},{"family":"Ojoache","given":"Idakwoji"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20084740","URL":"https://doi.org/10.5281/zenodo.20084740","source":"datacite"},{"id":"doi:10.5281/zenodo.20084741","type":"article-journal","title":"The Role of Cryptography in Network Security: A Systematic Review and Emerging Trends","abstract":"Cryptography is the backbone of modern network security, providing confidentiality, integrity, authentication, and non-repudiation for digital communication. However, the rapid evolution of cyber threats, particularly the looming arrival of large-scale quantum computers, poses serious challenges to the cryptographic algorithms that protect today's networks. This paper presents a systematic review of cryptography in network security, following the PRISMA 2020 guidelines. A total of 68 studies published between 2016 and 2025 were selected from five major academic databases: IEEE Xplore, ACM Digital Library, Scopus, Web of Science, and ScienceDirect. The review covers classical symmetric and asymmetric algorithms, widely deployed cryptographic protocols such as TLS 1.3, IPsec, and SSH, and the growing body of work on post-quantum cryptography (PQC). Key findings include the following: NIST finalized three post-quantum cryptographic standards (FIPS 203, 204, and 205) in August 2024; lightweight cryptography standards for IoT devices were published in 2025 with the selection of ASCON; and real-world deployment of hybrid classical/post-quantum schemes has already begun in major web browsers and messaging applications. This paper also examines emerging trends in homomorphic encryption, zero-knowledge proofs, and AI-driven cryptanalysis. Based on the findings, this review identifies critical gaps in PQC migration strategies, IoT security, and the integration of cryptography with artificial intelligence, and proposes directions for future research.","author":[{"family":"Makolo","given":"Daniel"},{"family":"Desmond","given":"Obafemi"},{"family":"Anibe","given":"Dauda"},{"family":"Ikoojo","given":"Ejiga"},{"family":"Adinoyi","given":"Lawal"},{"family":"Okoli","given":"Patience"},{"family":"Ojoache","given":"Idakwoji"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20084741","URL":"https://doi.org/10.5281/zenodo.20084741","source":"datacite"},{"id":"doi:10.5281/zenodo.20180625","type":"article-journal","title":"Ep. 2226: When Quantum Breaks Everything","abstract":"Episode summary: The threat from quantum computing isn't theoretical anymore. In August 2024, NIST finalized the first post-quantum cryptography standards—lattice-based algorithms designed to survive attacks from machines that don't yet exist. This episode explores what quantum computers actually do to modern encryption, why the \"harvest-now-decrypt-later\" attack is happening today, and how the internet's cryptographic foundation is being rebuilt. We also dig into the frontier: homomorphic encryption (computing on encrypted data), zero-knowledge proofs, and what it means when the computational substrate itself becomes the vulnerability. Show Notes # When Quantum Breaks Everything: Post-Quantum Cryptography and the Internet's Race Against Time The threat from quantum computing is often framed as distant and theoretical. But there's a problem happening right now that makes it urgent: nation-state adversaries are almost certainly recording encrypted internet traffic today, betting they'll be able to decrypt it in fifteen years when quantum computers mature. This \"harvest-now-decrypt-later\" attack is the real reason the cryptographic infrastructure of the internet is being overhauled. ## How Quantum Computers Break Current Encryption RSA and elliptic-curve cryptography—the algorithms that secure HTTPS, TLS, SSH, and certificate authorities—rely on mathematical problems that are believed to be computationally intractable. Factoring a 2,048-bit number would take classical computers longer than the age of the universe. But in 1994, Peter Shor published an algorithm that, running on a sufficiently powerful quantum computer, collapses that problem to hours or less. The same vulnerability exists in elliptic-curve cryptography, which uses the discrete logarithm problem over elliptic curves. Shor's algorithm breaks both with equal efficiency—which is why both underpin the internet's public-key infrastructure. The timeline, however, is uncertain. Current quantum systems like IBM's Heron and Google's Willow operate with hundreds to low thousands of physical qubits. Breaking RSA would require millions of high-fidelity physical qubits—most serious estimates place cryptographically relevant quantum computing at ten to twenty years out. But that timeline doesn't matter for data being encrypted today. Once it's stored, it can wait. ## NIST's Post-Quantum Standards Recognizing this urgency, NIST launched a post-quantum standardization process in 2016. They received eighty-two algorithm submissions from research teams worldwide and spent eight years evaluating them through mathematical analysis, cryptanalysis attempts, and performance benchmarking. In August 2024, they finalized the first three standards: - **ML-KEM**: A key encapsulation mechanism based on the CRYSTALS-Kyber scheme - **ML-DSA**: A digital signature algorithm based on CRYSTALS-Dilithium - **SLH-DSA**: A hash-based signature algorithm as a backup The first two are lattice-based, representing the field's consensus on what can survive quantum attacks. ## Lattice Cryptography and the Learning With Errors Problem Lattice-based cryptography operates on a deceptively simple premise: a lattice is a regular grid of points in high-dimensional space—imagine an infinitely extending checkerboard in five hundred dimensions. The hard problem at its core is the Learning With Errors (LWE) problem: given a secret vector multiplied by a public matrix plus a small random error term, recovering the secret is believed to be computationally hard—even for quantum computers. No known quantum algorithm provides a meaningful speedup against LWE. This is important to emphasize: we don't have a proof that LWE is hard. We have decades of cryptanalysis, connections to well-studied worst-case lattice problems, and strong evidence. But like RSA, it's a computational assumption, not a theorem. The NIST process mitigated this risk by standardizing algorithms from different mathematical families and including multip","author":[{"family":"Rosehill","given":"Daniel"},{"family":"Tts","given":"Chatterbox"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20180625","URL":"https://doi.org/10.5281/zenodo.20180625","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.22216","type":"manuscript","title":"Representation biases: will we achieve complete understanding by analyzing representations?","abstract":"A common approach in neuroscience is to study neural representations as a means to understand a system -- increasingly, by relating the neural representations to the internal representations learned by computational models. However, a recent work in machine learning (Lampinen, 2024) shows that learned feature representations may be biased to over-represent certain features, and represent others more weakly and less-consistently. For example, simple (linear) features may be more strongly and more consistently represented than complex (highly nonlinear) features. These biases could pose challenges for achieving full understanding of a system through representational analysis. In this perspective, we illustrate these challenges -- showing how feature representation biases can lead to strongly biased inferences from common analyses like PCA, regression, and RSA. We also present homomorphic encryption as a simple case study of the potential for strong dissociation between patterns of representation and computation. We discuss the implications of these results for representational comparisons between systems, and for neuroscience more generally.","author":[{"family":"Lampinen","given":"Andrew"},{"family":"Chan","given":"Stephanie"},{"family":"Li","given":"Yuxuan"},{"family":"Hermann","given":"Katherine"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.22216","URL":"https://doi.org/10.48550/arxiv.2507.22216","source":"datacite"},{"id":"doi:10.48550/arxiv.2504.13198","type":"manuscript","title":"Overcoming Bottlenecks in Homomorphic Encryption for the 2024 Mexican Federal Election","abstract":"On June 2, 2024, Mexico held its federal elections. The majority of Mexican citizens voted in person at the polls in this historic election. For the first time though, Mexican citizens living outside their country were able to vote online via a web app, either on a personal device or using an electronic voting kiosk at one of 23 embassies and consulates in the U.S., Canada, and Europe. In total, 144,734 people voted outside of Mexico: 122,496 on a personal device and 22,238 in-person at a kiosk. Voting was open for remote voting from 8PM, May 18, 2024 to 6PM, June 2, 2024 and was open for in-person voting from 8AM-6PM on June 2, 2024. This article describes the technical and cryptographic tools applied to secure the ex-patriate component of the election and to enable INE (Mexico's National Electoral Institute) to generate provable election results within minutes of the close of the election. This article will also describe how the solutions we present scale to elections on a national level.","author":[{"family":"Landquist","given":"Eric"},{"family":"Sawhney","given":"Nimit"},{"family":"Sawhney","given":"Simer"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2504.13198","URL":"https://doi.org/10.48550/arxiv.2504.13198","source":"datacite"},{"id":"doi:10.48550/arxiv.2502.16877","type":"manuscript","title":"APINT: A Full-Stack Framework for Acceleration of Privacy-Preserving Inference of Transformers based on Garbled Circuits","abstract":"As the importance of Privacy-Preserving Inference of Transformers (PiT) increases, a hybrid protocol that integrates Garbled Circuits (GC) and Homomorphic Encryption (HE) is emerging for its implementation. While this protocol is preferred for its ability to maintain accuracy, it has a severe drawback of excessive latency. To address this, existing protocols primarily focused on reducing HE latency, thus making GC the new latency bottleneck. Furthermore, previous studies only focused on individual computing layers, such as protocol or hardware accelerator, lacking a comprehensive solution at the system level. This paper presents APINT, a full-stack framework designed to reduce PiT's overall latency by addressing the latency problem of GC through both software and hardware solutions. APINT features a novel protocol that reallocates possible GC workloads to alternative methods (i.e., HE or standard matrix operation), substantially decreasing the GC workload. It also suggests GC-friendly circuit generation that reduces the number of AND gates at the most, which is the expensive operator in GC. Furthermore, APINT proposes an innovative netlist scheduling that combines coarse-grained operation mapping and fine-grained scheduling for maximal data reuse and minimal dependency. Finally, APINT's hardware accelerator, combined with its compiler speculation, effectively resolves the memory stall issue. Putting it all together, APINT achieves a remarkable end-to-end reduction in latency, outperforming the existing protocol on CPU platform by 12.2x online and 2.2x offline. Meanwhile, the APINT accelerator not only reduces its latency by 3.3x but also saves energy consumption by 4.6x while operating PiT compared to the state-of-the-art GC accelerator.","author":[{"family":"Cho","given":"Hyunjun"},{"family":"Jeon","given":"Jaeho"},{"family":"Heo","given":"Jaehoon"},{"family":"Kim","given":"Joo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2502.16877","URL":"https://doi.org/10.48550/arxiv.2502.16877","source":"datacite"},{"id":"doi:10.5281/zenodo.19808698","type":"article-journal","title":"A Multi-Layered Security Framework for Cloud-Native ERP Systems (2026)","abstract":"Organizational data management has undergone a fundamental transformation as a result of the migration of Enterprise Resource Planning (ERP) systems to cloud computing environments. This migration offers previously unheard-of scalability, but it also exposes vital operational assets to sophisticated cyber threats. Cloud-native ERPs necessitate a paradigm shift based on five fundamental security pillars: confidentiality, integrity, availability, accountability, and privacy. Traditional on-premise security mainly relied on perimeter defences. A thorough, multi-layered security framework that connects these theoretical foundations with cutting-edge enforcement techniques is presented in this paper. The study specifically looks into integrating AI-driven machine learning models for real-time anomaly detection in user behaviour and transaction logs with Zero Trust Architecture (ZTA) to remove implicit trust in multi-tenant SaaS environments. In order to protect data while processing, the paper also investigates sophisticated cryptographic methods, such as homomorphic encryption. This study shows how businesses can sustain ongoing data sovereignty and operational resilience by evaluating this suggested framework against modern attack vectors like sophisticated ransomware and advanced persistent threats (APTs). For security architects protecting next-generation cloud ERP ecosystems, the resulting blueprint offers practical, contemporary guidelines.","author":[{"family":"Reddy","given":"Som"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19808698","URL":"https://doi.org/10.5281/zenodo.19808698","source":"datacite"},{"id":"doi:10.5281/zenodo.19808699","type":"article-journal","title":"A Multi-Layered Security Framework for Cloud-Native ERP Systems (2026)","abstract":"Organizational data management has undergone a fundamental transformation as a result of the migration of Enterprise Resource Planning (ERP) systems to cloud computing environments. This migration offers previously unheard-of scalability, but it also exposes vital operational assets to sophisticated cyber threats. Cloud-native ERPs necessitate a paradigm shift based on five fundamental security pillars: confidentiality, integrity, availability, accountability, and privacy. Traditional on-premise security mainly relied on perimeter defences. A thorough, multi-layered security framework that connects these theoretical foundations with cutting-edge enforcement techniques is presented in this paper. The study specifically looks into integrating AI-driven machine learning models for real-time anomaly detection in user behaviour and transaction logs with Zero Trust Architecture (ZTA) to remove implicit trust in multi-tenant SaaS environments. In order to protect data while processing, the paper also investigates sophisticated cryptographic methods, such as homomorphic encryption. This study shows how businesses can sustain ongoing data sovereignty and operational resilience by evaluating this suggested framework against modern attack vectors like sophisticated ransomware and advanced persistent threats (APTs). For security architects protecting next-generation cloud ERP ecosystems, the resulting blueprint offers practical, contemporary guidelines.","author":[{"family":"Reddy","given":"Som"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19808699","URL":"https://doi.org/10.5281/zenodo.19808699","source":"datacite"},{"id":"doi:10.24406/publica-5189","type":"article-journal","title":"PRepChain: A versatile privacy-preserving reputation system for dynamic supply chain environments","abstract":"Despite their significant added value in the context of consumer-oriented e-commerce, reputation systems have seen limited adoption in other business settings and models these days. Yet, reliable reputation scores are essential in such settings for easing the establishment of new business relationships—an aspect that is particularly crucial in dynamic supply chain environments, where business partners change frequently. Existing approaches, however, usually target other application domains and fall short in addressing the specific challenges of dynamic supply chains—especially with respect to reliability (incl. availability) and privacy preservation (incl. confidentiality). To close this research gap and to support novel directions in this important research area, we propose PRepChain, our highly-configurable approach that leverages fully homomorphic encryption and distributed competences to provide businesses with a versatile reputation-enriched ecosystem. PRepChain is specifically designed to operate in dynamic environments by also offering a trade-off between data availability and confidentiality guarantees. We make contributions in four primary directions: (i) It offers performant privacy preservation even in large-scale settings, (ii) ensures availability of computed reputation scores, (iii) seamlessly integrates with existing supply chain information systems, and (iv) in addition to subjective reputation scores, it also supports reliably-calculated, i.e., objective, ones, thereby strengthening the reliability of third-party-sourced information. Our evaluation of PRepChain documents its performance - based on a real-world use case -, security, and privacy preservation, hence, its applicability. We conclude that it is indeed destined for practical deployments in modern supply networks.","author":[{"family":"Pennekamp","given":"Jan"},{"family":"Bader","given":"Lennart"},{"family":"Thevaraj","given":"Emildeon"},{"family":"Berninger","given":"Stefanie"},{"family":"Perau","given":"Martin"},{"family":"Schröer","given":"Tobias"},{"family":"Boos","given":"Wolfgang"},{"family":"Kanhere","given":"Salil"},{"family":"Wehrle","given":"Klaus"},{"family":"Unav"}],"issued":{"date-parts":[[2026]]},"DOI":"10.24406/publica-5189","URL":"https://doi.org/10.24406/publica-5189","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.12418","type":"manuscript","title":"High-Performance Pipelined NTT Accelerators with Homogeneous Digit-Serial Modulo Arithmetic","abstract":"The Number Theoretic Transform (NTT) is a fundamental operation in privacy-preserving technologies, particularly within fully homomorphic encryption (FHE). The efficiency of NTT computation directly impacts the overall performance of FHE, making hardware acceleration a critical technology that will enable realistic FHE applications. Custom accelerators, in FPGAs or ASICs, offer significant performance advantages due to their ability to exploit massive parallelism and specialized optimizations. However, the operation of NTT over large moduli requires large word-length modulo arithmetic that limits achievable clock frequencies in hardware and increases hardware area costs. To overcome such deficits, digit-serial arithmetic has been explored for modular multiplication and addition independently. The goal of this work is to leverage digit-serial modulo arithmetic combined with appropriate redundant data representation to design modular pipelined NTT accelerators that operate uniformly on arbitrary small digits, without the need for intermediate (de)serialization. The proposed architecture enables high clock frequencies through regular pipelining while maintaining parallelism. Experimental results demonstrate that the proposed approach outperforms state-of-the-art implementations and reduces hardware complexity under equal performance and input-output bandwidth constraints.","author":[{"family":"Alexakis","given":"George"},{"family":"Schoinianakis","given":"Dimitrios"},{"family":"Dimitrakopoulos","given":"Giorgos"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.12418","URL":"https://doi.org/10.48550/arxiv.2507.12418","source":"datacite"},{"id":"doi:10.5281/zenodo.18190467","type":"article-journal","title":"From Now On, Any AI Can Train on Everything and Memorize Nothing","abstract":"We present Y.I.N.-LLM, a privacy-preserving training architecture for Large Language Models that mathematically guarantees non-memorization of training data. The core innovation is the mandatory DP→ZK→HE ordering (Differential Privacy → Zero-Knowledge Proof → Homomorphic Encryption) applied to transformer gradients during training. Key results: (1) 2.3% accuracy loss at ε=1.0 privacy versus 15-40% with standard DP-SGD; (2) zero extractable training data across all tested attack vectors; (3) native GDPR Article 17 \"right to be forgotten\" compliance via cryptographic gradient subtraction; (4) EU AI Act Article 50 transparency compliance through verifiable privacy proofs. The Non-Memorization Theorem establishes that for any model M trained with Y.I.N.-LLM parameters (ε, δ), the probability of verbatim reproduction is bounded: P[M outputs y | x ∈ training] ≤ e^ε · P[M outputs y | x ∉ training]. This transforms copyright defense from argument to mathematics. Y.I.N.-LLM addresses the $10B+ memorization litigation crisis (NYT v. OpenAI, Getty v. Stability AI, Authors Guild v. OpenAI) by providing the first mathematically verifiable non-memorization guarantee with practical accuracy preservation. Patent Protected: U.S. Provisional Application 63/946,118 (filed December 21, 2025).","author":[{"family":"Mazari","given":"Ilyes"},{"family":"Mazari","given":"Yanis"},{"family":"Mazari","given":"Ilyan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18190467","URL":"https://doi.org/10.5281/zenodo.18190467","source":"datacite"},{"id":"doi:10.5281/zenodo.17955545","type":"article-journal","title":"DP→HE: A Temporal Architecture for Practical Homomorphic Encryption via Bidirectional Taint Analysis","abstract":"Fully Homomorphic Encryption (FHE) enables computation on encrypted data but imposes 100-10,000× performance overhead, rendering it impractical for most applications. Hardware Trusted Execution Environments (TEEs) offer an alternative but require specialized silicon and remain vulnerable to microarchitectural side-channel attacks. This paper introduces DP→HE (Deterministic Parsing to Homomorphic Encryption), a novel architectural framework that achieves 10-100× speedup over traditional FHE while maintaining equivalent cryptographic security. The key innovation is a mandatory temporal constraint: static program analysis must complete before any encryption operations commence. Combined with bidirectional taint analysis—intersecting forward propagation from secret sources with backward propagation from secret sinks—the architecture identifies a mathematically minimal encryption surface, typically comprising only 1-10% of program operations. This represents a paradigm shift from optimizing homomorphic operations to minimizing what requires encryption. The software-only implementation runs on commodity processors without hardware modifications, providing post-quantum security via lattice-based cryptography. U.S. Provisional Patent Application No. 63/941,743, filed December 16, 2025. Commercial implementation requires licensing. Contact: ilyesmazari@hotmail.com","author":[{"family":"Mazari","given":"Ilyes"},{"family":"Mazari","given":"Yanis"},{"family":"Mazari","given":"Ilyan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17955545","URL":"https://doi.org/10.5281/zenodo.17955545","source":"datacite"},{"id":"doi:10.5281/zenodo.17955546","type":"article-journal","title":"DP→HE: A Temporal Architecture for Practical Homomorphic Encryption via Bidirectional Taint Analysis","abstract":"Fully Homomorphic Encryption (FHE) enables computation on encrypted data but imposes 100-10,000× performance overhead, rendering it impractical for most applications. Hardware Trusted Execution Environments (TEEs) offer an alternative but require specialized silicon and remain vulnerable to microarchitectural side-channel attacks. This paper introduces DP→HE (Deterministic Parsing to Homomorphic Encryption), a novel architectural framework that achieves 10-100× speedup over traditional FHE while maintaining equivalent cryptographic security. The key innovation is a mandatory temporal constraint: static program analysis must complete before any encryption operations commence. Combined with bidirectional taint analysis—intersecting forward propagation from secret sources with backward propagation from secret sinks—the architecture identifies a mathematically minimal encryption surface, typically comprising only 1-10% of program operations. This represents a paradigm shift from optimizing homomorphic operations to minimizing what requires encryption. The software-only implementation runs on commodity processors without hardware modifications, providing post-quantum security via lattice-based cryptography. U.S. Provisional Patent Application No. 63/941,743, filed December 16, 2025. Commercial implementation requires licensing. Contact: ilyesmazari@hotmail.com","author":[{"family":"Mazari","given":"Ilyes"},{"family":"Mazari","given":"Yanis"},{"family":"Mazari","given":"Ilyan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17955546","URL":"https://doi.org/10.5281/zenodo.17955546","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.17341","type":"manuscript","title":"MetaFed: Advancing Privacy, Performance, and Sustainability in Federated Metaverse Systems","abstract":"The rapid expansion of immersive Metaverse applications introduces complex challenges at the intersection of performance, privacy, and environmental sustainability. Centralized architectures fall short in addressing these demands, often resulting in elevated energy consumption, latency, and privacy concerns. This paper proposes MetaFed, a decentralized federated learning (FL) framework that enables sustainable and intelligent resource orchestration for Metaverse environments. MetaFed integrates (i) multi-agent reinforcement learning for dynamic client selection, (ii) privacy-preserving FL using homomorphic encryption, and (iii) carbon-aware scheduling aligned with renewable energy availability. Evaluations on MNIST and CIFAR-10 using lightweight ResNet architectures demonstrate that MetaFed achieves up to 25% reduction in carbon emissions compared to conventional approaches, while maintaining high accuracy and minimal communication overhead. These results highlight MetaFed as a scalable solution for building environmentally responsible and privacy-compliant Metaverse infrastructures.","author":[{"family":"Yagiz","given":"Muhammet"},{"family":"Cengiz","given":"Zeynep"},{"family":"Goktas","given":"Polat"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.17341","URL":"https://doi.org/10.48550/arxiv.2508.17341","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.23034","type":"manuscript","title":"Efficient and Encrypted Inference using Binarized Neural Networks within In-Memory Computing Architectures","abstract":"Binarized Neural Networks (BNNs) are a class of deep neural networks designed to utilize minimal computational resources, which drives their popularity across various applications. Recent studies highlight the potential of mapping BNN model parameters onto emerging non-volatile memory technologies, specifically using crossbar architectures, resulting in improved inference performance compared to traditional CMOS implementations. However, the common practice of protecting model parameters from theft attacks by storing them in an encrypted format and decrypting them at runtime introduces significant computational overhead, thus undermining the core principles of in-memory computing, which aim to integrate computation and storage. This paper presents a robust strategy for protecting BNN model parameters, particularly within in-memory computing frameworks. Our method utilizes a secret key derived from a physical unclonable function to transform model parameters prior to storage in the crossbar. Subsequently, the inference operations are performed on the encrypted weights, achieving a very special case of Fully Homomorphic Encryption (FHE) with minimal runtime overhead. Our analysis reveals that inference conducted without the secret key results in drastically diminished performance, with accuracy falling below 15%. These results validate the effectiveness of our protection strategy in securing BNNs within in-memory computing architectures while preserving computational efficiency.","author":[{"family":"Rajendran","given":"Gokulnath"},{"family":"Deb","given":"Suman"},{"family":"Chattopadhyay","given":"Anupam"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.23034","URL":"https://doi.org/10.48550/arxiv.2510.23034","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.21086","type":"manuscript","title":"DictPFL: Efficient and Private Federated Learning on Encrypted Gradients","abstract":"Federated Learning (FL) enables collaborative model training across institutions without sharing raw data. However, gradient sharing still risks privacy leakage, such as gradient inversion attacks. Homomorphic Encryption (HE) can secure aggregation but often incurs prohibitive computational and communication overhead. Existing HE-based FL methods sit at two extremes: encrypting all gradients for full privacy at high cost, or partially encrypting gradients to save resources while exposing vulnerabilities. We present DictPFL, a practical framework that achieves full gradient protection with minimal overhead. DictPFL encrypts every transmitted gradient while keeping non-transmitted parameters local, preserving privacy without heavy computation. It introduces two key modules: Decompose-for-Partial-Encrypt (DePE), which decomposes model weights into a static dictionary and an updatable lookup table, only the latter is encrypted and aggregated, while the static dictionary remains local and requires neither sharing nor encryption; and Prune-for-Minimum-Encrypt (PrME), which applies encryption-aware pruning to minimize encrypted parameters via consistent, history-guided masks. Experiments show that DictPFL reduces communication cost by 402-748$\\times$ and accelerates training by 28-65$\\times$ compared to fully encrypted FL, while outperforming state-of-the-art selective encryption methods by 51-155$\\times$ in overhead and 4-19$\\times$ in speed. Remarkably, DictPFL's runtime is within 2$\\times$ of plaintext FL, demonstrating for the first time, that HE-based private federated learning is practical for real-world deployment. The code is publicly available at https://github.com/UCF-ML-Research/DictPFL.","author":[{"family":"Xue","given":"Jiaqi"},{"family":"Kumar","given":"Mayank"},{"family":"Shang","given":"Yuzhang"},{"family":"Gao","given":"Shangqian"},{"family":"Ning","given":"Rui"},{"family":"Zheng","given":"Mengxin"},{"family":"Jiang","given":"Xiaoqian"},{"family":"Lou","given":"Qian"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.21086","URL":"https://doi.org/10.48550/arxiv.2510.21086","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.01148","type":"manuscript","title":"Latency-Optimal Adaptive Split Inference for Privacy-Preserving Cloud-Edge-End Collaboration","abstract":"Internet of Things (IoT) end devices are increasingly expected to support privacy-sensitive batch inference, yet their limited computational resources often make full local execution of convolutional neural networks impractical. This paper presents a latency-optimal adaptive split inference framework for privacy-preserving cloud-edge-end collaboration. The end device acts as the trust anchor, executes the plaintext model prefix, encrypts the split activation using fully homomorphic encryption (FHE), and keeps the secret key locally, while the edge and cloud execute assigned model segments only on FHE ciphertexts. We formulate collaborative encrypted inference as a split-pair selection problem over an end-side split point and an edge-side termination point. The proposed planner jointly models plaintext prefix execution, encryption, communication, edge-side FHE execution, and cloud-side FHE completion, and supports both convolution-level and block-level split granularities. Experiments on CIFAR-10 and PathMNIST show that the proposed convolution-level collaborative scheme achieves amortized end-to-end speedups of approximately 12.9 times over full-cloud FHE and 3.9 times over the block-level alternative, while preserving the corresponding plaintext-model accuracy. Including modeled communication, the amortized latencies are 1033.279 s/sample on CIFAR-10 and 1023.429 s/sample on PathMNIST.","author":[{"family":"Li","given":"Yi"},{"family":"Zhang","given":"Peng"},{"family":"Au","given":"Man"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.01148","URL":"https://doi.org/10.48550/arxiv.2608.01148","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.26664","type":"manuscript","title":"TGHE: Template-based Graph Homomorphic Encryption for Privacy-Preserving GNN Inference in Edge-Cloud Systems","abstract":"Existing homomorphic encryption (HE)-based GNN systems adopt a graph-centric paradigm that couples per-query cost to global graph size, limiting evaluations to at most ~20k nodes and making them incompatible with dynamic, large-scale financial graphs. We propose TGHE (Template-based Graph Homomorphic Encryption), an ego-centric framework that resolves this by exploiting a template phenomenon: local computation trees in transaction graphs converge into a small set of structural shapes. TGHE canonicalizes ego-graphs at the edge and packs structurally identical trees into shared CKKS ciphertexts for SIMD-parallel encrypted inference, with two long-tail optimizers (Approximate Template Fitting and Topology Collapse) ensuring full SIMD coverage. On DGraphFin (3.7M nodes, 4.3M edges), TGHE-Collapse achieves a 66.9x speedup over the sequential encrypted baseline with less than 0.002 AUC loss.","author":[{"family":"Le","given":"Ngoc"},{"family":"Vu","given":"Thai"},{"family":"Le","given":"John"},{"family":"Cooper","given":"Heath"},{"family":"Shen","given":"Jun"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.26664","URL":"https://doi.org/10.48550/arxiv.2606.26664","source":"datacite"},{"id":"doi:10.48550/arxiv.2605.26903","type":"manuscript","title":"Practical Anonymous Two-Party Gradient Boosting Decision Tree","abstract":"Structured data is well handled by gradient-boosted decision trees (GBDT), which are usually trained on vertically partitioned features across mutually distrustful parties. High speed and interpretability make GBDTs popular in finance and healthcare, where neural networks may fall short. Enabling secure computation for GBDTs poses unique challenges, requiring secure record alignment for comparison. Relying on private set intersection (PSI) is a de facto approach. Mistaking PSI for a safety measure actually exposes which record identifiers (IDs) are shared between the datasets. Although circuit-PSI could help, it is costly for generic uses. New ideas are needed to efficiently train in a \"dark forest\". Aiming to hide the IDs, we initiate the study of anonymous GBDT training on split data held by two parties. Dual circuit-PSI in our design lets the parties alternate as receiver to run pick-then-sum over local features. Via oblivious programmable pseudorandom functions, we propagate circuit-PSI outputs as shared state across runs. Avoiding universal alignment, we resolve the neglected dilemma that ID hiding incurs a cost that scales with domain size. Next, we halve the cost of ciphertext packing used to convert single-instruction multiple-data homomorphic encryption from (ring) learning with errors in prior secure GBDT (Usenix Security' 23) and related secure machine-learning computations. Comparative experiments show our protocol remains competitive with leaky approaches in efficiency. Enabling ID-hiding aggregation, our techniques can extend to other vertically partitioned analytics.","author":[{"family":"Huang","given":"Chenyu"},{"family":"Zhang","given":"Fan"},{"family":"Du","given":"Minxin"},{"family":"Chow","given":"Sherman"},{"family":"Chen","given":"Huangxun"},{"family":"Rao","given":"Huaming"},{"family":"Huang","given":"Danqing"},{"family":"Qian","given":"Bo"},{"family":"Chen","given":"Peng"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.26903","URL":"https://doi.org/10.48550/arxiv.2605.26903","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.16359","type":"manuscript","title":"FEnc$^2$: Unifying Data Packing for Efficient Private Inference via Convolution and Architecture-Aware Fragment Encoding","abstract":"Fully Homomorphic Encryption (FHE) enables privacy-preserving machine learning but incurs extreme computational and memory overhead. These costs come not only from expensive low-level primitives, including Number Theoretic Transform (NTT), rotation, and key-switching, but also from inefficient ciphertext packing at the application level. Existing packing strategies typically preserve either neighboring data elements or feature grouping, but not both, leading to wasted ciphertext slots, excessive rotations, and inflated ciphertext counts. We propose FEnc2, a unified and principled fragment-based encoding framework for CKKS-based private convolutional neural network inference. FEnc2 optimizes slot utilization, rotation complexity, and ciphertext density through two components: 1)Conv-aware Encoding, which analytically selects an optimal fragment size to decouple spatial dependencies and jointly minimize inner-outer rotations across layers, and 2)Arch-aware Ct Compression, which restores ciphertext density after feature- or channel-reduction layers. Together, these transformations reshape encrypted workload structure and reduce homomorphic operations by one to two orders of magnitude. With full memory capacity utilized, i.e., at maximum batch size, FEnc2 achieves end-to-end latency speedups over the state-of-the-art Orion of up to 228.83x on GPU and 226.06x on CPU for LeNet on MNIST, and up to 4.55x on GPU and 9.43x on CPU for MobileNet on ImageNet. FEnc2 is hardware-agnostic yet architecturally transformative: by optimizing encrypted tensor layout before execution, it reduces ciphertext count and workload pressure on hardware, complementing primitive-level optimizations such as NTT and keyswitch accelerators. These results show that application-level data layout is a first-order architectural design dimension for encrypted inference and an important enabler for next-generation FHE systems.","author":[{"family":"Ran","given":"Ran"},{"family":"Gong","given":"Zhaoting"},{"family":"Xu","given":"Nuo"},{"family":"Xu","given":"Yuanchao"},{"family":"Yao","given":"Fan"},{"family":"Wen","given":"Wujie"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.16359","URL":"https://doi.org/10.48550/arxiv.2606.16359","source":"datacite"},{"id":"doi:10.48550/arxiv.2605.14230","type":"manuscript","title":"On the (non-)resilience of encrypted controllers to covert attacks","abstract":"The security of networked control systems (NCS) is receiving increasing attention from both cyber-security and system-theoretic perspectives. The former focuses on classical IT security goals such as confidentiality, integrity, and availability of process data, while the latter investigates tailored attacks (and detection schemes), including covert and zero-dynamics attacks. Confidentiality in control systems can, for instance, be achieved by securely outsourcing the evaluation of the controller to third-party platforms, such as cloud services. The underlying technology enabling such secure computation often is homomorphic encryption (HE). Recent works in encrypted control have proposed modifications to underlying HE schemes to achieve not only confidentiality but also resilience to certain types of integrity attacks. While extensions in this direction are desirable in principle, we show that the integrity problem in encrypted control cannot be solved by public-key HE schemes alone due to their inherent malleability. In other words, the same homomorphisms that enable encrypted control in the first place can be leveraged not only constructively but also destructively. More precisely, we demonstrate that NCS are vulnerable to covert attacks, even when encrypted control is employed. Remarkably, this remains possible without knowledge of an unencrypted model. Yet, resilience to such attacks can still be achieved through complementary techniques. We present an approach based on verifiable computation that integrates with modern homomorphic cryptosystems and is asymptotically secure while incurring no communication overhead.","author":[{"family":"Binfet","given":"Philipp"},{"family":"Adamek","given":"Janis"},{"family":"Darup","given":"Moritz"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.14230","URL":"https://doi.org/10.48550/arxiv.2605.14230","source":"datacite"},{"id":"doi:10.48550/arxiv.2605.13708","type":"manuscript","title":"DisAgg: Distributed Aggregators for Efficient Secure Aggregation in Federated Learning","abstract":"Federated learning enables collaborative model training across distributed clients, yet vanilla FL exposes client updates to the central server. Secure-aggregation schemes protect privacy against an honest-but-curious server, but existing approaches often suffer from many communication rounds, heavy public-key operations, or difficulty handling client dropouts. Recent methods like One-Shot Private Aggregation (OPA) cut rounds to a single server interaction per FL iteration, yet they impose substantial cryptographic and computational overhead on both server and clients. We propose a new protocol called DisAgg that leverages a small committee of clients called Aggregators to perform the aggregation itself: each client secret-shares its update vector to Aggregators, which locally compute partial sums and return only aggregated shares for server-side reconstruction. This design eliminates local masking and expensive homomorphic encryption, reducing endpoint computation while preserving privacy against a curious server and a limited fraction of colluding clients. By leveraging optimal trade-offs between communication and computation costs, DisAgg processes 100k-dimensional update vectors from 100k 5G clients with a 4.6x speedup compared to OPA, the previous best protocol.","author":[{"family":"Mehmood","given":"Haaris"},{"family":"Tatsis","given":"Giorgos"},{"family":"Alexopoulos","given":"Dimitrios"},{"family":"Saravanan","given":"Karthikeyan"},{"family":"Xu","given":"Jie"},{"family":"Drosou","given":"Anastasios"},{"family":"Ozay","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.13708","URL":"https://doi.org/10.48550/arxiv.2605.13708","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.23245","type":"manuscript","title":"Training Machine Learning Models on Encrypted Data: A Privacy-Preserving Framework using Homomorphic Encryption","abstract":"The use of Machine Learning (ML) for data-driven decision-making often relies on access to sensitive datasets, which introduces privacy challenges. Traditional encryption methods protect data at rest or in transit but fail to secure it during processing, exposing it to unauthorized access. Homomorphic encryption emerges as a transformative solution, enabling computations on encrypted data without decryption, thus preserving confidentiality throughout the ML pipeline. This paper addresses the challenge of training ML models on encrypted data while maintaining accuracy and efficiency by proposing a proof-of-concept for a privacy-preserving framework that leverages Cheon-Kim-Kim-Song (CKKS) for approximate real-number arithmetic. Also, it demonstrates the feasibility of training K-Nearest Neighbors (KNN) and linear regression models on encrypted data, and evaluates encrypted inference for a basic Multilayer Perceptron (MLP) architecture. Experimental results show that models trained under Homomorphic encryption achieve performance metrics comparable to plaintext-trained models, validating the approach. However, challenges such as computational overhead, noise management, and limited support for non-polynomial operations persist. This work lays the groundwork for broader adoption of privacy-preserving ML in real-world applications, balancing security with computational feasibility.","author":[{"family":"Marques","given":"Alexandre"},{"family":"Sá","given":"Beatriz"},{"family":"Botelho","given":"Rui"},{"family":"Pinto","given":"Pedro"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.23245","URL":"https://doi.org/10.48550/arxiv.2604.23245","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.09541","type":"manuscript","title":"Trans-RAG: Query-Centric Vector Transformation for Secure Cross-Organizational Retrieval","abstract":"Retrieval Augmented Generation (RAG) systems deployed across organizational boundaries face fundamental tensions between security, accuracy, and efficiency. Current encryption methods expose plaintext during decryption, while federated architectures prevent resource integration and incur substantial overhead. We introduce Trans-RAG, implementing a novel vector space language paradigm where each organization's knowledge exists in a mathematically isolated semantic space. At the core lies vector2Trans, a multi-stage transformation technique that enables queries to dynamically \"speak\" each organization's vector space \"language\" through query-centric transformations, eliminating decryption overhead while maintaining native retrieval efficiency. Security evaluations demonstrate near-orthogonal vector spaces with 89.90° angular separation and 99.81% isolation rates. Experiments across 8 retrievers, 3 datasets, and 3 LLMs show minimal accuracy degradation (3.5% decrease in nDCG@10) and significant efficiency improvements over homomorphic encryption.","author":[{"family":"Liu","given":"Yu"},{"family":"Peng","given":"Kun"},{"family":"Zhang","given":"Wenxiao"},{"family":"Yuan","given":"Fangfang"},{"family":"Cao","given":"Cong"},{"family":"Lu","given":"Wenxuan"},{"family":"Liu","given":"Yanbing"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.09541","URL":"https://doi.org/10.48550/arxiv.2604.09541","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.07047","type":"manuscript","title":"Leveraging ASIC AI Chips for Homomorphic Encryption","abstract":"Homomorphic Encryption (HE) provides strong data privacy for cloud services but at the cost of prohibitive computational overhead. While GPUs have emerged as a practical platform for accelerating HE, there remains an order-of-magnitude energy-efficiency gap compared to specialized (but expensive) HE ASICs. This paper explores an alternate direction: leveraging existing AI accelerators, like Google's TPUs with coarse-grained compute and memory architectures, to offer a path toward ASIC-level energy efficiency for HE. However, this architectural paradigm creates a fundamental mismatch with SoTA HE algorithms designed for GPUs. These algorithms rely heavily on: (1) high-precision (32-bit) integer arithmetic to now run on a TPU's low-throughput vector unit, leaving its high-throughput low-precision (8-bit) matrix engine (MXU) idle, and (2) fine-grained data permutations that are inefficient on the TPU's coarse-grained memory subsystem. Consequently, porting GPU-optimized HE libraries to TPUs results in severe resource under-utilization and performance degradation. To tackle above challenges, we introduce CROSS, a compiler framework that systematically transforms HE workloads to align with the TPU's architecture. CROSS makes two key contributions: (1) Basis-Aligned Transformation (BAT), a novel technique that converts high-precision modular arithmetic into dense, low-precision (INT8) matrix multiplications, unlocking and improving the utilization of TPU's MXU for HE, and (2) Memory-Aligned Transformation (MAT), which eliminates costly runtime data reordering by embedding reordering into compute kernels through offline parameter transformation. CROSS (TPU v6e) achieves higher throughput per watt on NTT and HE operators than WarpDrive, FIDESlib, FAB, HEAP, and Cheddar, establishing AI ASIC as the SotA efficient platform for HE operators. Code: https://github.com/EfficientPPML/CROSS","author":[{"family":"Tong","given":"Jianming"},{"family":"Huang","given":"Tianhao"},{"family":"Dang","given":"Jingtian"},{"family":"De Castro","given":"Leo"},{"family":"Itagi","given":"Anirudh"},{"family":"Golder","given":"Anupam"},{"family":"Ali","given":"Asra"},{"family":"Kun","given":"Jeremy"},{"family":"Jiang","given":"Jevin"},{"family":"Arvind"},{"family":"Suh","given":"GE"},{"family":"Krishna","given":"Tushar"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.07047","URL":"https://doi.org/10.48550/arxiv.2501.07047","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.03425","type":"manuscript","title":"AEGIS: Scaling Long-Sequence Homomorphic Encrypted Transformer Inference via Hybrid Parallelism on Multi-GPU Systems","abstract":"Fully Homomorphic Encryption (FHE) enables privacy-preserving Transformer inference, but long-sequence encrypted Transformers quickly exceed single-GPU memory capacity because encoded weights are already large and encrypted activations grow rapidly with sequence length. Multi-GPU execution therefore becomes unavoidable, yet scaling remains challenging because communication is jointly induced by application-level aggregation and encryption-level RNS coupling. Existing approaches either synchronize between devices frequently or replicate encrypted tensors across devices, leading to excessive communication and latency. We present AEGIS, an Application-Encryption Guided Inference System for scalable long-sequence encrypted Transformer inference on multi-GPU platforms. AEGIS derives device placement from ciphertext dependencies jointly induced by Transformer dataflow and CKKS polynomial coupling, co-locating modulus-coherent and token-coherent data so that communication is introduced only when application dependencies require it, while reordering polynomial operators to overlap the remaining collectives with computation. On 2048-token inputs, AEGIS reduces inter-GPU communication by up to 57.9% in feed-forward networks and 81.3% in self-attention versus prior state-of-the-art designs. On four GPUs, it achieves up to 96.62% scaling efficiency, 3.86x end-to-end speedup, and 69.1% per-device memory reduction. These results establish coordinated application-encryption parallelism as a practical foundation for scalable homomorphic Transformer inference.","author":[{"family":"Gong","given":"Zhaoting"},{"family":"Ran","given":"Ran"},{"family":"Yao","given":"Fan"},{"family":"Wen","given":"Wujie"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.03425","URL":"https://doi.org/10.48550/arxiv.2604.03425","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.20504","type":"manuscript","title":"Meeting in the Middle: A Co-Design Paradigm for FHE and AI Inference","abstract":"Modern cloud inference creates a two sided privacy problem where users reveal sensitive inputs to providers, while providers must execute proprietary model weights inside potentially leaky execution environments. Fully homomorphic encryption (FHE) offers cryptographic guarantees but remains prohibitively expensive for modern architectures. We argue that progress requires co-design where specializing FHE schemes/compilers for the static structure of inference circuits, while simultaneously constraining inference architectures to reduce dominant homomorphic cost drivers. We outline a meet in the middle agenda and concrete optimization targets on both axes.","author":[{"family":"Magri","given":"Bernardo"},{"family":"Marsh","given":"Benjamin"},{"family":"Gebheim","given":"Paul"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.20504","URL":"https://doi.org/10.48550/arxiv.2603.20504","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.18345","type":"manuscript","title":"Theodosian: A Deep Dive into Memory-Hierarchy-Centric FHE Acceleration","abstract":"Fully homomorphic encryption (FHE) enables secure computation on encrypted data, mitigating privacy concerns in cloud and edge environments. However, due to its high compute and memory demands, extensive acceleration research has been pursued across diverse hardware platforms, especially GPUs. In this paper, we perform a microarchitectural analysis of CKKS, a popular FHE scheme, on modern GPUs. Focusing on the memory hierarchy, we demonstrate that dominant kernels remain bound by the on-chip L2 cache despite its high bandwidth, exposing a persistent inner memory wall beyond the conventional off-chip DRAM bottleneck. Further, we reveal that the overall CKKS throughput is constrained by low per-kernel hardware utilization, caused by insufficient intra-kernel parallelism. Motivated by these findings, we introduce Theodosian, a set of complementary, memory-aware optimizations that improve cache efficiency and reduce runtime overheads. Theodosian achieves 1.45--1.83x performance improvements over a highly optimized baseline, Cheddar, across representative CKKS workloads. On an RTX 5090, we reduce the bootstrapping latency for 32,768 complex numbers from 22.1ms to 15.2ms, and further to 12.8ms with additional algorithmic optimizations, establishing a new state-of-the-art GPU performance to the best of our knowledge.","author":[{"family":"Choi","given":"Wonseok"},{"family":"Yu","given":"Hyunah"},{"family":"Kim","given":"Jongmin"},{"family":"Ji","given":"Hyesung"},{"family":"Park","given":"Jaiyoung"},{"family":"Ahn","given":"Jung"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.18345","URL":"https://doi.org/10.48550/arxiv.2512.18345","source":"datacite"},{"id":"doi:10.48550/arxiv.2505.21051","type":"manuscript","title":"SHE-LoRA: Selective Homomorphic Encryption for Federated Tuning with Heterogeneous LoRA","abstract":"Federated fine-tuning is critical for improving the performance of large language models (LLMs) in handling domain-specific tasks while keeping training data decentralized and private. However, prior work has shown that clients' private data can actually be recovered via gradient inversion attacks. Existing privacy preservation techniques against such attacks typically entail performance degradation and high costs, making them ill-suited for clients with heterogeneous data distributions and device capabilities. In this paper, we propose SHE-LoRA, which integrates selective homomorphic encryption (SHE) and low-rank adaptation (LoRA) to enable efficient and privacy-preserving federated tuning of LLMs in cross-device environments. Based on model parameter sensitivity assessment, heterogeneous clients adaptively negotiate and select a subset of model parameters for homomorphic encryption. To ensure accurate model aggregation, we design a column-aware secure aggregation method and customized reparameterization techniques to align the aggregation results with the heterogeneous device capabilities of clients. Extensive experiments demonstrate that SHE-LoRA maintains performance comparable to non-private baselines, achieves strong resistance to state-of-the-art attacks, and significantly reduces communication overhead by 99.71% and encryption time by 99.87%, compared to HE baselines.","author":[{"family":"Liu","given":"Jianmin"},{"family":"Yan","given":"Li"},{"family":"Li","given":"Borui"},{"family":"Yu","given":"Lei"},{"family":"Shen","given":"Chao"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2505.21051","URL":"https://doi.org/10.48550/arxiv.2505.21051","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.13024","type":"manuscript","title":"FedHENet: A Frugal Federated Learning Framework for Heterogeneous Environments","abstract":"Federated Learning (FL) enables collaborative training without centralizing data, essential for privacy compliance in real-world scenarios involving sensitive visual information. Most FL approaches rely on expensive, iterative deep network optimization, which still risks privacy via shared gradients. In this work, we propose FedHENet, extending the FedHEONN framework to image classification. By using a fixed, pre-trained feature extractor and learning only a single output layer, we avoid costly local fine-tuning. This layer is learned by analytically aggregating client knowledge in a single round of communication using homomorphic encryption (HE). Experiments show that FedHENet achieves competitive accuracy compared to iterative FL baselines while demonstrating superior stability performance and up to 70\\% better energy efficiency. Crucially, our method is hyperparameter-free, removing the carbon footprint associated with hyperparameter tuning in standard FL. Code available in https://github.com/AlejandroDopico2/FedHENet/","author":[{"family":"Dopico-Castro","given":"Alejandro"},{"family":"Fontenla-Romero","given":"Oscar"},{"family":"Guijarro-Berdiñas","given":"Bertha"},{"family":"Alonso-Betanzos","given":"Amparo"},{"family":"Digón","given":"Iván"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.13024","URL":"https://doi.org/10.48550/arxiv.2602.13024","source":"datacite"},{"id":"doi:10.48550/arxiv.2601.04912","type":"manuscript","title":"Decentralized Privacy-Preserving Federal Learning of Computer Vision Models on Edge Devices","abstract":"Collaborative training of a machine learning model comes with a risk of sharing sensitive or private data. Federated learning offers a way of collectively training a single global model without the need to share client data, by sharing only the updated parameters from each client's local model. A central server is then used to aggregate parameters from all clients and redistribute the aggregated model back to the clients. Recent findings have shown that even in this scenario, private data can be reconstructed only using information about model parameters. Current efforts to mitigate this are mainly focused on reducing privacy risks on the server side, assuming that other clients will not act maliciously. In this work, we analyzed various methods for improving the privacy of client data concerning both the server and other clients for neural networks. Some of these methods include homomorphic encryption, gradient compression, gradient noising, and discussion on possible usage of modified federated learning systems such as split learning, swarm learning or fully encrypted models. We have analyzed the negative effects of gradient compression and gradient noising on the accuracy of convolutional neural networks used for classification. We have shown the difficulty of data reconstruction in the case of segmentation networks. We have also implemented a proof of concept on the NVIDIA Jetson TX2 module used in edge devices and simulated a federated learning process.","author":[{"family":"Harenčák","given":"Damian"},{"family":"Gajdošech","given":"Lukáš"},{"family":"Madaras","given":"Martin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2601.04912","URL":"https://doi.org/10.48550/arxiv.2601.04912","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.01574","type":"manuscript","title":"IVE: An Accelerator for Single-Server Private Information Retrieval Using Versatile Processing Elements","abstract":"Private information retrieval (PIR) is an essential cryptographic protocol for privacy-preserving applications, enabling a client to retrieve a record from a server's database without revealing which record was requested. Single-server PIR based on homomorphic encryption has particularly gained immense attention for its ease of deployment and reduced trust assumptions. However, single-server PIR remains impractical due to its high computational and memory bandwidth demands. Specifically, reading the entirety of large databases from storage, such as SSDs, severely limits its performance. To address this, we propose IVE, an accelerator for single-server PIR with a systematic extension that enables practical retrieval from large databases using DRAM. Recent advances in DRAM capacity allow PIR for large databases to be served entirely from DRAM, removing its dependence on storage bandwidth. Although the memory bandwidth bottleneck still remains, multi-client batching effectively amortizes database access costs across concurrent requests to improve throughput. However, client-specific data remains a bottleneck, whose bandwidth requirements ultimately limits performance. IVE overcomes this by employing a large on-chip scratchpad with an operation scheduling algorithm that maximizes data reuse, further boosting throughput. Additionally, we introduce sysNTTU, a versatile functional unit that enhances area efficiency without sacrificing performance. We also propose a heterogeneous memory system architecture, which enables a linear scaling of database sizes without a throughput degradation. Consequently, IVE achieves up to 1,275x higher throughput compared to prior PIR hardware solutions.","author":[{"family":"Kim","given":"Sangpyo"},{"family":"Ji","given":"Hyesung"},{"family":"Kim","given":"Jongmin"},{"family":"Choi","given":"Wonseok"},{"family":"Park","given":"Jaiyoung"},{"family":"Ahn","given":"Jung"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.01574","URL":"https://doi.org/10.48550/arxiv.2512.01574","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.17642","type":"manuscript","title":"Quantum Federated Learning: Architectural Elements and Future Directions","abstract":"Federated learning (FL) focuses on collaborative model training without the need to move the private data silos to a central server. Despite its several benefits, the classical FL is plagued with several limitations, such as high computational power required for model training(which is critical for low-resource clients), privacy risks, large update traffic, and non-IID heterogeneity. This chapter surveys a hybrid paradigm - Quantum Federated Learning (QFL), which introduces quantum computation, that addresses multiple challenges of classical FL and offers rapid computing capability while keeping the classical orchestration intact. Firstly, we motivate QFL with a concrete presentation on pain points of classical FL, followed by a discussion on a general architecture of QFL frameworks specifying the roles of client and server, communication primitives and the quantum model placement. We classify the existing QFL systems based on four criteria - quantum architecture (pure QFL, hybrid QFL), data processing method (quantum data encoding, quantum feature mapping, and quantum feature selection &amp; dimensionality reduction), network topology (centralized, hierarchial, decentralized), and quantum security mechanisms (quantum key distribution, quantum homomorphic encryption, quantum differential privacy, blind quantum computing). We then describe applications of QFL in healthcare, vehicular networks, wireless networks, and network security, clearly highlighting where QFL improves communication efficiency, security, and performance compared to classical FL. We close with multiple challenges and future works in QFL, including extension of QFL beyond classification tasks, adversarial attacks, realistic hardware deployment, quantum communication protocols deployment, aggregation of different quantum models, and quantum split learning as an alternative to QFL.","author":[{"family":"Sai","given":"Siva"},{"family":"Sawaika","given":"Abhishek"},{"family":"Singh","given":"Prabhjot"},{"family":"Buyya","given":"Rajkumar"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.17642","URL":"https://doi.org/10.48550/arxiv.2510.17642","source":"datacite"},{"id":"doi:10.71886/bioem.2026.1224382","type":"article-journal","title":"Federated Correction of Batch Effects &amp; Heterogeneity in Single-cell and Multi-omics Genomics (privacy-preserving)","abstract":"The rapid proliferation of genomic data from large‑scale sequencing initiatives presents unprecedented opportunities for precision medicine, population genomics and biotechnology. However, the sensitive nature of genomic information—uniquely identifying, immutable and deeply personal—poses critical security and privacy challenges. Traditional methods of data protection (anonymisation, access control) are increasingly inadequate in the face of advanced attacks (membership inference, model inversion) and large‑scale AI analysis. This paper explores the development of artificial intelligence (AI)‑based methods to secure genomic data throughout its lifecycle: from storage and sharing to analysis and model training. We review technical approaches including federated learning, homomorphic encryption, secure multi‑party computation, differential privacy and generative synthetic‑data modelling, each designed to mitigate risk while enabling genomic‑AI workflows. We present a hypothetical benchmarking study where a federated‑learning pipeline augmented with differential‑privacy noise and encrypted aggregation reduced membership inference risk by ~45 % compared with naïve central models, while retaining &gt;90 % of predictive utility. Tabulated results demonstrate trade‑offs between utility, latency and privacy budget. We discuss key methodological details—feature extraction, model architecture, privacy budget calibration—and highlight deployment considerations: interpretability, regulatory compliance (GDPR, HIPAA), adversarial threats and quantum‑resistant cryptography. Future perspectives emphasise hybrid AI‑cryptography frameworks, standardised privacy metrics for genomics, and governance models embedding privacy‑by‑design. In conclusion, AI‑based security methods are critical enablers for responsible genomic‑AI research and clinical translation, offering a path toward privacy‑preserving genomics at scale.","author":[{"family":"Mohamed Sikkander","given":"Dr"},{"family":"Rodrigues","given":"Joel"},{"family":"Meena","given":"Manoharan"},{"family":"Abuelmakarem","given":"Hala"}],"issued":{"date-parts":[[2026]]},"DOI":"10.71886/bioem.2026.1224382","URL":"https://doi.org/10.71886/bioem.2026.1224382","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.22437","type":"manuscript","title":"mmFHE: mmWave Sensing with End-to-End Fully Homomorphic Encryption","abstract":"We present mmFHE, the first system that enables fully homomorphic encryption (FHE) for end-to-end mmWave radar sensing. mmFHE encrypts raw range profiles on a lightweight edge device and executes the entire mmWave signal-processing and ML inference pipeline homomorphically on an untrusted cloud that operates exclusively on ciphertexts. At the core of mmFHE is a library of seven composable, data-oblivious FHE kernels that replace standard DSP routines with fixed arithmetic circuits. These kernels can be flexibly composed into different application-specific pipelines. We demonstrate this approach on two representative tasks: vital-sign monitoring and gesture recognition. We formally prove two cryptographic guarantees for any pipeline assembled from this library: input privacy, the cloud learns nothing about the sensor data; and data obliviousness, the execution trace is identical on the cloud regardless of the data being processed. These guarantees effectively neutralize various supervised and unsupervised privacy attacks on raw data, including re-identification and data-dependent privacy leakage. Evaluation on three public radar datasets (270 vital-sign recordings, 600 gesture trials) shows that encryption introduces negligible error: HR/RR MAE &lt;10^-3 bpm versus plaintext, and 84.5% gesture accuracy (vs. 84.7% plaintext) with end-to-end cloud GPU latency of 103s for a 10s vital-sign window and 37s for a 3s gesture window. These results show that privacy-preserving end-to-end mmWave sensing is feasible on commodity hardware today.","author":[{"family":"Ahmed","given":"Tanvir"},{"family":"Gao","given":"Yixuan"},{"family":"Armouti","given":"Adnan"},{"family":"Nandakumar","given":"Rajalakshmi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.22437","URL":"https://doi.org/10.48550/arxiv.2603.22437","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.10182","type":"manuscript","title":"Secure Semantic Communication With Homomorphic Encryption","abstract":"In recent years, Semantic Communication (SemCom), which aims to achieve efficient and reliable transmission of meaning between agents, has garnered significant attention from both academia and industry. To ensure the security of communication systems, encryption techniques are employed to safeguard confidentiality and integrity. However, existing encryption schemes encounter obstacles when applied to SemCom. To address this issue, this paper explores the feasibility of applying homomorphic encryption (HE) to SemCom. Initially, we review the encryption algorithms utilized in mobile communication systems and analyze the challenges associated with their application to SemCom. Subsequently, we overview HE techniques and employ scale-invariant feature transform (SIFT) to demonstrate that the extractable semantic information can be preserved in homomorphic encrypted ciphertext. Based on this finding, we further propose the HE-joint source-channel coding (HE-JSCC) scheme, where the traditional JSCC model architecture is modified to support HE operations. Moreover, we present the simulation results for image classification and image generation tasks. Furthermore, we provide potential future research directions for homomorphic encrypted SemCom.","author":[{"family":"Meng","given":"Rui"},{"family":"Fan","given":"Dayu"},{"family":"Gao","given":"Haixiao"},{"family":"Yuan","given":"Yifan"},{"family":"Wang","given":"Bizhu"},{"family":"Xu","given":"Xiaodong"},{"family":"Sun","given":"Mengying"},{"family":"Dong","given":"Chen"},{"family":"Tao","given":"Xiaofeng"},{"family":"Zhang","given":"Ping"},{"family":"Niyato","given":"Dusit"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.10182","URL":"https://doi.org/10.48550/arxiv.2501.10182","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.03427","type":"manuscript","title":"Federated Learning: An approach with Hybrid Homomorphic Encryption","abstract":"Federated Learning (FL) is a distributed machine learning approach that promises privacy by keeping the data on the device. However, gradient reconstruction and membership-inference attacks show that model updates still leak information. Fully Homomorphic Encryption (FHE) can address those privacy concerns but it suffers from ciphertext expansion and requires prohibitive overhead on resource-constrained devices. We propose the first Hybrid Homomorphic Encryption (HHE) framework for FL that pairs the PASTA symmetric cipher with the BFV FHE scheme. Clients encrypt local model updates with PASTA and send both the lightweight ciphertexts and the PASTA key (itself BFV-encrypted) to the server, which performs a homomorphic evaluation of the decryption circuit of PASTA and aggregates the resulting BFV ciphertexts. A prototype implementation, developed on top of the Flower FL framework, shows that on independently and identically distributed MNIST dataset with 12 clients and 10 training rounds, the proposed HHE system achieves 97.6% accuracy, just 1.3% below plaintext, while reducing client upload bandwidth by over 2,000x and cutting client runtime by 30% compared to a system based solely on the BFV FHE scheme. However, server computational cost increases by roughly 15621x for each client participating in the training phase, a challenge to be addressed in future work.","author":[{"family":"Correia","given":"Pedro"},{"family":"Silva","given":"Ivan"},{"family":"Amorim","given":"Ivone"},{"family":"Maia","given":"Eva"},{"family":"Praça","given":"Isabel"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.03427","URL":"https://doi.org/10.48550/arxiv.2509.03427","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.03024","type":"manuscript","title":"Efficient Privacy-Preserving Recommendation on Sparse Data using Fully Homomorphic Encryption","abstract":"In today's data-driven world, recommendation systems personalize user experiences across industries but rely on sensitive data, raising privacy concerns. Fully homomorphic encryption (FHE) can secure these systems, but a significant challenge in applying FHE to recommendation systems is efficiently handling the inherently large and sparse user-item rating matrices. FHE operations are computationally intensive, and naively processing various sparse matrices in recommendation systems would be prohibitively expensive. Additionally, the communication overhead between parties remains a critical concern in encrypted domains. We propose a novel approach combining Compressed Sparse Row (CSR) representation with FHE-based matrix factorization that efficiently handles matrix sparsity in the encrypted domain while minimizing communication costs. Our experimental results demonstrate high recommendation accuracy with encrypted data while achieving the lowest communication costs, effectively preserving user privacy.","author":[{"family":"Chowdhury","given":"Moontaha"},{"family":"Bauer","given":"André"},{"family":"Zhou","given":"Minxuan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.03024","URL":"https://doi.org/10.48550/arxiv.2509.03024","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.14568","type":"manuscript","title":"Leuvenshtein: Efficient FHE-based Edit Distance Computation with Single Bootstrap per Cell","abstract":"This paper presents a novel approach to calculating the Levenshtein (edit) distance within the framework of Fully Homomorphic Encryption (FHE), specifically targeting third-generation schemes like TFHE. Edit distance computations are essential in applications across finance and genomics, such as DNA sequence alignment. We introduce an optimised algorithm that significantly reduces the cost of edit distance calculations called Leuvenshtein. This algorithm specifically reduces the number of programmable bootstraps (PBS) needed per cell of the calculation, lowering it from approximately 94 operations -- required by the conventional Wagner-Fisher algorithm -- to just 1. Additionally, we propose an efficient method for performing equality checks on characters, reducing ASCII character comparisons to only 2 PBS operations. Finally, we explore the potential for further performance improvements by utilising preprocessing when one of the input strings is unencrypted. Our Leuvenshtein achieves up to $278\\times$ faster performance compared to the best available TFHE implementation and up to $39\\times$ faster than an optimised implementation of the Wagner-Fisher algorithm. Moreover, when offline preprocessing is possible due to the presence of one unencrypted input on the server side, an additional $3\\times$ speedup can be achieved.","author":[{"family":"Legiest","given":"Wouter"},{"family":"D'anvers","given":"Jan"},{"family":"Spasic","given":"Bojan"},{"family":"Tran","given":"Nam"},{"family":"Verbauwhede","given":"Ingrid"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.14568","URL":"https://doi.org/10.48550/arxiv.2508.14568","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.12832","type":"manuscript","title":"Efficient and Verifiable Privacy-Preserving Convolutional Computation for CNN Inference with Untrusted Clouds","abstract":"The widespread adoption of convolutional neural networks (CNNs) in resource-constrained scenarios has driven the development of Machine Learning as a Service (MLaaS) system. However, this approach is susceptible to privacy leakage, as the data sent from the client to the untrusted cloud server often contains sensitive information. Existing CNN privacy-preserving schemes, while effective in ensuring data confidentiality through homomorphic encryption and secret sharing, face efficiency bottlenecks, particularly in convolution operations. In this paper, we propose a novel verifiable privacy-preserving scheme tailored for CNN convolutional layers. Our scheme enables efficient encryption and decryption, allowing resource-constrained clients to securely offload computations to the untrusted cloud server. Additionally, we present a verification mechanism capable of detecting the correctness of the results with a success probability of at least $1-\\frac{1}{\\left|Z\\right|}$. Extensive experiments conducted on 10 datasets and various CNN models demonstrate that our scheme achieves speedups ranging $26 \\times$ ~ $\\ 87\\times$ compared to the original plaintext model while maintaining accuracy.","author":[{"family":"Lu","given":"Jinyu"},{"family":"Sun","given":"Xinrong"},{"family":"Tao","given":"Yunting"},{"family":"Ji","given":"Tong"},{"family":"Kong","given":"Fanyu"},{"family":"Yang","given":"Guoqiang"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.12832","URL":"https://doi.org/10.48550/arxiv.2508.12832","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.13715","type":"manuscript","title":"Trans-XFed: An Explainable Federated Learning for Supply Chain Credit Assessment","abstract":"This paper proposes a Trans-XFed architecture that combines federated learning with explainable AI techniques for supply chain credit assessment. The proposed model aims to address several key challenges, including privacy, information silos, class imbalance, non-identically and independently distributed (Non-IID) data, and model interpretability in supply chain credit assessment. We introduce a performance-based client selection strategy (PBCS) to tackle class imbalance and Non-IID problems. This strategy achieves faster convergence by selecting clients with higher local F1 scores. The FedProx architecture, enhanced with homomorphic encryption, is used as the core model, and further incorporates a transformer encoder. The transformer encoder block provides insights into the learned features. Additionally, we employ the integrated gradient explainable AI technique to offer insights into decision-making. We demonstrate the effectiveness of Trans-XFed through experimental evaluations on real-world supply chain datasets. The obtained results show its ability to deliver accurate credit assessments compared to several baselines, while maintaining transparency and privacy.","author":[{"family":"Shi","given":"Jie"},{"family":"Siebes","given":"Arno"},{"family":"Mehrkanoon","given":"Siamak"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.13715","URL":"https://doi.org/10.48550/arxiv.2508.13715","source":"datacite"},{"id":"doi:10.48550/arxiv.2503.16233","type":"manuscript","title":"Empirical Analysis of Privacy-Fairness-Accuracy Trade-offs in Federated Learning: A Step Towards Responsible AI","abstract":"Federated Learning (FL) enables collaborative model training while preserving data privacy; however, balancing privacy preservation (PP) and fairness poses significant challenges. In this paper, we present the first unified large-scale empirical study of privacy-fairness-utility trade-offs in FL, advancing toward responsible AI deployment. Specifically, we systematically compare Differential Privacy (DP), Homomorphic Encryption (HE), and Secure Multi-Party Computation (SMC) with fairness-aware optimizers including q-FedAvg, q-MAML, Ditto, evaluating their performance under IID and non-IID scenarios using benchmark (MNIST, Fashion-MNIST) and real-world datasets (Alzheimer's MRI, credit-card fraud detection). Our analysis reveals HE and SMC significantly outperform DP in achieving equitable outcomes under data skew, although at higher computational costs. Remarkably, we uncover unexpected interactions: DP mechanisms can negatively impact fairness, and fairness-aware optimizers can inadvertently reduce privacy effectiveness. We conclude with practical guidelines for designing robust FL systems that deliver equitable, privacy-preserving, and accurate outcomes.","author":[{"family":"Wasif","given":"Dawood"},{"family":"Chen","given":"Dian"},{"family":"Madabushi","given":"Sindhuja"},{"family":"Alluru","given":"Nithin"},{"family":"Moore","given":"Terrence"},{"family":"Cho","given":"Jin"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2503.16233","URL":"https://doi.org/10.48550/arxiv.2503.16233","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.12050","type":"manuscript","title":"IDFace: Face Template Protection for Efficient and Secure Identification","abstract":"As face recognition systems (FRS) become more widely used, user privacy becomes more important. A key privacy issue in FRS is protecting the user's face template, as the characteristics of the user's face image can be recovered from the template. Although recent advances in cryptographic tools such as homomorphic encryption (HE) have provided opportunities for securing the FRS, HE cannot be used directly with FRS in an efficient plug-and-play manner. In particular, although HE is functionally complete for arbitrary programs, it is basically designed for algebraic operations on encrypted data of predetermined shape, such as a polynomial ring. Thus, a non-tailored combination of HE and the system can yield very inefficient performance, and many previous HE-based face template protection methods are hundreds of times slower than plain systems without protection. In this study, we propose IDFace, a new HE-based secure and efficient face identification method with template protection. IDFace is designed on the basis of two novel techniques for efficient searching on a (homomorphically encrypted) biometric database with an angular metric. The first technique is a template representation transformation that sharply reduces the unit cost for the matching test. The second is a space-efficient encoding that reduces wasted space from the encryption algorithm, thus saving the number of operations on encrypted templates. Through experiments, we show that IDFace can identify a face template from among a database of 1M encrypted templates in 126ms, showing only 2X overhead compared to the identification over plaintexts.","author":[{"family":"Kim","given":"Sunpill"},{"family":"Paik","given":"Seunghun"},{"family":"Hwang","given":"Chanwoo"},{"family":"Kim","given":"Dongsoo"},{"family":"Shin","given":"Junbum"},{"family":"Seo","given":"Jae"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.12050","URL":"https://doi.org/10.48550/arxiv.2507.12050","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.05649","type":"manuscript","title":"DESIGN: Encrypted GNN Inference via Server-Side Input Graph Pruning","abstract":"Graph Neural Networks (GNNs) have achieved state-of-the-art performance in various graph-based learning tasks. However, enabling privacy-preserving GNNs in encrypted domains, such as under Fully Homomorphic Encryption (FHE), typically incurs substantial computational overhead, rendering real-time and privacy-preserving inference impractical. In this work, we propose DESIGN (EncrypteD GNN Inference via sErver-Side Input Graph pruNing), a novel framework for efficient encrypted GNN inference. DESIGN tackles the critical efficiency limitations of existing FHE GNN approaches, which often overlook input data redundancy and apply uniform computational strategies. Our framework achieves significant performance gains through a hierarchical optimization strategy executed entirely on the server: first, FHE-compatible node importance scores (based on encrypted degree statistics) are computed from the encrypted graph. These scores then guide a homomorphic partitioning process, generating multi-level importance masks directly under FHE. This dynamically generated mask facilitates both input graph pruning (by logically removing unimportant elements) and a novel adaptive polynomial activation scheme, where activation complexity is tailored to node importance levels. Empirical evaluations demonstrate that DESIGN substantially accelerates FHE GNN inference compared to state-of-the-art methods while maintaining competitive model accuracy, presenting a robust solution for secure graph analytics. Our implementation is publicly available at https://github.com/LabRAI/DESIGN.","author":[{"family":"Zhao","given":"Kaixiang"},{"family":"Attalla","given":"Joseph"},{"family":"Lou","given":"Qian"},{"family":"Dong","given":"Yushun"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.05649","URL":"https://doi.org/10.48550/arxiv.2507.05649","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.12358","type":"manuscript","title":"Relative Entropy Regularized Reinforcement Learning for Efficient Encrypted Policy Synthesis","abstract":"We propose an efficient encrypted policy synthesis to develop privacy-preserving model-based reinforcement learning. We first demonstrate that the relative-entropy-regularized reinforcement learning framework offers a computationally convenient linear and ``min-free'' structure for value iteration, enabling a direct and efficient integration of fully homomorphic encryption with bootstrapping into policy synthesis. Convergence and error bounds are analyzed as encrypted policy synthesis propagates errors under the presence of encryption-induced errors including quantization and bootstrapping. Theoretical analysis is validated by numerical simulations. Results demonstrate the effectiveness of the RERL framework in integrating FHE for encrypted policy synthesis.","author":[{"family":"Suh","given":"Jihoon"},{"family":"Jang","given":"Yeongjun"},{"family":"Teranishi","given":"Kaoru"},{"family":"Tanaka","given":"Takashi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.12358","URL":"https://doi.org/10.48550/arxiv.2506.12358","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.08461","type":"manuscript","title":"ABC-FHE : A Resource-Efficient Accelerator Enabling Bootstrappable Parameters for Client-Side Fully Homomorphic Encryption","abstract":"As the demand for privacy-preserving computation continues to grow, fully homomorphic encryption (FHE)-which enables continuous computation on encrypted data-has become a critical solution. However, its adoption is hindered by significant computational overhead, requiring 10000-fold more computation compared to plaintext processing. Recent advancements in FHE accelerators have successfully improved server-side performance, but client-side computations remain a bottleneck, particularly under bootstrappable parameter configurations, which involve combinations of encoding, encrypt, decoding, and decrypt for large-sized parameters. To address this challenge, we propose ABC-FHE, an area- and power-efficient FHE accelerator that supports bootstrappable parameters on the client side. ABC-FHE employs a streaming architecture to maximize performance density, minimize area usage, and reduce off-chip memory access. Key innovations include a reconfigurable Fourier engine capable of switching between NTT and FFT modes. Additionally, an on-chip pseudo-random number generator and a unified on-the-fly twiddle factor generator significantly reduce memory demands, while optimized task scheduling enhances the CKKS client-side processing, achieving reduced latency. Overall, ABC-FHE occupies a die area of 28.638 mm2 and consumes 5.654 W of power in 28 nm technology. It delivers significant performance improvements, achieving a 1112x speed-up in encoding and encryption execution time compared to a CPU, and 214x over the state-of-the-art client-side accelerator. For decoding and decryption, it achieves a 963x speed-up over the CPU and 82x over the state-of-the-art accelerator.","author":[{"family":"Yune","given":"Sungwoong"},{"family":"Lee","given":"Hyojeong"},{"family":"Putra","given":"Adiwena"},{"family":"Cho","given":"Hyunjun"},{"family":"Manh","given":"Cuong"},{"family":"Jeon","given":"Jaeho"},{"family":"Kim","given":"Joo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.08461","URL":"https://doi.org/10.48550/arxiv.2506.08461","source":"datacite"},{"id":"doi:10.48550/arxiv.2503.15916","type":"manuscript","title":"ALLMod: Exploring $\\underline{\\mathbf{A}}$rea-Efficiency of $\\underline{\\mathbf{L}}$UT-based $\\underline{\\mathbf{L}}$arge Number $\\underline{\\mathbf{Mod}}$ular Reduction via Hybrid Workloads","abstract":"Modular arithmetic, particularly modular reduction, is widely used in cryptographic applications such as homomorphic encryption (HE) and zero-knowledge proofs (ZKP). High-bit-width operations are crucial for enhancing security; however, they are computationally intensive due to the large number of modular operations required. The lookup-table-based (LUT-based) approach, a ``space-for-time'' technique, reduces computational load by segmenting the input number into smaller bit groups, pre-computing modular reduction results for each segment, and storing these results in LUTs. While effective, this method incurs significant hardware overhead due to extensive LUT usage. In this paper, we introduce ALLMod, a novel approach that improves the area efficiency of LUT-based large-number modular reduction by employing hybrid workloads. Inspired by the iterative method, ALLMod splits the bit groups into two distinct workloads, achieving lower area costs without compromising throughput. We first develop a template to facilitate workload splitting and ensure balanced distribution. Then, we conduct design space exploration to evaluate the optimal timing for fusing workload results, enabling us to identify the most efficient design under specific constraints. Extensive evaluations show that ALLMod achieves up to $1.65\\times$ and $3\\times$ improvements in area efficiency over conventional LUT-based methods for bit-widths of $128$ and $8,192$, respectively.","author":[{"family":"Liu","given":"Fangxin"},{"family":"Li","given":"Haomin"},{"family":"Wang","given":"Zongwu"},{"family":"Zhang","given":"Bo"},{"family":"Zhang","given":"Mingzhe"},{"family":"Yan","given":"Shoumeng"},{"family":"Jiang","given":"Li"},{"family":"Guan","given":"Haibing"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2503.15916","URL":"https://doi.org/10.48550/arxiv.2503.15916","source":"datacite"},{"id":"doi:10.48550/arxiv.2504.15817","type":"manuscript","title":"EFFACT: A Highly Efficient Full-Stack FHE Acceleration Platform","abstract":"Fully Homomorphic Encryption (FHE) is a set of powerful cryptographic schemes that allows computation to be performed directly on encrypted data with an unlimited depth. Despite FHE's promising in privacy-preserving computing, yet in most FHE schemes, ciphertext generally blows up thousands of times compared to the original message, and the massive amount of data load from off-chip memory for bootstrapping and privacy-preserving machine learning applications (such as HELR, ResNet-20), both degrade the performance of FHE-based computation. Several hardware designs have been proposed to address this issue, however, most of them require enormous resources and power. An acceleration platform with easy programmability, high efficiency, and low overhead is a prerequisite for practical application. This paper proposes EFFACT, a highly efficient full-stack FHE acceleration platform with a compiler that provides comprehensive optimizations and vector-friendly hardware. We start by examining the computational overhead across different real-world benchmarks to highlight the potential benefits of reallocating computing resources for efficiency enhancement. Then we make a design space exploration to find an optimal SRAM size with high utilization and low cost. On the other hand, EFFACT features a novel optimization named streaming memory access which is proposed to enable high throughput with limited SRAMs. Regarding the software-side optimization, we also propose a circuit-level function unit reuse scheme, to substantially reduce the computing resources without performance degradation. Moreover, we design novel NTT and automorphism units that are suitable for a cost-sensitive and highly efficient architecture, leading to low area. For generality, EFFACT is also equipped with an ISA and a compiler backend that can support several FHE schemes like CKKS, BGV, and BFV.","author":[{"family":"Huang","given":"Yi"},{"family":"Gong","given":"Xinsheng"},{"family":"Kong","given":"Xiangyu"},{"family":"Chen","given":"Dibei"},{"family":"Zhu","given":"Jianfeng"},{"family":"Zhu","given":"Wenping"},{"family":"Li","given":"Liangwei"},{"family":"Gao","given":"Mingyu"},{"family":"Wei","given":"Shaojun"},{"family":"Zhang","given":"Aoyang"},{"family":"Liu","given":"Leibo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2504.15817","URL":"https://doi.org/10.48550/arxiv.2504.15817","source":"datacite"},{"id":"doi:10.48550/arxiv.2505.01273","type":"manuscript","title":"Anti-adversarial Learning: Desensitizing Prompts for Large Language Models","abstract":"With the widespread use of LLMs, preserving privacy in user prompts has become crucial, as prompts risk exposing privacy and sensitive data to the cloud LLMs. Traditional techniques like homomorphic encryption, secure multi-party computation, and federated learning face challenges due to heavy computational costs and user participation requirements, limiting their applicability in LLM scenarios. In this paper, we propose PromptObfus, a novel method for desensitizing LLM prompts. The core idea of PromptObfus is \"anti-adversarial\" learning, which perturbs privacy words in the prompt to obscure sensitive information while retaining the stability of model predictions. Specifically, PromptObfus frames prompt desensitization as a masked language modeling task, replacing privacy-sensitive terms with a [MASK] token. A desensitization model is trained to generate candidate replacements for each masked position. These candidates are subsequently selected based on gradient feedback from a surrogate model, ensuring minimal disruption to the task output. We demonstrate the effectiveness of our approach on three NLP tasks. Results show that PromptObfus effectively prevents privacy inference from remote LLMs while preserving task performance.","author":[{"family":"Li","given":"Xuan"},{"family":"Yin","given":"Zhe"},{"family":"Gu","given":"Xiaodong"},{"family":"Shen","given":"Beijun"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2505.01273","URL":"https://doi.org/10.48550/arxiv.2505.01273","source":"datacite"},{"id":"doi:10.48550/arxiv.2511.03341","type":"manuscript","title":"LaMoS: Enabling Efficient Large Number Modular Multiplication through SRAM-based CiM Acceleration","abstract":"Barrett's algorithm is one of the most widely used methods for performing modular multiplication, a critical nonlinear operation in modern privacy computing techniques such as homomorphic encryption (HE) and zero-knowledge proofs (ZKP). Since modular multiplication dominates the processing time in these applications, computational complexity and memory limitations significantly impact performance. Computing-in-Memory (CiM) is a promising approach to tackle this problem. However, existing schemes currently suffer from two main problems: 1) Most works focus on low bit-width modular multiplication, which is inadequate for mainstream cryptographic algorithms such as elliptic curve cryptography (ECC) and the RSA algorithm, both of which require high bit-width operations; 2) Recent efforts targeting large number modular multiplication rely on inefficient in-memory logic operations, resulting in high scaling costs for larger bit-widths and increased latency. To address these issues, we propose LaMoS, an efficient SRAM-based CiM design for large-number modular multiplication, offering high scalability and area efficiency. First, we analyze the Barrett's modular multiplication method and map the workload onto SRAM CiM macros for high bit-width cases. Additionally, we develop an efficient CiM architecture and dataflow to optimize large-number modular multiplication. Finally, we refine the mapping scheme for better scalability in high bit-width scenarios using workload grouping. Experimental results show that LaMoS achieves a $7.02\\times$ speedup and reduces high bit-width scaling costs compared to existing SRAM-based CiM designs.","author":[{"family":"Li","given":"Haomin"},{"family":"Liu","given":"Fangxin"},{"family":"Guan","given":"Chenyang"},{"family":"Wang","given":"Zongwu"},{"family":"Jiang","given":"Li"},{"family":"Guan","given":"Haibing"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2511.03341","URL":"https://doi.org/10.48550/arxiv.2511.03341","source":"datacite"},{"id":"doi:10.48550/arxiv.2511.00737","type":"manuscript","title":"EP-HDC: Hyperdimensional Computing with Encrypted Parameters for High-Throughput Privacy-Preserving Inference","abstract":"While homomorphic encryption (HE) provides strong privacy protection, its high computational cost has restricted its application to simple tasks. Recently, hyperdimensional computing (HDC) applied to HE has shown promising performance for privacy-preserving machine learning (PPML). However, when applied to more realistic scenarios such as batch inference, the HDC-based HE has still very high compute time as well as high encryption and data transmission overheads. To address this problem, we propose HDC with encrypted parameters (EP-HDC), which is a novel PPML approach featuring client-side HE, i.e., inference is performed on a client using a homomorphically encrypted model. Our EP-HDC can effectively mitigate the encryption and data transmission overhead, as well as providing high scalability with many clients while providing strong protection for user data and model parameters. In addition to application examples for our client-side PPML, we also present design space exploration involving quantization, architecture, and HE-related parameters. Our experimental results using the BFV scheme and the Face/Emotion datasets demonstrate that our method can improve throughput and latency of batch inference by orders of magnitude over previous PPML methods (36.52~1068x and 6.45~733x, respectively) with less than 1% accuracy degradation.","author":[{"family":"Park","given":"Jaewoo"},{"family":"Quan","given":"Chenghao"},{"family":"Lee","given":"Jongeun"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2511.00737","URL":"https://doi.org/10.48550/arxiv.2511.00737","source":"datacite"},{"id":"doi:10.5281/zenodo.17720023","type":"article-journal","title":"Review on Cloud Data Security Using VGG19-Deep Learning and Homomorphic Encryption","abstract":"Data is the new currency as lot of the user’s presence online is an upward trend. As a consequence the data storage on the various cloud platforms has been a new normal. Data security in cloud has turn formidable due to unique security issues and challenges. Conventional methods on security may not always cater a proper barter between computational efficiency and data security. This review paper discusses the blending facial key features of the face obtained from VGG19 deep learning with homomorphic encryption to enrich cloud data security. The VGG19 allows for robust feature extraction from face and facial key points for authentication, while homomorphic encryption scales computation in encrypted form; it acquire enhanced accuracy with scalability and preservation of privacy. Thus, this method ensure a better approach in next-generation cloud security frameworks.","author":[{"family":"Chandrasekhar","given":"Tadi"},{"family":"Basanta","given":"Th"},{"family":"Swaminathan","given":"JN"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17720023","URL":"https://doi.org/10.5281/zenodo.17720023","source":"datacite"},{"id":"doi:10.5281/zenodo.17720024","type":"article-journal","title":"Review on Cloud Data Security Using VGG19-Deep Learning and Homomorphic Encryption","abstract":"Data is the new currency as lot of the user’s presence online is an upward trend. As a consequence the data storage on the various cloud platforms has been a new normal. Data security in cloud has turn formidable due to unique security issues and challenges. Conventional methods on security may not always cater a proper barter between computational efficiency and data security. This review paper discusses the blending facial key features of the face obtained from VGG19 deep learning with homomorphic encryption to enrich cloud data security. The VGG19 allows for robust feature extraction from face and facial key points for authentication, while homomorphic encryption scales computation in encrypted form; it acquire enhanced accuracy with scalability and preservation of privacy. Thus, this method ensure a better approach in next-generation cloud security frameworks.","author":[{"family":"Chandrasekhar","given":"Tadi"},{"family":"Basanta","given":"Th"},{"family":"Swaminathan","given":"JN"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17720024","URL":"https://doi.org/10.5281/zenodo.17720024","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.21483","type":"manuscript","title":"Introducing GRAFHEN: Group-based Fully Homomorphic Encryption without Noise","abstract":"We present GRAFHEN, a new cryptographic scheme which offers Fully Homomorphic Encryption without the need for bootstrapping (or in other words, without noise). Building on the work of Nuida and others, we achieve this using encodings in groups. The groups are represented on a machine using rewriting systems. In this way the subgroup membership problem, which an attacker would have to solve in order to break the scheme, becomes maximally hard, while performance is preserved. In fact we include a simple benchmark demonstrating that our implementation runs several orders of magnitude faster than existing standards. We review many possible attacks against our protocol and explain how to protect the scheme in each case.","author":[{"family":"Guillot","given":"Pierre"},{"family":"Duc","given":"Auguste"},{"family":"Koskas","given":"Michel"},{"family":"Méhats","given":"Florian"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.21483","URL":"https://doi.org/10.48550/arxiv.2510.21483","source":"datacite"},{"id":"doi:10.48550/arxiv.2504.16091","type":"manuscript","title":"Post-Quantum Homomorphic Encryption: A Case for Code-Based Alternatives","abstract":"Homomorphic Encryption (HE) allows secure and privacy-protected computation on encrypted data without the need to decrypt it. Since Shor's algorithm rendered prime factorisation and discrete logarithm-based ciphers insecure with quantum computations, researchers have been working on building post-quantum homomorphic encryption (PQHE) algorithms. Most of the current PQHE algorithms are secured by Lattice-based problems and there have been limited attempts to build ciphers based on error-correcting code-based problems. This review presents an overview of the current approaches to building PQHE schemes and justifies code-based encryption as a novel way to diversify post-quantum algorithms. We present the mathematical underpinnings of existing code-based cryptographic frameworks and their security and efficiency guarantees. We compare lattice-based and code-based homomorphic encryption solutions identifying challenges that have inhibited the progress of code-based schemes. We finally propose five new research directions to advance post-quantum code-based homomorphic encryption.","author":[{"family":"Bhoi","given":"Siddhartha"},{"family":"Arakala","given":"Arathi"},{"family":"Corman","given":"Amy"},{"family":"Rao","given":"Asha"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2504.16091","URL":"https://doi.org/10.48550/arxiv.2504.16091","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.28198","type":"manuscript","title":"Performative Privacy: When Differential Privacy Maximizes Utility","abstract":"Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and thus participation, improving utility in the long term. However, this claim has not been formalized so far. In parallel, performative learning provides a framework for studying learning systems whose deployment affects the data they later observe. In this work, we bring these two perspectives together and introduce \\emph{performative privacy}, where data leakage reduces future participation. We study a simple model where agents repeatedly contribute data for mean estimation but may leave the system when their data is leaked. Privacy is implemented through differentially private mechanisms, creating a trade-off between estimation noise and future participation. We show, through a theoretical study of the dynamics and numerical experiments, that a finite privacy budget can outperform non-private estimation in the long term when the feedback loop between leakage and participation is sufficiently strong. This provides first evidence that differential privacy can be optimal not only as a protection mechanism, but also from the perspective of long-term utility.","author":[{"family":"Mukherjee","given":"Uddalak"},{"family":"Cyffers","given":"Edwige"},{"family":"Chevaleyre","given":"Yann"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.28198","URL":"https://doi.org/10.48550/arxiv.2608.28198","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.27782","type":"manuscript","title":"Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots","abstract":"Memorization in large language models is measured through a zoo of definitions whose formal relations are unknown, and differential privacy (DP) is treated as a proxy against all of them at once. We pin down the exact DP constant for the two that carry the practical weight, counterfactual memorization and adaptive extraction, and show that they do not control each other. Under $f$-DP, every adaptive extraction protocol with list budget $m$ succeeds with probability at most $1-f(κ)$ for the oblivious baseline $κ$, and the bound is tight on a dense set of baselines: DP uniformly controls extraction exactly up to a threshold in how well the secret can be guessed a priori. Min-entropy certifies that baseline distribution-free, since $H_\\infty\\geε\\log_2 e+\\log_2(m/τ)$ holds extraction below a risk level $τ\\le1/2$ under pure $ε$-DP for every prior, and is exact on uniform priors. On the memorization side, $f$-DP caps the counterfactual memorization of any bounded score at an advantage functional $η(f)$, equal to $\\tanh(ε/2)$ under pure DP; for $k\\ge2$ duplicated copies the naive $ε\\mapsto kε$ bound $\\tanh(kε/2)$ is unattainable, the exact constant being a closed-form staircase attained by geometric noisy counting. That cap is attained inside the local score class used in practice, and it is there that the two measures separate: one mechanism is memorized yet unextractable, another fully extractable yet exactly invisible to every loss-based score. The two-sided blind spot this opens for loss-based auditing and unlearning verification survives on billion-parameter models: a reserved-trigger release is recovered verbatim from one prompt while the audits practitioners deploy certify it clean.","author":[{"family":"Che","given":"Xujun"},{"family":"Xu","given":"Depeng"},{"family":"Yuan","given":"Shuhan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.27782","URL":"https://doi.org/10.48550/arxiv.2608.27782","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.01607","type":"manuscript","title":"Minimax optimal differentially private synthetic data for smooth queries","abstract":"Differentially private synthetic data enables the sharing and analysis of sensitive datasets while providing rigorous privacy guarantees for individual contributors. A central challenge is to achieve strong utility guarantees for meaningful downstream analysis. Many existing methods ensure uniform accuracy over broad query classes, such as all Lipschitz functions, but this level of generality often leads to suboptimal rates for statistics of practical interest. Since many common data analysis queries exhibit smoothness beyond what worst-case Lipschitz bounds capture, we ask whether exploiting this additional structure can yield improved utility. We study the problem of generating $(\\varepsilon,δ)$-differentially private synthetic data from a dataset of size $n$ supported on the hypercube $[-1,1]^d$, with utility guarantees uniformly for all smooth queries having bounded derivatives up to order $k$. We propose a polynomial-time algorithm that achieves a minimax error rate of $O_{k,d}(n^{-\\min \\{1, \\frac{k}{d}\\}})$, up to a $\\log(n)$ factor. This characterization uncovers a phase transition at $k=d$. Our results generalize the Chebyshev moment matching framework of (Musco et al., 2025; Wang et al., 2016) and strictly improve the error rates for $k$-smooth queries established in \\citep{wang2016differentially}. Moreover, we establish the first minimax lower bound for the utility of $(\\varepsilon,δ)$-differentially private synthetic data with respect to $k$-smooth queries, extending the Wasserstein lower bound for $\\varepsilon$-differential privacy in (Boedihardjo et al., 2024).","author":[{"family":"Ding","given":"Rundong"},{"family":"He","given":"Yiyun"},{"family":"Zhu","given":"Yizhe"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.01607","URL":"https://doi.org/10.48550/arxiv.2602.01607","source":"datacite"},{"id":"doi:10.48550/arxiv.2605.20069","type":"manuscript","title":"Smooth Partial Lotteries for Stable Randomized Selection","abstract":"Competitive selection processes, from scientific funding to admissions and hiring, use evaluations to score candidates, and eventually choose a subset of them based on those scores. Recently, many organizations have adopted partial lotteries, which randomize selection based on evaluation scores. However, existing lottery designs are inherently unstable, as a small change to a single candidate's score can cause large shifts in their selection probabilities. This instability undermines a key goal of lotteries: reducing the influence of fine-grained score distinctions near the decision boundary. We propose smoothness as a design principle for partial lotteries, formalizing it as a Lipschitz condition on the mapping from review scores over candidates to selection probabilities. We introduce the Clipped Linear Lottery, a simple mechanism in which selection probabilities scale linearly with estimated quality between an upper threshold, above which we always accept, and a lower threshold, below which we always reject. We prove that the Clipped Linear Lottery's worst-case regret matches a lower bound for any smooth selection rule up to a factor of $(1 - k/n)$, where $k/n$ is the acceptance rate. We compare smooth selection to other stability notions like Individual Fairness and Differential Privacy, showing that the Clipped Linear Lottery achieves a better smoothness-regret tradeoff than alternatives. Experiments on real peer review data from ICLR 2025, NeurIPS 2024, and the Swiss National Science Foundation demonstrate that existing lottery designs are highly unstable in practice even under perturbations to a single score. Our experiments also confirm the tightness of our theoretical analysis and show that our proposed Clipped Linear Lottery achieves a better smoothness-utility tradeoff than alternatives in practice.","author":[{"family":"Goldberg","given":"Alexander"},{"family":"Fanti","given":"Giulia"},{"family":"Shah","given":"Nihar"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.20069","URL":"https://doi.org/10.48550/arxiv.2605.20069","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.15588","type":"manuscript","title":"Memory-Efficient Differentially Private Training with Gradient Random Projection","abstract":"Differential privacy (DP) protects sensitive data during neural network training, but standard methods like DP-Adam suffer from high memory overhead due to per-sample gradient clipping, limiting scalability. We introduce DP-GRAPE (Gradient RAndom ProjEction), a DP training method that significantly reduces memory usage while maintaining utility on par with first-order DP approaches. DP-GRAPE is motivated by our finding that privatization flattens the gradient singular value spectrum, making SVD-based projections (as in GaLore (Zhao et al., 2024)) unnecessary. Consequently, DP-GRAPE employs three key components: (1) random Gaussian matrices replace SVD-based subspaces, (2) gradients are privatized after projection, and (3) projection is applied during backpropagation. These contributions eliminate the need for costly SVD computations, enable substantial memory savings, and lead to improved utility. Despite operating in lower-dimensional subspaces, our theoretical analysis shows that DP-GRAPE achieves a privacy-utility tradeoff comparable to DP-SGD. Our extensive empirical experiments show that DP-GRAPE can significantly reduce the memory footprint of DP training without sacrificing accuracy or training time. In particular, DP-GRAPE reduces memory usage by over 63% when pre-training Vision Transformers and over 70% when fine-tuning RoBERTa-Large as compared to DP-Adam, while achieving similar performance. We further demonstrate that DP-GRAPE scales to fine-tuning large models such as OPT with up to 6.7 billion parameters, a scale at which DP-Adam fails due to memory constraints. Our code is available at https://github.com/alexmul1114/DP_GRAPE.","author":[{"family":"Mulrooney","given":"Alex"},{"family":"Gupta","given":"Devansh"},{"family":"Flemings","given":"James"},{"family":"Zhang","given":"Huanyu"},{"family":"Annavaram","given":"Murali"},{"family":"Razaviyayn","given":"Meisam"},{"family":"Zhang","given":"Xinwei"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.15588","URL":"https://doi.org/10.48550/arxiv.2506.15588","source":"datacite"},{"id":"doi:10.48550/arxiv.2504.00919","type":"manuscript","title":"Nonparametric spectral density estimation using interactive mechanisms under local differential privacy","abstract":"We study the problem of estimating the spectral density of a centered stationary Gaussian time series under local differential privacy constraints. Specifically, we propose new interactive privacy mechanisms for three tasks: recovering a single covariance coefficient, recovering the spectral density at a fixed frequency, and global recovery. Our approach achieves faster rates through a two-stage process: we first apply the Laplace mechanism to the truncated value, and then use the resulting privatized sample to learn about the dependence mechanism in the time series. For spectral densities belonging to Hölder and Sobolev smoothness classes, we demonstrate that our algorithms improve upon the non-interactive mechanism of Kroll (2024) for small privacy parameter $α$, since the pointwise rates depend on $nα^2$ instead of $nα^4$. Moreover, we show that the rate $(nα^4)^{-1}$ is optimal for estimating a covariance coefficient with non-interactive mechanisms. However, the $L_2$ rate of our interactive estimator is slower than the pointwise rate. We show how to use these procedures to provide a bona fide locally differentially private estimator of the entire covariance matrix. A simulation study validates our findings.","author":[{"family":"Butucea","given":"Cristina"},{"family":"Klockmann","given":"Karolina"},{"family":"Krivobokova","given":"Tatyana"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2504.00919","URL":"https://doi.org/10.48550/arxiv.2504.00919","source":"datacite"},{"id":"doi:10.5061/dryad.zcrjdfnqs","type":"article-journal","title":"Refining impact assessment in undergraduate STEM education: Differential item functioning analysis of field-based learning interventions","abstract":"This dataset contains self-efficacy survey responses from undergraduate biology students enrolled in three different course formats (lecture, introductory field, and intensive field) at the University of California, Santa Cruz from 2016-2019. The data structure includes pre/post survey responses measuring students' self-efficacy across four skill areas (species identification, experimental design, oral presentation, and field research), demographic information (URM status, first-generation status, gender, and Educational Opportunity Program status), and calculated change scores for 564 students. The dataset demonstrates how Differential Item Functioning (DIF) analysis can quantify both the magnitude and demographic patterns of educational interventions with greater precision than traditional assessment methods. Analysis revealed field course students were significantly more likely to report higher self-efficacy ratings compared to lecture course students (odds ratios ranging from 2-167 times higher), with historically minoritized students showing greater gains in field settings. This dataset has significant reuse potential for researchers studying educational interventions, assessment methodology, field-based learning, and equity in STEM education. All data were collected under IRB approval (UCSC #HS3230) with student identifiers anonymized to ensure ethical compliance and privacy protection.RetryClaude can make mistakes. Please double-check responses.","author":[{"family":"Bhatti","given":"Haider"},{"family":"Arcila Hernández","given":"Lina"},{"family":"Balachandran","given":"Lalitha"},{"family":"Kouba","given":"Paige"},{"family":"Croll","given":"Donald"},{"family":"Dayton","given":"Gage"},{"family":"Marnocha","given":"Erin"},{"family":"Beltran","given":"Roxanne"},{"family":"Zavaleta","given":"Erika"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5061/dryad.zcrjdfnqs","URL":"https://doi.org/10.5061/dryad.zcrjdfnqs","source":"datacite"},{"id":"oa:W4410959142","type":"article-journal","title":"Phase-Adaptive Federated Learning for Privacy-Preserving Personalized Travel Itinerary Generation","abstract":"We propose Phase-Adaptive Federated Learning (PAFL), a novel framework for privacy-preserving personalized travel itinerary generation that dynamically balances privacy and utility through a phase-dependent aggregation mechanism inspired by phase-change materials. (1) PAFL’s primary objective is to dynamically optimize the privacy–utility trade-off in federated travel recommendation systems through phase-adaptive anonymization. The phase parameter φ ∈ [0, 1] operates as a tunable control variable that continuously adjusts the latent space geometry between differentially private (φ→1) and utility-optimized (φ→0) representations via a thermodynamic-inspired transformation. Conventional federated learning approaches often rely on static privacy-preserving techniques, which either degrade recommendation quality or inadequately protect sensitive user data; PAFL addresses this limitation through three key innovations: a latent-space phase transformer, a differential privacy-gradient inverter with mathematically provable reconstruction bounds (εt ≤ 1.0), and a lightweight sequential transformer. (2) PAFL’s core innovation lies in its phase-adaptive mechanism that dynamically balances privacy preservation through differential privacy and utility maintenance via gradient inversion, governed by the tunable phase parameter φ. Experimental results demonstrate statistically significant improvements, with 18.7% higher HR@10 (p < 0.01) and 62% lower membership inference risk compared to state-of-the-art methods, while maintaining εtotal < 2.3 over 100 training rounds. The framework advances federated learning for sensitive recommendation tasks by establishing a new paradigm for adaptive privacy–utility optimization.","author":[{"family":"Chen","given":"Xiaolong"},{"family":"Zhang","given":"Hongfeng"},{"family":"Wong","given":"Cora"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/tourhosp6020100","URL":"https://doi.org/10.3390/tourhosp6020100","source":"openalex"},{"id":"oa:W4410984179","type":"article-journal","title":"ADVANCEMENTS IN FEDERATED LEARNING FOR SECURE DATA SHARING IN FINANCIAL SERVICES","abstract":"This paper explores the application of Federated Learning (FL) in the financial sector, focusing on enhancing security and privacy in key areas such as fraud detection, Anti-Money Laundering (AML) compliance, and biometric authentication systems. FL enables collaborative model training across multiple financial institutions without sharing sensitive transaction data, thereby preserving privacy while improving the accuracy of fraud detection models. In AML compliance, FL facilitates the development of robust models by leveraging diverse datasets, enhancing the ability to detect suspicious activities. Moreover, FL strengthens biometric authentication systems by decentralizing model training, reducing the risks of data breaches, and ensuring compliance with privacy regulations. The paper also evaluates the performance of a loan default prediction model trained using FL, highlighting challenges with class imbalance and model bias toward the majority class. The classification report indicates high recall (98%) but also shows a potential for misclassifying non-default cases, leading to a moderate precision (81%) and an F1-score of 89%. The model's AUC of 0.69 suggests moderate discriminatory power, with room for improvement in its ability to differentiate between default and non-default cases. The model achieves an overall accuracy of 80%. Despite these challenges, it demonstrates good generalization capabilities while maintaining the privacy of client data, presenting a promising approach to secure financial transaction analysis.","author":[{"family":"Unuigbokhai","given":"Nkem"},{"family":"Oise","given":"Godfrey"},{"family":"Akilo","given":"Babalola"},{"family":"Nwabuokei","given":"Onyemaechi"},{"family":"Odimayomi","given":"Joy"},{"family":"Bakare","given":"Sofiat"},{"family":"Atake","given":"Onoriode"}],"issued":{"date-parts":[[2025]]},"DOI":"10.33003/fjs-2025-0905-3207","URL":"https://doi.org/10.33003/fjs-2025-0905-3207","source":"openalex"},{"id":"oa:W4411874056","type":"article-journal","title":"Enhancing agricultural commodity price forecasting with deep learning","abstract":"Accurate forecasting of agricultural commodity prices is essential for market planning and policy formulation, especially in agriculture-dependent economies like India. Price volatility, driven by factors such as weather variability and market demand fluctuations, poses significant forecasting challenges. This study evaluates the performance of traditional stochastic models, machine learning techniques, and deep learning approaches in forecasting the prices of 23 commodities using daily wholesale price data from January 2010 to June 2024. Models assessed include Autoregressive Integrated Moving Average, Support Vector Regression, Extreme Gradient Boosting, Multilayer Perceptron, Recurrent Neural Networks, Long Short-Term Memory Networks, Gated Recurrent Units, and Echo State Networks. Results show that deep learning models, particularly Long Short-Term Memory and Gated Recurrent Units, outperform others in capturing complex temporal patterns, achieving superior accuracy across error metrics. The results indicate that deep learning models, particularly Long Short-Term Memory (LSTM) networks and Gated Recurrent Units (GRU), demonstrate superior performance in capturing complex temporal patterns. For instance, the GRU model achieved a Root Mean Squared Error (RMSE) of 369.54 for onions and 210.35 for tomatoes, significantly outperforming the ARIMA model, which recorded RMSE values of 1564.62 and 1298.60, respectively. Furthermore, the Mean Absolute Percentage Error (MAPE) for GRU was notably lower, at 14.59% for onions and 10.58% for tomatoes. These results underscore the efficacy of deep learning approaches in addressing the inherent volatility and nonlinear dynamics of agricultural commodity prices. These findings offer valuable insights for policymakers, traders, and farmers, enabling better market interventions, crop planning, and risk management. The study recommends exploring hybrid models and incorporating external factors like weather data to further enhance forecasting reliability.","author":[{"family":"Manogna","given":"RL"},{"family":"Dharmaji","given":"Vijay"},{"family":"Sarang","given":"S"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-05103-z","URL":"https://doi.org/10.1038/s41598-025-05103-z","source":"openalex"},{"id":"oa:W4409429890","type":"article-journal","title":"FedMEM: Adaptive Personalized Federated Learning Framework for Heterogeneous Mobile Edge Environments","abstract":"With the growth of the Internet of Things (IoT) and communication technologies, edge devices have become more diverse. This diversity has increased the computational load on these systems and led to differences between devices. In mobile edge computing, variations in communication and computing resources can prevent some devices from updating models quickly. This delay affects overall performance. In addition, in federated learning, data that is not independently and identically distributed (non-IID) makes it hard for clients to maintain personalized models.To address these issues, this paper introduces a personalized federated learning framework. This framework enhances the resource allocation optimization algorithm by dynamically adjusting the depth of model inference and the bandwidth allocation strategy, which assists devices with limited computational capabilities in completing inference tasks promptly. Furthermore, it divide the client models into global and personalized layers. Only the global layers are combined, which helps manage the diversity in data distributions. Simulation results show that the proposed FedMEM method is superior to other state-of-the-art methods, and can drastically reduce system latency.","author":[{"family":"Ximing","given":"Chen"},{"family":"Xilong","given":"He"},{"family":"Du","given":"Cheng"},{"family":"Tie-Jun","given":"Wu"},{"family":"Qingyu","given":"Tian"},{"family":"Chen","given":"Rongrong"},{"family":"Qiu","given":"Jing"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s44196-025-00814-7","URL":"https://doi.org/10.1007/s44196-025-00814-7","source":"openalex"},{"id":"oa:W4408750008","type":"article-journal","title":"Scale-MIA: A Scalable Model Inversion Attack against Secure Federated Learning via Latent Space Reconstruction","abstract":"Federated learning is known for its capability to safeguard the participants' data privacy.However, recently emerged model inversion attacks (MIAs) have shown that a malicious parameter server can reconstruct individual users' local data samples from model updates.The state-of-the-art attacks either rely on computation-intensive iterative optimization methods to reconstruct each input batch, making scaling difficult, or involve the malicious parameter server adding extra modules before the global model architecture, rendering the attacks too conspicuous and easily detectable.To overcome these limitations, we propose Scale-MIA, a novel MIA capable of efficiently and accurately reconstructing local training samples from the aggregated model updates, even when the system is protected by a robust secure aggregation (SA) protocol.Scale-MIA utilizes the inner architecture of models and identifies the latent space as the critical layer for breaching privacy.Scale-MIA decomposes the complex reconstruction task into an innovative two-step process.The first step is to reconstruct the latent space representations (LSRs) from the aggregated model updates using a closed-form inversion mechanism, leveraging specially crafted linear layers.Then in the second step, the LSRs are fed into a fine-tuned generative decoder to reconstruct the whole input batch.We implemented Scale-MIA on commonly used machine learning models and conducted comprehensive experiments across various settings.The results demonstrate that Scale-MIA achieves excellent performance on different datasets, exhibiting high reconstruction rates, accuracy, and attack efficiency on a larger scale compared to state-of-the-art MIAs.Our code is available at https://github.com/unknown123489/Scale-MIA.Break Secure Attacker's Attack Attack Need Auxiliary Model Attack Aggregation?Capability Overhead Scale Dataset?Agnostic?DLG [3], iDLG [4] No Weak (Curious) Large Single image No Yes Inverting Grad [5] No Weak (Curious) Large 8 No Yes GradInversion [7] No Weak (Curious) Large 48 No No (ResNet) GradViT [9] No Weak (Curious) Large 8 No No (ViT) APRIL-Optim [8] No Weak (Curious) Large Single-image No No (ViT) APRIL-Analytic [8] No Weak (Curious) Small Single-image No No (ViT) R-GAP [6] No Weak (Curious) Small Single-image No Yes Leak in FA [25] No Weak (Curious) Small 50 No Yes Fishing for data [17] Yes Medium (Modify params) Large 256 Yes Yes Eluding SecureAgg [16] Yes Medium (Modify params) Large 512 Yes Yes Robbing the fed [18] Yes Strong (Change architect) Small 1024+ Yes Yes LOKI [19] Yes Strong (Change architect) Small 1024+ Yes Yes","author":[{"family":"Shi","given":"Shanghao"},{"family":"Wang","given":"Ning"},{"family":"Yang","given":"Xiao"},{"family":"Zhang","given":"Chaoyu"},{"family":"Shi","given":"Yi"},{"family":"Hou","given":"YT"},{"family":"Lou","given":"Wenjing"}],"issued":{"date-parts":[[2025]]},"DOI":"10.14722/ndss.2025.240644","URL":"https://doi.org/10.14722/ndss.2025.240644","source":"openalex"},{"id":"oa:W4415151187","type":"article-journal","title":"Quantum-Driven Reinforcement Learning for Spectral Energy Optimization in Massive MIMO Hybrid Beamforming for 6G","abstract":"Abstract The evolution of 6G wireless networks demands highly efficient beamforming strategies to optimize spectral and energy efficiency in massive MIMO systems. This study introduces a Quantum-Driven Reinforcement Learning (QDRL) framework for Spectral Energy Optimization in Massive MIMO Hybrid Beamforming for 6G, leveraging Quantum Deep Q-Networks (Q-DQN), Quantum Policy Gradient (QPG), and Quantum Approximate Optimization Algorithm (QAOA). The framework integrates mruby-based lightweight scripting for efficient deployment in edge-AI environments, enhancing computational flexibility and resource efficiency. Performance evaluations demonstrate that the Hybrid Quantum Model achieves 11.21 bps/Hz spectral efficiency, 97% resource utilization efficiency, and reduces energy consumption to 0.50 Joules/bit, outperforming classical models. The Bit Error Rate (BER) is minimized to 0.0025, and the convergence time is 48.7 s, significantly improving computational efficiency. Comparative analysis with conventional Deep Reinforcement Learning (DRL) techniques shows that the proposed quantum-enhanced model provides a 32% improvement in energy efficiency and a 21% reduction in computational complexity. The integration of mruby enhances the adaptability of the system in low-power and embedded environments, making it a viable solution for real-time 6G hybrid beamforming. This research highlights the transformative potential of quantum-assisted AI frameworks for scalable, high-speed, and energy-efficient wireless communication.","author":[{"family":"Krishnamoorthy","given":"R"},{"family":"Begum","given":"MA"},{"family":"Maguluri","given":"Lakshmana"},{"family":"Abdelhaq","given":"Maha"},{"family":"Alsaqour","given":"Raed"},{"family":"Selvarajan","given":"Shitharth"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s11277-025-11855-8","URL":"https://doi.org/10.1007/s11277-025-11855-8","source":"openalex"},{"id":"oa:W4412429308","type":"article-journal","title":"Emerging role of generative AI in renewable energy forecasting and system optimization","abstract":"• Generative AI improves RES forecasting accuracy by up to 25 %. • GANs and VAEs optimize microgrid and storage operations. • Federated learning enables privacy-preserving energy model training. • Black-box models pose interpretability and regulatory challenges. • AI–IoT convergence supports real-time energy system optimization. The rapid integration of renewable energy sources (RES) into modern power systems introduces significant challenges in forecasting accuracy, grid stability, and energy optimization. Generative Artificial Intelligence (Gen-AI), including architectures such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and transformers, offers new capabilities to overcome data sparsity, nonlinearity, and uncertainty in renewable-dominant systems. This study aims to comprehensively review the emerging role of Gen-AI in improving solar and wind forecasting, load prediction, energy storage management, and smart grid optimization. Using a comparative and synthesis-based methodology, this review analyses findings from high-impact publications between 2023 and 2025. Results indicate that GAN-based models reduce root mean square error (RMSE) by 15–20 % in solar irradiance forecasting and significantly enhance spatial-temporal wind simulations. Time-series GAN-LSTM hybrids enhance demand forecasting accuracy under nonlinear conditions, while VAE-driven dispatch models achieve gains of 9–12 % in energy efficiency and curtailment reduction. The novelty of this review lies in mapping Gen-AI's integration with digital twins, federated learning, and AI–IoT frameworks, which enables the real-time, privacy-preserving optimisation of complex energy systems. The principal conclusion is that Gen-AI serves as a transformative tool to enhance system resilience, forecasting precision, and operational flexibility in renewable energy networks. For sustainable implementation, future developments must address challenges in model explainability, data privacy, and scalability. These findings support the journal’s scope by highlighting AI-driven advancements for the reliable, efficient, and sustainable transformation of energy systems.","author":[{"family":"Erdiwansyah","given":"Erdiwansyah"},{"family":"Mamat","given":"Rizalman"},{"family":"Syafrizal","given":"Syafrizal"},{"family":"Ghazali","given":"Mohd"},{"family":"Basrawi","given":"Firdaus"},{"family":"Rosdi","given":"SM"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1016/j.scca.2025.100099","URL":"https://doi.org/10.1016/j.scca.2025.100099","source":"openalex"},{"id":"oa:W4414538246","type":"article-journal","title":"UAV-Assisted Federated Learning With Robust Resource and Trajectory Optimization Under Location Uncertainties","abstract":"Federated learning (FL) has emerged as a promising solution to facilitate the deployment of artificial intelligence (AI) on wireless devices. However, heterogeneity of wireless devices, including disparities in computation capabilities, data sizes, and energy constraints, introduces delays in the FL completion time, particularly due to inefficient communication and slow updates from resource-constrained devices. To address this issue, we propose an unmanned aerial vehicle (UAV)-assisted FL framework that integrates UAV as a central server, collaborating with the devices to facilitate the model training process. Accordingly, we jointly consider computation and transmission strategies, as well as the task assignment and UAV trajectory to minimize the completion time of the FL process. Particularly, we consider the location uncertainties associated with the devices, along with the consequent chance-constrained aggregation process, to achieve a robust learning process. We employ the Bernstein-type inequalities to reformulate the probabilistic-form optimization into its deterministic counterpart. Then we solve the problem under a block coordinate descent framework. Simulation results demonstrate that the proposed approach significantly reduces the completion time of FL and achieves robust performance guarantee in the presence of location deviations.","author":[{"family":"Wang","given":"Chen"},{"family":"Tang","given":"Xiao"},{"family":"Xiong","given":"Zehui"},{"family":"Zhai","given":"Daosen"},{"family":"Zhang","given":"Ruonan"},{"family":"Niyato","given":"Dusit"},{"family":"Han","given":"Zhu"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/tccn.2025.3614635","URL":"https://doi.org/10.1109/tccn.2025.3614635","source":"openalex"},{"id":"oa:W4416409999","type":"article-journal","title":"Secure blockchain integrated deep learning framework for federated risk-adaptive and privacy-preserving IoT edge intelligence sets","abstract":"An enormous demand for a secure, scalable, intelligent edge computing framework has emerged for the exponentially increasing number of Internet of Things (IoT) devices for any substrate of modern digital infrastructure. These edge nodes distributed across heterogeneous environments serve as primary interfaces for sensing, computation, and actuations. Their physical deployment in unattended scenarios puts them at risk of being targets for resource manipulation. One widely accepted IoT architecture with traditional notions of edge may consider a threat to its centralized knowledge with an unbounded attack surface that includes anything that can remotely connect to the edge from the cloud-like domain. Existing strategies either forget the dynamic risk context of edge nodes or do not achieve a reasonable trade-off between security and resource constraints, essentially degrading the robustness and trustworthiness of solutions intended for real-life scenarios. To address the existing gaps, the work presents a novel Blockchain Integrated Deep Learning Framework for secure IoT edge computing, introducing a hybrid architecture where the transparency of blockchain meets deep learning flexibility. The proposed system incorporates five specialized components: Blockchain-Orchestrated Federated Curriculum Learning (BOFCL), which ensures risk-prioritized training using threat indices derived from blockchain logs; this adaptive sequencing enhances responsiveness to high-risk edge scenarios. Zero-Knowledge Proof Enabled Secure Inference Engine (ZK-SIE) provides verifiable privacy-preserving inference, ensuring model integrity without exposing input data or model internals in process. Blockchain Indexed Adversarial Attack Simulator (BI-AAS) focuses on testing the models in edge environments against attack scenarios drawn from common adversarial profiles and thereby facilitates a model defensive retraining. Energy-Aware Lightweight Consensus with Adaptive Synchronization (ELCAS) avoids overhead by seeking energy-efficient participants for global model synchronization in constrained environments. Trust Indexed Model Provenance and Deployment Ledger (TIMPDL) ensures model lineage tracking and deploy ability in a transparent manner by providing composite trust scores computed from data quality, node reputation, and validation metrics. Altogether, the framework combines the data integrity, adversarial robustness, and trust-aware deployment, shortening training latency, synchronization energy, and privacy leakage. It is a foundational advancement supporting secure decentralized edge intelligence for next-generation IoT ecosystems.","author":[{"family":"Swathi","given":"K"},{"family":"Durga","given":"Putta"},{"family":"Prasad","given":"KV"},{"family":"Chaitanya","given":"AK"},{"family":"Santhi","given":"Kuraganti"},{"family":"Vidyullatha","given":"P"},{"family":"Rao","given":"Sannidhi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-24895-8","URL":"https://doi.org/10.1038/s41598-025-24895-8","source":"openalex"},{"id":"oa:W7124702472","type":"article-journal","title":"Blockchain-based personalized federated learning framework for drug recommendation systems resilient to model poisoning","abstract":"Abstract Federated learning enables multiple healthcare entities to collaboratively train a global model while ensuring patient data privacy through local model training without sharing raw data. However, FL remains vulnerable to adversarial attacks such as model poisoning, data injection, and model inversion that compromise model integrity. To address these challenges, this paper presents a blockchain-based personalized federated learning (FL) framework designed to enhance the security, privacy, and efficiency of decentralized model training in healthcare environments. It integrates Practical Byzantine Fault Tolerance (PBFT) for tamper-resistant aggregation, L2-norm anomaly filtering for lightweight adversarial defense, and a Neural Architecture Search (NAS)-optimized hybrid Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) model with attention to enable efficient, personalized modeling of non-IID healthcare data. Together, these components address key FL challenges, including robustness to model poisoning, accuracy, and deployment on resource-constrained devices. To evaluate its effectiveness, we apply the proposed framework to drug recommendation tasks using three real-world medical datasets, namely Symptom2Disease, UCL Drug, and Dermo Questions, achieving F1-scores of 0.97, 0.70, and 0.83, respectively. The framework demonstrates competitive performance compared to conventional and state-of-the-art methods while significantly reducing the number of trainable parameters, highlighting its suitability for real-time, on-device healthcare applications. These results validate the framework’s ability to deliver secure, personalized, and privacy-preserving recommendations in intelligent healthcare systems.","author":[{"family":"Apak","given":"Sina"},{"family":"Değim","given":"İsmail"},{"family":"Zahertar","given":"Samaneh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1007/s00521-025-11828-9","URL":"https://doi.org/10.1007/s00521-025-11828-9","source":"openalex"},{"id":"oa:W4413441047","type":"article-journal","title":"Privacy-Preserving Diabetes and Heart Disease Prediction via Federated Learning and WCO","abstract":"Diabetes, afflicting 537 million worldwide, is a prevalent and lethal non-communicable ailment. Its onset, influenced by factors like obesity and family history, manifests symptoms such as frequent urination. Long-term complications encompass heart, kidney, and nerve ailments. Early prediction mitigates risks. All these encompassing strategies are designed to improve prediction precision and facilitate proactive diabetes control. This research employed SMOTE methods to tackle imbalanced classes, utilizing various classification algorithms such as Random Forest, XGBoost, Multilayer Perceptron, Gradient Boost, and AdaBoost. Following extensive training and evaluation, the AdaBoost classifier delivered superior outcomes, achieving a 94.02% accuracy rate, an F1 score of 93.32%, and an AUC of 0.95. In the healthcare industry, accurately forecasting diabetes mellitus is crucial; however, privacy laws hinder the transfer of medical information from the Internet of Medical Things (IoMT), causing delays in diagnosis. This study introduces the Federated Learning with Weighted Conglomeration Optimization (FLWCO) model as a solution to these challenges. In Centralized Learning, AdaBoost with WCO achieves an accuracy of 95.32% when tested on a Kaggle dataset consisting of 96,146 instances. During the second stage, FLWCO achieves a superior 97.27% accuracy rate compared to other federated learning techniques. The method not only guarantees privacy conformity but also decreases communication expenses. FLWCO demonstrates superiority over existing federated learning algorithms in real-world heart illness prediction. Furthermore, the proposed model can be employed to estimate the likelihood of heart disease in individuals with diabetes. This highlights the potential of federated learning, especially FLWCO, in leveraging distributed data while preserving privacy, facilitating accurate diabetes mellitus diagnosis, and addressing challenges in sharing medical information securely and efficiently.","author":[{"family":"Dash","given":"Sachikanta"},{"family":"Padhy","given":"Sasmita"},{"family":"Suman","given":"Preetam"},{"family":"Mal","given":"Sandip"},{"family":"Malviya","given":"Lokesh"},{"family":"Suman","given":"Amrit"},{"family":"Kishore","given":"Jaydeep"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s44196-025-00956-8","URL":"https://doi.org/10.1007/s44196-025-00956-8","source":"openalex"},{"id":"oa:W4416650196","type":"article-journal","title":"Certifying the Right to Be Forgotten: Primal–Dual Optimization for Sample and Label Unlearning in Vertical Federated Learning","abstract":"Federated unlearning has become an attractive approach to address privacy concerns in collaborative machine learning, for situations when sensitive data are remembered by AI models during the machine learning process. It enables the removal of specific data influences from trained models, aligning with the growing emphasis on the “right to be forgotten.” While extensively studied in horizontal federated learning, unlearning in vertical federated learning (VFL) remains challenging due to the distributed feature architecture. VFL unlearning includes sample unlearning that removes specific data points’ influence and label unlearning that removes entire classes. Since different parties hold complementary features of the same samples, unlearning tasks require cross-party coordination, creating computational overhead and feature interdependencies. To address such challenges, we propose FedORA (Federated Optimization for data Removal via primal-dual Algorithm), designed for sample and label unlearning in VFL. FedORA formulates the removal of certain samples or labels as a constrained optimization problem solved using a primal-dual framework. Our approach introduces a new unlearning loss function that promotes classification uncertainty rather than misclassification. An adaptive step size enhances convergence, while an asymmetric batch design handles unlearning and retained data efficiently to reduce computational costs, considering the prior influence of the remaining data on the model. We provide theoretical analysis proving that the model difference between FedORA and Train-from-scratch is bounded, establishing guarantees for unlearning effectiveness. Experiments on tabular and image datasets demonstrate that FedORA achieves unlearning effectiveness and utility preservation comparable to Train-from-scratch with reduced computation and communication overhead.","author":[{"family":"Jiang","given":"Yu"},{"family":"Tong","given":"Xindi"},{"family":"Liu","given":"Ziyao"},{"family":"Zhang","given":"Xiaoxi"},{"family":"Lam","given":"Kwok‐yan"},{"family":"Tan","given":"Chee"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/tifs.2025.3636788","URL":"https://doi.org/10.1109/tifs.2025.3636788","source":"openalex"},{"id":"oa:W4415878254","type":"article-journal","title":"DFCA: Decentralized Federated Clustering Algorithm","abstract":"Clustered Federated Learning has emerged as an effective approach for handling heterogeneous data across clients by partitioning them into clusters with similar or identical data distributions. However, most existing methods, including the Iterative Federated Clustering Algorithm (IFCA), rely on a central server to coordinate model updates, typically requiring stable connectivity, synchronous communication rounds, and global aggregation of client models. These assumptions are difficult to satisfy in decentralized and heterogeneous environments, where clients may only have limited, local communication with a small subset of peers. As a result, such methods create a bottleneck and a single point of failure, limiting their applicability in realistic decentralized learning settings. This limitation is particularly severe in Internet of Things settings, where large numbers of resource-constrained devices, intermittent or sparse connectivity, and dynamic participation make reliance on a central server impractical. In this work, we introduce the Decentralized Federated Clustering Algorithm (DFCA), a fully decentralized clustered federated learning algorithm that enables clients to collaboratively train cluster-specific models without central coordination. DFCA uses a sequential running average to aggregate models from neighbors as updates arrive, providing a communication-efficient alternative to batch aggregation while maintaining clustering performance. Our experiments on various datasets demonstrate that DFCA outperforms other decentralized algorithms and performs comparably to centralized IFCA, even under sparse connectivity, highlighting its robustness and practicality for dynamic real-world decentralized networks.","author":[{"family":"Kirch","given":"Jonas"},{"family":"Becker","given":"S"},{"family":"Rodrigues","given":"Tiago"},{"family":"Harmeling","given":"Stefan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/jiot.2026.3669440","URL":"https://doi.org/10.1109/jiot.2026.3669440","source":"openalex"},{"id":"oa:W4412840942","type":"article-journal","title":"Novel Federated Graph Contrastive Learning for IoMT Security: Protecting Data Poisoning and Inference Attacks","abstract":"Malware evolution presents growing security threats for resource-constrained Internet of Medical Things (IoMT) devices. Conventional federated learning (FL) often suffers from slow convergence, high communication overhead, and fairness issues in dynamic IoMT environments. In this paper, we propose FedGCL, a secure and efficient FL framework integrating contrastive graph representation learning for enhanced feature discrimination, a Jain-index-based fairness-aware aggregation mechanism, an adaptive synchronization scheduler to optimize communication rounds, and secure aggregation via homomorphic encryption within a Trusted Execution Environment. We evaluate FedGCL on four benchmark malware datasets (Drebin, Malgenome, Kronodroid, and TUANDROMD) using 5 to 15 graph neural network clients over 20 communication rounds. Our experiments demonstrate that FedGCL achieves 96.3% global accuracy within three rounds and converges to 98.9% by round twenty—reducing required training rounds by 45% compared to FedAvg—while incurring only approximately 10% additional computational overhead. By preserving patient data privacy at the edge, FedGCL enhances system resilience without sacrificing model performance. These results indicate FedGCL’s promise as a secure, efficient, and fair federated malware detection solution for IoMT ecosystems.","author":[{"family":"Daulay","given":"Amarudin"},{"family":"Ramli","given":"Kalamullah"},{"family":"Harwahyu","given":"Ruki"},{"family":"Hidayat","given":"Taufik"},{"family":"Pranggono","given":"Bernardi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/math13152471","URL":"https://doi.org/10.3390/math13152471","source":"openalex"},{"id":"oa:W4413140400","type":"article-journal","title":"FedNolowe: A normalized loss-based weighted aggregation strategy for robust federated learning in heterogeneous environments","abstract":"Federated Learning supports collaborative model training across distributed clients while keeping sensitive data decentralized. Still, non-independent and identically distributed data pose challenges like unstable convergence and client drift. We propose Federated Normalized Loss-based Weighted Aggregation (FedNolowe) (Code is available at https://github.com/dongld-2020/fednolowe), a new method that weights client contributions using normalized training losses, favoring those with lower losses to improve global model stability. Unlike prior methods tied to dataset sizes or resource-heavy techniques, FedNolowe employs a two-stage L1 normalization, reducing computational complexity by 40% in floating-point operations while matching state-of-the-art performance. A detailed sensitivity analysis shows our two-stage weighting maintains stability in heterogeneous settings by mitigating extreme loss impacts while remaining effective in independent and identically distributed scenarios.","author":[{"family":"Le","given":"Duy"},{"family":"Tuong","given":"Nguyen"},{"family":"Tran","given":"Anh"},{"family":"Dao","given":"Minh"},{"family":"Bao","given":"Pham"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1371/journal.pone.0322766","URL":"https://doi.org/10.1371/journal.pone.0322766","source":"openalex"},{"id":"oa:W4417224611","type":"article-journal","title":"Blockchain-Enabled Hierarchical Federated Learning Framework for Anomaly Detection in IoT Systems","abstract":"The rapid expansion of the Internet of Things (IoT) across domains such as industrial automation, smart healthcare, and intelligent transportation has intensified security challenges, particularly in terms of detecting anomalies across large-scale, heterogeneous networks. To address these challenges, this study introduces a blockchain-enabled hierarchical federated learning (Block-HFL) approach that combines federated model aggregation with blockchain-based authentication and immutable storage. This approach has enhanced scalability, reduced communication latency, and ensured trustworthy model management while preserving data privacy. In comparison with existing hierarchical and non-hierarchical FL approaches, the proposed Block-HFL framework introduces an accuracy-based leader election mechanism that enhances fairness and improves global model convergence. Experimental evaluations on the Edge-IIoTset dataset show that Block-HFL consistently maintains detection accuracy above 94% as the number of clients increases from 4 to 16, outperforming baseline FL models under similar non-IID conditions. Moreover, blockchain integration ensures secure, transparent, and tamper-proof global model management with minimal computational cost, confirming that the proposed framework provides an efficient and trustworthy solution for distributed anomaly detection in IoT systems.","author":[{"family":"Alharthi","given":"Haya"},{"family":"Alshehri","given":"Suhair"},{"family":"Kalkatawi","given":"Manal"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/app152413037","URL":"https://doi.org/10.3390/app152413037","source":"openalex"},{"id":"oa:W4412078570","type":"article-journal","title":"Towards fair decentralized benchmarking of healthcare AI algorithms with the Federated Tumor Segmentation (FeTS) challenge","abstract":"Computational competitions are the standard for benchmarking medical image analysis algorithms, but they typically use small curated test datasets acquired at a few centers, leaving a gap to the reality of diverse multicentric patient data. To this end, the Federated Tumor Segmentation (FeTS) Challenge represents the paradigm for real-world algorithmic performance evaluation. The FeTS challenge is a competition to benchmark (i) federated learning aggregation algorithms and (ii) state-of-the-art segmentation algorithms, across multiple international sites. Weight aggregation and client selection techniques were compared using a multicentric brain tumor dataset in realistic federated learning simulations, yielding benefits for adaptive weight aggregation, and efficiency gains through client sampling. Quantitative performance evaluation of state-of-the-art segmentation algorithms on data distributed internationally across 32 institutions yielded good generalization on average, albeit the worst-case performance revealed data-specific modes of failure. Similar multi-site setups can help validate the real-world utility of healthcare AI algorithms in the future.","author":[{"family":"Zenk","given":"Maximilian"},{"family":"Baid","given":"Ujjwal"},{"family":"Pati","given":"Sarthak"},{"family":"Linardos","given":"Akis"},{"family":"Edwards","given":"Brandon"},{"family":"Sheller","given":"Micah"},{"family":"Foley","given":"Patrick"},{"family":"Aristizábal","given":"Alejandro"},{"family":"Zimmerer","given":"David"},{"family":"Груздев","given":"АД"},{"family":"Martin","given":"Jason"},{"family":"Shinohara","given":"Russell"},{"family":"Reinke","given":"Annika"},{"family":"Isensee","given":"Fabian"},{"family":"Parampottupadam","given":"Santhosh"},{"family":"Parekh","given":"Kaushal"},{"family":"Floca","given":"Ralf"},{"family":"Kassem","given":"Hasan"},{"family":"Baheti","given":"Bhakti"},{"family":"Thakur","given":"Siddhesh"},{"family":"Chung","given":"Verena"},{"family":"Kushibar","given":"Kaisar"},{"family":"Lekadir","given":"Karim"},{"family":"Jiang","given":"Meirui"},{"family":"Yin","given":"Youtan"},{"family":"Yang","given":"Hongzheng"},{"family":"Liu","given":"Quande"},{"family":"Chen","given":"Cheng"},{"family":"Dou","given":"Qi"},{"family":"Heng","given":"Pheng‐ann"},{"family":"Zhang","given":"Xiaofan"},{"family":"Zhang","given":"Shaoting"},{"family":"Khan","given":"Muhammad"},{"family":"Azeem","given":"Mohammad"},{"family":"Jafaritadi","given":"Mojtaba"},{"family":"Alhoniemi","given":"Esa"},{"family":"Kontio","given":"Elina"},{"family":"Khan","given":"Suleiman"},{"family":"Mächler","given":"Leon"},{"family":"Ezhov","given":"Ivan"},{"family":"Kofler","given":"Florian"},{"family":"Shit","given":"Suprosanna"},{"family":"Paetzold","given":"Johannes"},{"family":"Loehr","given":"Timo"},{"family":"Wiestler","given":"Benedikt"},{"family":"Peiris","given":"Himashi"},{"family":"Pawar","given":"Kamlesh"},{"family":"Zhong","given":"Shenjun"},{"family":"Chen","given":"Zhaolin"},{"family":"Hayat","given":"Munawar"},{"family":"Egan","given":"Gary"},{"family":"Harandi","given":"Mehrtash"},{"family":"Isik-Polat","given":"Ece"},{"family":"Polat","given":"Görkem"},{"family":"Koçyiğit","given":"Altan"},{"family":"Temizel","given":"Alptekin"},{"family":"Tuladhar","given":"Anup"},{"family":"Tyagi","given":"Lakshay"},{"family":"Souza","given":"Raissa"},{"family":"Forkert","given":"Nils"},{"family":"Mouchès","given":"Pauline"},{"family":"Wilms","given":"Matthias"},{"family":"Shambhat","given":"Vishruth"},{"family":"Maurya","given":"Akansh"},{"family":"Danannavar","given":"Shubham"},{"family":"Kalla","given":"Rohit"},{"family":"Anand","given":"Vikas"},{"family":"Krishnamurthi","given":"Ganapathy"},{"family":"Nalawade","given":"Sahil"},{"family":"Ganesh","given":"Chandan"},{"family":"Wagner","given":"Benjamin"},{"family":"Reddy","given":"Divya"},{"family":"Das","given":"Yudhajit"},{"family":"Yu","given":"Fang"},{"family":"Fei","given":"Baowei"},{"family":"Madhuranthakam","given":"Ananth"},{"family":"Maldjian","given":"Joseph"},{"family":"Singh","given":"Gaurav"},{"family":"Ren","given":"Jianxun"},{"family":"Zhang","given":"Wei"},{"family":"An","given":"Ning"},{"family":"Hu","given":"Qingyu"},{"family":"Zhang","given":"Youjia"},{"family":"Zhou","given":"Ying"},{"family":"Siomos","given":"Vasilis"},{"family":"Tarroni","given":"Giacomo"},{"family":"Passeratpalmbach","given":"Jonathan"},{"family":"Rawat","given":"Ambrish"},{"family":"Zizzo","given":"Giulio"},{"family":"Kadhe","given":"Swanand"},{"family":"Epperlein","given":"Jonathan"},{"family":"Braghin","given":"Stefano"},{"family":"Wang","given":"Yuan"},{"family":"Kanagavelu","given":"Renuga"},{"family":"Wei","given":"Qingsong"},{"family":"Yang","given":"Yechao"},{"family":"Liu","given":"Yong"},{"family":"Kotowski","given":"Krzysztof"},{"family":"Adamski","given":"Szymon"},{"family":"Machura","given":"Bartosz"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41467-025-60466-1","URL":"https://doi.org/10.1038/s41467-025-60466-1","source":"openalex"},{"id":"oa:W4412423696","type":"article-journal","title":"Network-based intrusion detection using deep learning technique","abstract":"A high growth rate in network traffic and the complexity of cyber threats have made it necessary to create more effective and flexible intrusion detection systems. Most traditional Network-based Intrusion Detection Systems (NIDS) can become weak at detecting new patterns of attacks due to the use of obsolete data or traditional machine learning models. To overcome the mentioned constraints, the current research presents a new deep learning solution that combines Sequential Deep Neural Networks (DNN) and Rectified Linear Unit (ReLU) activation unit with an Extra Tree Classifier feature selection procedure. The proposed model is trained and tested on the new rich and up-to-date UNSW-NB15 set, which provides a realistic reflection of the real-life network traffic and attack vectors. The interesting novelty of this study is the tactical use of ReLU-based DNN combined with feature optimization through the Extra Tree Classifier, which not only overcomes general problems like vanishing gradients and overfitting but also greatly increases the interpretability of the model and the efficiency of its computation. This dimensional reduction of the feature space (43 to only 8 highly relevant features) retains the high accuracy of the model but with better inference speed, which is a crucial aspect of the real-time deployment of NIDS. The results show that with the Sequential DNN approach, the binary class (0 for normal and 1 for attack records) achieved 97.93% accuracy, 97% Precision, 97% Recall and 97% F1-score. Furthermore, the detailed experimental testing, such as ROC curves and Confusion Matrices, confirmed that the Sequential DNN performed well in comparison to other Existing Studies. These findings underscore the effectiveness of deep learning architectures enhanced with optimized feature selection in detecting network intrusions, making the proposed system a promising solution for securing critical infrastructure in sectors such as finance, healthcare, and government networks.","author":[{"family":"Farhan","given":"Muhammad"},{"family":"Din","given":"Hafiz"},{"family":"Ullah","given":"SMW"},{"family":"Hussain","given":"Muhammad"},{"family":"Khan","given":"Muhammad"},{"family":"Mazhar","given":"Tehseen"},{"family":"Khattak","given":"Umar"},{"family":"Jaghdam","given":"Ines"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-08770-0","URL":"https://doi.org/10.1038/s41598-025-08770-0","source":"openalex"},{"id":"oa:W4407548040","type":"article-journal","title":"Advances in Neuroimaging and Deep Learning for Emotion Detection: A Systematic Review of Cognitive Neuroscience and Algorithmic Innovations","abstract":"Background/Objectives: The following systematic review integrates neuroimaging techniques with deep learning approaches concerning emotion detection. It, therefore, aims to merge cognitive neuroscience insights with advanced algorithmic methods in pursuit of an enhanced understanding and applications of emotion recognition. Methods: The study was conducted following PRISMA guidelines, involving a rigorous selection process that resulted in the inclusion of 64 empirical studies that explore neuroimaging modalities such as fMRI, EEG, and MEG, discussing their capabilities and limitations in emotion recognition. It further evaluates deep learning architectures, including neural networks, CNNs, and GANs, in terms of their roles in classifying emotions from various domains: human-computer interaction, mental health, marketing, and more. Ethical and practical challenges in implementing these systems are also analyzed. Results: The review identifies fMRI as a powerful but resource-intensive modality, while EEG and MEG are more accessible with high temporal resolution but limited by spatial accuracy. Deep learning models, especially CNNs and GANs, have performed well in classifying emotions, though they do not always require large and diverse datasets. Combining neuroimaging data with behavioral and cognitive features improves classification performance. However, ethical challenges, such as data privacy and bias, remain significant concerns. Conclusions: The study has emphasized the efficiencies of neuroimaging and deep learning in emotion detection, while various ethical and technical challenges were also highlighted. Future research should integrate behavioral and cognitive neuroscience advances, establish ethical guidelines, and explore innovative methods to enhance system reliability and applicability.","author":[{"family":"Halkiopoulos","given":"Constantinos"},{"family":"Gkintoni","given":"Evgenia"},{"family":"Aroutzidis","given":"Anthimos"},{"family":"Antonopoulou","given":"Hera"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/diagnostics15040456","URL":"https://doi.org/10.3390/diagnostics15040456","source":"openalex"},{"id":"oa:W4410061972","type":"article-journal","title":"Adversarial machine learning: a review of methods, tools, and critical industry sectors","abstract":"Abstract The rapid advancement of Artificial Intelligence (AI), particularly Machine Learning (ML) and Deep Learning (DL), has produced high-performance models widely used in various applications, ranging from image recognition and chatbots to autonomous driving and smart grid systems. However, security threats arise from the vulnerabilities of ML models to adversarial attacks and data poisoning, posing risks such as system malfunctions and decision errors. Meanwhile, data privacy concerns arise, especially with personal data being used in model training, which can lead to data breaches. This paper surveys the Adversarial Machine Learning (AML) landscape in modern AI systems, while focusing on the dual aspects of robustness and privacy. Initially, we explore adversarial attacks and defenses using comprehensive taxonomies. Subsequently, we investigate robustness benchmarks alongside open-source AML technologies and software tools that ML system stakeholders can use to develop robust AI systems. Lastly, we delve into the landscape of AML in four industry fields –automotive, digital healthcare, electrical power and energy systems (EPES), and Large Language Model (LLM)-based Natural Language Processing (NLP) systems– analyzing attacks, defenses, and evaluation concepts, thereby offering a holistic view of the modern AI-reliant industry and promoting enhanced ML robustness and privacy preservation in the future.","author":[{"family":"Pelekis","given":"Sotiris"},{"family":"Koutroubas","given":"Thanos"},{"family":"Blika","given":"Afroditi"},{"family":"Berdelis","given":"Anastasis"},{"family":"Karakolis","given":"Evangelos"},{"family":"Ntanos","given":"Christos"},{"family":"Spiliotis","given":"Evangelos"},{"family":"Askounis","given":"Dimitris"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s10462-025-11147-4","URL":"https://doi.org/10.1007/s10462-025-11147-4","source":"openalex"},{"id":"oa:W4412480786","type":"article-journal","title":"An Asynchronous Federated Learning Aggregation Method Based on Adaptive Differential Privacy","abstract":"Federated learning is a distributed machine learning technique that allows multiple devices to collaborate on learning a shared model without exchanging data. It can be used to improve model accuracy while protecting user privacy. However, traditional federated learning is vulnerable to attacks from generative adversarial networks (GANs). As a new privacy protection method, differential privacy enhances privacy protection capabilities by sacrificing some data accuracy. To optimize the privacy budget allocation scheme in traditional differential privacy, we propose a differential privacy method called ADP-FL, which dynamically adjusts the privacy budget based on Newton’s Law of Cooling. While maintaining the overall privacy budget, it dynamically tunes adaptive parameters to improve training accuracy. Additionally, we propose an asynchronous federated learning aggregation scheme that combines privacy budget with data freshness, thereby reducing the impact of differential privacy on accuracy. We conducted extensive experiments on differential privacy algorithms based on Gaussian mechanisms and Laplace mechanisms. The experimental results show that, under the same privacy budget, our algorithm achieves higher accuracy and lower communication overhead compared to the baseline algorithm.","author":[{"family":"Wu","given":"Jiawen"},{"family":"Xia","given":"Geming"},{"family":"Huang","given":"Hongwei"},{"family":"Yu","given":"Chaodong"},{"family":"Zhang","given":"Yuze"},{"family":"Li","given":"Hongfeng"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/electronics14142847","URL":"https://doi.org/10.3390/electronics14142847","source":"openalex"},{"id":"oa:W4409310894","type":"article-journal","title":"DRL-Based Joint Aggregation Frequency and Edge Association for Energy-Efficient Hierarchical Federated Learning","abstract":"Hierarchical Federated Learning (HFL) has been proposed to achieve large-scale model training and more efficient communication, surpassing conventional Federated Learning (FL). However, inappropriate aggregation frequency and edge association in HFL result in excessive energy consumption for users with poor channels or hinder its convergence performance due to stochastic gradient descent (SGD) and Non-Independent and Identical Distribution (NIID) data, which is particularly challenging for energy-limited users. Motivated by this, a joint aggregation frequency and edge association optimization problem is proposed to minimize the long-term energy consumption during HFL training process. The problem can be formulated by incorporating computation, communication model and convergence analysis together. Due to the coupling between control variables, we decompose it into two sub-problems and adopt an iterative algorithm to approximate their optimal solutions. Specifically, the aggregation frequency is optimized under a given edge association by convex optimization to trade-off the computation and communication energy consumption, considering the convergence characteristic and SGD noise. Then, Deep Reinforcement Learning (DRL) is adopted to optimize edge association based on data distribution, dynamic channels and the derived aggregation frequency. Simulation results demonstrate that our proposed strategy achieves the lowest energy consumption while attaining the required model accuracy, outperforming other benchmarks.","author":[{"family":"Ren","given":"Yijing"},{"family":"Wu","given":"Changxiang"},{"family":"So","given":"Daniel"},{"family":"Tang","given":"Jie"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/twc.2025.3556514","URL":"https://doi.org/10.1109/twc.2025.3556514","source":"openalex"},{"id":"oa:W4410008025","type":"article-journal","title":"When Fine-Tuning LLMs Meets Data Privacy: An Empirical Study of Federated Learning in LLM-Based Program Repair","abstract":"Software systems have been evolving rapidly and inevitably introducing bugs at an increasing rate, leading to significant maintenance costs. While large language models (LLMs) have demonstrated remarkable potential in enhancing software development and maintenance practices, particularly in automated program repair (APR), they rely heavily on high-quality code repositories. Most code repositories are proprietary assets that capture the diversity and nuances of real-world industry software practices, which public datasets cannot fully represent. However, obtaining such data from various industries is hindered by data privacy concerns, as companies are reluctant to share their proprietary codebases. There has also been no in-depth investigation of collaborative software development by learning from private and decentralized data while preserving data privacy for program repair. To address the gap, we investigate federated learning as a privacy-preserving method for fine-tuning LLMs on proprietary and decentralized data to boost collaborative software development and maintenance. We use the private industrial dataset TutorCode for fine-tuning and the EvalRepair-Java benchmark for evaluation, and assess whether federated fine-tuning enhances program repair. We then further explore how code heterogeneity (i.e., variations in coding style, complexity, and embedding) and different federated learning algorithms affect bug fixing to provide practical implications for real-world software development collaboration. Our evaluation reveals that federated fine-tuning can significantly enhance program repair, achieving increases of up to 16.67% for Top@10 and 18.44% for Pass@10, even comparable to the bug-fixing capabilities of centralized learning. Moreover, the negligible impact of code heterogeneity implies that industries can effectively collaborate despite diverse data distributions. Different federated algorithms also demonstrate unique strengths across LLMs, suggesting that tailoring the optimization process to specific LLM characteristics can further improve program repair.","author":[{"family":"Luo","given":"Wenqiang"},{"family":"Keung","given":"Jacky"},{"family":"Yang","given":"Boyang"},{"family":"Ye","given":"He"},{"family":"Goues","given":"Claire"},{"family":"Bissyandé","given":"Tegawendé"},{"family":"Tian","given":"Haoye"},{"family":"Le","given":"Xuan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1145/3733599","URL":"https://doi.org/10.1145/3733599","source":"openalex"},{"id":"oa:W4413381171","type":"article-journal","title":"FLUID: Dynamic Model-Agnostic Federated Learning with Pruning and Knowledge Distillation for Maritime Predictive Maintenance","abstract":"Predictive maintenance (PdM) is vital to maritime operations; however, the traditional deep learning solutions currently offered heavily depend on centralized data aggregation, which is impractical under the limited connectivity, privacy concerns, and resource constraints found in maritime vessels. Federated Learning addresses privacy by training models locally, yet most FL methods assume homogeneous client architectures and exchange full model weights, leading to heavy communication overhead and sensitivity to system heterogeneity. To overcome these challenges, we introduce FLUID, a dynamic, model-agnostic FL framework that combines client clustering, structured pruning, and student–teacher knowledge distillation. FLUID first groups vessels into resource tiers and calibrates pruning strategies on the most capable client to determine optimal sparsity levels. In subsequent FL rounds, clients exchange logits over a small reference set, decoupling global aggregation from specific model architectures. We evaluate FLUID on a real-world heavy-fuel-oil purifier dataset under realistic heterogeneous deployment. With mixed pruning across clients, FLUID achieves a global R2 of 0.9352, compared with 0.9757 for a centralized baseline. Predictive consistency also remains high for client-based data, with a mean per-client MAE of 0.02575 ± 0.0021 and a mean RMSE of 0.0419 ± 0.0036. These results demonstrate FLUID’s ability to deliver accurate, efficient, and privacy-preserving PdM in heterogeneous maritime fleets.","author":[{"family":"Kalafatelis","given":"Alexandros"},{"family":"Pitsiakou","given":"Angeliki"},{"family":"Νομικός","given":"Νικόλαος"},{"family":"Tsoulakos","given":"Nikolaos"},{"family":"Syriopoulos","given":"Theodore"},{"family":"Trakadas","given":"Panagiotis"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/jmse13081569","URL":"https://doi.org/10.3390/jmse13081569","source":"openalex"},{"id":"oa:W4413794989","type":"article-journal","title":"Comparative analysis of deep learning architectures in solar power prediction","abstract":"Integrating renewable energy sources into the electricity grid requires accurate forecasts of solar power production. With the aim of enhancing the accuracy and reliability of forecasts, this study presents a comprehensive comparative analysis of eight state-of-the-art Deep Learning (DL) architectures-Autoencoder, Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), Simple Recurrent Neural Network (SimpleRNN), Convolutional Neural Network (CNN), Temporal Convolutional Network (TCN), Transformer, and Lightweight Informer for Long Sequence Time-Series Forecasting (InformerLite)-applied to solar power prediction using a dataset with 4,200 historical records and 20 meteorological and astronomical features. A comprehensive assessment of Root Mean Squared Error [Formula: see text], Mean Absolute Error [Formula: see text], Mean Absolute Percentage Error [Formula: see text], and Coefficient of Determination [Formula: see text] metrics was performed on the training, validation, and test datasets. The TCN model had the greatest performance across all models, achieving a test R² of 0.7786, an [Formula: see text] of 429.4863, and a balanced relative standard deviation ([Formula: see text]) of 0.6827, so exhibiting an exceptional capacity to capture temporal patterns. The Autoencoder achieved a [Formula: see text] of 0.7648 and had the greatest overall performance on the entire dataset, resulting in a Whole [Formula: see text] of 0.8437. In contrast, the Transformer model demonstrated significantly poorer performance (Test [Formula: see text] = 0.0714), underscoring its limitations in this context without any architectural modifications. This study not only demonstrates the best DL models for solar power forecasting as qualified by useful statistical metrics, but also provides a scalable, interpretable, and extensible forecasting framework for real-world energy systems. The findings verify the informed DL integration to smart grid scenarios, laying the foundations for further developments in hybrid modeling, multi-horizon prediction, and deployment in resource-constrained environments with limited computational power and resources.","author":[{"family":"Abdelsattar","given":"Montaser"},{"family":"Azim","given":"Mohamed"},{"family":"Abdelmoety","given":"Ahmed"},{"family":"Emad-Eldeen","given":"Ahmed"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-14908-x","URL":"https://doi.org/10.1038/s41598-025-14908-x","source":"openalex"},{"id":"oa:W4408272724","type":"article-journal","title":"Decentralized AI at the Edge: Federated Learning, Quantum Optimization and IoT Scalability","abstract":"Decentralized artificial intelligence (AI) at the edge marks a revolutionary evolution in computing, enabling efficient, privacy-preserving, and scalable solutions tailored for the Internet of Things (IoT). This paper integrates cutting-edge advancements in federated learning (FL), quantum optimization, and scalable IoT architectures to propose a cohesive framework for next-generation edge AI systems. We conducted an extensive literature review covering privacy-focused decentralized AI, quantum-enhanced optimization methods, and IoT system scalability. Our research highlights significant enhancements in model accuracy, resource efficiency, and data privacy through detailed comparative analysis and simulation-based experiments. Federated learning ensures local data processing, mitigating privacy risks, while quantum optimization accelerates complex computations, boosting system performance. However, challenges persist, including device heterogeneity, communication bottlenecks, and nascent quantum security risks. Our findings indicate that combining FL with quantum techniques can substantially improve edge AI scalability and effectiveness. Nonetheless, real-world deployment requires overcoming practical hurdles like interoperability and energy constraints. This paper thoroughly synthesizes the current landscape and charts a forward-looking agenda for research and innovation in decentralized edge AI.","author":[{"family":"Kiran","given":"Surya"},{"family":"Kumar","given":"Arjun"},{"family":"Chukkala","given":"Swathi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.30574/ijsra.2025.14.3.0633","URL":"https://doi.org/10.30574/ijsra.2025.14.3.0633","source":"openalex"},{"id":"oa:W4409988825","type":"article-journal","title":"A Survey of Clustering Federated Learning in Heterogeneous Data Scenarios","abstract":"Federated learning, as a collaborative training paradigm that preserves raw data privacy, offers an effective solution for data protection concerns. However, its practical implementation faces significant challenges due to data heterogeneity. This heterogeneity manifests as non-independent and identically distributed (non-IID) data across participating entities, resulting in degraded model performance, slower convergence rates, and training instability. While conventional federated learning approaches—including parameter averaging, knowledge distillation, and personalization techniques—offer certain advantages, their efficacy remains limited in severely heterogeneous environments. This survey systematically examines research advancements in clustered federated learning for addressing data heterogeneity challenges, encompassing fundamental principles, model architecture development, and algorithmic implementations. We provide a detailed analysis of innovative algorithms ranging from IFCA to FedGroup, and from FCL-GNN to FedAC, highlighting their technical contributions and applicable scenarios. Furthermore, we explore emerging research directions including clustering interpretability, multi-source heterogeneous information fusion, dynamic clustering mechanisms, and resource-aware optimization. Clustered federated learning effectively enhances model performance and convergence efficiency while maintaining privacy by grouping participants with similar data distributions into clusters and training specialized models for each cluster. With ongoing technological progress, clustered federated learning shows promise for achieving an optimal balance between privacy preservation and learning efficiency in critical domains such as healthcare and finance, thereby contributing to the sustainable development of artificial intelligence technologies.","author":[{"family":"Liu","given":"Entuo"},{"family":"Yang","given":"Wentong"},{"family":"Gu","given":"Yonggen"},{"family":"Long","given":"Wei"},{"family":"Istvan","given":"Szabo"},{"family":"Jiang","given":"Lin‐hua"}],"issued":{"date-parts":[[2025]]},"DOI":"10.54097/v7wcad61","URL":"https://doi.org/10.54097/v7wcad61","source":"openalex"},{"id":"oa:W4406102782","type":"article-journal","title":"Ultra-Short-Term Distributed Photovoltaic Power Probabilistic Forecasting Method Based on Federated Learning and Joint Probability Distribution Modeling","abstract":"The accurate probabilistic forecasting of ultra-short-term power generation from distributed photovoltaic (DPV) systems is of great significance for optimizing electricity markets and managing energy on the user side. Existing methods regarding cluster information sharing tend to easily trigger issues of data privacy leakage during information sharing, or they suffer from insufficient information sharing while protecting data privacy, leading to suboptimal forecasting performance. To address these issues, this paper proposes a privacy-preserving deep federated learning method for the probabilistic forecasting of ultra-short-term power generation from DPV systems. Firstly, a collaborative feature federated learning framework is established. For the central server, information sharing among clients is realized through the interaction of global models and features while avoiding the direct interaction of raw data to ensure the security of client data privacy. For local clients, a Transformer autoencoder is used as the forecasting model to extract local temporal features, which are combined with global features to form spatiotemporal correlation features, thereby deeply exploring the spatiotemporal correlations between different power stations and improving the accuracy of forecasting. Subsequently, a joint probability distribution model of forecasting values and errors is constructed, and the distribution patterns of errors are finely studied based on the dependencies between data to enhance the accuracy of probabilistic forecasting. Finally, the effectiveness of the proposed method was validated through real datasets.","author":[{"family":"Wang","given":"Yübo"},{"family":"Huo","given":"Chao"},{"family":"Xu","given":"Fei"},{"family":"Zheng","given":"Libin"},{"family":"Hao","given":"Ling"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/en18010197","URL":"https://doi.org/10.3390/en18010197","source":"openalex"},{"id":"oa:W4410089534","type":"article-journal","title":"Empowering Federated Graph Rationale Learning with Latent Environments","abstract":"The success of Graph Neural Networks (GNNs) in graph classification has heightened interest in explainable GNNs, particularly through graph rationalization. This method aims to enhance GNNs explainability by identifying subgraph structures (i.e., rationales) that support model predictions. However, existing methods often rely on centralized datasets, posing challenges in scenarios where data privacy is crucial, such as in molecular property prediction. Federated Learning (FL) offers a solution by enabling collaborative model training without sharing raw data. In this context, Federated Graph Rationalization emerges as a promising research direction. However, in each client, the rationalization methods often rely on client-specific shortcuts to compose rationales and make task predictions. Data heterogeneity, characterized by non-IID data across clients, exacerbates this problem, leading to poor prediction performance. To address these challenges, we propose the Environment-aware Data Augmentation (EaDA) method for Federated Graph Rationalization. EaDA comprises two main components: the Environment-aware Rationale Extraction (ERE) module and the Local-Global Alignment (LGA) module. The ERE module employs prototype learning to infer and share abstract environment information across clients, which are then aggregated to form a global environment. This information is used to generate counterfactual samples for local clients, enhancing the robustness of task predictions. The LGA module uses contrastive learning methods to align local and global rationale representations, mitigating performance degradation due to data heterogeneity. Comprehensive experiments on benchmark datasets demonstrate the effectiveness of our approaches. Code is available at https://github.com/yuelinan/Codes-of-EaDA.","author":[{"family":"Yue","given":"Linan"},{"family":"Liu","given":"Qi"},{"family":"Li","given":"Yawen"},{"family":"Yao","given":"Fangzhou"},{"family":"Gao","given":"Weibo"},{"family":"Du","given":"Junping"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1145/3696410.3714929","URL":"https://doi.org/10.1145/3696410.3714929","source":"openalex"},{"id":"oa:W4407391569","type":"article-journal","title":"Participant Selection for Efficient and Trusted Federated Learning in Blockchain-Assisted Hierarchical Federated Learning Architectures","abstract":"Federated learning has attracted widespread attention due to its strong capabilities of privacy protection, making it a powerful supporting technology for addressing data silos in the future. However, federated learning still lags significantly behind traditional centralized learning in terms of learning efficiency and system security. In this paper, we first construct a hierarchical federated learning architecture integrated with blockchain based on the cooperation of the cloud, edge, and terminal, which has the ability to enhance the security of federated learning while reducing the introduction costs of blockchain. Under this architecture, we propose a semi-asynchronous aggregation scheme at the edge layer and introduce a hierarchical aggregation scheme that combines it with synchronous aggregation at the cloud end to improve system efficiency. Furthermore, we present a multi-objective node selection scheme that considers various influencing factors such as security and efficiency. We formulate the node selection problem as a Markov Decision Process (MDP) and propose a solution based on deep reinforcement learning to address it more efficiently. The experimental results show that the proposed scheme can effectively improve system efficiency and enhance system security. In addition, the proposed DQN-based node selection algorithm can efficiently realize the selection of the optimal policy.","author":[{"family":"Liu","given":"Peng"},{"family":"Jia","given":"Lili"},{"family":"Xiao","given":"Yang"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/fi17020075","URL":"https://doi.org/10.3390/fi17020075","source":"openalex"},{"id":"oa:W4414051559","type":"article-journal","title":"Privacy-Preserving Federated Learning for Predictive Maintenance in Smart Manufacturing Networks","abstract":"Smart manufacturing environments (digitalized production systems with integrated sensor networks and data analytics capabilities) require advanced predictive maintenance capabilities, yet implementation faces significant barriers due to data privacy concerns and proprietary knowledge protection requirements.Traditional machine learning approaches necessitate centralized data repositories, creating obstacles for collaborative maintenance optimization across organizational boundaries.This research develops and evaluates a federated learning framework that enables effective predictive maintenance while preserving data privacy in manufacturing networks.The study implemented a horizontal federated learning architecture with secure aggregation protocols and differential privacy techniques across multiple aerospace manufacturing facilities.System performance was evaluated through comparative analysis against centralized and standalone approaches across multiple predictive maintenance use cases.The federated approach achieved 93.7% of centralized model accuracy while eliminating cross-facility data sharing, with failure prediction lead times approaching centralized performance while substantially outperforming standalone models.Computational overhead increased modestly, but network data transfer requirements decreased by 94%.Privacy analysis confirmed that proprietary process parameters could not be reconstructed from shared model updates.This research advances smart manufacturing capabilities by providing a practical implementation framework for privacy-preserving predictive maintenance across organizational boundaries, enabling industry collaboration while maintaining intellectual property protection.","author":[{"family":"Escorcia","given":"Yulineth"},{"family":"Sabirov","given":"Sardor"},{"family":"Saydullayev","given":"Bakhodir"},{"family":"Umarov","given":"AV"},{"family":"Atamuratova","given":"Zukhra"},{"family":"Alsayah","given":"Ahmed"},{"family":"Tulekov","given":"Yerzhan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.24867/ijiem-390","URL":"https://doi.org/10.24867/ijiem-390","source":"openalex"},{"id":"oa:W4412877087","type":"article-journal","title":"HtFLlib: A Comprehensive Heterogeneous Federated Learning Library and Benchmark","abstract":"As AI evolves, collaboration among heterogeneous models helps overcome data scarcity by enabling knowledge transfer across institutions and devices.Traditional Federated Learning (FL) only supports homogeneous models, limiting collaboration among clients with heterogeneous model architectures.To address this, Heterogeneous Federated Learning (HtFL) methods are developed to enable collaboration across diverse heterogeneous models while tackling the data heterogeneity issue at the same time.However, a comprehensive benchmark for standardized evaluation and analysis of the rapidly growing HtFL methods is lacking.Firstly, the highly varied datasets, model heterogeneity scenarios, and different method implementations become hurdles to making easy and fair comparisons among HtFL methods.Secondly, the effectiveness and robustness of HtFL methods are under-explored in various scenarios, such as the medical domain and sensor signal modality.To fill this gap, we introduce the first Heterogeneous Federated Learning Library (HtFLlib), an easy-to-use and extensible framework that integrates multiple datasets and model heterogeneity scenarios, offering a robust benchmark for research and practical applications.Specifically, HtFLlib integrates (1) 12 datasets spanning various domains, modalities, and data heterogeneity scenarios; (2) 40 model architectures, ranging from small to large, across three modalities;(3) a modularized and easy-to-extend HtFL codebase with implementations of 10 representative HtFL methods; and (4) systematic evaluations in terms of accuracy, convergence, computation costs, and communication costs.We emphasize the advantages and potential of state-of-the-art HtFL methods and hope that HtFLlib will catalyze advancing HtFL research and enable its broader applications.The code is released at https://github.com/TsingZ0/HtFLlib.","author":[{"family":"Zhang","given":"Jianqing"},{"family":"Wu","given":"Xinghao"},{"family":"Zhou","given":"Yanbing"},{"family":"Sun","given":"Xiaoting"},{"family":"Cai","given":"Qiqi"},{"family":"Liu","given":"Yang"},{"family":"Yang","given":"Hua"},{"family":"Zheng","given":"Zhenzhe"},{"family":"Cao","given":"Jian"},{"family":"Yang","given":"Qiang"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1145/3711896.3737379","URL":"https://doi.org/10.1145/3711896.3737379","source":"openalex"},{"id":"oa:W4410453221","type":"article-journal","title":"Quantization-based chained privacy-preserving federated learning","abstract":"Federated Learning (FL) is an advanced distributed machine learning framework crucial in protecting data privacy and security. By enabling multiple participants to train models while keeping their data local collaboratively, FL effectively mitigates the risks associated with centralized storage and sharing of raw data. However, traditional FL schemes face significant challenges regarding communication efficiency, computational costs, and privacy preservation. For instance, its communication and computational overhead in edge computing scenarios is often excessively high, hindering real-time applications. This paper proposes an innovative federated learning framework, Q-Chain FL, integrating quantization compression techniques into a chained FL architecture. This Q-Chain FL scheme adopts efficient compression and transmission of model parameter differences at the user node and executes seamless decompression and aggregation at the server node. Experiments on several publicly available datasets, including MNIST, CIFAR-10, and CelebA, demonstrate low communication and computational overhead, fast convergence speed, and high security of Q-Chain FL. Compared to traditional FedAvg and Chain-PPFL, Q-Chain FL reduces communication overhead by approximately 62.5% and 44.7%, respectively. These results underscore the robustness and adaptability of Q-Chain FL in various datasets and real-world learning scenarios.","author":[{"family":"Liu","given":"Ya"},{"family":"Wu","given":"Shumin"},{"family":"Li","given":"Yibo"},{"family":"Zhao","given":"Fengyu"},{"family":"Ren","given":"Yanli"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-01420-5","URL":"https://doi.org/10.1038/s41598-025-01420-5","source":"openalex"},{"id":"oa:W4413240396","type":"article-journal","title":"Federated Learning for Semantic Communication Based on CNNs and Transformer","abstract":"This study focuses on the latest research advancements in the field of semantic communication. Traditional communication systems prioritize the transmission of raw data, whilst semantic communication emphasizes conveying the meaning represented by the data. However, the extracted semantic information is often ambiguous and subject to subjective evaluation. To address this problem, this study proposes a model that combines a convolutional neural network (CNN) with a Transformer, called DeepSC‐CT. The model utilizes a CNN to extract semantic information from the data, followed by a Transformer model to capture spatial relationships and contextual information within the semantic content. We utilize federated learning to train the model and propose an adaptive aggregation algorithm to accelerate the convergence process. Moreover, we expand the single‐modality semantic communication model to encompass multiple modalities, such as texts, audio, and images. Furthermore, this study introduces a learnable position‐encoding method for the Transformer. The experimental results and visual effects of audio and image restoration demonstrate that the proposed method exhibits impressive performance and that the proposed model shows robust data restoration capabilities under various signal‐to‐noise ratio conditions.","author":[{"family":"Li","given":"Shufeng"},{"family":"Cai","given":"Yujun"},{"family":"Deng","given":"Zhaokai"},{"family":"Ba","given":"Xinran"},{"family":"Zheng","given":"Qinghe"},{"family":"Zhang","given":"Xinruo"},{"family":"Su","given":"Baoxin"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1155/int/3750087","URL":"https://doi.org/10.1155/int/3750087","source":"openalex"},{"id":"oa:W4413885161","type":"article-journal","title":"Federated Multi-Agent DRL for Task Offloading in Vehicular Edge Computing","abstract":"With the expansion of vehicle-to-everything (V2X) networks and the rising demand for intelligent services, vehicle edge computing encounters heightened requirements for more efficient task offloading. This study proposes a task offloading technique that utilizes federated collaboration and multi-agent deep reinforcement learning to reduce system latency and energy consumption. The task offloading issue is formulated as a Markov decision process (MDP), and a framework utilizing the Multi-Agent Dueling Double Deep Q-Network (MAD3QN) is developed to facilitate agents in making optimal offloading decisions inside intricate environments. Secondly, Federated Learning (FL) is implemented during the training phase, leveraging local training outcomes from many vehicles to enhance the global model, thus augmenting the learning efficiency of the agents. Experimental results indicate that, compared to conventional baseline algorithms, the proposed method decreases latency and energy consumption by at least 10% and 9%, respectively, while enhancing the average reward by at least 21%.","author":[{"family":"Zhao","given":"Hongwei"},{"family":"Li","given":"Yu"},{"family":"Pang","given":"Zhixi"},{"family":"Ma","given":"Zihan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/electronics14173501","URL":"https://doi.org/10.3390/electronics14173501","source":"openalex"},{"id":"oa:W4415774286","type":"article-journal","title":"Privacy and Trust in Blockchain-Federated Intrusion Detection Systems: Taxonomy, Challenges and Perspectives","abstract":"Intrusion Detection Systems (IDS) play a critical role in protecting modern networks, but traditional centralized designs raise serious concerns regarding data privacy, trust, and scalability. Federated Learning (FL) reduces privacy risks through decentralized model training, and blockchain enhances trust by providing immutability and transparency. Combining these technologies creates a promising paradigm for secure and trustworthy IDS. This paper presents a comprehensive survey of blockchain-federated IDS with a particular focus on privacy and trust. The key contribution is a multi-dimensional taxonomy that integrates IDS architectures, FL strategies, blockchain types, and consensus mechanisms, providing a clear and structured view of this emerging field. We categorize threats into data, communication, and model levels, and map representative defense mechanisms to each. We also review applications in vehicular networks, industrial and medical Internet of Things (IoT), and metaverse scenarios. Finally, we highlight key challenges, including non-IID data, lightweight consensus, incentive mechanisms, and poisoning-resilient aggregation, and outline future research directions.","author":[{"family":"Yuan","given":"Cao"},{"family":"Ku","given":"Chin"},{"family":"Kumar","given":"Rahul"},{"family":"Khan","given":"Arshad"}],"issued":{"date-parts":[[2025]]},"DOI":"10.62762/jrsc.2025.399812","URL":"https://doi.org/10.62762/jrsc.2025.399812","source":"openalex"},{"id":"oa:W4409912927","type":"article-journal","title":"Managing Supply and Value Chains in a New Era of Global Trade: The Promise of Federated Learning","abstract":"As global trade shifts from an era of efficiency-driven globalisation to a new compliance-centred paradigm, customs administrations face mounting challenges – ranging from forced labour and environmental enforcement to fractured supply chain visibility and escalating transaction volumes, particularly in e-commerce. This article introduces federated learning as a practical, privacy-preserving solution for enabling secure data collaboration across public and private actors without commingling or centralising sensitive information. We trace the structural failures of Globalisation 1.0 and propose a modernised model of border management built on federated system architecture and trusted networks. These systems allow customs authorities to apply actionable intelligence across multi-tier value chains, strengthen enforcement capabilities, and expedite legitimate trade. The article outlines key steps towards implementation, including legal, technical and institutional reforms, and argues that federated architectures can form the foundation for next-generation risk management and trade facilitation strategies. The future of effective customs governance will depend on embracing secure, data-driven collaboration within and across borders.","author":[{"family":"Bersin","given":"Alan"},{"family":"Swartz","given":"Peter"},{"family":"Karlsson","given":"Lars"}],"issued":{"date-parts":[[2025]]},"DOI":"10.55596/001c.133998","URL":"https://doi.org/10.55596/001c.133998","source":"openalex"},{"id":"oa:W4414359239","type":"article-journal","title":"Federated Low-Rank Adaptation for Foundation Models: A Survey","abstract":"Effectively leveraging private datasets remains a significant challenge in developing foundation models. Federated Learning (FL) has recently emerged as a collaborative framework that enables multiple users to fine-tune these models while mitigating data privacy risks. Meanwhile, Low-Rank Adaptation (LoRA) offers a resource-efficient alternative for fine-tuning foundation models by dramatically reducing the number of trainable parameters. This survey examines how LoRA has been integrated into federated fine-tuning for foundation models—an area we term FedLoRA—by focusing on three key challenges: distributed learning, heterogeneity, and efficiency. We further categorize existing work based on the specific methods used to address each challenge. Finally, we discuss open research questions and highlight promising directions for future investigation, outlining the next steps for advancing FedLoRA.","author":[{"family":"Yang","given":"Yiyuan"},{"family":"Long","given":"Guodong"},{"family":"Lu","given":"Qinghua"},{"family":"Zhu","given":"Liming"},{"family":"Jiang","given":"Jing"},{"family":"Zhang","given":"Chengqi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.24963/ijcai.2025/1196","URL":"https://doi.org/10.24963/ijcai.2025/1196","source":"openalex"},{"id":"oa:W4410961669","type":"article-journal","title":"FedMDKGE: Multi-granularity Dynamic Knowledge Graph Embedding in Federated Learning","abstract":"As knowledge is time-sensitive, some researchers have started to focus on dynamic knowledge graphs to provide time-dimensioned knowledge content thus reflecting richer information. But they have not yet combined temporal information at different granularities. Also, in the case of multiple knowledge graphs distributed across different clients, it is of interest to ensure that the knowledge graph embedding representations are learned without exposing data and collaboratively. Therefore, in this paper, we propose a framework for multi-granularity dynamic knowledge graph embedding in federated learning (FedMDKGE), which allows multiple parties to interact securely with temporal information at different granularities. In the client, we present a multi-granularity dynamic knowledge graph embedding model that improves the capability of dynamic knowledge graph embedding representation by focusing on multi-granularity temporal facts from the perspective of knowledge utilization. On the server, we design a multi-granularity aggregation rule to accommodate multi-party information aggregation at different granularities. Finally, we conduct extensive experiments to demonstrate the superior performance of our model. The results on these real datasets show that FedMDKGE considering multi-granularity temporal information performs better than all comparative baselines and interconnect information for multi-party dynamic knowledge graph embedding without exposing data.","author":[{"family":"Huang","given":"Wei"},{"family":"Chen","given":"Junling"},{"family":"Wang","given":"Dexian"},{"family":"Zhang","given":"Pengfei"},{"family":"Liu","given":"Jia"},{"family":"Li","given":"Tianrui"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s44196-025-00878-5","URL":"https://doi.org/10.1007/s44196-025-00878-5","source":"openalex"},{"id":"oa:W7116103563","type":"article-journal","title":"Federated Learning for Multi-Disease Ophthalmic Diagnostics Using OCT Angiography","abstract":"Purpose: To conduct a comprehensive systematic evaluation of federated learning (FL) strategies for multi-disease retinal classification using OCT angiography (OCTA), implementing a 2-part experimental framework to establish foundational feasibility and optimize performance under realistic heterogeneous conditions while ensuring privacy preservation. Design: = 0.5). Participants: A total of 456 OCTA images from patients with 7 retinal pathologies, with diabetic retinopathy (31.1%) and normal cases (25.2%) comprising the majority, sourced from the public OCTA-500 data set (n = 300) and a private collection from the University of Illinois Chicago (n = 156). Methods: Five FL aggregation strategies (federated averaging [FedAvg], federated proximal [FedProx], federated magnetic resonance imaging [FedMRI], federated Adagrad, and federated Yogi) were systematically evaluated across multiple optimization dimensions: 7 architecture configurations spanning vision transformers, established convolutional neural networks, and hybrid models; 5 transfer learning freezing strategies; 3 local epoch configurations (2, 5, and 10); and scalability analysis across 2, 3, and 5-client federations. Security mechanisms including differential privacy (ε = 1.0-8.0) and secure aggregation were integrated and evaluated. Performance was assessed across 3 classification scenarios: 7-class, 4-class modified, and 4-class streamlined. Main Outcome Measures: Classification accuracy, receiver-operating-characteristic area under the curve (ROC-AUC), and macro-averaged F1-score with comprehensive privacy-utility analysis and computational efficiency metrics. Results: Under controlled conditions, FL achieved superior performance in simplified classifications, with FedAvg, FedProx, and FedMRI reaching 72.09% accuracy versus 69.77% centralized training. Comprehensive optimization identified DenseNet121 as optimal architecture (79.55% accuracy, 89.68% ROC-AUC), with \"most\" freezing strategy (75% frozen layers) providing 60% training time reduction while maintaining superior performance. Federated proximal demonstrated exceptional resilience to heterogeneity (-11.7% degradation). Bonawitz secure aggregation achieved optimal privacy-utility balance (63.64% accuracy with cryptographic guarantees), whereas differential privacy maintained clinical utility under moderate constraints (ε ≈ 4-6). Conclusions: This systematic evaluation establishes FL as a comprehensive solution for privacy-preserving multi-institutional OCTA-based disease classification, with careful architectural selection, optimization strategies, and security mechanisms enabling performance that matches or exceeds centralized approaches while maintaining regulatory compliance and clinical utility. Financial Disclosures: The authors have no proprietary or commercial interest in any materials discussed in this article.","author":[{"family":"Nabil","given":"Ahammed"},{"family":"Gholami","given":"Sina"},{"family":"Leng","given":"Theodore"},{"family":"Lim","given":"Jennifer"},{"family":"Alam","given":"Minhaj"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1016/j.xops.2025.101030","URL":"https://doi.org/10.1016/j.xops.2025.101030","source":"openalex"},{"id":"oa:W4416873395","type":"article-journal","title":"Federated Deep Reinforcement Learning for Privacy-Preserving Robotic-Assisted Surgery","abstract":"The integration of Reinforcement Learning (RL) into robotic-assisted surgery (RAS) holds significant promise for advancing surgical precision, adaptability, and autonomous decision-making. However, the development of robust RL models in clinical settings is hindered by key challenges, including stringent patient data privacy regulations, limited access to diverse surgical datasets, and high procedural variability. To address these limitations, this paper presents a Federated Deep Reinforcement Learning (FDRL) framework that enables decentralised training of RL models across multiple healthcare institutions without exposing sensitive patient information. A central innovation of the proposed framework is its dynamic policy adaptation mechanism, which allows surgical robots to select and tailor patient-specific policies in real-time, thereby ensuring personalised and optimised interventions. To uphold rigorous privacy standards while facilitating collaborative learning, the FDRL framework incorporates secure aggregation, differential privacy, and homomorphic encryption techniques. Experimental results demonstrate a $60 \\%$ reduction in privacy leakage compared to conventional methods, with surgical precision maintained within a $1.5 \\%$ margin of a centralised baseline. This work establishes a foundational approach for adaptive, secure, and patient-centric AI-driven surgical robotics, offering a pathway toward clinical translation and scalable deployment across diverse healthcare environments.","author":[{"family":"Hafeez","given":"Sana"},{"family":"Mulkana","given":"Sundas"},{"family":"Imran","given":"Muhammad"},{"family":"Sevegnani","given":"Michele"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/icdcsw63273.2025.00128","URL":"https://doi.org/10.1109/icdcsw63273.2025.00128","source":"openalex"},{"id":"oa:W4409895099","type":"article-journal","title":"Personalized Federated Transfer Learning for Building Energy Forecasting via Model Ensemble with Multi-Level Masking in Heterogeneous Sensing Environment","abstract":"Effective building energy prediction is essential for optimizing energy management, but existing models struggle with data scarcity and sensor heterogeneity across different buildings. Conventional approaches, including centralized and transfer learning methods, fail to generalize well due to varying sensor configurations and inconsistent data availability. To overcome these challenges, this study proposes a Personalized Federated Learning (pFL) framework that integrates multi-level feature masking, model ensemble techniques, and knowledge transfer to enhance predictive performance across diverse buildings. The proposed feature masking strategy extracts the most relevant time-series features, while model ensemble learning improves generalization, and knowledge transfer enables adaptive fine-tuning for each building. These techniques allow pFL to retain global knowledge while personalizing to local energy consumption patterns, making it more effective than traditional FL methods. Experiments conducted on a campus energy dataset demonstrate that pFL consistently outperforms FedAvg, FedProx, and standalone models in energy prediction accuracy. The most significant improvements are observed in buildings with highly fluctuating consumption patterns, validating the effectiveness of the proposed approach in handling heterogeneous sensing environments. This study highlights the potential of Federated Learning for scalable and adaptive energy prediction. Future work will focus on refining multi-horizon forecasting and developing strategies to enhance knowledge sharing among buildings for improved long-term performance.","author":[{"family":"Kim","given":"Hakjae"},{"family":"Dorjgochoo","given":"Sarangerel"},{"family":"Park","given":"Hansaem"},{"family":"Lee","given":"Sung"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/electronics14091790","URL":"https://doi.org/10.3390/electronics14091790","source":"openalex"},{"id":"oa:W4416909357","type":"article-journal","title":"Federated Deep Learning for Robust Multi-Modal Biometric Authentication Based on Facial and Eye-Blink Cues","abstract":"The increasing demand for secure and user-friendly authentication mechanisms has led to the exploration of biometric systems that leverage unique physiological traits. Among these, face recognition and eye blink detection have emerged as effective and non-intrusive modalities. However, traditional biometric systems typically rely on centralized data storage and processing, raising significant concerns about user privacy, data security, and potential breaches. To address these challenges, this paper proposes a federated learning-based framework that combines face and eye blink recognition for robust user authentication.The proposed system utilizes OpenCV for real-time image capture and processing, enabling users to register by submitting facial images and customized eye blink patterns. These biometric features are used to train local models that remain on the user's device, ensuring that raw biometric data is never transmitted to external servers. Instead, model parameters are shared and aggregated at a centralized server using federated learning techniques, resulting in a global model that benefits from decentralized data sources while maintaining user privacy.The system is divided into key modules: face registration, eye blink training, federated model updating, and multi-modal authentication. Each module plays a vital role in establishing a secure and user-specific identity. The integration of eye blink recognition as a secondary verification layer significantly enhances the system's resistance to spoofing attacks and impersonation. Experimental evaluations demonstrate the system’s effectiveness in accurately identifying users while preserving privacy and reducing server dependency.This research offers a novel contribution to biometric security by combining federated learning with multi-modal authentication, paving the way for privacy-preserving, scalable, and intelligent user verification systems in real-world applications.","author":[{"family":"Balaji","given":"A"},{"family":"Balanjali","given":"Doppalapudi"},{"family":"Subbaiah","given":"Guntu"},{"family":"Reddy","given":"Avula"},{"family":"Karthik","given":"Daggubati"}],"issued":{"date-parts":[[2025]]},"DOI":"10.65521/ijacect.v14i1.167","URL":"https://doi.org/10.65521/ijacect.v14i1.167","source":"openalex"},{"id":"oa:W4415819639","type":"article-journal","title":"On-Device Federated Learning for Energy-Efficient Smart Irrigation","abstract":"This study presents a novel federated learning (FL) methodology implemented directly on STM32-based microcontrollers (MCUs) for energy-efficient smart irrigation. To the best of our knowledge, this is the first work to demonstrate end-to-end FL training and aggregation on real STM32 MCU clients (STM32F722ZE), under realistic energy and memory constraints. Unlike most prior studies that rely on simulated clients or high-power edge devices, our framework deploys lightweight neural networks trained locally on MCUs and synchronized via message queuing telemetry transport (MQTT) communication. Using a smart agriculture (SA) dataset partitioned by soil type, 7 clients collaboratively trained a model over 3 federated rounds. Experimental results show that MCU clients achieved competitive accuracy (70–82%) compared to PC clients (80–85%) while consuming orders of magnitude less energy. Specifically, MCU inference required only 0.95 mJ per sample versus 60–70 mJ on PCs, and training consumed ∼70 mJ per epoch versus nearly 20 J. Latency remained modest, with MCU inference averaging 3.2 ms per sample compared to sub-millisecond execution on PCs, a negligible overhead in irrigation scenarios. The evaluation also considers the payoff between accuracy, energy consumption, and latency through the Energy Latency Accuracy Index (ELAI). This integrated perspective highlights the trade-offs inherent in deploying FL on heterogeneous devices and demonstrates the efficiency advantages of MCU-based training in energy-constrained smart irrigation settings.","author":[{"family":"Dakhia","given":"Zohra"},{"family":"Lazzaro","given":"Alessia"},{"family":"Sebti","given":"Mohamed"},{"family":"Russo","given":"Mariateresa"},{"family":"Merenda","given":"Massimo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/electronics14214311","URL":"https://doi.org/10.3390/electronics14214311","source":"openalex"},{"id":"oa:W4416018425","type":"article-journal","title":"Collision avoidance in UAV swarms: A learning-centric perspective on collaborative intelligence","abstract":"As UAV swarm deployments become more prevalent in mission critical domains, collision avoidance remains a key challenge in ensuring safety, coordination, and autonomy at scale. This survey investigates the state of the art in learning based collision avoidance strategies enabled through collaborative intelligence in UAV swarms. We introduce a six dimensional taxonomy that classifies approaches across decision making paradigms, swarm coordination models, communication architectures, learning methodologies, execution strategies, and safety assurance mechanisms. The survey places particular emphasis on learning based methodologies, which we categorize into four prominent techniques: reinforcement learning, federated learning, neuro inspired models, and hybrid approaches. For each, we provide a detailed review of training architectures, scalability, robustness, and real-time feasibility. Drawing on peer-reviewed publications (2019 till early 2025), we synthesize comparative insights into their application contexts, including trajectory planning, vision-based navigation, decentralized coordination, and multi-agent conflict resolution, while assessing trade-offs in deployment complexity and operational safety. Beyond method specific analysis, the survey highlights key distinctions, practical challenges, and enabling technologies, concluding with open challenges and future directions for scalable and verifiable UAV swarm intelligence.","author":[{"family":"Khargharia","given":"Himadri"},{"family":"Ouali","given":"Anis"},{"family":"Shakya","given":"Siddhartha"},{"family":"Ahmad","given":"S"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1016/j.neucom.2025.132020","URL":"https://doi.org/10.1016/j.neucom.2025.132020","source":"openalex"},{"id":"oa:W4412688173","type":"article-journal","title":"Review: machine learning approaches for diverse alloy systems","abstract":"Abstract The integration of machine learning (ML) into alloy design has revolutionized the discovery and optimization of advanced materials by enabling high-throughput, data-driven methodologies. This review systematically examines recent advancements in ML applications across diverse alloy systems, including steels, aluminum alloys, magnesium alloys, nickel-based superalloys, high-entropy alloys (HEAs), shape memory alloys, and metallic glasses. We categorize ML approaches into supervised, unsupervised, and reinforcement learning paradigms, detailing their specific implementations for property prediction, phase stability analysis, and composition optimization. Advanced techniques, such as inverse design frameworks and physics-informed ML models, have demonstrated substantial improvements in predictive accuracy and interpretability by integrating domain knowledge with data-driven approaches. The review further explores the synergy between ML and traditional computational methods, including CALPHAD-based thermodynamic modeling and density functional theory (DFT), enhancing the reliability of property predictions. We highlight case studies where ML-driven strategies have successfully accelerated alloy discovery, optimized mechanical properties, and identified novel compositions with tailored performance metrics. Additionally, we address key challenges in ML-driven alloy design, including data scarcity, feature selection, model interpretability, and the necessity for standardized benchmarking datasets. By providing a comprehensive evaluation of current methodologies and emerging trends, this review underscores the transformative role of ML in advancing next-generation alloy design and manufacturing, ultimately enabling the rapid development of high-performance materials for aerospace, energy, biomedical, and structural applications.","author":[{"family":"Rahman","given":"Arafat"},{"family":"Hossain","given":"Md"},{"family":"Siddique","given":"Abdullah"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s10853-025-11154-4","URL":"https://doi.org/10.1007/s10853-025-11154-4","source":"openalex"},{"id":"oa:W4410269797","type":"article-journal","title":"Federated transfer learning for distributed drought stage prediction","abstract":"Abstract Due to the uncertain nature of drought, it is one of the most menacing natural disasters. Drought modeling (Prediction, Detection, Forecasting, and Stage Prediction) is very essential for efficient policy making. But one of the key problems with drought modeling is the limited availability of centralized datasets. To address this problem, we are a novel proposing federated learning based transfer learning models for the prediction of drought stages. In this study, satellite image dataset was collected from the Tharparkar district (prone to drought) of Pakistan. We trained the dataset using traditional and federated learning approaches, comparing centralized ML models, pre-trained models, and their respective federated learning models (FL-ResNet, FL-DenseNet, FL-MobileNet). The development of these models is the novel aspect of the study specifically for the use case of drought stage prediction. Based on the final evaluation, FL-MobileNet achieved 82% precision while baseline MobileNet scored 68%. The results show the effectiveness of novelty (federated learning), that our proposed framework improves the performance of the drought stage classification task.","author":[{"family":"Raza","given":"Muhammad"},{"family":"Umar","given":"Aqsa"},{"family":"Rasheed","given":"Jawad"},{"family":"Aşuroğlu","given":"Tunç"},{"family":"Alsubai","given":"Shtwai"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s44163-025-00288-8","URL":"https://doi.org/10.1007/s44163-025-00288-8","source":"openalex"},{"id":"oa:W4409328620","type":"article-journal","title":"Federated LeViT-ResUNet for Scalable and Privacy-Preserving Agricultural Monitoring Using Drone and Internet of Things Data","abstract":"Precision agriculture is necessary for dealing with problems like pest outbreaks, a lack of water, and declining crop health. Manual inspections and broad-spectrum pesticide application are inefficient, time-consuming, and dangerous. New drone photography and IoT sensors offer quick, high-resolution, multimodal agricultural data collecting. Regional diversity, data heterogeneity, and privacy problems make it hard to conclude these data. This study proposes a lightweight, hybrid deep learning architecture called federated LeViT-ResUNet that combines the spatial efficiency of LeViT transformers with ResUNet’s exact pixel-level segmentation to address these issues. The system uses multispectral drone footage and IoT sensor data to identify real-time insect hotspots, crop health, and yield prediction. The dynamic relevance and sparsity-based feature selector (DRS-FS) improves feature ranking and reduces redundancy. Spectral normalization, spatial–temporal alignment, and dimensionality reduction provide reliable input representation. Unlike centralized models, our platform trains over-dispersed client datasets using federated learning to preserve privacy and capture regional trends. A huge, open-access agricultural dataset from varied environmental circumstances was used for simulation experiments. The suggested approach improves on conventional models like ResNet, DenseNet, and the vision transformer with a 98.9% classification accuracy and 99.3% AUC. The LeViT-ResUNet system is scalable and sustainable for privacy-preserving precision agriculture because of its high generalization, low latency, and communication efficiency. This study lays the groundwork for real-time, intelligent agricultural monitoring systems in diverse, resource-constrained farming situations.","author":[{"family":"Aldossary","given":"Mohammad"},{"family":"Almutairi","given":"Jaber"},{"family":"Alzamil","given":"Ibrahim"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/agronomy15040928","URL":"https://doi.org/10.3390/agronomy15040928","source":"openalex"},{"id":"oa:W4412800085","type":"article-journal","title":"Dual prompt personalized federated learning in foundation models","abstract":"Personalized federated learning (PFL) has garnered significant attention for its ability to address heterogeneous client data distributions while preserving data privacy. However, when local client data is limited, deep learning models often suffer from insufficient training, leading to suboptimal performance. Foundation models, such as CLIP (Contrastive Language-Image Pretraining), exhibit strong feature extraction capabilities and can alleviate this issue by fine-tuning on limited local data. Despite their potential, foundation models are rarely utilized in federated learning scenarios, and challenges related to integrating new clients remain largely unresolved. To address these challenges, we propose the Dual Prompt Personalized Federated Learning (DP 2 FL) framework, which introduces dual prompts and an adaptive aggregation strategy. DP 2 FL combines global task awareness with local data-driven insights, enabling local models to achieve effective generalization while remaining adaptable to specific data distributions. Moreover, DP 2 FL introduces a global model that enables prediction on new data sources and seamlessly integrates newly added clients without requiring retraining. Experimental results in highly heterogeneous environments validate the effectiveness of DP 2 FL’s prompt design and aggregation strategy, underscoring the advantages of prediction on novel data sources and demonstrating the seamless integration of new clients into the federated learning framework.","author":[{"family":"Chang","given":"Ying"},{"family":"Shi","given":"Xiaohu"},{"family":"Xiaohui","given":"Zhao"},{"family":"Chen","given":"Zhaohuang"},{"family":"Ma","given":"DC"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-11864-4","URL":"https://doi.org/10.1038/s41598-025-11864-4","source":"openalex"},{"id":"oa:W4409067086","type":"article-journal","title":"Design of an Immersive Basketball Tactical Training System Based on Digital Twins and Federated Learning","abstract":"To address the challenges of dynamic adversarial scenario modeling distortion, insufficient cross-institutional data privacy protection, and simplistic evaluation systems in collegiate basketball tactical education, this study proposes and validates an immersive instructional system integrating digital twin and federated learning technologies. The four-tier architecture (sensing layer, digital twin layer, federated layer, and interaction layer) synthesizes multimodal data (motion trajectories and physiological signals) with Multi-Agent Reinforcement Learning (MARL) to enable virtual–physical integrated tactical simulation and real-time error correction. Experimental results demonstrate that the experimental group achieved 35.2% higher tactical execution accuracy (TEA) (p < 0.01), 1.8 s faster decision making (p < 0.05), and 47% improved team coordination efficiency compared to the controls. The hierarchical federated learning framework (trajectory ε = 0.8; physiology ε = 0.3) maintained model precision loss at 2.4% while optimizing communication efficiency by 23%, ensuring privacy preservation. A novel three-dimensional “Skill–Creativity–Load” evaluation system revealed a 22% increase in unconventional tactical applications (p = 0.013) through the Tactical Creativity Index (TCI). By implementing lightweight federated architecture with dynamic cognitive offloading mechanisms, the system enables resource-constrained institutions to achieve 87% of the pedagogical effectiveness observed in elite programs, offering an innovative solution to reconcile educational equity with technological ethics. Future research should focus on long-term skill transfer, multimodal adaptive learning, and ethical framework development to advance intelligent sports education from efficiency-oriented paradigms to competency-based transformation.","author":[{"family":"Lv","given":"Xiongce"},{"family":"Tao","given":"Ye"},{"family":"Zhang","given":"Yifan"},{"family":"Xue","given":"Yang"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/app15073831","URL":"https://doi.org/10.3390/app15073831","source":"openalex"},{"id":"oa:W4410732892","type":"article-journal","title":"The role of explainable AI in enhancing breast cancer diagnosis using machine learning and deep learning models","abstract":"Breast cancer is still a big health issue around the world, and it needs to be found quickly and perfectly to improve patient outcomes and lower death rates. Although artificial intelligence (AI) has showed amazing promise in breast cancer prediction mainly machine learning (ML) algorithms as well as deep learning (DL), practical use of these models is greatly hampered by their lack of interpretability and transparency. By giving complicated AI models interpretability, explainable artificial intelligence (XAI) becomes an essential tool to improve trust and transparency. XAI's efficacy in clinical environments is yet perfectly unidentified however, and its proper implementation into breast cancer diagnostics is hence ignored. Focussing on their interpretability, clinical application, and influence on decision-making, this paper systematically reviews machine learning, deep learning impact on breast cancer diagnosis and current XAI approaches used to breast cancer detection, prognosis, and treatment. This work presents a thorough assessment of XAI approaches classified by data kinds (imaging, genomic, and clinical), a comparative analysis of LIME (Local Interpretable Model-agnostic Explanations), SHAP (SHapley Additive exPlanations), and Grad-CAM, and highlights important issues and future directions of research. This work highlights the possibility of XAI to enhance clinical decision-making and patient confidence by closing the gap between great diagnosis accuracy and interpretability. The results support multidisciplinary cooperation among medical experts, scientists in artificial intelligence, and legislators to guarantee the responsible and ethical integration of artificial intelligence in society.","author":[{"family":"Ansari","given":"Zulfikar"},{"family":"Tripathi","given":"Manish"},{"family":"Ahmed","given":"Rafeeq"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s44163-025-00307-8","URL":"https://doi.org/10.1007/s44163-025-00307-8","source":"openalex"},{"id":"oa:W4412929609","type":"article-journal","title":"Ensuring Zero Trust in GDPR-Compliant Deep Federated Learning Architecture","abstract":"Deep Federated Learning (DFL) revolutionizes machine learning (ML) by enabling collaborative model training across diverse, decentralized data sources without direct data sharing, emphasizing user privacy and data sovereignty. Despite its potential, DFL’s application in sensitive sectors is hindered by challenges in meeting rigorous standards like the GDPR, with traditional setups struggling to ensure compliance and maintain trust. Addressing these issues, our research introduces an innovative Zero Trust-based DFL architecture designed for GDPR compliant systems, integrating advanced security and privacy mechanisms to ensure safe and transparent cross-node data processing. Our base paper proposed the basic GDPR-Compliant DFL Architecture. Now we validate the previously proposed architecture by formally verifying it using High-Level Petri Nets (HLPNs). This Zero Trust-based framework facilitates secure, decentralized model training without direct data sharing. Furthermore, we have also implemented a case study using the MNIST and CIFAR-10 datasets to evaluate the existing approach with the proposed Zero Trust-based DFL methodology. Our experiments confirmed its effectiveness in enhancing trust, complying with GDPR, and promoting DFL adoption in privacy-sensitive areas, achieving secure, ethical Artificial Intelligence (AI) with transparent and efficient data processing.","author":[{"family":"Abbas","given":"Zahra"},{"family":"Ahmad","given":"Sunila"},{"family":"Anjum","given":"Adeel"},{"family":"Syed","given":"Madiha"},{"family":"Malik","given":"Saif"},{"family":"Rehman","given":"Semeen"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/computers14080317","URL":"https://doi.org/10.3390/computers14080317","source":"openalex"},{"id":"oa:W4412708560","type":"article-journal","title":"Intelligent waste sorting for urban sustainability using deep learning","abstract":"Smart cities’ have experienced an increasingly higher rate of urbanization and increase of the population leading to strengthening the pressing needs in waste management. In this paper, we present an intelligent waste classification system that utilises Convolutional Neural Networks (CNNs) for automatic segregation into twelve categories of waste, employing image data. The model is trained on 15,535 images from a publicly available dataset using preprocessing and data augmentation to increase generalisation and mitigate class imbalance. A performance comparison in terms of precision, recall, F1 score, and accuracy shows that the proposed ResNet-based model yields a classification accuracy of 98.16%, outperforming previous work on conventional deep learning architectures. Experimental results demonstrate that the model is a robust framework for handling various types of waste (organic, recyclable, and hazardous) and is a very general model, as confirmed by cross-validation and real-world tests. The proposed system demonstrates great promise for upscaling in automatic waste management towards long-term urban sustainability, improved recycling, and reduced environmental threats.","author":[{"family":"Ahmad","given":"GF"},{"family":"Aleem","given":"Fizza"},{"family":"Alyas","given":"Tahir"},{"family":"Abbas","given":"Qaiser"},{"family":"Nawaz","given":"Waqas"},{"family":"Ghazal","given":"Taher"},{"family":"Aziz","given":"Abdul"},{"family":"Aleem","given":"Shady"},{"family":"Tabassum","given":"Nadia"},{"family":"Ibrahim","given":"Aidarus"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-08461-w","URL":"https://doi.org/10.1038/s41598-025-08461-w","source":"openalex"},{"id":"oa:W4413358968","type":"article-journal","title":"Deep learning for intrusion detection in emerging technologies: a comprehensive survey and new perspectives","abstract":"Abstract Intrusion Detection Systems (IDS) can help cybersecurity analysts detect malicious activities in computational environments. Recently, Deep Learning (DL) methods in IDS have demonstrated notable performance, revealing new underlying cybersecurity patterns in systems’ operations. Conversely, issues such as low performance in real systems, high false positive rates, and lack of explainability hinder its real-world deployment. In addition, the adoption of many new emerging technologies, such as cloud, edge computing, and the Internet of Things (IoT) introduces new forms of vulnerabilities. Therefore, the improvement of intrusion detection in emerging technologies depends on the clear definitions of challenging security problems and the limitations of existing solutions. The main goal of this research is to conduct a literature review of DL solutions for intrusion detection in emerging technologies to understand the state-of-the-art solutions and their limitations. Specifically, we conduct a comprehensive review of IDS-based automated threat defense methods, with the objective of identifying the landscape of, and opportunities for, incorporating DL methods into IDS. To accomplish this, a thorough review of IDS methods is conducted for multiple platforms and technologies, focusing on the use of common DL techniques. To expand on the study, several widely used IDS datasets are evaluated to assess their ability to train DL models and support researchers in understanding their characteristics and limitations. The analysis of attack vectors in emerging technologies is conducted, enabling an in-depth evaluation of security solutions in the future. Our findings show many clear opportunities for future research, including addressing the gap between solutions for controlled/simulated environments versus real systems, overcoming trustworthiness issues, including lack of explainability, and further exploring operationalization issues such as deployable solutions and continuous detection. Our analysis highlights that the operationalization of DL for intrusion detection in emerging technologies represents a key challenge to be addressed in the next few years.","author":[{"family":"Neto","given":"Euclides"},{"family":"Iqbal","given":"Shahrear"},{"family":"Buffett","given":"Scott"},{"family":"Sultana","given":"Madeena"},{"family":"Taylor","given":"Adrian"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s10462-025-11346-z","URL":"https://doi.org/10.1007/s10462-025-11346-z","source":"openalex"},{"id":"oa:W4414380024","type":"article-journal","title":"Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings","abstract":"MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.","author":[{"family":"Montalvo","given":"Nigreisy"},{"family":"Silvente","given":"Francisco"},{"family":"Capriotti","given":"Emidio"},{"family":"Rausell","given":"Antonio"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1093/bioinformatics/btaf523","URL":"https://doi.org/10.1093/bioinformatics/btaf523","source":"openalex"},{"id":"oa:W4412539477","type":"article-journal","title":"Blockchain-based federated learning framework for malicious node detection in internet of vehicles (IoV) networks using fog and cloud computing","abstract":"Due to the continuous digitalization, IoV networks are vulnerable to various communication attacks by malicious network nodes. In these attacks, the malicious entities disseminate faulty information in the network, which affects quick and intelligent decision-making in the network. Many deep learning and machine learning techniques are proposed for the classification of legitimate and malicious vehicular entities. These techniques have a centralized model training structure, which has low classification accuracy and is vulnerable to privacy leakage. To address these issues, we propose a blockchain-based federated learning framework for distributed classification of malicious and legitimate vehicles. The proposed model uses the capabilities of Long short-term memory (LSTM) and Naive Bayes (NB) for efficient and reliable malicious node detection. In our proposed model, the distributed models are trained on each locally installed virtual machine with a federated learning mechanism and then a unified model is generated at the centralized cloud server. The proposed model not only enhances the accuracy and privacy preservation but also solves the issues of centralized Internet of Vehicles (IoV) networks such as single point of failure and performance bottlenecks by utilizing the capabilities of blockchain. We used the Vehicular Reference Misbehavior (VeReMi) dataset for evaluation of our proposed model. The results show that our proposed LSTM and NB-based model outperforms centralized benchmark classification methods in malicious node detection. With an accuracy of 95%, the LSTM-based model demonstrates superior performance in identifying both malicious and legitimate vehicles, achieving a precision of 0.96 and a recall of 0.97. The high value of precision and recall shows that our model can efficiently discriminate between malicious and legitimate vehicles in the IoV network.","author":[{"family":"Bandarapu","given":"Srinivas"},{"family":"Bilal","given":"Muhammad"},{"family":"Chatterjee","given":"Pushpalika"},{"family":"Cheema","given":"Adnan"},{"family":"Rashid","given":"Junaid"},{"family":"Kim","given":"Jungeun"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s44443-025-00134-y","URL":"https://doi.org/10.1007/s44443-025-00134-y","source":"openalex"},{"id":"oa:W4413120471","type":"article-journal","title":"Realistic Urban Traffic Generator Using Decentralized Federated Learning for the SUMO Simulator","abstract":"Realistic urban traffic simulation is essential for sustainable urban planning and the development of intelligent transportation systems. However, generating high-fidelity, time-varying traffic profiles that accurately reflect real-world conditions, especially in large-scale scenarios, remains a major challenge. Existing methods often suffer from limitations in accuracy, scalability, or raise privacy concerns due to centralized data processing. This work introduces DesRUTGe (Decentralized Realistic Urban Traffic Generator), a novel framework that integrates Deep Reinforcement Learning (DRL) agents with the SUMO simulator to generate realistic 24-hour traffic patterns. A key innovation of DesRUTGe is its use of Decentralized Federated Learning (DFL), wherein each traffic detector and its corresponding urban zone function as an independent learning node. These nodes train local DRL models using minimal historical data and collaboratively refine their performance by exchanging model parameters with selected peers (e.g., geographically adjacent zones), without requiring a central coordinator. Evaluated using real-world data from the city of Barcelona, DesRUTGe outperforms standard SUMO-based tools such as RouteSampler, as well as other centralized learning approaches, by delivering more accurate traffic pattern generation.","author":[{"family":"Bazán-Guillén","given":"Alberto"},{"family":"Beis-Penedo","given":"Carlos"},{"family":"Cajaraville-Aboy","given":"Diego"},{"family":"Bautista","given":"Pablo"},{"family":"Redondo","given":"Rebeca"},{"family":"Llopis","given":"Luis"},{"family":"Vilas","given":"Ana"},{"family":"Igartua","given":"Mónica"},{"family":"Fernándezveiga","given":"Manuel"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/ojcoms.2025.3597019","URL":"https://doi.org/10.1109/ojcoms.2025.3597019","source":"openalex"},{"id":"oa:W4414281561","type":"article-journal","title":"Federated Learning with Adversarial Optimisation for Secure and Efficient 5G Edge Computing Networks","abstract":"With the evolution of 5G edge computing networks, privacy-aware applications are gaining significant attention due to their decentralised processing capabilities. However, these networks face substantial challenges to ensure privacy and security, specifically in a Federated Learning (FL) setup, where adversarial attacks can potentially influence the model integrity. Conventional privacy-preserving FL mechanisms are often susceptible to such attacks, leading to degraded model performance and severe security vulnerabilities. To address this issue, we propose FL with adversarial optimisation framework to improve adversarial robustness in 5G edge computing networks while ensuring privacy preservation. The proposed framework considers two models; a classifier model and an adversary model, where the classifier model is integrated with the adversary model, trained jointly considering Fast Gradient Sign Method (FGSM) for generation of adversarial perturbations. This adversarial optimisation enhances classifier’s resilience to attacks, thereby improving both privacy preservation and model accuracy. Experimental analysis reveals that the proposed model achieves up to 99.44% accuracy on adversarial test data, while improving robustness and sustaining high precision and recall across varying client scenarios. The experimental results further ensure the effectiveness of the proposed model in terms of communication efficiency and computational efficiency while reducing inference time and FLOPs making it ideal for secure 5G edge computing applications.","author":[{"family":"Zafar","given":"Saniya"},{"family":"White","given":"Jonathan"},{"family":"Legg","given":"Phil"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/bdcc9090238","URL":"https://doi.org/10.3390/bdcc9090238","source":"openalex"},{"id":"oa:W4415241570","type":"article-journal","title":"Federated Learning for Privacy-Preserving Apparel Supply Chain Analytics","abstract":"The apparel industry operates through highly complex and globalized supply chains, where effective data analytics plays a critical role in improving demand forecasting, inventory management, logistics coordination, and sustainability practices. However, organizations within the supply chain are often reluctant to share sensitive data due to concerns about privacy, security, compliance, and competitive risks. Traditional centralized analytics approaches exacerbate these concerns by requiring raw data aggregation, thereby increasing the likelihood of breaches and loss of confidentiality. Federated Learning (FL) has emerged as a transformative paradigm that addresses these challenges by enabling decentralized model training without the need to exchange raw data. In this study, we investigate the application of federated learning to apparel supply chain analytics, with a focus on balancing data utility and privacy preservation. We present a framework that integrates federated optimization, secure aggregation, and differential privacy to allow suppliers, manufacturers, distributors, and retailers to collaboratively train robust predictive models while maintaining strict data sovereignty. Our experimental evaluation demonstrates that federated models achieve comparable or superior forecasting accuracy relative to centralized approaches, while significantly reducing privacy risks. Moreover, results indicate notable improvements in demand forecasting, trend identification, and cost optimization tasks across heterogeneous datasets. By reducing data silos, federated learning fosters stronger collaboration, enhances supply chain resilience, and supports sustainability objectives. Overall, this work provides a practical pathway for implementing privacy-preserving analytics in apparel supply chains through federated learning.","author":[{"family":"Rahman","given":"Mizanur"},{"family":"Haque","given":"Samsul"},{"family":"Sany","given":"SMAA"}],"issued":{"date-parts":[[2025]]},"DOI":"10.30574/wjaets.2025.17.1.1386","URL":"https://doi.org/10.30574/wjaets.2025.17.1.1386","source":"openalex"},{"id":"oa:W4413291117","type":"article-journal","title":"Federated Learning for Medical Image Analysis: Privacy-Preserving Paradigms and Clinical Challenges","abstract":"Federated Learning (FL) has emerged as a transformative paradigm in medical image analysis, addressing the critical challenges of data scarcity and patient privacy. By enabling collaborative model training across decentralized datasets without requiring data sharing, FL aligns with stringent privacy regulations like HIPAA and GDPR. However, existing surveys on FL for medical image analysis often focus narrowly on aspects like privacy and security or fail to categorize methods within a clear taxonomy. Our survey bridges these gaps by systematically organizing FL methodologies for medical image analysis around three core pillars: training, architecture, and unlearning. We emphasize the unique demands of the medical domain, such as handling heterogeneous imaging modalities and annotations. Unlike prior works, our survey strikes a balance between technical rigor and clinical practicality, covering approaches not only for privacy and security but also for accuracy and efficiency. By synthesizing insights from various studies, we provide a comprehensive roadmap to guide researchers and practitioners in leveraging FL's potential to advance AI-driven healthcare.","author":[{"family":"Hu","given":"Juntao"},{"family":"Yang","given":"Zhengjie"},{"family":"Wang","given":"Peng"},{"family":"Zhao","given":"Guanyi"},{"family":"Huang","given":"Hong"},{"family":"Zong","given":"Zhimin"},{"family":"Wu","given":"Dapeng"}],"issued":{"date-parts":[[2025]]},"DOI":"10.53941/tai.2025.100010","URL":"https://doi.org/10.53941/tai.2025.100010","source":"openalex"},{"id":"oa:W4412987399","type":"article-journal","title":"Federated Learning for Early Cardiac Anomaly Prediction in Cross-Silo IoMT Environments","abstract":"Early detection of cardiovascular anomalies remains critical for proactive patient care, especially within the growing ecosystem of Internet of Medical Things (IoMT) devices. This study explores the application of Federated Learning (FL) to predict early cardiac events using electrocardiogram (ECG) signals across heterogeneous IoMT silos without centralized data sharing. We focus on Premature Ventricular Contraction (PVC) as an example of early event prediction. Using three realworld ECG datasets (PTB-XL, Chapman-Shaoxing, and MITBIH), we simulate cross-silo environments where local models are trained independently and aggregated through FL. Our experiments demonstrate that local models can already achieve high classification performance, but global models obtained via FL lead to consistent improvements in macro precision, recall, and F1-scores across datasets. Visual analysis of early ECG segments further highlights inter-dataset variability, emphasizing the importance of silo-specific characteristics. The results validate that FL is a promising strategy to enable scalable, privacypreserving, and accurate early cardiovascular event prediction in IoMT systems, bridging clinical silos while safeguarding sensitive patient data.","author":[{"family":"Georgiades","given":"Michael"},{"family":"Christodoulou","given":"Lakis"},{"family":"Chari","given":"Andreas"},{"family":"Wang","given":"Kezhi"},{"family":"Ho","given":"Kin"},{"family":"Hou","given":"Yun"},{"family":"Chai","given":"Wei"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/dcoss-iot65416.2025.00087","URL":"https://doi.org/10.1109/dcoss-iot65416.2025.00087","source":"openalex"},{"id":"oa:W4412599322","type":"article-journal","title":"An optimized oversampling-based federated transfer learning approach for rotating machinery cluster fault diagnosis","abstract":"Abstract With the development of distributed industrial systems, rotating machinery as the core power and transmission unit of complex distributed industrial systems, its fault diagnosis is very necessary and faces the serious challenge of Non-Independent and Identically Distributed (Non-IID). Although federated transfer learning (FTL) provides decentralized solutions, existing methods do not adequately address the poor classification results caused by data imbalance within the client. This study integrates optimized oversampling techniques into a federated transfer learning framework, proposes an optimized oversampling-based federated transfer learning approach. Firstly, annular region sample optimization (ARSO) is proposed to tackle ambiguous class boundaries from arbitrary sample selection in synthetic minority oversampling technique (SMOTE) by optimizing the sample selection strategy through annular regions. Then ARSO is integrated into a federated transfer learning framework with a One-Dimensional Convolutional Neural Network (1D-CNN), balance the amount of data among clients by extending a few classes of data before federated transfer learning, screen high-quality source clients for knowledge transfer based on a privacy-preserving transfer mechanism selects source clients via category-completeness metadata, and aligns domains using encrypted feature embeddings, proposed the annular region sample optimization federated transfer learning (ARSO-FTL). Experiments demonstrate ARSO-FTL achieves leading performance, recording 96.65% accuracy and an AUC of 0.96. It outperforms distributed baselines and effectively addresses intra-client imbalance and Non-IID challenges within federated transfer learning.","author":[{"family":"Xu","given":"Zhao"},{"family":"Yu","given":"Liya"},{"family":"Li","given":"Shaobo"},{"family":"Li","given":"Chuanjiang"},{"family":"Feng","given":"Yixiong"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1093/jcde/qwaf068","URL":"https://doi.org/10.1093/jcde/qwaf068","source":"openalex"},{"id":"oa:W4410192464","type":"article-journal","title":"Evaluating machine learning algorithms for energy consumption prediction in electric vehicles: A comparative study","abstract":"An accurate energy consumption prediction becomes crucial with increasing electric vehicle usage for effective power grid management. This research examined the performance of eleven machine learning models for this purpose: Ridge Regression, Lasso Regression, K-Nearest Neighbors, Gradient Boosting, Support Vector Regression, Multi-Layer Perceptron, XGBoost, CatBoost, LightGBM, Gaussian Processes for Regression(GPR) and Extra Trees Regressor, considering real historical data from Colorado. The models were evaluated using different metrics: Mean Absolute Error (MAE), Mean Squared Error (MSE), R², Root Mean Squared Error(RMSE) and Normalized Root Mean Squared Error(NRMSE), with visual analyses through scatter plots and time series plots. The best model observed was the Extra Trees Regressor, which had an MAE of 0.5888, an MSE of 3.2683, R² value of 0.9592, RMSE of 1.8078 and NRMSE of 0.020. Gradient Boosting and KNN also returned good results, although they were slightly more dispersed. Nevertheless, while non-linear models like MLP, XGBoost, CatBoost, LightGBM and linear models such as Ridge and Lasso Regression offer valuable insights, they exhibit shortcomings in estimating energy, especially at extreme levels, highlighting limitations in capturing complex non-linear interactions. This study focuses on their applicability to energy projections to demonstrate how well ensemble and non-linear models may capture intricate patterns in time series. These cutting-edge machine learning techniques might greatly enhance energy demand predictions.","author":[{"family":"Hussain","given":"Izhar"},{"family":"Ching","given":"Kok"},{"family":"Uttraphan","given":"Chessda"},{"family":"Tay","given":"Kim"},{"family":"Noor","given":"Adil"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-94946-7","URL":"https://doi.org/10.1038/s41598-025-94946-7","source":"openalex"},{"id":"oa:W7118956619","type":"article-journal","title":"SplitML: A Unified Privacy-Preserving Architecture for Federated Split-Learning in Heterogeneous Environments","abstract":"While Federated Learning (FL) and Split Learning (SL) aim to uphold data confidentiality by localized training, they remain susceptible to adversarial threats such as model poisoning and sophisticated inference attacks. To mitigate these vulnerabilities, we propose SplitML, a secure and privacy-preserving framework for Federated Split Learning (FSL). By integrating IND−CPAD secure Fully Homomorphic Encryption (FHE) with Differential Privacy (DP), SplitML establishes a defense-in-depth strategy that minimizes information leakage and thwarts reconstructive inference attempts. The framework accommodates heterogeneous model architectures by allowing clients to collaboratively train only the common top layers while keeping their bottom layers exclusive to each participant. This partitioning strategy ensures that the layers closest to the sensitive input data are never exposed to the centralized server. During the training phase, participants utilize multi-key CKKS FHE to facilitate secure weight aggregation, which ensures that no single entity can access individual updates in plaintext. For collaborative inference, clients exchange activations protected by single-key CKKS FHE to achieve a consensus derived from Total Labels (TL) or Total Predictions (TP). This consensus mechanism enhances decision reliability by aggregating decentralized insights while obfuscating soft-label confidence scores that could be exploited by attackers. Our empirical evaluation demonstrates that SplitML provides substantial defense against Membership Inference (MI) attacks, reduces temporal training costs compared to standard encrypted FL, and improves inference precision via its consensus mechanism, all while maintaining a negligible impact on federation overhead.","author":[{"family":"Trivedi","given":"Devharsh"},{"family":"Boudguiga","given":"Aymen"},{"family":"Kaaniche","given":"Nesrine"},{"family":"Triandopoulos","given":"Nikos"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/electronics15020267","URL":"https://doi.org/10.3390/electronics15020267","source":"openalex"},{"id":"oa:W4410748034","type":"article-journal","title":"FedCVG: a two-stage robust federated learning optimization algorithm","abstract":"Federated learning provides an effective solution to the data privacy issue in distributed machine learning. However, distributed federated learning systems are inherently susceptible to data poisoning attacks and data heterogeneity. Under conditions of high data heterogeneity, the gradient conflict problem in federated learning becomes more pronounced, making traditional defense mechanisms against poisoning attacks less adaptable between scenarios with and without attacks. To address this challenge, we design a two-stage federated learning framework for defending against poisoning attacks-FedCVG. During implementation, FedCVG first removes malicious clients using a reputation-based clustering method, and then optimizes communication overhead through a virtual aggregation mechanism. Extensive experimental results show that, compared to other baseline methods, FedCVG improves average accuracy by 4.2% and reduces communication overhead by approximately 50% while defending against poisoning attacks.","author":[{"family":"Zhang","given":"Runze"},{"family":"Zhang","given":"Yang"},{"family":"Zhao","given":"Yating"},{"family":"Jia","given":"Bin"},{"family":"Lian","given":"Wenjuan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-02722-4","URL":"https://doi.org/10.1038/s41598-025-02722-4","source":"openalex"},{"id":"oa:W4409409952","type":"article-journal","title":"Strategies to Improve the Robustness and Generalizability of Deep Learning Segmentation and Classification in Neuroimaging","abstract":"Artificial Intelligence (AI) and deep learning models have revolutionized diagnosis, prognostication, and treatment planning by extracting complex patterns from medical images, enabling more accurate, personalized, and timely clinical decisions. Despite its promise, challenges such as image heterogeneity across different centers, variability in acquisition protocols and scanners, and sensitivity to artifacts hinder the reliability and clinical integration of deep learning models. Addressing these issues is critical for ensuring accurate and practical AI-powered neuroimaging applications. We reviewed and summarized the strategies for improving the robustness and generalizability of deep learning models for the segmentation and classification of neuroimages. This review follows a structured protocol, comprehensively searching Google Scholar, PubMed, and Scopus for studies on neuroimaging, task-specific applications, and model attributes. Peer-reviewed, English-language studies on brain imaging were included. The extracted data were analyzed to evaluate the implementation and effectiveness of these techniques. The study identifies key strategies to enhance deep learning in neuroimaging, including regularization, data augmentation, transfer learning, and uncertainty estimation. These approaches address major challenges such as data variability and domain shifts, improving model robustness and ensuring consistent performance across diverse clinical settings. The technical strategies summarized in this review can enhance the robustness and generalizability of deep learning models for segmentation and classification to improve their reliability for real-world clinical practice.","author":[{"family":"Tran","given":"Anh"},{"family":"Zeevi","given":"Tal"},{"family":"Payabvash","given":"Seyedmehdi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/biomedinformatics5020020","URL":"https://doi.org/10.3390/biomedinformatics5020020","source":"openalex"},{"id":"oa:W7139924667","type":"article-journal","title":"A systematic literature review on federated cyber-attack detection for edge intelligence: Challenges, approaches, and future directions","abstract":"The rapid expansion of Edge Computing (EC) and Internet of Things devices has introduced significant cybersecurity challenges, necessitating advanced and privacy-preserving attack detection strategies. Traditional cyber-attack detection methods and centralized machine learning solutions face critical limitations in addressing privacy concerns, resource constraints, and the evolving nature of cyber threats in edge environments. Federated Learning (FL) offers a transformative solution by enabling distributed model training across edge devices while preserving data privacy. This systematic literature review investigates FL for cyber-attack detection in EC environments using the PRISMA methodology, analyzing 131 primary studies from 2020–2025 across five major databases. Our contributions include: (1) a comprehensive PRISMA-compliant framework encompassing seven thematic areas with detailed comparative analysis, (2) an in-depth gap analysis with actionable recommendations for privacy-performance trade-offs, scalability, and standardization challenges, and (3) a forward-looking research agenda addressing generative models, collaborative defense, 6G-enabled intelligence, and zero-trust architectures. Unlike existing surveys, this work provides the most comprehensive scope with bibliometric analysis, multi-perspective evaluation, and practical deployment guidelines, serving as a foundational reference for advancing federated cyber-attack detection in edge computing environments.","author":[{"family":"Sharmin","given":"Zeseya"},{"family":"Uddin","given":"Md"},{"family":"Xiang","given":"Yong"},{"family":"Chen","given":"Feifei"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.cosrev.2026.100965","URL":"https://doi.org/10.1016/j.cosrev.2026.100965","source":"openalex"},{"id":"oa:W4415688581","type":"article-journal","title":"Robust aggregation algorithms for federated learning in unreliable network environments","abstract":"Federated learning (FL) allows joint model training on distributed devices without losing data locality, but its results are significantly worse in unreliable network systems where packets are dropped, clients fail, resources are heterogeneous, and adversarial (Byzantine) agents exist. The viability of FL to withstand these unfavorable conditions is keyed on the robust aggregation algorithms. The paper meticulously examines powerful methods of aggregation, which include: geometric-median methods (RFA), Krum/Multi-Krum, trimmed-mean/coordinate-wise defenses, g-divergence estimators, trust-based aggregators (FLTrust), and layer-wise aggregation methods (FedRoLA) and compares their performance on simulated unreliable networks, which model packet loss, communication delay, and malicious client actions (McMahan et al., 2017; Blanchard et al., 2017; P We examine accuracy, convergence speed, communication cost and resilience in the face of model-poisoning attacks with the help of benchmark image tasks and a set of network unreliability scenarios. We find that robust aggregators combining statistical outlier resistance and structural (layer-wise) aggregation or trust calibration (especially RFA and FedRoLA) are more accurate and converge more quickly than naive FedAvg in high packet-loss and moderate Byzantine contamination and we observe up to 12 percent improvement in test accuracy with 30 percent simulated packet loss. We also talk about the trade offs between robustness, communication overhead and privacy (secure aggregation) and present a hybrid design pattern that incorporates robust aggregation and adaptive client selection with secure aggregation so as to address both unreliable links as well as adversarial updates. The results provide prescriptive advice on the use of FL in mobile, IoT, and vehicular networks that have limited reliability and security requirements and provide future directions such as privacy-conscious robust aggregation, fairness-conscious weighting, and testbed implementation.","author":[{"family":"Zeng","given":"Ziyang"},{"family":"Yang","given":"Shiyu"},{"family":"Ding","given":"Guanyu"}],"issued":{"date-parts":[[2025]]},"DOI":"10.54097/n0dpaf43","URL":"https://doi.org/10.54097/n0dpaf43","source":"openalex"},{"id":"oa:W4412947457","type":"article-journal","title":"Client to Server: Heterogeneous Distribution Knowledge Transfer for Federated Learning","abstract":"Federated learning (FL) is an emerging distributed machine learning paradigm that provides privacy guarantees for training robust models on distributed clients. The primary challenge of FL is data heterogeneity,which slows down model convergence and degrades model performance. Knowledge Distillation has recently demonstrated effectiveness in addressing this challenge. However, these approaches neglect the statistical heterogeneity in local models and the uncertainty of the data distribution in the global model, which results in the ensemble knowledge cannot be fully utilized to guide local model learning. In this work, we propose an unsupervised knowledge distillation method migrating the local class-level pseudo-data sample scheme in the server for fine-tuning the global model. Specifically, we provide the conditional autoencoder for each client to maintain a dynamic generator in the server, which ensembles the client’s class-level information. The proposal produces an auxiliary dataset representing the global class-level distribution to regulate the local model as an inductive knowledge bias and employs unsupervised knowledge distillation to enhance the aggregated model’s performance. The extensive experiments show that our proposal significantly outperforms the current state-of-theart FL algorithms and can be integrated as a flexible plugin into existing FL optimization algorithms to enhance model performance.","author":[{"family":"Zhao","given":"Rui"},{"family":"Yang","given":"Xiao"},{"family":"Zhi","given":"Peng"},{"family":"Zhang","given":"Zhihe"},{"family":"Liu","given":"Gang"},{"family":"Di","given":"Changyan"},{"family":"Zhou","given":"Qingguo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.26599/tst.2025.9010047","URL":"https://doi.org/10.26599/tst.2025.9010047","source":"openalex"},{"id":"oa:W4407237697","type":"article-journal","title":"A hybrid machine learning model for intrusion detection in wireless sensor networks leveraging data balancing and dimensionality reduction","abstract":"Intrusion detection systems are essential for securing wireless sensor networks (WSNs) and Internet of Things (IoT) environments against various threats. This study presents a novel hybrid machine learning (ML) model that integrates KMeans-SMOTE (KMS) for data balancing and principal component analysis (PCA) for dimensionality reduction, evaluated using the WSN-DS and TON-IoT datasets. The model employs classifiers such as Decision Tree Classifier, Random Forest Classifier (RFC), and gradient boosting techniques like XGBoost (XGBC) to enhance detection accuracy and efficiency. The proposed hybrid (KMS + PCA + RFC) approach achieves remarkable performance, with an accuracy of 99.94% and an f1-score of 99.94% on the WSN-DS dataset. For the TON-IoT dataset, it achieves 99.97% accuracy and an f1-score of 99.97%, outperforming traditional SMOTE TomekLink and Generative Adversarial Network-based data balancing techniques. This hybrid approach addresses class imbalance and high-dimensionality challenges, providing scalable and robust intrusion detection. Complexity analysis reveals that the proposed model reduces training and prediction times, making it suitable for real-time applications.","author":[{"family":"Talukder","given":"Md"},{"family":"Khalid","given":"Majdi"},{"family":"Sultana","given":"Nasrin"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-87028-1","URL":"https://doi.org/10.1038/s41598-025-87028-1","source":"openalex"},{"id":"oa:W4415481241","type":"article-journal","title":"ZTID-IoV: Zero-Trust Intrusion Detection in IoV Using Neurosymbolic AI Approach With Federated Meta-Learning","abstract":"The rapid growth of the Internet of Vehicles (IoVs) and smart consumer electronics has generated cybersecurity concerns that require an intelligent, adaptable, and privacy-preserving Intrusion Detection System (IDS). This study introduces ZTID-IoV, a novel neurosymbolic AI framework that integrates federated learning, lightweight transformers, and meta-learning to improve threat detection while preserving user privacy in consumer IoVs. Our approach leverages neural components such as a transformer model for recognizing patterns in network traffic, combined with symbolic AI techniques such as self-organizing maps for interpretable client clustering and rule-guided reasoning, to achieve robust cybersecurity in distributed environments. A lightweight transformer architecture optimizes performance for resource-constrained edge devices, and SOM-based clustering enhances model aggregation by grouping devices with similar behavioral patterns. The proposed system employs Model-Agnostic Meta-Learning (MAML) to enable rapid adaptation to emerging threats across diverse consumer devices, while federated learning ensures decentralized model training without exposing sensitive user data. Experiments on real-world IoT intrusion datasets demonstrate that our framework achieves higher detection accuracy compared to centralized and pure neural approaches while maintaining low computational overhead. Additionally, the neurosymbolic design provides interpretable threat explanations, crucial for consumer applications where transparency is essential. The results highlight the potential of ZTID-IoV in enabling zero-trust security for IoV and other connected consumer electronics. This work contributes to the evolving landscape of AI-driven cybersecurity by addressing critical challenges in privacy and adaptability, making ZTID-IoV particularly suitable for next-generation IoV ecosystems.","author":[{"family":"Ullah","given":"Farhan"},{"family":"Srivastava","given":"Gautam"},{"family":"Mostarda","given":"Leonardo"},{"family":"Raza","given":"Umar"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/tce.2025.3625081","URL":"https://doi.org/10.1109/tce.2025.3625081","source":"openalex"},{"id":"oa:W4408372351","type":"article-journal","title":"Using Homomorphic Proxy Re‐Encryption to Enhance Security and Privacy of Federated Learning‐Based Intelligent Connected Vehicles","abstract":"Intelligent connected vehicles (ICVs) are one of the fast‐growing directions that plays a significant role in the area of autonomous driving. To realize collaborative computation among ICVs, federated learning (FL) or federated‐based large language model (FedLLM) as a promising distributed approach has been used to support various collaborative application computations in ICVs scenarios, for example, analyzing vehicle driving information to realize trajectory prediction, voice‐activated controls, conversational AI assistants. Unfortunately, recent research reveals that FL systems are still faced with privacy challenges from honest‐but‐curious server, honest‐but‐curious distributed participants, or the collusion between participants and the server. These threats can lead to the leakage of sensitive private data, such as location information and driving conditions. Homomorphic encryption (HE) is one of the typical mitigation that has few effects on the model accuracy and has been studied before. However, single‐key HE cannot resist collusion between participants and the server, multikey HE is not suitable for ICVs scenarios. In this work, we proposed a novel approach that combines FL with homomorphic proxy re‐encryption (PRE) which is based on participants’ ID information. By doing so, the FL‐based ICVs can be able to successfully defend against privacy threats. In addition, we analyze the security and performance of our method, and the theoretical analysis and the experiment results show that our defense framework with ID‐based homomorphic PRE can achieve a high‐security level and efficient computation. We anticipate that our approach can serve as a fundamental point to support the extensive research on FedLLMs privacy‐preserving.","author":[{"family":"Bai","given":"Yang"},{"family":"Rao","given":"YS"},{"family":"Wu","given":"Hongyan"},{"family":"Wang","given":"Juan"},{"family":"Yang","given":"Wentao"},{"family":"Xing","given":"Gaojie"},{"family":"Yang","given":"Jiawei"},{"family":"Yuan","given":"Xiaoshu"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1049/ise2/4632786","URL":"https://doi.org/10.1049/ise2/4632786","source":"openalex"},{"id":"oa:W4412806734","type":"article-journal","title":"Breakthroughs in Brain Tumor Detection: Leveraging Deep Learning and Transfer Learning for MRI-Based Classification","abstract":"Identifying and classifying brain tumors play a pivotal role in gaining insights into their underlying mechanisms. In contemporary medical practice, the integration of Computer-assisted Diagnosis (CAD) and machine learning, particularly deep learning, has significantly enhanced the radiologist's ability to accurately identify brain tumors. Unlike traditional machine learning methods, which often rely on manual feature engineering for classification, deep learning models can be structured to prevent the need for manual feature extraction, yielding highly accurate classification outcomes. This paper customizes advanced deep learning models including VGG19, ResNet50, InceptionV3, and EfficientNetV2 as the most powerful deep learning models aimed at the identification of both binary (normal and abnormal) and multiclass: 17 classes including Glioma, Meningioma, Neurocytoma, and other types of injuries such as Abscesses and Cysts. We utilize a publicly available dataset containing 4449 MRI images. Subsequently, we conduct a comprehensive comparative analysis of our proposed models against existing models in the literature. Our experimental findings indicates that EfficientNetV2 outperforms other state-of-the-art deep-learning models.","author":[{"family":"Golkarieh","given":"Alireza"},{"family":"Boroujeni","given":"Sajjad"},{"family":"Kiashemshaki","given":"Kiana"},{"family":"Deldadehasl","given":"Maryam"},{"family":"Aghayarzadeh","given":"Hamed"},{"family":"Ramezani","given":"Azita"}],"issued":{"date-parts":[[2025]]},"DOI":"10.59543/comdem.v2i.14243","URL":"https://doi.org/10.59543/comdem.v2i.14243","source":"openalex"},{"id":"oa:W4415594398","type":"article-journal","title":"RewardChain: A Blockchain-Based Incentive Mechanism for Federated Learning in Consumer-Centric Internet of Medical Things","abstract":"Federated learning is a promising approach that enables collaborative machine learning (ML) in distributed environments, such as the Internet of Medical Things (IoMT) while preserving consumer privacy. It allows multiple consumers to collaboratively train a model using their own data, sharing only the locally trained model rather than the raw data. Most existing federated learning systems assume a high level of trust in participating nodes, which is unrealistic in real-world consumer-centric scenarios. Involving untrusted nodes can compromise the integrity of the training process and result in potential data breaches. To address these challenges, this paper presents REWARDCHAIN, a novel federated learning framework that leverages blockchain technology to ensure trust and accountability among untrusted IoMT consumers. By recording all model updates and client contributions on an immutable blockchain ledger, REWARDCHAIN allows auditing of the entire training process and attributing any malicious behaviour to specific nodes. Moreover, we design an incentive mechanism that evaluates contributions based on data quality and participant reputation. This system motivates participants to contribute high-quality data through a reputation-constrained reward allocation. Our evaluations show that REWARDCHAIN effectively balances trust, security, and model performance, facilitating a more secure and effective federated learning ecosystem.","author":[{"family":"Alsharidah","given":"Ahmad"},{"family":"Jha","given":"Devki"},{"family":"Solaiman","given":"Ellis"},{"family":"Wei","given":"Bo"},{"family":"Aujla","given":"Gagangeet"},{"family":"Ranjan","given":"Rajiv"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/tce.2025.3626199","URL":"https://doi.org/10.1109/tce.2025.3626199","source":"openalex"},{"id":"oa:W4415648046","type":"article-journal","title":"A Multi-View-Based Federated Learning Approach for Intrusion Detection","abstract":"Intrusion detection aims to identify the unauthorized activities within computer networks or systems by classifying events into normal or abnormal categories. As modern scenarios often involve multi-source data, multi-view fusion deep learning methods are employed to leverage diverse viewpoints for enhancing security threat detection. This paper introduces a novel intrusion detection approach using multi-view fusion within a federated learning framework, proposing an integrated AE Neural SVM (AE-NSVM) model that combines auto-encoder (AE) multi-view feature extraction and Support Vector Machine (SVM) classification. This approach simultaneously learns representative features from multiple views and classifies network samples into normal or seven attack categories while employing federated learning across clients to ensure adaptability and robustness in diverse network environments. The experimental results obtained from two benchmark datasets validate its superiority: on TON_IoT, the CAE-NSVM model achieves a highest F1-measure of 0.792 (1.4% higher than traditional pipeline systems); on UNSW-NB15, it delivers an F1-score of 0.829 with a 73% reduced training time and an 89% faster inference compared to baseline models. These results demonstrate the advantages of multi-view fusion in federated learning for balancing accuracy and efficiency in distributed intrusion detection systems.","author":[{"family":"Yu","given":"Jia"},{"family":"Wang","given":"Guoqiang"},{"family":"Shi","given":"Nianfeng"},{"family":"Saxena","given":"Raghav"},{"family":"Lee","given":"Brian"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/electronics14214166","URL":"https://doi.org/10.3390/electronics14214166","source":"openalex"},{"id":"doi:10.5281/zenodo.21552485","type":"article-journal","title":"A Conceptual Model for Intelligent Automation Loops in High-Throughput, Multi-Phase Crude Processing Units","abstract":"The increasing complexity and scale of high-throughput, multi-phase crude processing units (CUs) in modern refineries necessitate the development of advanced automation strategies beyond conventional control loops. Traditional systems, typically based on fixed-logic PID controllers and static setpoint optimization, are often inadequate in dealing with nonlinear process dynamics, rapid phase changes, and operational variability inherent in multi-phase crude streams. This proposes a conceptual model for intelligent automation loops that leverage real-time data, adaptive control strategies, and artificial intelligence (AI) to optimize the operation of such units. The proposed model integrates several key components: a distributed sensor network for high-fidelity, multi-phase flow data acquisition; dynamic process modeling using hybrid techniques (first-principles and machine learning); and intelligent control algorithms capable of self-tuning, fault detection, and optimization under changing operating conditions. Digital twin technology is incorporated to simulate and validate control actions in real-time, while edge computing infrastructure supports low-latency decision-making and reduces reliance on centralized systems. Use cases such as adaptive separator control, slug flow management, and real-time energy efficiency optimization are used to demonstrate the model's applicability and effectiveness. The intelligent loop framework enables predictive behavior, reduces manual intervention, and enhances operational resilience leading to reduced downtime, improved throughput, and better resource utilization. Importantly, the model supports integration with existing Distributed Control Systems (DCS) and Supervisory Control and Data Acquisition (SCADA) systems, enabling phased implementation and minimal disruption to ongoing operations. This conceptual framework addresses key challenges in refinery automation, including legacy integration, cybersecurity, and data quality. Future research will focus on full autonomy, federated AI deployment across refinery assets, and standardization of intelligent loop architectures. The proposed model represents a strategic step toward smarter, safer, and more efficient crude processing in the era of Industry 4.0.","author":[{"family":"Ofoedu","given":"Andrew"},{"family":"Ozor","given":"Joshua"},{"family":"Sofoluwe","given":"Oludayo"},{"family":"Jambol","given":"Dazok"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.21552485","URL":"https://doi.org/10.5281/zenodo.21552485","source":"datacite"},{"id":"doi:10.5281/zenodo.21552486","type":"article-journal","title":"A Conceptual Model for Intelligent Automation Loops in High-Throughput, Multi-Phase Crude Processing Units","abstract":"The increasing complexity and scale of high-throughput, multi-phase crude processing units (CUs) in modern refineries necessitate the development of advanced automation strategies beyond conventional control loops. Traditional systems, typically based on fixed-logic PID controllers and static setpoint optimization, are often inadequate in dealing with nonlinear process dynamics, rapid phase changes, and operational variability inherent in multi-phase crude streams. This proposes a conceptual model for intelligent automation loops that leverage real-time data, adaptive control strategies, and artificial intelligence (AI) to optimize the operation of such units. The proposed model integrates several key components: a distributed sensor network for high-fidelity, multi-phase flow data acquisition; dynamic process modeling using hybrid techniques (first-principles and machine learning); and intelligent control algorithms capable of self-tuning, fault detection, and optimization under changing operating conditions. Digital twin technology is incorporated to simulate and validate control actions in real-time, while edge computing infrastructure supports low-latency decision-making and reduces reliance on centralized systems. Use cases such as adaptive separator control, slug flow management, and real-time energy efficiency optimization are used to demonstrate the model's applicability and effectiveness. The intelligent loop framework enables predictive behavior, reduces manual intervention, and enhances operational resilience leading to reduced downtime, improved throughput, and better resource utilization. Importantly, the model supports integration with existing Distributed Control Systems (DCS) and Supervisory Control and Data Acquisition (SCADA) systems, enabling phased implementation and minimal disruption to ongoing operations. This conceptual framework addresses key challenges in refinery automation, including legacy integration, cybersecurity, and data quality. Future research will focus on full autonomy, federated AI deployment across refinery assets, and standardization of intelligent loop architectures. The proposed model represents a strategic step toward smarter, safer, and more efficient crude processing in the era of Industry 4.0.","author":[{"family":"Ofoedu","given":"Andrew"},{"family":"Ozor","given":"Joshua"},{"family":"Sofoluwe","given":"Oludayo"},{"family":"Jambol","given":"Dazok"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.21552486","URL":"https://doi.org/10.5281/zenodo.21552486","source":"datacite"},{"id":"doi:10.48550/arxiv.2401.04472","type":"manuscript","title":"A Survey on Efficient Federated Learning Methods for Foundation Model Training","abstract":"Federated Learning (FL) has become an established technique to facilitate privacy-preserving collaborative training across a multitude of clients. However, new approaches to FL often discuss their contributions involving small deep-learning models only and focus on training full models on clients. In the wake of Foundation Models (FM), the reality is different for many deep learning applications. Typically, FMs have already been pre-trained across a wide variety of tasks and can be fine-tuned to specific downstream tasks over significantly smaller datasets than required for full model training. However, access to such datasets is often challenging. By its design, FL can help to open data silos. With this survey, we introduce a novel taxonomy focused on computational and communication efficiency, the vital elements to make use of FMs in FL systems. We discuss the benefits and drawbacks of parameter-efficient fine-tuning (PEFT) for FL applications, elaborate on the readiness of FL frameworks to work with FMs, and provide future research opportunities on how to evaluate generative models in FL as well as the interplay of privacy and PEFT.","author":[{"family":"Woisetschläger","given":"Herbert"},{"family":"Isenko","given":"Alexander"},{"family":"Wang","given":"Shiqiang"},{"family":"Mayer","given":"Ruben"},{"family":"Jacobsen","given":"Hans"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2401.04472","URL":"https://doi.org/10.48550/arxiv.2401.04472","source":"datacite"},{"id":"doi:10.48550/arxiv.2311.18741","type":"manuscript","title":"VREM-FL: Mobility-Aware Computation-Scheduling Co-Design for Vehicular Federated Learning","abstract":"Assisted and autonomous driving are rapidly gaining momentum and will soon become a reality. Artificial intelligence and machine learning are regarded as key enablers thanks to the massive amount of data that smart vehicles will collect from onboard sensors. Federated learning is one of the most promising techniques for training global machine learning models while preserving data privacy of vehicles and optimizing communications resource usage. In this article, we propose vehicular radio environment map federated learning (VREM-FL), a computation-scheduling co-design for vehicular federated learning that combines mobility of vehicles with 5G radio environment maps. VREM-FL jointly optimizes learning performance of the global model and wisely allocates communication and computation resources. This is achieved by orchestrating local computations at the vehicles in conjunction with transmission of their local models in an adaptive and predictive fashion, by exploiting radio channel maps. The proposed algorithm can be tuned to trade training time for radio resource usage. Experimental results demonstrate that VREM-FL outperforms literature benchmarks for both a linear regression model (learning time reduced by 28%) and a deep neural network for semantic image segmentation (doubling the number of model updates within the same time window).","author":[{"family":"Ballotta","given":"Luca"},{"family":"Fabbro","given":"Nicolò"},{"family":"Perin","given":"Giovanni"},{"family":"Schenato","given":"Luca"},{"family":"Rossi","given":"Michele"},{"family":"Piro","given":"Giuseppe"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2311.18741","URL":"https://doi.org/10.48550/arxiv.2311.18741","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.15632","type":"manuscript","title":"Federated Behavioural Planes: Explaining the Evolution of Client Behaviour in Federated Learning","abstract":"Federated Learning (FL), a privacy-aware approach in distributed deep learning environments, enables many clients to collaboratively train a model without sharing sensitive data, thereby reducing privacy risks. However, enabling human trust and control over FL systems requires understanding the evolving behaviour of clients, whether beneficial or detrimental for the training, which still represents a key challenge in the current literature. To address this challenge, we introduce Federated Behavioural Planes (FBPs), a novel method to analyse, visualise, and explain the dynamics of FL systems, showing how clients behave under two different lenses: predictive performance (error behavioural space) and decision-making processes (counterfactual behavioural space). Our experiments demonstrate that FBPs provide informative trajectories describing the evolving states of clients and their contributions to the global model, thereby enabling the identification of clusters of clients with similar behaviours. Leveraging the patterns identified by FBPs, we propose a robust aggregation technique named Federated Behavioural Shields to detect malicious or noisy client models, thereby enhancing security and surpassing the efficacy of existing state-of-the-art FL defense mechanisms. Our code is publicly available on GitHub.","author":[{"family":"Fenoglio","given":"Dario"},{"family":"Dominici","given":"Gabriele"},{"family":"Barbiero","given":"Pietro"},{"family":"Tonda","given":"Alberto"},{"family":"Gjoreski","given":"Martin"},{"family":"Langheinrich","given":"Marc"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.15632","URL":"https://doi.org/10.48550/arxiv.2405.15632","source":"datacite"},{"id":"doi:10.48550/arxiv.2410.03070","type":"manuscript","title":"FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization","abstract":"Federated Learning (FL) is a method for training machine learning models using distributed data sources. It ensures privacy by allowing clients to collaboratively learn a shared global model while storing their data locally. However, a significant challenge arises when dealing with missing modalities in clients' datasets, where certain features or modalities are unavailable or incomplete, leading to heterogeneous data distribution. While previous studies have addressed the issue of complete-modality missing, they fail to tackle partial-modality missing on account of severe heterogeneity among clients at an instance level, where the pattern of missing data can vary significantly from one sample to another. To tackle this challenge, this study proposes a novel framework named FedMAC, designed to address multi-modality missing under conditions of partial-modality missing in FL. Additionally, to avoid trivial aggregation of multi-modal features, we introduce contrastive-based regularization to impose additional constraints on the latent representation space. The experimental results demonstrate the effectiveness of FedMAC across various client configurations with statistical heterogeneity, outperforming baseline methods by up to 26% in severe missing scenarios, highlighting its potential as a solution for the challenge of partially missing modalities in federated systems. Our source code is provided at https://github.com/nmduonggg/PEPSY","author":[{"family":"Nguyen","given":"Manh"},{"family":"Nguyen","given":"Trung"},{"family":"Pham","given":"Huy"},{"family":"Hoang","given":"Trong"},{"family":"Nguyen","given":"Phi"},{"family":"Huynh","given":"Thanh"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2410.03070","URL":"https://doi.org/10.48550/arxiv.2410.03070","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.20821","type":"manuscript","title":"Pursuing Overall Welfare in Federated Learning through Sequential Decision Making","abstract":"In traditional federated learning, a single global model cannot perform equally well for all clients. Therefore, the need to achieve the client-level fairness in federated system has been emphasized, which can be realized by modifying the static aggregation scheme for updating the global model to an adaptive one, in response to the local signals of the participating clients. Our work reveals that existing fairness-aware aggregation strategies can be unified into an online convex optimization framework, in other words, a central server's sequential decision making process. To enhance the decision making capability, we propose simple and intuitive improvements for suboptimal designs within existing methods, presenting AAggFF. Considering practical requirements, we further subdivide our method tailored for the cross-device and the cross-silo settings, respectively. Theoretical analyses guarantee sublinear regret upper bounds for both settings: $\\mathcal{O}(\\sqrt{T \\log{K}})$ for the cross-device setting, and $\\mathcal{O}(K \\log{T})$ for the cross-silo setting, with $K$ clients and $T$ federation rounds. Extensive experiments demonstrate that the federated system equipped with AAggFF achieves better degree of client-level fairness than existing methods in both practical settings. Code is available at https://github.com/vaseline555/AAggFF","author":[{"family":"Hahn","given":"Seok"},{"family":"Kim","given":"Gi"},{"family":"Lee","given":"Junghye"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.20821","URL":"https://doi.org/10.48550/arxiv.2405.20821","source":"datacite"},{"id":"doi:10.48550/arxiv.2407.15402","type":"manuscript","title":"Tackling Selfish Clients in Federated Learning","abstract":"Federated Learning (FL) is a distributed machine learning paradigm facilitating participants to collaboratively train a model without revealing their local data. However, when FL is deployed into the wild, some intelligent clients can deliberately deviate from the standard training process to make the global model inclined toward their local model, thereby prioritizing their local data distribution. We refer to this novel category of misbehaving clients as selfish. In this paper, we propose a Robust aggregation strategy for FL server to mitigate the effect of Selfishness (in short RFL-Self). RFL-Self incorporates an innovative method to recover (or estimate) the true updates of selfish clients from the received ones, leveraging robust statistics (median of norms) of the updates at every round. By including the recovered updates in aggregation, our strategy offers strong robustness against selfishness. Our experimental results, obtained on MNIST and CIFAR-10 datasets, demonstrate that just 2% of clients behaving selfishly can decrease the accuracy by up to 36%, and RFL-Self can mitigate that effect without degrading the global model performance.","author":[{"family":"Augello","given":"Andrea"},{"family":"Gupta","given":"Ashish"},{"family":"Re","given":"Giuseppe"},{"family":"Das","given":"Sajal"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2407.15402","URL":"https://doi.org/10.48550/arxiv.2407.15402","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.06312","type":"manuscript","title":"FedGCS: A Generative Framework for Efficient Client Selection in Federated Learning via Gradient-based Optimization","abstract":"Federated Learning faces significant challenges in statistical and system heterogeneity, along with high energy consumption, necessitating efficient client selection strategies. Traditional approaches, including heuristic and learning-based methods, fall short of addressing these complexities holistically. In response, we propose FedGCS, a novel generative client selection framework that innovatively recasts the client selection process as a generative task. Drawing inspiration from the methodologies used in large language models, FedGCS efficiently encodes abundant decision-making knowledge within a continuous representation space, enabling efficient gradient-based optimization to search for optimal client selection that will be finally output via generation. The framework comprises four steps: (1) automatic collection of diverse \"selection-score\" pair data using classical client selection methods; (2) training an encoder-evaluator-decoder framework on this data to construct a continuous representation space; (3) employing gradient-based optimization in this space for optimal client selection; (4) generating the final optimal client selection via using beam search for the well-trained decoder. FedGCS outperforms traditional methods by being more comprehensive, generalizable, and efficient, simultaneously optimizing for model performance, latency, and energy consumption. The effectiveness of FedGCS is proven through extensive experimental analyses.","author":[{"family":"Ning","given":"Zhiyuan"},{"family":"Tian","given":"Chunlin"},{"family":"Xiao","given":"Meng"},{"family":"Fan","given":"Wei"},{"family":"Wang","given":"Pengyang"},{"family":"Li","given":"Li"},{"family":"Wang","given":"Pengfei"},{"family":"Zhou","given":"Yuanchun"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.06312","URL":"https://doi.org/10.48550/arxiv.2405.06312","source":"datacite"},{"id":"doi:10.48550/arxiv.2409.06067","type":"manuscript","title":"MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning","abstract":"Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In light of the recent advances in multimodal large language models (MLLMs), such as GPT-4v and LLaVA, which demonstrate their exceptional proficiency in multimodal tasks, such as image captioning and multimodal question answering. We introduce a novel federated learning framework, named Multimodal Large Language Model Assisted Federated Learning (MLLM-LLaVA-FL), which employs powerful MLLMs at the server end to address the heterogeneous and long-tailed challenges. Owing to the advanced cross-modality representation capabilities and the extensive open-vocabulary prior knowledge of MLLMs, our framework is adept at harnessing the extensive, yet previously underexploited, open-source data accessible from websites and powerful server-side computational resources. Hence, the MLLM-LLaVA-FL not only enhances the performance but also avoids increasing the risk of privacy leakage and the computational burden on local devices, distinguishing it from prior methodologies. Our framework has three key stages. Initially, we conduct global visual-text pretraining of the model. This pretraining is facilitated by utilizing the extensive open-source data available online, with the assistance of MLLMs. Subsequently, the pretrained model is distributed among various clients for local training. Finally, once the locally trained models are transmitted back to the server, a global alignment is carried out under the supervision of MLLMs to further enhance the performance. Experimental evaluations on established benchmarks, show that our framework delivers promising performance in the typical scenarios with data heterogeneity and long-tail distribution across different clients in FL.","author":[{"family":"Zhang","given":"Jianyi"},{"family":"Yang","given":"Hao"},{"family":"Li","given":"Ang"},{"family":"Guo","given":"Xin"},{"family":"Wang","given":"Pu"},{"family":"Wang","given":"Haiming"},{"family":"Chen","given":"Yiran"},{"family":"Li","given":"Hai"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2409.06067","URL":"https://doi.org/10.48550/arxiv.2409.06067","source":"datacite"},{"id":"doi:10.48550/arxiv.2411.02115","type":"manuscript","title":"FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation","abstract":"Federated learning (FL) is a collaborative machine learning approach that enables multiple clients to train models without sharing their private data. With the rise of deep learning, large-scale models have garnered significant attention due to their exceptional performance. However, a key challenge in FL is the limitation imposed by clients with constrained computational and communication resources, which hampers the deployment of these large models. The Mixture of Experts (MoE) architecture addresses this challenge with its sparse activation property, which reduces computational workload and communication demands during inference and updates. Additionally, MoE facilitates better personalization by allowing each expert to specialize in different subsets of the data distribution. To alleviate the communication burdens between the server and clients, we propose FedMoE-DA, a new FL model training framework that leverages the MoE architecture and incorporates a novel domain-aware, fine-grained aggregation strategy to enhance the robustness, personalizability, and communication efficiency simultaneously. Specifically, the correlation between both intra-client expert models and inter-client data heterogeneity is exploited. Moreover, we utilize peer-to-peer (P2P) communication between clients for selective expert model synchronization, thus significantly reducing the server-client transmissions. Experiments demonstrate that our FedMoE-DA achieves excellent performance while reducing the communication pressure on the server.","author":[{"family":"Zhan","given":"Ziwei"},{"family":"Zhao","given":"Wenkuan"},{"family":"Li","given":"Yuanqing"},{"family":"Liu","given":"Weijie"},{"family":"Zhang","given":"Xiaoxi"},{"family":"Tan","given":"Chee"},{"family":"Wu","given":"Chuan"},{"family":"Guo","given":"Deke"},{"family":"Chen","given":"Xu"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2411.02115","URL":"https://doi.org/10.48550/arxiv.2411.02115","source":"datacite"},{"id":"doi:10.7302/25148","type":"article-journal","title":"Data Harmonization, Standardization, and Collaboration for Diabetic Retinal Disease (DRD) Research: Report From the 2024 Mary Tyler Moore Vision Initiative Workshop on Data","abstract":"The 2024 Mary Tyler Moore Vision Initiative (MTM Vision) Workshop on Data convened to discuss best practices and specific considerations for building a comprehensive, shareable MTM Vision data lake. The workshop aimed to accelerate the development of new indications, therapies, and regulatory pathways for diabetic retinal disease (DRD) by standardizing and harmonizing clinical data and ocular ’omics analyses. Standardization of data collection, the use of common data elements, and data interoperability were emphasized, alongside federated learning approaches to promote data sharing and collaboration while maintaining data privacy and security. The integration of molecular data with other multimodal data types was recognized as a promising strategy for leveraging machine learning and AI approaches to advancing therapeutics development and improving treatment outcomes for DRD patients. Partnerships with entities such as the National Eye Institute, part of the National Institutes of Health, foundations, and industry were deemed vital for the successful implementation of these initiatives.","author":[{"family":"Domalpally","given":"A"},{"family":"Fickweiler","given":"W"},{"family":"Levine"},{"family":"Goetz","given":"Ke"},{"family":"Vanderbeek","given":"Bl"},{"family":"Lee","given":"A"},{"family":"Sundstrom","given":"Jm"},{"family":"Markel","given":"D"},{"family":"Sun","given":"Jk"}],"issued":{"date-parts":[[2024]]},"DOI":"10.7302/25148","URL":"https://doi.org/10.7302/25148","source":"datacite"},{"id":"doi:10.48550/arxiv.2408.01765","type":"manuscript","title":"Joint Model Pruning and Resource Allocation for Wireless Time-triggered Federated Learning","abstract":"Time-triggered federated learning, in contrast to conventional event-based federated learning, organizes users into tiers based on fixed time intervals. However, this network still faces challenges due to a growing number of devices and limited wireless bandwidth, increasing issues like stragglers and communication overhead. In this paper, we apply model pruning to wireless Time-triggered systems and jointly study the problem of optimizing the pruning ratio and bandwidth allocation to minimize training loss under communication latency constraints. To solve this joint optimization problem, we perform a convergence analysis on the gradient $l_2$-norm of the asynchronous multi-tier federated learning (FL) model with adaptive model pruning. The convergence upper bound is derived and a joint optimization problem of pruning ratio and wireless bandwidth is defined to minimize the model training loss under a given communication latency constraint. The closed-form solutions for wireless bandwidth and pruning ratio by using KKT conditions are then formulated. As indicated in the simulation experiments, our proposed TT-Prune demonstrates a 40% reduction in communication cost, compared with the asynchronous multi-tier FL without model pruning, while maintaining the model convergence at the same level.","author":[{"family":"Zhang","given":"Xinlu"},{"family":"Deng","given":"Yansha"},{"family":"Mahmoodi","given":"Toktam"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2408.01765","URL":"https://doi.org/10.48550/arxiv.2408.01765","source":"datacite"},{"id":"doi:10.48550/arxiv.2402.10092","type":"manuscript","title":"Workflow Optimization for Parallel Split Learning","abstract":"Split learning (SL) has been recently proposed as a way to enable resource-constrained devices to train multi-parameter neural networks (NNs) and participate in federated learning (FL). In a nutshell, SL splits the NN model into parts, and allows clients (devices) to offload the largest part as a processing task to a computationally powerful helper. In parallel SL, multiple helpers can process model parts of one or more clients, thus, considerably reducing the maximum training time over all clients (makespan). In this paper, we focus on orchestrating the workflow of this operation, which is critical in highly heterogeneous systems, as our experiments show. In particular, we formulate the joint problem of client-helper assignments and scheduling decisions with the goal of minimizing the training makespan, and we prove that it is NP-hard. We propose a solution method based on the decomposition of the problem by leveraging its inherent symmetry, and a second one that is fully scalable. A wealth of numerical evaluations using our testbed's measurements allow us to build a solution strategy comprising these methods. Moreover, we show that this strategy finds a near-optimal solution, and achieves a shorter makespan than the baseline scheme by up to 52.3%.","author":[{"family":"Tirana","given":"Joana"},{"family":"Tsigkari","given":"Dimitra"},{"family":"Iosifidis","given":"George"},{"family":"Chatzopoulos","given":"Dimitris"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2402.10092","URL":"https://doi.org/10.48550/arxiv.2402.10092","source":"datacite"},{"id":"doi:10.48550/arxiv.2401.12012","type":"manuscript","title":"TurboSVM-FL: Boosting Federated Learning through SVM Aggregation for Lazy Clients","abstract":"Federated learning is a distributed collaborative machine learning paradigm that has gained strong momentum in recent years. In federated learning, a central server periodically coordinates models with clients and aggregates the models trained locally by clients without necessitating access to local data. Despite its potential, the implementation of federated learning continues to encounter several challenges, predominantly the slow convergence that is largely due to data heterogeneity. The slow convergence becomes particularly problematic in cross-device federated learning scenarios where clients may be strongly limited by computing power and storage space, and hence counteracting methods that induce additional computation or memory cost on the client side such as auxiliary objective terms and larger training iterations can be impractical. In this paper, we propose a novel federated aggregation strategy, TurboSVM-FL, that poses no additional computation burden on the client side and can significantly accelerate convergence for federated classification task, especially when clients are \"lazy\" and train their models solely for few epochs for next global aggregation. TurboSVM-FL extensively utilizes support vector machine to conduct selective aggregation and max-margin spread-out regularization on class embeddings. We evaluate TurboSVM-FL on multiple datasets including FEMNIST, CelebA, and Shakespeare using user-independent validation with non-iid data distribution. Our results show that TurboSVM-FL can significantly outperform existing popular algorithms on convergence rate and reduce communication rounds while delivering better test metrics including accuracy, F1 score, and MCC.","author":[{"family":"Wang","given":"Mengdi"},{"family":"Bodonhelyi","given":"Anna"},{"family":"Bozkir","given":"Efe"},{"family":"Kasneci","given":"Enkelejda"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2401.12012","URL":"https://doi.org/10.48550/arxiv.2401.12012","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.15474","type":"manuscript","title":"Unlearning during Learning: An Efficient Federated Machine Unlearning Method","abstract":"In recent years, Federated Learning (FL) has garnered significant attention as a distributed machine learning paradigm. To facilitate the implementation of the right to be forgotten, the concept of federated machine unlearning (FMU) has also emerged. However, current FMU approaches often involve additional time-consuming steps and may not offer comprehensive unlearning capabilities, which renders them less practical in real FL scenarios. In this paper, we introduce FedAU, an innovative and efficient FMU framework aimed at overcoming these limitations. Specifically, FedAU incorporates a lightweight auxiliary unlearning module into the learning process and employs a straightforward linear operation to facilitate unlearning. This approach eliminates the requirement for extra time-consuming steps, rendering it well-suited for FL. Furthermore, FedAU exhibits remarkable versatility. It not only enables multiple clients to carry out unlearning tasks concurrently but also supports unlearning at various levels of granularity, including individual data samples, specific classes, and even at the client level. We conducted extensive experiments on MNIST, CIFAR10, and CIFAR100 datasets to evaluate the performance of FedAU. The results demonstrate that FedAU effectively achieves the desired unlearning effect while maintaining model accuracy. Our code is availiable at https://github.com/Liar-Mask/FedAU.","author":[{"family":"Gu","given":"Hanlin"},{"family":"Zhu","given":"Gongxi"},{"family":"Zhang","given":"Jie"},{"family":"Zhao","given":"Xinyuan"},{"family":"Han","given":"Yuxing"},{"family":"Fan","given":"Lixin"},{"family":"Yang","given":"Qiang"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.15474","URL":"https://doi.org/10.48550/arxiv.2405.15474","source":"datacite"},{"id":"doi:10.48550/arxiv.2406.08267","type":"manuscript","title":"A deep cut into Split Federated Self-supervised Learning","abstract":"Collaborative self-supervised learning has recently become feasible in highly distributed environments by dividing the network layers between client devices and a central server. However, state-of-the-art methods, such as MocoSFL, are optimized for network division at the initial layers, which decreases the protection of the client data and increases communication overhead. In this paper, we demonstrate that splitting depth is crucial for maintaining privacy and communication efficiency in distributed training. We also show that MocoSFL suffers from a catastrophic quality deterioration for the minimal communication overhead. As a remedy, we introduce Momentum-Aligned contrastive Split Federated Learning (MonAcoSFL), which aligns online and momentum client models during training procedure. Consequently, we achieve state-of-the-art accuracy while significantly reducing the communication overhead, making MonAcoSFL more practical in real-world scenarios.","author":[{"family":"Przewięźlikowski","given":"Marcin"},{"family":"Osial","given":"Marcin"},{"family":"Zieliński","given":"Bartosz"},{"family":"Śmieja","given":"Marek"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2406.08267","URL":"https://doi.org/10.48550/arxiv.2406.08267","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.11525","type":"manuscript","title":"Overcoming Data and Model Heterogeneities in Decentralized Federated Learning via Synthetic Anchors","abstract":"Conventional Federated Learning (FL) involves collaborative training of a global model while maintaining user data privacy. One of its branches, decentralized FL, is a serverless network that allows clients to own and optimize different local models separately, which results in saving management and communication resources. Despite the promising advancements in decentralized FL, it may reduce model generalizability due to lacking a global model. In this scenario, managing data and model heterogeneity among clients becomes a crucial problem, which poses a unique challenge that must be overcome: How can every client's local model learn generalizable representation in a decentralized manner? To address this challenge, we propose a novel Decentralized FL technique by introducing Synthetic Anchors, dubbed as DeSA. Based on the theory of domain adaptation and Knowledge Distillation (KD), we theoretically and empirically show that synthesizing global anchors based on raw data distribution facilitates mutual knowledge transfer. We further design two effective regularization terms for local training: 1) REG loss that regularizes the distribution of the client's latent embedding with the anchors and 2) KD loss that enables clients to learn from others. Through extensive experiments on diverse client data distributions, we showcase the effectiveness of DeSA in enhancing both inter- and intra-domain accuracy of each client.","author":[{"family":"Huang","given":"Chun"},{"family":"Srinivas","given":"Kartik"},{"family":"Zhang","given":"Xin"},{"family":"Li","given":"Xiaoxiao"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.11525","URL":"https://doi.org/10.48550/arxiv.2405.11525","source":"datacite"},{"id":"doi:10.48550/arxiv.2410.08892","type":"manuscript","title":"Federated Learning in Practice: Reflections and Projections","abstract":"Federated Learning (FL) is a machine learning technique that enables multiple entities to collaboratively learn a shared model without exchanging their local data. Over the past decade, FL systems have achieved substantial progress, scaling to millions of devices across various learning domains while offering meaningful differential privacy (DP) guarantees. Production systems from organizations like Google, Apple, and Meta demonstrate the real-world applicability of FL. However, key challenges remain, including verifying server-side DP guarantees and coordinating training across heterogeneous devices, limiting broader adoption. Additionally, emerging trends such as large (multi-modal) models and blurred lines between training, inference, and personalization challenge traditional FL frameworks. In response, we propose a redefined FL framework that prioritizes privacy principles rather than rigid definitions. We also chart a path forward by leveraging trusted execution environments and open-source ecosystems to address these challenges and facilitate future advancements in FL.","author":[{"family":"Daly","given":"Katharine"},{"family":"Eichner","given":"Hubert"},{"family":"Kairouz","given":"Peter"},{"family":"Mcmahan","given":"HB"},{"family":"Ramage","given":"Daniel"},{"family":"Xu","given":"Zheng"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2410.08892","URL":"https://doi.org/10.48550/arxiv.2410.08892","source":"datacite"},{"id":"doi:10.48550/arxiv.2407.14154","type":"manuscript","title":"Where is the Testbed for my Federated Learning Research?","abstract":"Progressing beyond centralized AI is of paramount importance, yet, distributed AI solutions, in particular various federated learning (FL) algorithms, are often not comprehensively assessed, which prevents the research community from identifying the most promising approaches and practitioners from being convinced that a certain solution is deployment-ready. The largest hurdle towards FL algorithm evaluation is the difficulty of conducting real-world experiments over a variety of FL client devices and different platforms, with different datasets and data distribution, all while assessing various dimensions of algorithm performance, such as inference accuracy, energy consumption, and time to convergence, to name a few. In this paper, we present CoLExT, a real-world testbed for FL research. CoLExT is designed to streamline experimentation with custom FL algorithms in a rich testbed configuration space, with a large number of heterogeneous edge devices, ranging from single-board computers to smartphones, and provides real-time collection and visualization of a variety of metrics through automatic instrumentation. According to our evaluation, porting FL algorithms to CoLExT requires minimal involvement from the developer, and the instrumentation introduces minimal resource usage overhead. Furthermore, through an initial investigation involving popular FL algorithms running on CoLExT, we reveal previously unknown trade-offs, inefficiencies, and programming bugs.","author":[{"family":"Božič","given":"Janez"},{"family":"Faustino","given":"Amândio"},{"family":"Radovič","given":"Boris"},{"family":"Canini","given":"Marco"},{"family":"Pejović","given":"Veljko"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2407.14154","URL":"https://doi.org/10.48550/arxiv.2407.14154","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.04146","type":"manuscript","title":"pFedLVM: A Large Vision Model (LVM)-Driven and Latent Feature-Based Personalized Federated Learning Framework in Autonomous Driving","abstract":"Deep learning-based Autonomous Driving (AD) models often exhibit poor generalization due to data heterogeneity in an ever domain-shifting environment. While Federated Learning (FL) could improve the generalization of an AD model (known as FedAD system), conventional models often struggle with under-fitting as the amount of accumulated training data progressively increases. To address this issue, instead of conventional small models, employing Large Vision Models (LVMs) in FedAD is a viable option for better learning of representations from a vast volume of data. However, implementing LVMs in FedAD introduces three challenges: (I) the extremely high communication overheads associated with transmitting LVMs between participating vehicles and a central server; (II) lack of computing resource to deploy LVMs on each vehicle; (III) the performance drop due to LVM focusing on shared features but overlooking local vehicle characteristics. To overcome these challenges, we propose pFedLVM, a LVM-Driven, Latent Feature-Based Personalized Federated Learning framework. In this approach, the LVM is deployed only on central server, which effectively alleviates the computational burden on individual vehicles. Furthermore, the exchange between central server and vehicles are the learned features rather than the LVM parameters, which significantly reduces communication overhead. In addition, we utilize both shared features from all participating vehicles and individual characteristics from each vehicle to establish a personalized learning mechanism. This enables each vehicle's model to learn features from others while preserving its personalized characteristics, thereby outperforming globally shared models trained in general FL. Extensive experiments demonstrate that pFedLVM outperforms the existing state-of-the-art approaches.","author":[{"family":"Kou","given":"Wei"},{"family":"Lin","given":"Qingfeng"},{"family":"Tang","given":"Ming"},{"family":"Xu","given":"Sheng"},{"family":"Ye","given":"Rongguang"},{"family":"Leng","given":"Yang"},{"family":"Wang","given":"Shuai"},{"family":"Li","given":"Guofa"},{"family":"Chen","given":"Zhenyu"},{"family":"Zhu","given":"Guangxu"},{"family":"Wu","given":"Yik"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.04146","URL":"https://doi.org/10.48550/arxiv.2405.04146","source":"datacite"},{"id":"doi:10.48550/arxiv.2404.10728","type":"manuscript","title":"Randomized Exploration in Cooperative Multi-Agent Reinforcement Learning","abstract":"We present the first study on provably efficient randomized exploration in cooperative multi-agent reinforcement learning (MARL). We propose a unified algorithm framework for randomized exploration in parallel Markov Decision Processes (MDPs), and two Thompson Sampling (TS)-type algorithms, CoopTS-PHE and CoopTS-LMC, incorporating the perturbed-history exploration (PHE) strategy and the Langevin Monte Carlo exploration (LMC) strategy, respectively, which are flexible in design and easy to implement in practice. For a special class of parallel MDPs where the transition is (approximately) linear, we theoretically prove that both CoopTS-PHE and CoopTS-LMC achieve a $\\widetilde{\\mathcal{O}}(d^{3/2}H^2\\sqrt{MK})$ regret bound with communication complexity $\\widetilde{\\mathcal{O}}(dHM^2)$, where $d$ is the feature dimension, $H$ is the horizon length, $M$ is the number of agents, and $K$ is the number of episodes. This is the first theoretical result for randomized exploration in cooperative MARL. We evaluate our proposed method on multiple parallel RL environments, including a deep exploration problem (i.e., $N$-chain), a video game, and a real-world problem in energy systems. Our experimental results support that our framework can achieve better performance, even under conditions of misspecified transition models. Additionally, we establish a connection between our unified framework and the practical application of federated learning.","author":[{"family":"Hsu","given":"Hao"},{"family":"Wang","given":"Weixin"},{"family":"Pajic","given":"Miroslav"},{"family":"Xu","given":"Pan"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2404.10728","URL":"https://doi.org/10.48550/arxiv.2404.10728","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.13879","type":"manuscript","title":"FACT or Fiction: Can Truthful Mechanisms Eliminate Federated Free Riding?","abstract":"Standard federated learning (FL) approaches are vulnerable to the free-rider dilemma: participating agents can contribute little to nothing yet receive a well-trained aggregated model. While prior mechanisms attempt to solve the free-rider dilemma, none have addressed the issue of truthfulness. In practice, adversarial agents can provide false information to the server in order to cheat its way out of contributing to federated training. In an effort to make free-riding-averse federated mechanisms truthful, and consequently less prone to breaking down in practice, we propose FACT. FACT is the first federated mechanism that: (1) eliminates federated free riding by using a penalty system, (2) ensures agents provide truthful information by creating a competitive environment, and (3) encourages agent participation by offering better performance than training alone. Empirically, FACT avoids free-riding when agents are untruthful, and reduces agent loss by over 4x.","author":[{"family":"Bornstein","given":"Marco"},{"family":"Bedi","given":"Amrit"},{"family":"Mohamed","given":"Abdirisak"},{"family":"Huang","given":"Furong"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.13879","URL":"https://doi.org/10.48550/arxiv.2405.13879","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.02140","type":"manuscript","title":"An Information Theoretic Perspective on Conformal Prediction","abstract":"Conformal Prediction (CP) is a distribution-free uncertainty estimation framework that constructs prediction sets guaranteed to contain the true answer with a user-specified probability. Intuitively, the size of the prediction set encodes a general notion of uncertainty, with larger sets associated with higher degrees of uncertainty. In this work, we leverage information theory to connect conformal prediction to other notions of uncertainty. More precisely, we prove three different ways to upper bound the intrinsic uncertainty, as described by the conditional entropy of the target variable given the inputs, by combining CP with information theoretical inequalities. Moreover, we demonstrate two direct and useful applications of such connection between conformal prediction and information theory: (i) more principled and effective conformal training objectives that generalize previous approaches and enable end-to-end training of machine learning models from scratch, and (ii) a natural mechanism to incorporate side information into conformal prediction. We empirically validate both applications in centralized and federated learning settings, showing our theoretical results translate to lower inefficiency (average prediction set size) for popular CP methods.","author":[{"family":"Correia","given":"Alvaro"},{"family":"Massoli","given":"Fabio"},{"family":"Louizos","given":"Christos"},{"family":"Behboodi","given":"Arash"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.02140","URL":"https://doi.org/10.48550/arxiv.2405.02140","source":"datacite"},{"id":"doi:10.48550/arxiv.2409.06123","type":"manuscript","title":"Contrastive Federated Learning with Tabular Data Silos","abstract":"Learning from vertical partitioned data silos is challenging due to the segmented nature of data, sample misalignment, and strict privacy concerns. Federated learning has been proposed as a solution. However, sample misalignment across silos often hinders optimal model performance and suggests data sharing within the model, which breaks privacy. Our proposed solution is Contrastive Federated Learning with Tabular Data Silos (CFL), which offers a solution for data silos with sample misalignment without the need for sharing original or representative data to maintain privacy. CFL begins with local acquisition of contrastive representations of the data within each silo and aggregates knowledge from other silos through the federated learning algorithm. Our experiments demonstrate that CFL solves the limitations of existing algorithms for data silos and outperforms existing tabular contrastive learning. CFL provides performance improvements without loosening privacy.","author":[{"family":"Ginanjar","given":"Achmad"},{"family":"Li","given":"Xue"},{"family":"Hua","given":"Wen"},{"family":"Pei","given":"Jiaming"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2409.06123","URL":"https://doi.org/10.48550/arxiv.2409.06123","source":"datacite"},{"id":"doi:10.48550/arxiv.2406.03519","type":"manuscript","title":"Noise-Aware Algorithm for Heterogeneous Differentially Private Federated Learning","abstract":"High utility and rigorous data privacy are of the main goals of a federated learning (FL) system, which learns a model from the data distributed among some clients. The latter has been tried to achieve by using differential privacy in FL (DPFL). There is often heterogeneity in clients privacy requirements, and existing DPFL works either assume uniform privacy requirements for clients or are not applicable when server is not fully trusted (our setting). Furthermore, there is often heterogeneity in batch and/or dataset size of clients, which as shown, results in extra variation in the DP noise level across clients model updates. With these sources of heterogeneity, straightforward aggregation strategies, e.g., assigning clients aggregation weights proportional to their privacy parameters will lead to lower utility. We propose Robust-HDP, which efficiently estimates the true noise level in clients model updates and reduces the noise-level in the aggregated model updates considerably. Robust-HDP improves utility and convergence speed, while being safe to the clients that may maliciously send falsified privacy parameter to server. Extensive experimental results on multiple datasets and our theoretical analysis confirm the effectiveness of Robust-HDP. Our code can be found here.","author":[{"family":"Malekmohammadi","given":"Saber"},{"family":"Yu","given":"Yaoliang"},{"family":"Cao","given":"Yang"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2406.03519","URL":"https://doi.org/10.48550/arxiv.2406.03519","source":"datacite"},{"id":"doi:10.48550/arxiv.2403.03333","type":"manuscript","title":"Federated Learning over Connected Modes","abstract":"Statistical heterogeneity in federated learning poses two major challenges: slow global training due to conflicting gradient signals, and the need of personalization for local distributions. In this work, we tackle both challenges by leveraging recent advances in \\emph{linear mode connectivity} -- identifying a linearly connected low-loss region in the parameter space of neural networks, which we call solution simplex. We propose federated learning over connected modes (\\textsc{Floco}), where clients are assigned local subregions in this simplex based on their gradient signals, and together learn the shared global solution simplex. This allows personalization of the client models to fit their local distributions within the degrees of freedom in the solution simplex and homogenizes the update signals for the global simplex training. Our experiments show that \\textsc{Floco} accelerates the global training process, and significantly improves the local accuracy with minimal computational overhead in cross-silo federated learning settings.","author":[{"family":"Grinwald","given":"Dennis"},{"family":"Wiesner","given":"Philipp"},{"family":"Nakajima","given":"Shinichi"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2403.03333","URL":"https://doi.org/10.48550/arxiv.2403.03333","source":"datacite"},{"id":"doi:10.48550/arxiv.2407.17754","type":"manuscript","title":"DualFed: Enjoying both Generalization and Personalization in Federated Learning via Hierachical Representations","abstract":"In personalized federated learning (PFL), it is widely recognized that achieving both high model generalization and effective personalization poses a significant challenge due to their conflicting nature. As a result, existing PFL methods can only manage a trade-off between these two objectives. This raises an interesting question: Is it feasible to develop a model capable of achieving both objectives simultaneously? Our paper presents an affirmative answer, and the key lies in the observation that deep models inherently exhibit hierarchical architectures, which produce representations with various levels of generalization and personalization at different stages. A straightforward approach stemming from this observation is to select multiple representations from these layers and combine them to concurrently achieve generalization and personalization. However, the number of candidate representations is commonly huge, which makes this method infeasible due to high computational costs.To address this problem, we propose DualFed, a new method that can directly yield dual representations correspond to generalization and personalization respectively, thereby simplifying the optimization task. Specifically, DualFed inserts a personalized projection network between the encoder and classifier. The pre-projection representations are able to capture generalized information shareable across clients, and the post-projection representations are effective to capture task-specific information on local clients. This design minimizes the mutual interference between generalization and personalization, thereby achieving a win-win situation. Extensive experiments show that DualFed can outperform other FL methods. Code is available at https://github.com/GuogangZhu/DualFed.","author":[{"family":"Zhu","given":"Guogang"},{"family":"Liu","given":"Xuefeng"},{"family":"Niu","given":"Jianwei"},{"family":"Tang","given":"Shaojie"},{"family":"Wu","given":"Xinghao"},{"family":"Zhang","given":"Jiayuan"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2407.17754","URL":"https://doi.org/10.48550/arxiv.2407.17754","source":"datacite"},{"id":"doi:10.48550/arxiv.2409.09727","type":"manuscript","title":"From Challenges and Pitfalls to Recommendations and Opportunities: Implementing Federated Learning in Healthcare","abstract":"Federated learning holds great potential for enabling large-scale healthcare research and collaboration across multiple centres while ensuring data privacy and security are not compromised. Although numerous recent studies suggest or utilize federated learning based methods in healthcare, it remains unclear which ones have potential clinical utility. This review paper considers and analyzes the most recent studies up to May 2024 that describe federated learning based methods in healthcare. After a thorough review, we find that the vast majority are not appropriate for clinical use due to their methodological flaws and/or underlying biases which include but are not limited to privacy concerns, generalization issues, and communication costs. As a result, the effectiveness of federated learning in healthcare is significantly compromised. To overcome these challenges, we provide recommendations and promising opportunities that might be implemented to resolve these problems and improve the quality of model development in federated learning with healthcare.","author":[{"family":"Li","given":"Ming"},{"family":"Xu","given":"Pengcheng"},{"family":"Hu","given":"Junjie"},{"family":"Tang","given":"Zeyu"},{"family":"Yang","given":"Guang"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2409.09727","URL":"https://doi.org/10.48550/arxiv.2409.09727","source":"datacite"},{"id":"doi:10.48550/arxiv.2409.04637","type":"manuscript","title":"Enhancing Quantum Security over Federated Learning via Post-Quantum Cryptography","abstract":"Federated learning (FL) has become one of the standard approaches for deploying machine learning models on edge devices, where private training data are distributed across clients, and a shared model is learned by aggregating locally computed updates from each client. While this paradigm enhances communication efficiency by only requiring updates at the end of each training epoch, the transmitted model updates remain vulnerable to malicious tampering, posing risks to the integrity of the global model. Although current digital signature algorithms can protect these communicated model updates, they fail to ensure quantum security in the era of large-scale quantum computing. Fortunately, various post-quantum cryptography algorithms have been developed to address this vulnerability, especially the three NIST-standardized algorithms - Dilithium, FALCON, and SPHINCS+. In this work, we empirically investigate the impact of these three NIST-standardized PQC algorithms for digital signatures within the FL procedure, covering a wide range of models, tasks, and FL settings. Our results indicate that Dilithium stands out as the most efficient PQC algorithm for digital signature in federated learning. Additionally, we offer an in-depth discussion of the implications of our findings and potential directions for future research.","author":[{"family":"Li","given":"Pingzhi"},{"family":"Chen","given":"Tianlong"},{"family":"Liu","given":"Junyu"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2409.04637","URL":"https://doi.org/10.48550/arxiv.2409.04637","source":"datacite"},{"id":"doi:10.3929/ethz-b-000719997","type":"article-journal","title":"Hiding in Plain Sight: Disguising Data Stealing Attacks in Federated Learning","abstract":"Malicious server (MS) attacks have enabled the scaling of data stealing in federated learning to large batch sizes and secure aggregation, settings previously considered private. However, many concerns regarding the client-side detectability of MS attacks were raised, questioning their practicality. In this work, for the first time, we thoroughly study client-side detectability. We first demonstrate that all prior MS attacks are detectable by principled checks, and formulate a necessary set of requirements that a practical MS attack must satisfy. Next, we propose SEER, a novel attack framework that satisfies these requirements. The key insight of SEER is the use of a secret decoder, jointly trained with the shared model. We show that SEER can steal user data from gradients of realistic networks, even for large batch sizes of up to 512 and under secure aggregation. Our work is a promising step towards assessing the true vulnerability of federated learning in real-world settings.","author":[{"family":"Garov","given":"Kostadin"},{"family":"Dimitrov","given":"Dimitar"},{"family":"Jovanović","given":"Nikola"},{"family":"Vechev","given":"Martin"}],"issued":{"date-parts":[[2024]]},"DOI":"10.3929/ethz-b-000719997","URL":"https://doi.org/10.3929/ethz-b-000719997","source":"datacite"},{"id":"doi:10.48550/arxiv.2401.10070","type":"manuscript","title":"Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks","abstract":"To protect privacy and meet legal regulations, federated learning (FL) has gained significant attention for training speech-to-text (S2T) systems, including automatic speech recognition (ASR) and speech translation (ST). However, the commonly used FL approach (i.e., \\textsc{FedAvg}) in S2T tasks typically suffers from extensive communication overhead due to multi-round interactions based on the whole model and performance degradation caused by data heterogeneity among clients.To address these issues, we propose a personalized federated S2T framework that introduces \\textsc{FedLoRA}, a lightweight LoRA module for client-side tuning and interaction with the server to minimize communication overhead, and \\textsc{FedMem}, a global model equipped with a $k$-nearest-neighbor ($k$NN) classifier that captures client-specific distributional shifts to achieve personalization and overcome data heterogeneity. Extensive experiments based on Conformer and Whisper backbone models on CoVoST and GigaSpeech benchmarks show that our approach significantly reduces the communication overhead on all S2T tasks and effectively personalizes the global model to overcome data heterogeneity.","author":[{"family":"Du","given":"Yichao"},{"family":"Zhang","given":"Zhirui"},{"family":"Yue","given":"Linan"},{"family":"Huang","given":"Xu"},{"family":"Zhang","given":"Yuqing"},{"family":"Xu","given":"Tong"},{"family":"Xu","given":"Linli"},{"family":"Chen","given":"Enhong"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2401.10070","URL":"https://doi.org/10.48550/arxiv.2401.10070","source":"datacite"},{"id":"doi:10.48550/arxiv.2404.14389","type":"manuscript","title":"Poisoning Attacks on Federated Learning-based Wireless Traffic Prediction","abstract":"Federated Learning (FL) offers a distributed framework to train a global control model across multiple base stations without compromising the privacy of their local network data. This makes it ideal for applications like wireless traffic prediction (WTP), which plays a crucial role in optimizing network resources, enabling proactive traffic flow management, and enhancing the reliability of downstream communication-aided applications, such as IoT devices, autonomous vehicles, and industrial automation systems. Despite its promise, the security aspects of FL-based distributed wireless systems, particularly in regression-based WTP problems, remain inadequately investigated. In this paper, we introduce a novel fake traffic injection (FTI) attack, designed to undermine the FL-based WTP system by injecting fabricated traffic distributions with minimal knowledge. We further propose a defense mechanism, termed global-local inconsistency detection (GLID), which strategically removes abnormal model parameters that deviate beyond a specific percentile range estimated through statistical methods in each dimension. Extensive experimental evaluations, performed on real-world wireless traffic datasets, demonstrate that both our attack and defense strategies significantly outperform existing baselines.","author":[{"family":"Zhang","given":"Zifan"},{"family":"Fang","given":"Minghong"},{"family":"Huang","given":"Jiayuan"},{"family":"Liu","given":"Yuchen"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2404.14389","URL":"https://doi.org/10.48550/arxiv.2404.14389","source":"datacite"},{"id":"doi:10.48550/arxiv.2402.15166","type":"manuscript","title":"Convergence Analysis of Split Federated Learning on Heterogeneous Data","abstract":"Split federated learning (SFL) is a recent distributed approach for collaborative model training among multiple clients. In SFL, a global model is typically split into two parts, where clients train one part in a parallel federated manner, and a main server trains the other. Despite the recent research on SFL algorithm development, the convergence analysis of SFL is missing in the literature, and this paper aims to fill this gap. The analysis of SFL can be more challenging than that of federated learning (FL), due to the potential dual-paced updates at the clients and the main server. We provide convergence analysis of SFL for strongly convex and general convex objectives on heterogeneous data. The convergence rates are $O(1/T)$ and $O(1/\\sqrt[3]{T})$, respectively, where $T$ denotes the total number of rounds for SFL training. We further extend the analysis to non-convex objectives and the scenario where some clients may be unavailable during training. Experimental experiments validate our theoretical results and show that SFL outperforms FL and split learning (SL) when data is highly heterogeneous across a large number of clients.","author":[{"family":"Han","given":"Pengchao"},{"family":"Huang","given":"Chao"},{"family":"Tian","given":"Geng"},{"family":"Tang","given":"Ming"},{"family":"Liu","given":"Xin"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2402.15166","URL":"https://doi.org/10.48550/arxiv.2402.15166","source":"datacite"},{"id":"doi:10.48550/arxiv.2412.20253","type":"manuscript","title":"Election of Collaborators via Reinforcement Learning for Federated Brain Tumor Segmentation","abstract":"Federated learning (FL) enables collaborative model training across decentralized datasets while preserving data privacy. However, optimally selecting participating collaborators in dynamic FL environments remains challenging. We present RL-HSimAgg, a novel reinforcement learning (RL) and similarity-weighted aggregation (simAgg) algorithm using harmonic mean to manage outlier data points. This paper proposes applying multi-armed bandit algorithms to improve collaborator selection and model generalization. By balancing exploration-exploitation trade-offs, these RL methods can promote resource-efficient training with diverse datasets. We demonstrate the effectiveness of Epsilon-greedy (EG) and upper confidence bound (UCB) algorithms for federated brain lesion segmentation. In simulation experiments on internal and external validation sets, RL-HSimAgg with UCB collaborator outperformed the EG method across all metrics, achieving higher Dice scores for Enhancing Tumor (0.7334 vs 0.6797), Tumor Core (0.7432 vs 0.6821), and Whole Tumor (0.8252 vs 0.7931) segmentation. Therefore, for the Federated Tumor Segmentation Challenge (FeTS 2024), we consider UCB as our primary client selection approach in federated Glioblastoma lesion segmentation of multi-modal MRIs. In conclusion, our research demonstrates that RL-based collaborator management, e.g. using UCB, can potentially improve model robustness and flexibility in distributed learning environments, particularly in domains like brain tumor segmentation.","author":[{"family":"Khan","given":"Muhammad"},{"family":"Kontio","given":"Elina"},{"family":"Khan","given":"Suleiman"},{"family":"Jafaritadi","given":"Mojtaba"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2412.20253","URL":"https://doi.org/10.48550/arxiv.2412.20253","source":"datacite"},{"id":"doi:10.48550/arxiv.2412.20250","type":"manuscript","title":"Recommender Engine Driven Client Selection in Federated Brain Tumor Segmentation","abstract":"This study presents a robust and efficient client selection protocol designed to optimize the Federated Learning (FL) process for the Federated Tumor Segmentation Challenge (FeTS 2024). In the evolving landscape of FL, the judicious selection of collaborators emerges as a critical determinant for the success and efficiency of collective learning endeavors, particularly in domains requiring high precision. This work introduces a recommender engine framework based on non-negative matrix factorization (NNMF) and a hybrid aggregation approach that blends content-based and collaborative filtering. This method intelligently analyzes historical performance, expertise, and other relevant metrics to identify the most suitable collaborators. This approach not only addresses the cold start problem where new or inactive collaborators pose selection challenges due to limited data but also significantly improves the precision and efficiency of the FL process. Additionally, we propose harmonic similarity weight aggregation (HSimAgg) for adaptive aggregation of model parameters. We utilized a dataset comprising 1,251 multi-parametric magnetic resonance imaging (mpMRI) scans from individuals diagnosed with glioblastoma (GBM) for training purposes and an additional 219 mpMRI scans for external evaluations. Our federated tumor segmentation approach achieved dice scores of 0.7298, 0.7424, and 0.8218 for enhancing tumor (ET), tumor core (TC), and whole tumor (WT) segmentation tasks respectively on the external validation set. In conclusion, this research demonstrates that selecting collaborators with expertise aligned to specific tasks, like brain tumor segmentation, improves the effectiveness of FL networks.","author":[{"family":"Khan","given":"Muhammad"},{"family":"Kontio","given":"Elina"},{"family":"Khan","given":"Suleiman"},{"family":"Jafaritadi","given":"Mojtaba"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2412.20250","URL":"https://doi.org/10.48550/arxiv.2412.20250","source":"datacite"},{"id":"doi:10.48550/arxiv.2406.14898","type":"manuscript","title":"Safely Learning with Private Data: A Federated Learning Framework for Large Language Model","abstract":"Private data, being larger and quality-higher than public data, can greatly improve large language models (LLM). However, due to privacy concerns, this data is often dispersed in multiple silos, making its secure utilization for LLM training a challenge. Federated learning (FL) is an ideal solution for training models with distributed private data, but traditional frameworks like FedAvg are unsuitable for LLM due to their high computational demands on clients. An alternative, split learning, offloads most training parameters to the server while training embedding and output layers locally, making it more suitable for LLM. Nonetheless, it faces significant challenges in security and efficiency. Firstly, the gradients of embeddings are prone to attacks, leading to potential reverse engineering of private data. Furthermore, the server's limitation of handle only one client's training request at a time hinders parallel training, severely impacting training efficiency. In this paper, we propose a Federated Learning framework for LLM, named FL-GLM, which prevents data leakage caused by both server-side and peer-client attacks while improving training efficiency. Specifically, we first place the input block and output block on local client to prevent embedding gradient attacks from server. Secondly, we employ key-encryption during client-server communication to prevent reverse engineering attacks from peer-clients. Lastly, we employ optimization methods like client-batching or server-hierarchical, adopting different acceleration methods based on the actual computational capabilities of the server. Experimental results on NLU and generation tasks demonstrate that FL-GLM achieves comparable metrics to centralized chatGLM model, validating the effectiveness of our federated learning framework.","author":[{"family":"Zheng","given":"Jiaying"},{"family":"Zhang","given":"Hainan"},{"family":"Wang","given":"Lingxiang"},{"family":"Qiu","given":"Wangjie"},{"family":"Zheng","given":"Hongwei"},{"family":"Zheng","given":"Zhiming"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2406.14898","URL":"https://doi.org/10.48550/arxiv.2406.14898","source":"datacite"},{"id":"doi:10.5281/zenodo.21053046","type":"article-journal","title":"A Comparative Analysis of Artificial Intelligence and Business Intelligence using Big Data Analytics","abstract":"This research presents a comparative analysis of Artificial Intelligence (AI) and Business Intelligence (BI) using Big Data Analytics. The study evaluates multiple machine learning algorithms, including Random Forest, Neural Networks, Support Vector Machines, Decision Trees, Logistic Regression, Naïve Bayes, AdaBoost, Gradient Boosting, and K-Nearest Neighbors, for talent recruitment and business intelligence applications. Experimental results demonstrate that Random Forest and Neural Networks provide the highest prediction accuracy, enabling organizations to improve recruitment decisions, operational efficiency, and data-driven business intelligence while highlighting future directions for explainable AI, federated learning, and real-time analytics.","author":[{"family":"Vegineni","given":"Gopi"},{"family":"Addanki","given":"Sireesha"},{"family":"Mandal","given":"Rishabh"},{"family":"Ellahi","given":"Ehsan"},{"family":"Marella","given":"Bhagath"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.21053046","URL":"https://doi.org/10.5281/zenodo.21053046","source":"datacite"},{"id":"doi:10.5281/zenodo.21053047","type":"article-journal","title":"A Comparative Analysis of Artificial Intelligence and Business Intelligence using Big Data Analytics","abstract":"This research presents a comparative analysis of Artificial Intelligence (AI) and Business Intelligence (BI) using Big Data Analytics. The study evaluates multiple machine learning algorithms, including Random Forest, Neural Networks, Support Vector Machines, Decision Trees, Logistic Regression, Naïve Bayes, AdaBoost, Gradient Boosting, and K-Nearest Neighbors, for talent recruitment and business intelligence applications. Experimental results demonstrate that Random Forest and Neural Networks provide the highest prediction accuracy, enabling organizations to improve recruitment decisions, operational efficiency, and data-driven business intelligence while highlighting future directions for explainable AI, federated learning, and real-time analytics.","author":[{"family":"Vegineni","given":"Gopi"},{"family":"Addanki","given":"Sireesha"},{"family":"Mandal","given":"Rishabh"},{"family":"Ellahi","given":"Ehsan"},{"family":"Marella","given":"Bhagath"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.21053047","URL":"https://doi.org/10.5281/zenodo.21053047","source":"datacite"},{"id":"doi:10.48550/arxiv.2406.14429","type":"manuscript","title":"CollaFuse: Collaborative Diffusion Models","abstract":"In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic images. However, the application of diffusion models poses numerous challenges, particularly concerning data availability, computational requirements, and privacy. Traditional approaches to address these shortcomings, like federated learning, often impose significant computational burdens on individual clients, especially those with constrained resources. In response to these challenges, we introduce the novel approach CollaFuse for distributed collaborative diffusion models inspired by split learning. Our approach facilitates collaborative training of diffusion models while alleviating client computational burdens during image synthesis. This reduced computational burden is achieved by retaining data and computationally inexpensive processes locally at each client while outsourcing the computationally expensive processes to shared, more efficient server resources. Through experiments on the common datasets CelebA, CIFAR-10, and Animals-with-Attributes2, our approach demonstrates enhanced performance while decreasing information disclosure as it reduces the necessity for sharing raw data. These capabilities hold significant potential across various application areas, including the design of edge computing solutions. Thus, our work advances distributed machine learning by contributing to the evolution of collaborative diffusion models.","author":[{"family":"Allmendinger","given":"Simeon"},{"family":"Zipperling","given":"Domenique"},{"family":"Struppek","given":"Lukas"},{"family":"Kühl","given":"Niklas"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2406.14429","URL":"https://doi.org/10.48550/arxiv.2406.14429","source":"datacite"},{"id":"doi:10.48550/arxiv.2408.00243","type":"manuscript","title":"A Survey on the Applications of Zero-Knowledge Proofs","abstract":"Zero-knowledge proofs (ZKPs) enable computational integrity and privacy by allowing one party to prove the truth of a statement without revealing underlying data. Compared with alternatives such as homomorphic encryption and secure multiparty computation, ZKPs offer distinct advantages in universality and minimal trust assumptions, with applications spanning blockchain systems and confidential verification of computational tasks. This survey provides a technical overview of ZKPs with a focus on an increasingly relevant subset called zkSNARKs. Unlike prior surveys emphasizing algorithmic and theoretical aspects, we take a broader view of practical deployments and recent use cases across multiple domains including blockchain privacy, scaling, storage, and interoperability, as well as non-blockchain applications such as voting, authentication, timelocks, and machine learning. To support consistent comparison, we provide (i) a taxonomy of application areas, (ii) evaluation criteria including proof size, prover and verifier time, memory, and setup assumptions, and (iii) comparative tables summarizing key tradeoffs and representative systems. The survey also covers supporting infrastructure, including zero-knowledge virtual machines, domain-specific languages, libraries, and frameworks. While emphasizing zkSNARKs for their prevalence in deployed systems, we compare them with zkSTARKs and Bulletproofs to clarify transparency and performance tradeoffs. We conclude with future research and application directions.","author":[{"family":"Lavin","given":"Ryan"},{"family":"Liu","given":"Xuekai"},{"family":"Mohanty","given":"Hardhik"},{"family":"Norman","given":"Logan"},{"family":"Zaarour","given":"Giovanni"},{"family":"Krishnamachari","given":"Bhaskar"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2408.00243","URL":"https://doi.org/10.48550/arxiv.2408.00243","source":"datacite"},{"id":"doi:10.48550/arxiv.2408.05629","type":"manuscript","title":"Quantum-secure multiparty deep learning","abstract":"Secure multiparty computation enables the joint evaluation of multivariate functions across distributed users while ensuring the privacy of their local inputs. This field has become increasingly urgent due to the exploding demand for computationally intensive deep learning inference. These computations are typically offloaded to cloud computing servers, leading to vulnerabilities that can compromise the security of the clients' data. To solve this problem, we introduce a linear algebra engine that leverages the quantum nature of light for information-theoretically secure multiparty computation using only conventional telecommunication components. We apply this linear algebra engine to deep learning and derive rigorous upper bounds on the information leakage of both the deep neural network weights and the client's data via the Holevo and the Cramér-Rao bounds, respectively. Applied to the MNIST classification task, we obtain test accuracies exceeding $96\\%$ while leaking less than $0.1$ bits per weight symbol and $0.01$ bits per data symbol. This weight leakage is an order of magnitude below the minimum bit precision required for accurate deep learning using state-of-the-art quantization techniques. Our work lays the foundation for practical quantum-secure computation and unlocks secure cloud deep learning as a field.","author":[{"family":"Sulimany","given":"Kfir"},{"family":"Vadlamani","given":"Sri"},{"family":"Hamerly","given":"Ryan"},{"family":"Iyengar","given":"Prahlad"},{"family":"Englund","given":"Dirk"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2408.05629","URL":"https://doi.org/10.48550/arxiv.2408.05629","source":"datacite"},{"id":"doi:10.48550/arxiv.2402.08956","type":"manuscript","title":"Seagull: Privacy preserving network verification system","abstract":"The Internet relies on routing protocols to direct traffic efficiently across interconnected networks, with the Border Gateway Protocol (BGP) serving as the core mechanism managing routing between autonomous systems. However, BGP configurations are largely manual, making them susceptible to human errors that can lead to outages or security vulnerabilities. Verifying the correctness and convergence of BGP configurations is therefore essential for maintaining a stable and secure Internet. Yet, this verification process faces two key challenges: preserving the privacy of proprietary routing information and ensuring scalability across large, distributed networks. This paper introduces a privacy-preserving verification framework that leverages multiparty computation (MPC) to validate BGP configurations without exposing sensitive routing data. Our approach overcomes both privacy and scalability challenges by ensuring that no information beyond the verification outcome is revealed. Through formal analysis, we show that the proposed method achieves strong privacy guarantees and practical scalability, providing a secure and efficient foundation for verifying BGP-based routing in the Internet backbone.","author":[{"family":"Daneshamooz","given":"Jaber"},{"family":"Yu","given":"Melody"},{"family":"Maddury","given":"Sucheer"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2402.08956","URL":"https://doi.org/10.48550/arxiv.2402.08956","source":"datacite"},{"id":"doi:10.5281/zenodo.14054751","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.14054751","URL":"https://doi.org/10.5281/zenodo.14054751","source":"datacite"},{"id":"doi:10.5281/zenodo.11243884","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.11243884","URL":"https://doi.org/10.5281/zenodo.11243884","source":"datacite"},{"id":"doi:10.5281/zenodo.11121487","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.11121487","URL":"https://doi.org/10.5281/zenodo.11121487","source":"datacite"},{"id":"doi:10.5281/zenodo.11632895","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.11632895","URL":"https://doi.org/10.5281/zenodo.11632895","source":"datacite"},{"id":"doi:10.5281/zenodo.11503803","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.11503803","URL":"https://doi.org/10.5281/zenodo.11503803","source":"datacite"},{"id":"doi:10.5281/zenodo.11243825","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.11243825","URL":"https://doi.org/10.5281/zenodo.11243825","source":"datacite"},{"id":"doi:10.5281/zenodo.14392692","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.14392692","URL":"https://doi.org/10.5281/zenodo.14392692","source":"datacite"},{"id":"doi:10.5281/zenodo.11263148","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.11263148","URL":"https://doi.org/10.5281/zenodo.11263148","source":"datacite"},{"id":"doi:10.5281/zenodo.12749258","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.12749258","URL":"https://doi.org/10.5281/zenodo.12749258","source":"datacite"},{"id":"doi:10.5281/zenodo.11442543","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.11442543","URL":"https://doi.org/10.5281/zenodo.11442543","source":"datacite"},{"id":"doi:10.5281/zenodo.12760291","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.12760291","URL":"https://doi.org/10.5281/zenodo.12760291","source":"datacite"},{"id":"doi:10.5281/zenodo.14438944","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.14438944","URL":"https://doi.org/10.5281/zenodo.14438944","source":"datacite"},{"id":"doi:10.5281/zenodo.12749076","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.12749076","URL":"https://doi.org/10.5281/zenodo.12749076","source":"datacite"},{"id":"doi:10.5281/zenodo.14205172","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.14205172","URL":"https://doi.org/10.5281/zenodo.14205172","source":"datacite"},{"id":"doi:10.5281/zenodo.12750133","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.12750133","URL":"https://doi.org/10.5281/zenodo.12750133","source":"datacite"},{"id":"doi:10.5281/zenodo.14094513","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.14094513","URL":"https://doi.org/10.5281/zenodo.14094513","source":"datacite"},{"id":"doi:10.5281/zenodo.13132856","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.13132856","URL":"https://doi.org/10.5281/zenodo.13132856","source":"datacite"},{"id":"doi:10.5281/zenodo.14093182","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.14093182","URL":"https://doi.org/10.5281/zenodo.14093182","source":"datacite"},{"id":"doi:10.5281/zenodo.13347752","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.13347752","URL":"https://doi.org/10.5281/zenodo.13347752","source":"datacite"},{"id":"doi:10.5281/zenodo.13866996","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.13866996","URL":"https://doi.org/10.5281/zenodo.13866996","source":"datacite"},{"id":"doi:10.5281/zenodo.13321314","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.13321314","URL":"https://doi.org/10.5281/zenodo.13321314","source":"datacite"},{"id":"doi:10.5281/zenodo.12751530","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.12751530","URL":"https://doi.org/10.5281/zenodo.12751530","source":"datacite"},{"id":"doi:10.5281/zenodo.12793562","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.12793562","URL":"https://doi.org/10.5281/zenodo.12793562","source":"datacite"},{"id":"doi:10.5281/zenodo.11447242","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.11447242","URL":"https://doi.org/10.5281/zenodo.11447242","source":"datacite"},{"id":"doi:10.5281/zenodo.12795334","type":"article-journal","title":"vantage6","abstract":"Vantage6 stands for privacy preserving federated learning infrastructure for secure insight exchange. The project is inspired by the Personal Health Train (PHT) concept. In this analogy vantage6 is the tracks and stations. Compatible algorithms are the trains, and computation tasks are the journey. Vantage6 is completely open source under the Apache License. What vantage6 does: delivering algorithms to data stations and collecting their results managing users, organizations, collaborations, computation tasks and their results providing control (security) at the data-stations to their owners The vantage6 infrastructure is designed with three fundamental functional aspects of federated learning. Autonomy. All involved parties should remain independent and autonomous. Heterogeneity. Parties should be allowed to have differences in hardware and operating systems. Flexibility. Related to the latter, a federated learning infrastructure should not limit the use of relevant data.","author":[{"family":"Martin","given":"Frank"},{"family":"Beusekom","given":"Bart"},{"family":"Leurs","given":"Richard"},{"family":"Sieswerda","given":"Melle"},{"family":"Soest","given":"Johan"},{"family":"Alradhi","given":"Hasan"},{"family":"Moncada-Torres","given":"Arturo"},{"family":"Baccinelli","given":"Walter"},{"family":"Smits","given":"Djura"},{"family":"Sanchez Gomez","given":"Luis"},{"family":"Harms","given":"Alexander"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.12795334","URL":"https://doi.org/10.5281/zenodo.12795334","source":"datacite"},{"id":"doi:10.69987/aimlr.2024.50304","type":"article-journal","title":"Research on Cross-Platform Digital Advertising User Behavior Analysis Framework Based on Federated Learning","abstract":"This information is released over the digital movement based on behavioral courses for user assessment. The framework is important to challenge the privacy-keeping the data division and manipulation throughout the platform announced. The network network architecture is designed to detect users of behavioral behavior when keeping personal information from secure. The framework implements an adaptive model aggregation strategy with dynamic weight adjustment mechanisms to optimize cross-platform model performance. Special protection, including special data and homomorphic encryption, has been integrated with security data during training and competition level. Tests have followed our greatest datase in the world, completed for a pre-commitment, the proposed efforts Over 200 million users across 5 million users when maintaining strategic warranty. Assessmental evaluation of significant improvements in advertisement, including the pronouncement (CTR), when minimized time %. The framework procedures for privacy personally used to investigate the characteristics of new ecosystem.","author":[{"family":"Zhang","given":"Kai"},{"family":"Xing","given":"Suchuan"},{"family":"Chen","given":"Yizhe"},{"family":"Chen","given":"YH"}],"issued":{"date-parts":[[2024]]},"DOI":"10.69987/aimlr.2024.50304","URL":"https://doi.org/10.69987/aimlr.2024.50304","source":"openalex"},{"id":"doi:10.5281/zenodo.21684175","type":"article-journal","title":"Developing AI-Augmented Intrusion Detection Systems for Cloud-Based Financial Platforms with Real-Time Risk Analysis","abstract":"Cloud-based financial platforms are increasingly targeted by sophisticated cyber threats due to the high-value transactions and sensitive data they manage. Traditional intrusion detection systems (IDS) often struggle to provide timely and accurate threat detection in dynamic and distributed cloud environments. This paper reviews the development of AI-augmented intrusion detection systems specifically tailored for financial services operating in cloud infrastructures. It explores how artificial intelligence—particularly machine learning, deep learning, and hybrid models—can enhance threat detection accuracy, reduce false positives, and support real-time risk analysis. The study also evaluates the role of federated learning, behavioral analytics, and anomaly detection in detecting insider threats, zero-day vulnerabilities, and fraud in financial transactions. Furthermore, the paper discusses the integration of explainable AI (XAI) to ensure transparency and regulatory compliance in threat assessment processes. Key implementation challenges such as data privacy, model drift, scalability, and integration with cloud-native architectures are also critically analyzed. The review concludes by proposing a layered, adaptive AI-IDS framework designed to safeguard cloud-based financial ecosystems through continuous learning, context-aware threat modeling, and real-time risk prioritization.","author":[{"family":"Ayanbode","given":"Noah"},{"family":"Cadet","given":"Emmanuel"},{"family":"Etim","given":"Edima"},{"family":"Essien","given":"Iboro"},{"family":"Ajayi","given":"Joshua"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.21684175","URL":"https://doi.org/10.5281/zenodo.21684175","source":"datacite"},{"id":"doi:10.5281/zenodo.21684176","type":"article-journal","title":"Developing AI-Augmented Intrusion Detection Systems for Cloud-Based Financial Platforms with Real-Time Risk Analysis","abstract":"Cloud-based financial platforms are increasingly targeted by sophisticated cyber threats due to the high-value transactions and sensitive data they manage. Traditional intrusion detection systems (IDS) often struggle to provide timely and accurate threat detection in dynamic and distributed cloud environments. This paper reviews the development of AI-augmented intrusion detection systems specifically tailored for financial services operating in cloud infrastructures. It explores how artificial intelligence—particularly machine learning, deep learning, and hybrid models—can enhance threat detection accuracy, reduce false positives, and support real-time risk analysis. The study also evaluates the role of federated learning, behavioral analytics, and anomaly detection in detecting insider threats, zero-day vulnerabilities, and fraud in financial transactions. Furthermore, the paper discusses the integration of explainable AI (XAI) to ensure transparency and regulatory compliance in threat assessment processes. Key implementation challenges such as data privacy, model drift, scalability, and integration with cloud-native architectures are also critically analyzed. The review concludes by proposing a layered, adaptive AI-IDS framework designed to safeguard cloud-based financial ecosystems through continuous learning, context-aware threat modeling, and real-time risk prioritization.","author":[{"family":"Ayanbode","given":"Noah"},{"family":"Cadet","given":"Emmanuel"},{"family":"Etim","given":"Edima"},{"family":"Essien","given":"Iboro"},{"family":"Ajayi","given":"Joshua"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.21684176","URL":"https://doi.org/10.5281/zenodo.21684176","source":"datacite"},{"id":"doi:10.48550/arxiv.2409.15723","type":"manuscript","title":"Federated Large Language Models: Current Progress and Future Directions","abstract":"Large Language Models have achieved impressive performance across diverse applications, yet their training typically depends on centralized data collection, raising serious privacy and governance concerns. Federated Learning offers a decentralized alternative by enabling multiple clients to collaboratively train shared models without exposing raw local data. However, integrating FL with LLMs introduces new challenges, including data heterogeneity, convergence instability, communication overhead, and computational constraints. This survey provides a comprehensive and up-to-date overview of Federated Learning for Large Language Models (FedLLM). We systematically review recent advances, with particular emphasis on federated fine-tuning and federated prompt learning, and analyze how existing methods address efficiency, personalization, and security challenges. We further summarize emerging directions such as federated pre-training and federated agents. Our goal is to offer a structured perspective on this rapidly evolving field and to highlight promising avenues for future research.","author":[{"family":"Yao","given":"Yuhang"},{"family":"Zhang","given":"Jianyi"},{"family":"Wu","given":"Junda"},{"family":"Huang","given":"Chengkai"},{"family":"Xia","given":"Yu"},{"family":"Yu","given":"Tong"},{"family":"Zhang","given":"Ruiyi"},{"family":"Kim","given":"Sungchul"},{"family":"Rossi","given":"Ryan"},{"family":"Li","given":"Ang"},{"family":"Yao","given":"Lina"},{"family":"Mcauley","given":"Julian"},{"family":"Chen","given":"Yiran"},{"family":"Joe-Wong","given":"Carlee"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2409.15723","URL":"https://doi.org/10.48550/arxiv.2409.15723","source":"datacite"},{"id":"doi:10.48550/arxiv.2402.09059","type":"manuscript","title":"I can't see it but I can Fine-tune it: On Encrypted Fine-tuning of Transformers using Fully Homomorphic Encryption","abstract":"In today's machine learning landscape, fine-tuning pretrained transformer models has emerged as an essential technique, particularly in scenarios where access to task-aligned training data is limited. However, challenges surface when data sharing encounters obstacles due to stringent privacy regulations or user apprehension regarding personal information disclosure. Earlier works based on secure multiparty computation (SMC) and fully homomorphic encryption (FHE) for privacy-preserving machine learning (PPML) focused more on privacy-preserving inference than privacy-preserving training. In response, we introduce BlindTuner, a privacy-preserving fine-tuning system that enables transformer training exclusively on homomorphically encrypted data for image classification. Our extensive experimentation validates BlindTuner's effectiveness by demonstrating comparable accuracy to non-encrypted models. Notably, our findings highlight a substantial speed enhancement of 1.5x to 600x over previous work in this domain.","author":[{"family":"Panzade","given":"Prajwal"},{"family":"Takabi","given":"Daniel"},{"family":"Cai","given":"Zhipeng"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2402.09059","URL":"https://doi.org/10.48550/arxiv.2402.09059","source":"datacite"},{"id":"doi:10.48550/arxiv.2410.08864","type":"manuscript","title":"The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses","abstract":"We formalize and analyze the trade-off between backdoor-based watermarks and adversarial defenses, framing it as an interactive protocol between a verifier and a prover. While previous works have primarily focused on this trade-off, our analysis extends it by identifying transferable attacks as a third, counterintuitive, but necessary option. Our main result shows that for all learning tasks, at least one of the three exists: a watermark, an adversarial defense, or a transferable attack. By transferable attack, we refer to an efficient algorithm that generates queries indistinguishable from the data distribution and capable of fooling all efficient defenders. Using cryptographic techniques, specifically fully homomorphic encryption, we construct a transferable attack and prove its necessity in this trade-off. Finally, we show that tasks of bounded VC-dimension allow adversarial defenses against all attackers, while a subclass allows watermarks secure against fast adversaries.","author":[{"family":"Głuch","given":"Grzegorz"},{"family":"Turan","given":"Berkant"},{"family":"Nagarajan","given":"Sai"},{"family":"Pokutta","given":"Sebastian"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2410.08864","URL":"https://doi.org/10.48550/arxiv.2410.08864","source":"datacite"},{"id":"doi:10.48550/arxiv.2312.04356","type":"manuscript","title":"NeuJeans: Private Neural Network Inference with Joint Optimization of Convolution and FHE Bootstrapping","abstract":"Fully homomorphic encryption (FHE) is a promising cryptographic primitive for realizing private neural network inference (PI) services by allowing a client to fully offload the inference task to a cloud server while keeping the client data oblivious to the server. This work proposes NeuJeans, an FHE-based solution for the PI of deep convolutional neural networks (CNNs). NeuJeans tackles the critical problem of the enormous computational cost for the FHE evaluation of CNNs. We introduce a novel encoding method called Coefficients-in-Slot (CinS) encoding, which enables multiple convolutions in one HE multiplication without costly slot permutations. We further observe that CinS encoding is obtained by conducting the first several steps of the Discrete Fourier Transform (DFT) on a ciphertext in conventional Slot encoding. This property enables us to save the conversion between CinS and Slot encodings as bootstrapping a ciphertext starts with DFT. Exploiting this, we devise optimized execution flows for various two-dimensional convolution (conv2d) operations and apply them to end-to-end CNN implementations. NeuJeans accelerates the performance of conv2d-activation sequences by up to 5.68 times compared to state-of-the-art FHE-based PI work and performs the PI of a CNN at the scale of ImageNet within a mere few seconds.","author":[{"family":"Ju","given":"Jae"},{"family":"Park","given":"Jaiyoung"},{"family":"Kim","given":"Jongmin"},{"family":"Kang","given":"Minsik"},{"family":"Kim","given":"Donghwan"},{"family":"Cheon","given":"Jung"},{"family":"Ahn","given":"Jung"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2312.04356","URL":"https://doi.org/10.48550/arxiv.2312.04356","source":"datacite"},{"id":"doi:10.48550/arxiv.2411.07468","type":"manuscript","title":"Privacy-Preserving Verifiable Neural Network Inference Service","abstract":"Machine learning has revolutionized data analysis and pattern recognition, but its resource-intensive training has limited accessibility. Machine Learning as a Service (MLaaS) simplifies this by enabling users to delegate their data samples to an MLaaS provider and obtain the inference result using a pre-trained model. Despite its convenience, leveraging MLaaS poses significant privacy and reliability concerns to the client. Specifically, sensitive information from the client inquiry data can be leaked to an adversarial MLaaS provider. Meanwhile, the lack of a verifiability guarantee can potentially result in biased inference results or even unfair payment issues. While existing trustworthy machine learning techniques, such as those relying on verifiable computation or secure computation, offer solutions to privacy and reliability concerns, they fall short of simultaneously protecting the privacy of client data and providing provable inference verifiability. In this paper, we propose vPIN, a privacy-preserving and verifiable CNN inference scheme that preserves privacy for client data samples while ensuring verifiability for the inference. vPIN makes use of partial homomorphic encryption and commit-and-prove succinct non-interactive argument of knowledge techniques to achieve desirable security properties. In vPIN, we develop various optimization techniques to minimize the proving circuit for homomorphic inference evaluation thereby, improving the efficiency and performance of our technique. We fully implemented and evaluated our vPIN scheme on standard datasets (e.g., MNIST, CIFAR-10). Our experimental results show that vPIN achieves high efficiency in terms of proving time, verification time, and proof size, while providing client data privacy guarantees and provable verifiability.","author":[{"family":"Riasi","given":"Arman"},{"family":"Guajardo","given":"Jorge"},{"family":"Hoang","given":"Thang"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2411.07468","URL":"https://doi.org/10.48550/arxiv.2411.07468","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.14569","type":"manuscript","title":"PrivCirNet: Efficient Private Inference via Block Circulant Transformation","abstract":"Homomorphic encryption (HE)-based deep neural network (DNN) inference protects data and model privacy but suffers from significant computation overhead. We observe transforming the DNN weights into circulant matrices converts general matrix-vector multiplications into HE-friendly 1-dimensional convolutions, drastically reducing the HE computation cost. Hence, in this paper, we propose \\method, a protocol/network co-optimization framework based on block circulant transformation. At the protocol level, PrivCirNet customizes the HE encoding algorithm that is fully compatible with the block circulant transformation and reduces the computation latency in proportion to the block size. At the network level, we propose a latency-aware formulation to search for the layer-wise block size assignment based on second-order information. PrivCirNet also leverages layer fusion to further reduce the inference cost. We compare PrivCirNet with the state-of-the-art HE-based framework Bolt (IEEE S\\&amp;P 2024) and the HE-friendly pruning method SpENCNN (ICML 2023). For ResNet-18 and Vision Transformer (ViT) on Tiny ImageNet, PrivCirNet reduces latency by $5.0\\times$ and $1.3\\times$ with iso-accuracy over Bolt, respectively, and improves accuracy by $4.1\\%$ and $12\\%$ over SpENCNN, respectively. For MobileNetV2 on ImageNet, PrivCirNet achieves $1.7\\times$ lower latency and $4.2\\%$ better accuracy over Bolt and SpENCNN, respectively. Our code and checkpoints are available on Git Hub.","author":[{"family":"Xu","given":"Tianshi"},{"family":"Wu","given":"Lemeng"},{"family":"Wang","given":"Runsheng"},{"family":"Li","given":"Meng"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.14569","URL":"https://doi.org/10.48550/arxiv.2405.14569","source":"datacite"},{"id":"doi:10.48550/arxiv.2410.21192","type":"manuscript","title":"On Homomorphic Encryption Based Strategies for Class Imbalance in Federated Learning","abstract":"Class imbalance in training datasets can lead to bias and poor generalization in machine learning models. While pre-processing of training datasets can efficiently address both these issues in centralized learning environments, it is challenging to detect and address these issues in a distributed learning environment such as federated learning. In this paper, we propose FLICKER, a privacy preserving framework to address issues related to global class imbalance in federated learning. At the heart of our contribution lies the popular CKKS homomorphic encryption scheme, which is used by the clients to privately share their data attributes, and subsequently balance their datasets before implementing the FL scheme. Extensive experimental results show that our proposed method significantly improves the FL accuracy numbers when used along with popular datasets and relevant baselines.","author":[{"family":"Guleria","given":"Arpit"},{"family":"Harshan","given":"J"},{"family":"Prasad","given":"Ranjitha"},{"family":"Bharath","given":"BN"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2410.21192","URL":"https://doi.org/10.48550/arxiv.2410.21192","source":"datacite"},{"id":"doi:10.48550/arxiv.2408.06167","type":"manuscript","title":"Blind-Match: Efficient Homomorphic Encryption-Based 1:N Matching for Privacy-Preserving Biometric Identification","abstract":"We present Blind-Match, a novel biometric identification system that leverages homomorphic encryption (HE) for efficient and privacy-preserving 1:N matching. Blind-Match introduces a HE-optimized cosine similarity computation method, where the key idea is to divide the feature vector into smaller parts for processing rather than computing the entire vector at once. By optimizing the number of these parts, Blind-Match minimizes execution time while ensuring data privacy through HE. Blind-Match achieves superior performance compared to state-of-the-art methods across various biometric datasets. On the LFW face dataset, Blind-Match attains a 99.63% Rank-1 accuracy with a 128-dimensional feature vector, demonstrating its robustness in face recognition tasks. For fingerprint identification, Blind-Match achieves a remarkable 99.55% Rank-1 accuracy on the PolyU dataset, even with a compact 16-dimensional feature vector, significantly outperforming the state-of-the-art method, Blind-Touch, which achieves only 59.17%. Furthermore, Blind-Match showcases practical efficiency in large-scale biometric identification scenarios, such as Naver Cloud's FaceSign, by processing 6,144 biometric samples in 0.74 seconds using a 128-dimensional feature vector.","author":[{"family":"Choi","given":"Hyunmin"},{"family":"Kim","given":"Jiwon"},{"family":"Song","given":"Chiyoung"},{"family":"Woo","given":"Simon"},{"family":"Kim","given":"Hyoungshick"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2408.06167","URL":"https://doi.org/10.48550/arxiv.2408.06167","source":"datacite"},{"id":"doi:10.48550/arxiv.2402.06609","type":"manuscript","title":"You Still See Me: How Data Protection Supports the Architecture of AI Surveillance","abstract":"Data forms the backbone of artificial intelligence (AI). Privacy and data protection laws thus have strong bearing on AI systems. Shielded by the rhetoric of compliance with data protection and privacy regulations, privacy-preserving techniques have enabled the extraction of more and new forms of data. We illustrate how the application of privacy-preserving techniques in the development of AI systems--from private set intersection as part of dataset curation to homomorphic encryption and federated learning as part of model computation--can further support surveillance infrastructure under the guise of regulatory permissibility. Finally, we propose technology and policy strategies to evaluate privacy-preserving techniques in light of the protections they actually confer. We conclude by highlighting the role that technologists could play in devising policies that combat surveillance AI technologies.","author":[{"family":"Yew","given":"Rui"},{"family":"Qin","given":"Lucy"},{"family":"Venkatasubramanian","given":"Suresh"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2402.06609","URL":"https://doi.org/10.48550/arxiv.2402.06609","source":"datacite"},{"id":"doi:10.48550/arxiv.2307.06554","type":"manuscript","title":"TPU as Cryptographic Accelerator","abstract":"Cryptographic schemes like Fully Homomorphic Encryption (FHE) and Zero-Knowledge Proofs (ZKPs), while offering powerful privacy-preserving capabilities, are often hindered by their computational complexity. Polynomial multiplication, a core operation in these schemes, is a major performance bottleneck. While algorithmic advancements and specialized hardware like GPUs and FPGAs have shown promise in accelerating these computations, the recent surge in AI accelerators (TPUs/NPUs) presents a new opportunity. This paper explores the potential of leveraging TPUs/NPUs to accelerate polynomial multiplication, thereby enhancing the performance of FHE and ZKP schemes. We present techniques to adapt polynomial multiplication to these AI-centric architectures and provide a preliminary evaluation of their effectiveness. We also discuss current limitations and outline future directions for further performance improvements, paving the way for wider adoption of advanced cryptographic tools.","author":[{"family":"Karanjai","given":"Rabimba"},{"family":"Shin","given":"Sangwon"},{"family":"Xiong","given":"And"},{"family":"Fan","given":"Xinxin"},{"family":"Chen","given":"Lin"},{"family":"Zhang","given":"Tianwei"},{"family":"Suh","given":"Taeweon"},{"family":"Shi","given":"Weidong"},{"family":"Kuchta","given":"Veronika"},{"family":"Sica","given":"Francesco"},{"family":"Xu","given":"Lei"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2307.06554","URL":"https://doi.org/10.48550/arxiv.2307.06554","source":"datacite"},{"id":"doi:10.48550/arxiv.2409.16675","type":"manuscript","title":"CryptoTrain: Fast Secure Training on Encrypted Dataset","abstract":"Secure training, while protecting the confidentiality of both data and model weights, typically incurs significant training overhead. Traditional Fully Homomorphic Encryption (FHE)-based non-inter-active training models are heavily burdened by computationally demanding bootstrapping. To develop an efficient secure training system, we established a foundational framework, CryptoTrain-B, utilizing a hybrid cryptographic protocol that merges FHE with Oblivious Transfer (OT) for handling linear and non-linear operations, respectively. This integration eliminates the need for costly bootstrapping. Although CryptoTrain-B sets a new baseline in performance, reducing its training overhead remains essential. We found that ciphertext-ciphertext multiplication (CCMul) is a critical bottleneck in operations involving encrypted inputs and models. Our solution, the CCMul-Precompute technique, involves precomputing CCMul offline and resorting to the less resource-intensive ciphertext-plaintext multiplication (CPMul) during private training. Furthermore, conventional polynomial convolution in FHE systems tends to encode irrelevant and redundant values into polynomial slots, necessitating additional polynomials and ciphertexts for input representation and leading to extra multiplications. Addressing this, we introduce correlated polynomial convolution, which encodes only related input values into polynomials, thus drastically reducing the number of computations and overheads. By integrating CCMul-Precompute and correlated polynomial convolution into CryptoTrain-B, we facilitate a rapid and efficient secure training framework, CryptoTrain. Extensive experiments demonstrate that CryptoTrain achieves a ~5.3X training time reduction compared to prior methods.","author":[{"family":"Xue","given":"Jiaqi"},{"family":"Zhang","given":"Yancheng"},{"family":"Wang","given":"Yanshan"},{"family":"Wang","given":"Xueqiang"},{"family":"Zheng","given":"Hao"},{"family":"Lou","given":"Qian"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2409.16675","URL":"https://doi.org/10.48550/arxiv.2409.16675","source":"datacite"},{"id":"doi:10.4230/lipics.itp.2024.9","type":"article-journal","title":"Verifying Peephole Rewriting in SSA Compiler IRs","abstract":"There is an increasing need for domain-specific reasoning in modern compilers. This has fueled the use of tailored intermediate representations (IRs) based on static single assignment (SSA), like in the MLIR compiler framework. Interactive theorem provers (ITPs) provide strong guarantees for the end-to-end verification of compilers (e.g., CompCert). However, modern compilers and their IRs evolve at a rate that makes proof engineering alongside them prohibitively expensive. Nevertheless, well-scoped push-button automated verification tools such as the Alive peephole verifier for LLVM-IR gained recognition in domains where SMT solvers offer efficient (semi) decision procedures. In this paper, we aim to combine the convenience of automation with the versatility of ITPs for verifying peephole rewrites across domain-specific IRs. We formalize a core calculus for SSA-based IRs that is generic over the IR and covers so-called regions (nested scoping used by many domain-specific IRs in the MLIR ecosystem). Our mechanization in the Lean proof assistant provides a user-friendly frontend for translating MLIR syntax into our calculus. We provide scaffolding for defining and verifying peephole rewrites, offering tactics to eliminate the abstraction overhead of our SSA calculus. We prove correctness theorems about peephole rewriting, as well as two classical program transformations. To evaluate our framework, we consider three use cases from the MLIR ecosystem that cover different levels of abstractions: (1) bitvector rewrites from LLVM, (2) structured control flow, and (3) fully homomorphic encryption. We envision that our mechanization provides a foundation for formally verified rewrites on new domain-specific IRs.","author":[{"family":"Bhat","given":"Siddharth"},{"family":"Keizer","given":"Alex"},{"family":"Hughes","given":"Chris"},{"family":"Goens","given":"Andrés"},{"family":"Grosser","given":"Tobias"}],"issued":{"date-parts":[[2024]]},"DOI":"10.4230/lipics.itp.2024.9","URL":"https://doi.org/10.4230/lipics.itp.2024.9","source":"datacite"},{"id":"doi:10.4230/lipics.itc.2024.11","type":"article-journal","title":"Fast Secure Computations on Shared Polynomials and Applications to Private Set Operations","abstract":"Secure multi-party computation aims to allow a set of players to compute a given function on their secret inputs without revealing any other information than the result of the computation. In this work, we focus on the design of secure multi-party protocols for shared polynomial operations. We consider the classical model where the adversary is honest-but-curious, and where the coefficients (or any secret values) are either encrypted using an additively homomorphic encryption scheme or shared using a threshold linear secret-sharing scheme. Our protocols terminate after a constant number of rounds and minimize the number of secure multiplications. In their seminal article at PKC 2006, Mohassel and Franklin proposed constant-rounds protocols for the main operations on (shared) polynomials. In this work, we improve the fan-in multiplication of nonzero polynomials, the multi-point polynomial evaluation and the polynomial interpolation (on secret points) to reach a quasi-linear complexity (instead of quadratic in Mohassel and Franklin’s work) in the degree of shared input/output polynomials. Computing with shared polynomials is a core component of several multi-party protocols for privacy-preserving operations on private sets, like the private disjointness test or the private set intersection. Using our new protocols, we are able to improve the complexity of such protocols and to design the first variants which always return a correct result.","author":[{"family":"Giorgi","given":"Pascal"},{"family":"Laguillaumie","given":"Fabien"},{"family":"Ottow","given":"Lucas"},{"family":"Vergnaud","given":"Damien"}],"issued":{"date-parts":[[2024]]},"DOI":"10.4230/lipics.itc.2024.11","URL":"https://doi.org/10.4230/lipics.itc.2024.11","source":"datacite"},{"id":"doi:10.48550/arxiv.2407.19871","type":"manuscript","title":"Fast Private Location-based Information Retrieval Over the Torus","abstract":"Location-based services offer immense utility, but also pose significant privacy risks. In response, we propose LocPIR, a novel framework using homomorphic encryption (HE), specifically the TFHE scheme, to preserve user location privacy when retrieving data from public clouds. Our system employs TFHE's expertise in non-polynomial evaluations, crucial for comparison operations. LocPIR showcases minimal client-server interaction, reduced memory overhead, and efficient throughput. Performance tests confirm its computational speed, making it a viable solution for practical scenarios, demonstrated via application to a COVID-19 alert model. Thus, LocPIR effectively addresses privacy concerns in location-based services, enabling secure data sharing from the public cloud.","author":[{"family":"Yoo","given":"Joon"},{"family":"Hong","given":"Mi"},{"family":"Heo","given":"Ji"},{"family":"Lee","given":"Kang"},{"family":"Yoon","given":"Ji"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2407.19871","URL":"https://doi.org/10.48550/arxiv.2407.19871","source":"datacite"},{"id":"doi:10.48550/arxiv.2407.07308","type":"manuscript","title":"BoostCom: Towards Efficient Universal Fully Homomorphic Encryption by Boosting the Word-wise Comparisons","abstract":"Fully Homomorphic Encryption (FHE) allows for the execution of computations on encrypted data without the need to decrypt it first, offering significant potential for privacy-preserving computational operations. Emerging arithmetic-based FHE schemes (ar-FHE), like BGV, demonstrate even better performance in word-wise comparison operations over non-arithmetic FHE (na-FHE) schemes, such as TFHE, especially for basic tasks like comparing values, finding maximums, and minimums. This shows the universality of ar-FHE in effectively handling both arithmetic and non-arithmetic operations without the expensive conversion between arithmetic and non-arithmetic FHEs. We refer to universal arithmetic Fully Homomorphic Encryption as uFHE. The arithmetic operations in uFHE remain consistent with those in the original arithmetic FHE, which have seen significant acceleration. However, its non-arithmetic comparison operations differ, are slow, and have not been as thoroughly studied or accelerated. In this paper, we introduce BoostCom, a scheme designed to speed up word-wise comparison operations, enhancing the efficiency of uFHE systems. BoostCom involves a multi-prong optimizations including infrastructure acceleration (Multi-level heterogeneous parallelization and GPU-related improvements), and algorithm-aware optimizations (slot compaction, non-blocking comparison semantic). Together, BoostCom achieves an end-to-end performance improvement of more than an order of magnitude (11.1x faster) compared to the state-of-the-art CPU-based uFHE systems, across various FHE parameters and tasks.","author":[{"family":"Yudha","given":"Ardhi"},{"family":"Xue","given":"Jiaqi"},{"family":"Lou","given":"Qian"},{"family":"Zhou","given":"Huiyang"},{"family":"Solihin","given":"Yan"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2407.07308","URL":"https://doi.org/10.48550/arxiv.2407.07308","source":"datacite"},{"id":"doi:10.48550/arxiv.2407.03685","type":"manuscript","title":"Verifying Peephole Rewriting In SSA Compiler IRs","abstract":"There is an increasing need for domain-specific reasoning in modern compilers. This has fueled the use of tailored intermediate representations (IRs) based on static single assignment (SSA), like in the MLIR compiler framework. Interactive theorem provers (ITPs) provide strong guarantees for the end-to-end verification of compilers (e.g., CompCert). However, modern compilers and their IRs evolve at a rate that makes proof engineering alongside them prohibitively expensive. Nevertheless, well-scoped push-button automated verification tools such as the Alive peephole verifier for LLVM-IR gained recognition in domains where SMT solvers offer efficient (semi) decision procedures. In this paper, we aim to combine the convenience of automation with the versatility of ITPs for verifying peephole rewrites across domain-specific IRs. We formalize a core calculus for SSA-based IRs that is generic over the IR and covers so-called regions (nested scoping used by many domain-specific IRs in the MLIR ecosystem). Our mechanization in the Lean proof assistant provides a user-friendly frontend for translating MLIR syntax into our calculus. We provide scaffolding for defining and verifying peephole rewrites, offering tactics to eliminate the abstraction overhead of our SSA calculus. We prove correctness theorems about peephole rewriting, as well as two classical program transformations. To evaluate our framework, we consider three use cases from the MLIR ecosystem that cover different levels of abstractions: (1) bitvector rewrites from LLVM, (2) structured control flow, and (3) fully homomorphic encryption. We envision that our mechanization provides a foundation for formally verified rewrites on new domain-specific IRs.","author":[{"family":"Bhat","given":"Siddharth"},{"family":"Keizer","given":"Alex"},{"family":"Hughes","given":"Chris"},{"family":"Goens","given":"Andrés"},{"family":"Grosser","given":"Tobias"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2407.03685","URL":"https://doi.org/10.48550/arxiv.2407.03685","source":"datacite"},{"id":"doi:10.48550/arxiv.2406.06808","type":"manuscript","title":"Fast White-Box Adversarial Streaming Without a Random Oracle","abstract":"Recently, the question of adversarially robust streaming, where the stream is allowed to depend on the randomness of the streaming algorithm, has gained a lot of attention. In this work, we consider a strong white-box adversarial model (Ajtai et al. PODS 2022), in which the adversary has access to all past random coins and the parameters used by the streaming algorithm. We focus on the sparse recovery problem and extend our result to other tasks such as distinct element estimation and low-rank approximation of matrices and tensors. The main drawback of previous work is that it requires a random oracle, which is especially problematic in the streaming model since the amount of randomness is counted in the space complexity of a streaming algorithm. Also, the previous work suffers from large update time. We construct a near-optimal solution for the sparse recovery problem in white-box adversarial streams, based on the subexponentially secure Learning with Errors assumption. Importantly, our solution does not require a random oracle and has a polylogarithmic per item processing time. We also give results in a related white-box adversarially robust distributed model. Our constructions are based on homomorphic encryption schemes satisfying very mild structural properties that are currently satisfied by most known schemes.","author":[{"family":"Feng","given":"Ying"},{"family":"Jain","given":"Aayush"},{"family":"Woodruff","given":"David"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2406.06808","URL":"https://doi.org/10.48550/arxiv.2406.06808","source":"datacite"},{"id":"doi:10.48550/arxiv.2404.06216","type":"manuscript","title":"Privacy-preserving Scanpath Comparison for Pervasive Eye Tracking","abstract":"As eye tracking becomes pervasive with screen-based devices and head-mounted displays, privacy concerns regarding eye-tracking data have escalated. While state-of-the-art approaches for privacy-preserving eye tracking mostly involve differential privacy and empirical data manipulations, previous research has not focused on methods for scanpaths. We introduce a novel privacy-preserving scanpath comparison protocol designed for the widely used Needleman-Wunsch algorithm, a generalized version of the edit distance algorithm. Particularly, by incorporating the Paillier homomorphic encryption scheme, our protocol ensures that no private information is revealed. Furthermore, we introduce a random processing strategy and a multi-layered masking method to obfuscate the values while preserving the original order of encrypted editing operation costs. This minimizes communication overhead, requiring a single communication round for each iteration of the Needleman-Wunsch process. We demonstrate the efficiency and applicability of our protocol on three publicly available datasets with comprehensive computational performance analyses and make our source code publicly accessible.","author":[{"family":"Ozdel","given":"Suleyman"},{"family":"Bozkir","given":"Efe"},{"family":"Kasneci","given":"Enkelejda"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2404.06216","URL":"https://doi.org/10.48550/arxiv.2404.06216","source":"datacite"},{"id":"doi:10.48550/arxiv.2312.05264","type":"manuscript","title":"All Rivers Run to the Sea: Private Learning with Asymmetric Flows","abstract":"Data privacy is of great concern in cloud machine-learning service platforms, when sensitive data are exposed to service providers. While private computing environments (e.g., secure enclaves), and cryptographic approaches (e.g., homomorphic encryption) provide strong privacy protection, their computing performance still falls short compared to cloud GPUs. To achieve privacy protection with high computing performance, we propose Delta, a new private training and inference framework, with comparable model performance as non-private centralized training. Delta features two asymmetric data flows: the main information-sensitive flow and the residual flow. The main part flows into a small model while the residuals are offloaded to a large model. Specifically, Delta embeds the information-sensitive representations into a low-dimensional space while pushing the information-insensitive part into high-dimension residuals. To ensure privacy protection, the low-dimensional information-sensitive part is secured and fed to a small model in a private environment. On the other hand, the residual part is sent to fast cloud GPUs, and processed by a large model. To further enhance privacy and reduce the communication cost, Delta applies a random binary quantization technique along with a DP-based technique to the residuals before sharing them with the public platform. We theoretically show that Delta guarantees differential privacy in the public environment and greatly reduces the complexity in the private environment. We conduct empirical analyses on CIFAR-10, CIFAR-100 and ImageNet datasets and ResNet-18 and ResNet-34, showing that Delta achieves strong privacy protection, fast training, and inference without significantly compromising the model utility.","author":[{"family":"Niu","given":"Yue"},{"family":"Ali","given":"Ramy"},{"family":"Prakash","given":"Saurav"},{"family":"Avestimehr","given":"Salman"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2312.05264","URL":"https://doi.org/10.48550/arxiv.2312.05264","source":"datacite"},{"id":"doi:10.48550/arxiv.2308.04890","type":"manuscript","title":"CiFHER: A Chiplet-Based FHE Accelerator with a Resizable Structure","abstract":"Fully homomorphic encryption (FHE) is in the spotlight as a definitive solution for privacy, but the high computational overhead of FHE poses a challenge to its practical adoption. Although prior studies have attempted to design ASIC accelerators to mitigate the overhead, their designs require excessive chip resources (e.g., areas) to contain and process massive data for FHE operations. We propose CiFHER, a chiplet-based FHE accelerator with a resizable structure, to tackle the challenge with a cost-effective multi-chip module (MCM) design. First, we devise a flexible core architecture whose configuration is adjustable to conform to the global organization of chiplets and design constraints. Its distinctive feature is a composable functional unit providing varying computational throughput for the number-theoretic transform, the most dominant function in FHE. Then, we establish generalized data mapping methodologies to minimize the interconnect overhead when organizing the chips into the MCM package in a tiled manner, which becomes a significant bottleneck due to the packaging constraints. This study demonstrates that a CiFHER package composed of a number of compact chiplets provides performance comparable to state-of-the-art monolithic ASIC accelerators while significantly reducing the package-wide power consumption and manufacturing cost.","author":[{"family":"Kim","given":"Sangpyo"},{"family":"Kim","given":"Jongmin"},{"family":"Choi","given":"Jaeyoung"},{"family":"Ahn","given":"Jung"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2308.04890","URL":"https://doi.org/10.48550/arxiv.2308.04890","source":"datacite"},{"id":"doi:10.48550/arxiv.2401.09604","type":"manuscript","title":"MedBlindTuner: Towards Privacy-preserving Fine-tuning on Biomedical Images with Transformers and Fully Homomorphic Encryption","abstract":"Advancements in machine learning (ML) have significantly revolutionized medical image analysis, prompting hospitals to rely on external ML services. However, the exchange of sensitive patient data, such as chest X-rays, poses inherent privacy risks when shared with third parties. Addressing this concern, we propose MedBlindTuner, a privacy-preserving framework leveraging fully homomorphic encryption (FHE) and a data-efficient image transformer (DEiT). MedBlindTuner enables the training of ML models exclusively on FHE-encrypted medical images. Our experimental evaluation demonstrates that MedBlindTuner achieves comparable accuracy to models trained on non-encrypted images, offering a secure solution for outsourcing ML computations while preserving patient data privacy. To the best of our knowledge, this is the first work that uses data-efficient image transformers and fully homomorphic encryption in this domain.","author":[{"family":"Panzade","given":"Prajwal"},{"family":"Takabi","given":"Daniel"},{"family":"Cai","given":"Zhipeng"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2401.09604","URL":"https://doi.org/10.48550/arxiv.2401.09604","source":"datacite"},{"id":"doi:10.5281/zenodo.8087573","type":"article-journal","title":"An Efficient Protocol For Computations Delegation Using Rerandomizable Garbled Circuits","abstract":"We consider the problem of delegating computations, in which a client may outsource the computation of a function, represented as a boolean circuit, to an external entity, then receives the result and a proof of correctness that can be verified efficiently. Most existing noninteractive solutions rely on the use of fully homomorphic encryption, which is practically unaffordable. As it turns out, making a probabilistic assumption about the honesty of the external entity eliminates the need for FHE. This paper describes an efficient protocol for delegating computations that also guarantees input and output privacy. The protocol is a new variant of the multi-server model and the cycle architecture from Ananth et al., which is based on the assumption of the existence of at least one honest server. We carefully utilize an alternative re-randomizable variant of Yao’s garbled circuits based on the ElGamal encryption scheme, which leads to a secure, private, and efficient protocol based on the DDH assumption. Furthermore, we describe an example implementation of an application of the protocol: a fully decentralized verifiable system for outsourcing computations, VDCS, which serves as a proof-of-concept prototype.","author":[{"family":"Ibrahim","given":"Ahmed"},{"family":"Gouhar","given":"Amr"},{"family":"Ghazy","given":"Ahmed"},{"family":"Ahmed","given":"Mohamed"},{"family":"Youssif","given":"George"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.8087573","URL":"https://doi.org/10.5281/zenodo.8087573","source":"datacite"},{"id":"doi:10.5281/zenodo.8087574","type":"article-journal","title":"An Efficient Protocol For Computations Delegation Using Rerandomizable Garbled Circuits","abstract":"We consider the problem of delegating computations, in which a client may outsource the computation of a function, represented as a boolean circuit, to an external entity, then receives the result and a proof of correctness that can be verified efficiently. Most existing noninteractive solutions rely on the use of fully homomorphic encryption, which is practically unaffordable. As it turns out, making a probabilistic assumption about the honesty of the external entity eliminates the need for FHE. This paper describes an efficient protocol for delegating computations that also guarantees input and output privacy. The protocol is a new variant of the multi-server model and the cycle architecture from Ananth et al., which is based on the assumption of the existence of at least one honest server. We carefully utilize an alternative re-randomizable variant of Yao’s garbled circuits based on the ElGamal encryption scheme, which leads to a secure, private, and efficient protocol based on the DDH assumption. Furthermore, we describe an example implementation of an application of the protocol: a fully decentralized verifiable system for outsourcing computations, VDCS, which serves as a proof-of-concept prototype.","author":[{"family":"Ibrahim","given":"Ahmed"},{"family":"Gouhar","given":"Amr"},{"family":"Ghazy","given":"Ahmed"},{"family":"Ahmed","given":"Mohamed"},{"family":"Youssif","given":"George"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.8087574","URL":"https://doi.org/10.5281/zenodo.8087574","source":"datacite"},{"id":"doi:10.48550/arxiv.2412.01858","type":"manuscript","title":"MQFL-FHE: Multimodal Quantum Federated Learning Framework with Fully Homomorphic Encryption","abstract":"The integration of fully homomorphic encryption (FHE) in federated learning (FL) has led to significant advances in data privacy. However, during the aggregation phase, it often results in performance degradation of the aggregated model, hindering the development of robust representational generalization. In this work, we propose a novel multimodal quantum federated learning framework that utilizes quantum computing to counteract the performance drop resulting from FHE. For the first time in FL, our framework combines a multimodal quantum mixture of experts (MQMoE) model with FHE, incorporating multimodal datasets for enriched representation and task-specific learning. Our MQMoE framework enhances performance on multimodal datasets and combined genomics and brain MRI scans, especially for underrepresented categories. Our results also demonstrate that the quantum-enhanced approach mitigates the performance degradation associated with FHE and improves classification accuracy across diverse datasets, validating the potential of quantum interventions in enhancing privacy in FL.","author":[{"family":"Dutta","given":"Siddhant"},{"family":"Innan","given":"Nouhaila"},{"family":"Yahia","given":"Sadok"},{"family":"Shafique","given":"Muhammad"},{"family":"Neira","given":"David"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2412.01858","URL":"https://doi.org/10.48550/arxiv.2412.01858","source":"datacite"},{"id":"doi:10.26181/22357747.v1","type":"article-journal","title":"A Review of Homomorphic Encryption for Privacy-Preserving Biometrics","abstract":"The advancement of biometric technology has facilitated wide applications of biometrics in law enforcement, border control, healthcare and financial identification and verification. Given the peculiarity of biometric features (e.g., unchangeability, permanence and uniqueness), the security of biometric data is a key area of research. Security and privacy are vital to enacting integrity, reliability and availability in biometric-related applications. Homomorphic encryption (HE) is concerned with data manipulation in the cryptographic domain, thus addressing the security and privacy issues faced by biometrics. This survey provides a comprehensive review of state-of-the-art HE research in the context of biometrics. Detailed analyses and discussions are conducted on various HE approaches to biometric security according to the categories of different biometric traits. Moreover, this review presents the perspective of integrating HE with other emerging technologies (e.g., machine/deep learning and blockchain) for biometric security. Finally, based on the latest development of HE in biometrics, challenges and future research directions are put forward.","author":[{"family":"Yang","given":"Wencheng"},{"family":"Wang","given":"Song"},{"family":"Cui","given":"Hui"},{"family":"Tang","given":"Zhaohui"},{"family":"Li","given":"Yan"}],"issued":{"date-parts":[[2023]]},"DOI":"10.26181/22357747.v1","URL":"https://doi.org/10.26181/22357747.v1","source":"datacite"},{"id":"doi:10.26181/22357747","type":"article-journal","title":"A Review of Homomorphic Encryption for Privacy-Preserving Biometrics","abstract":"The advancement of biometric technology has facilitated wide applications of biometrics in law enforcement, border control, healthcare and financial identification and verification. Given the peculiarity of biometric features (e.g., unchangeability, permanence and uniqueness), the security of biometric data is a key area of research. Security and privacy are vital to enacting integrity, reliability and availability in biometric-related applications. Homomorphic encryption (HE) is concerned with data manipulation in the cryptographic domain, thus addressing the security and privacy issues faced by biometrics. This survey provides a comprehensive review of state-of-the-art HE research in the context of biometrics. Detailed analyses and discussions are conducted on various HE approaches to biometric security according to the categories of different biometric traits. Moreover, this review presents the perspective of integrating HE with other emerging technologies (e.g., machine/deep learning and blockchain) for biometric security. Finally, based on the latest development of HE in biometrics, challenges and future research directions are put forward.","author":[{"family":"Yang","given":"Wencheng"},{"family":"Wang","given":"Song"},{"family":"Cui","given":"Hui"},{"family":"Tang","given":"Zhaohui"},{"family":"Li","given":"Yan"}],"issued":{"date-parts":[[2023]]},"DOI":"10.26181/22357747","URL":"https://doi.org/10.26181/22357747","source":"datacite"},{"id":"doi:10.48550/arxiv.2412.01650","type":"manuscript","title":"Privacy-Preserving Federated Learning via Homomorphic Adversarial Networks","abstract":"Privacy-preserving federated learning (PPFL) aims to train a global model for multiple clients while maintaining their data privacy. However, current PPFL protocols exhibit one or more of the following insufficiencies: considerable degradation in accuracy, the requirement for sharing keys, and cooperation during the key generation or decryption processes. As a mitigation, we develop the first protocol that utilizes neural networks to implement PPFL, as well as incorporating an Aggregatable Hybrid Encryption scheme tailored to the needs of PPFL. We name these networks as Homomorphic Adversarial Networks (HANs) which demonstrate that neural networks are capable of performing tasks similar to multi-key homomorphic encryption (MK-HE) while solving the problems of key distribution and collaborative decryption. Our experiments show that HANs are robust against privacy attacks. Compared with non-private federated learning, experiments conducted on multiple datasets demonstrate that HANs exhibit a negligible accuracy loss (at most 1.35%). Compared to traditional MK-HE schemes, HANs increase encryption aggregation speed by 6,075 times while incurring a 29.2 times increase in communication overhead.","author":[{"family":"Dong","given":"Wenhan"},{"family":"Lin","given":"Chao"},{"family":"He","given":"Xinlei"},{"family":"Xu","given":"Shengmin"},{"family":"Huang","given":"Xinyi"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2412.01650","URL":"https://doi.org/10.48550/arxiv.2412.01650","source":"datacite"},{"id":"doi:10.48550/arxiv.2412.07187","type":"manuscript","title":"A New Federated Learning Framework Against Gradient Inversion Attacks","abstract":"Federated Learning (FL) aims to protect data privacy by enabling clients to collectively train machine learning models without sharing their raw data. However, recent studies demonstrate that information exchanged during FL is subject to Gradient Inversion Attacks (GIA) and, consequently, a variety of privacy-preserving methods have been integrated into FL to thwart such attacks, such as Secure Multi-party Computing (SMC), Homomorphic Encryption (HE), and Differential Privacy (DP). Despite their ability to protect data privacy, these approaches inherently involve substantial privacy-utility trade-offs. By revisiting the key to privacy exposure in FL under GIA, which lies in the frequent sharing of model gradients that contain private data, we take a new perspective by designing a novel privacy preserve FL framework that effectively ``breaks the direct connection'' between the shared parameters and the local private data to defend against GIA. Specifically, we propose a Hypernetwork Federated Learning (HyperFL) framework that utilizes hypernetworks to generate the parameters of the local model and only the hypernetwork parameters are uploaded to the server for aggregation. Theoretical analyses demonstrate the convergence rate of the proposed HyperFL, while extensive experimental results show the privacy-preserving capability and comparable performance of HyperFL. Code is available at https://github.com/Pengxin-Guo/HyperFL.","author":[{"family":"Guo","given":"Pengxin"},{"family":"Zeng","given":"Shuang"},{"family":"Chen","given":"Wenhao"},{"family":"Zhang","given":"Xiaodan"},{"family":"Ren","given":"Weihong"},{"family":"Zhou","given":"Yuyin"},{"family":"Qu","given":"Liangqiong"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2412.07187","URL":"https://doi.org/10.48550/arxiv.2412.07187","source":"datacite"},{"id":"doi:10.48550/arxiv.2310.02563","type":"manuscript","title":"Practical, Private Assurance of the Value of Collaboration via Fully Homomorphic Encryption","abstract":"Two parties wish to collaborate on their datasets. However, before they reveal their datasets to each other, the parties want to have the guarantee that the collaboration would be fruitful. We look at this problem from the point of view of machine learning, where one party is promised an improvement on its prediction model by incorporating data from the other party. The parties would only wish to collaborate further if the updated model shows an improvement in accuracy. Before this is ascertained, the two parties would not want to disclose their models and datasets. In this work, we construct an interactive protocol for this problem based on the fully homomorphic encryption scheme over the Torus (TFHE) and label differential privacy, where the underlying machine learning model is a neural network. Label differential privacy is used to ensure that computations are not done entirely in the encrypted domain, which is a significant bottleneck for neural network training according to the current state-of-the-art FHE implementations. We formally prove the security of our scheme assuming honest-but-curious parties, but where one party may not have any expertise in labelling its initial dataset. Experiments show that we can obtain the output, i.e., the accuracy of the updated model, with time many orders of magnitude faster than a protocol using entirely FHE operations.","author":[{"family":"Asghar","given":"Hassan"},{"family":"Lu","given":"Zhigang"},{"family":"Zhao","given":"Zhongrui"},{"family":"Kaafar","given":"Dali"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2310.02563","URL":"https://doi.org/10.48550/arxiv.2310.02563","source":"datacite"},{"id":"doi:10.48550/arxiv.2407.13055","type":"manuscript","title":"Cheddar: A Swift Fully Homomorphic Encryption Library Designed for GPU Architectures","abstract":"Fully homomorphic encryption (FHE) frees cloud computing from privacy concerns by enabling secure computation on encrypted data. However, its substantial computational and memory overhead results in significantly slower performance compared to unencrypted processing. To mitigate this overhead, we present Cheddar, a high-performance FHE library for GPUs, achieving substantial speedups over previous GPU implementations. We systematically enable 32-bit FHE execution, leveraging the 32-bit integer datapath within GPUs. We optimize GPU kernels using efficient low-level primitives and algorithms tailored to specific GPU architectures. Further, we alleviate the memory bandwidth burden by adjusting common FHE operational sequences and extensively applying kernel fusion. Cheddar delivers performance improvements of 2.18--4.45$\\times$ for representative FHE workloads compared to state-of-the-art GPU implementations.","author":[{"family":"Choi","given":"Wonseok"},{"family":"Kim","given":"Jongmin"},{"family":"Ahn","given":"Jung"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2407.13055","URL":"https://doi.org/10.48550/arxiv.2407.13055","source":"datacite"},{"id":"doi:10.48550/arxiv.2308.03734","type":"manuscript","title":"Labeling without Seeing? Blind Annotation for Privacy-Preserving Entity Resolution","abstract":"The entity resolution problem requires finding pairs across datasets that belong to different owners but refer to the same entity in the real world. To train and evaluate solutions (either rule-based or machine-learning-based) to the entity resolution problem, generating a ground truth dataset with entity pairs or clusters is needed. However, such a data annotation process involves humans as domain oracles to review the plaintext data for all candidate record pairs from different parties, which inevitably infringes the privacy of data owners, especially in privacy-sensitive cases like medical records. To the best of our knowledge, there is no prior work on privacy-preserving ground truth dataset generation, especially in the domain of entity resolution. We propose a novel blind annotation protocol based on homomorphic encryption that allows domain oracles to collaboratively label ground truths without sharing data in plaintext with other parties. In addition, we design a domain-specific easy-to-use language that hides the sophisticated underlying homomorphic encryption layer. Rigorous proof of the privacy guarantee is provided and our empirical experiments via an annotation simulator indicate the feasibility of our privacy-preserving protocol (f-measure on average achieves more than 90\\% compared with the real ground truths).","author":[{"family":"Yao","given":"Yixiang"},{"family":"Jin","given":"Weizhao"},{"family":"Ravi","given":"Srivatsan"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2308.03734","URL":"https://doi.org/10.48550/arxiv.2308.03734","source":"datacite"},{"id":"doi:10.48550/arxiv.2402.15738","type":"manuscript","title":"Privacy-Preserving State Estimation in the Presence of Eavesdroppers: A Survey","abstract":"Networked systems are increasingly the target of cyberattacks that exploit vulnerabilities within digital communications, embedded hardware, and software. Arguably, the simplest class of attacks -- and often the first type before launching destructive integrity attacks -- are eavesdropping attacks, which aim to infer information by collecting system data and exploiting it for malicious purposes. A key technology of networked systems is state estimation, which leverages sensing and actuation data and first-principles models to enable trajectory planning, real-time monitoring, and control. However, state estimation can also be exploited by eavesdroppers to identify models and reconstruct states with the aim of, e.g., launching integrity (stealthy) attacks and inferring sensitive information. It is therefore crucial to protect disclosed system data to avoid an accurate state estimation by eavesdroppers. This survey presents a comprehensive review of existing literature on privacy-preserving state estimation methods, while also identifying potential limitations and research gaps. Our primary focus revolves around three types of methods: cryptography, data perturbation, and transmission scheduling, with particular emphasis on Kalman-like filters. Within these categories, we delve into the concepts of homomorphic encryption and differential privacy, which have been extensively investigated in recent years in the context of privacy-preserving state estimation. Finally, we shed light on several technical and fundamental challenges surrounding current methods and propose potential directions for future research.","author":[{"family":"Yan","given":"Xinhao"},{"family":"Zhou","given":"Guanzhong"},{"family":"Quevedo","given":"Daniel"},{"family":"Murguia","given":"Carlos"},{"family":"Chen","given":"Bo"},{"family":"Huang","given":"Hailong"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2402.15738","URL":"https://doi.org/10.48550/arxiv.2402.15738","source":"datacite"},{"id":"doi:10.48550/arxiv.2310.12401","type":"manuscript","title":"Privacy-Preserving Hierarchical Anonymization Framework over Encrypted Data","abstract":"Smart cities, which can monitor the real world and provide smart services in a variety of fields, have improved people's living standards as urbanization has accelerated. However, there are security and privacy concerns because smart city applications collect large amounts of privacy-sensitive information from people and their social circles. Anonymization, which generalizes data and reduces data uniqueness is an important step in preserving the privacy of sensitive information. However, anonymization methods frequently require large datasets and rely on untrusted third parties to collect and manage data, particularly in a cloud environment. In this case, private data leakage remains a critical issue, discouraging users from sharing their data and impeding the advancement of smart city services. This problem can be solved if the computational entity can perform the anonymization process without obtaining the original plain text. This study proposed a hierarchical k-anonymization framework using homomorphic encryption and secret sharing composed of two types of domains. Different computing methods are selected flexibly, and two domains are connected hierarchically to obtain higher-level anonymization results in an efficient manner. The experimental results show that connecting two domains can accelerate the anonymization process, indicating that the proposed secure hierarchical architecture is practical and efficient.","author":[{"family":"Jia","given":"Jing"},{"family":"Saito","given":"Kenta"},{"family":"Nishi","given":"Hiroaki"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2310.12401","URL":"https://doi.org/10.48550/arxiv.2310.12401","source":"datacite"},{"id":"doi:10.48550/arxiv.2308.14725","type":"manuscript","title":"Applications of Finite non-Abelian Simple Groups to Cryptography in the Quantum Era","abstract":"The theory of finite simple groups is a (rather unexplored) area likely to provide interesting computational problems and modelling tools useful in a cryptographic context. In this note, we review some applications of finite non-abelian simple groups to cryptography and discuss different scenarios in which this theory is clearly central, providing the relevant definitions to make the material accessible to both cryptographers and group theorists, in the hope of stimulating further interaction between these two (non-disjoint) communities. In particular, we look at constructions based on various group-theoretic factorization problems, review group theoretical hash functions, and discuss fully homomorphic encryption using simple groups. The Hidden Subgroup Problem is also briefly discussed in this context.","author":[{"family":"Vasco","given":"María"},{"family":"Kahrobaei","given":"Delaram"},{"family":"Mckemmie","given":"Eilidh"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2308.14725","URL":"https://doi.org/10.48550/arxiv.2308.14725","source":"datacite"},{"id":"doi:10.48550/arxiv.2305.02225","type":"manuscript","title":"Data Privacy with Homomorphic Encryption in Neural Networks Training and Inference","abstract":"The use of Neural Networks (NNs) for sensitive data processing is becoming increasingly popular, raising concerns about data privacy and security. Homomorphic Encryption (HE) has the potential to be used as a solution to preserve data privacy in NN. This study provides a comprehensive analysis on the use of HE for NN training and classification, focusing on the techniques and strategies used to enhance data privacy and security. The current state-of-the-art in HE for NNs is analysed, and the challenges and limitations that need to be addressed to make it a reliable and efficient approach for privacy preservation are identified. Also, the different categories of HE schemes and their suitability for NNs are discussed, as well as the techniques used to optimize the accuracy and efficiency of encrypted models. The review reveals that HE has the potential to provide strong data privacy guarantees for NNs, but several challenges need to be addressed, such as limited support for advanced NN operations, scalability issues, and performance trade-offs.","author":[{"family":"Amorim","given":"Ivone"},{"family":"Maia","given":"Eva"},{"family":"Barbosa","given":"Pedro"},{"family":"Praça","given":"Isabel"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2305.02225","URL":"https://doi.org/10.48550/arxiv.2305.02225","source":"datacite"},{"id":"doi:10.48550/arxiv.2411.14639","type":"manuscript","title":"Differentially Private Adaptation of Diffusion Models via Noisy Aggregated Embeddings","abstract":"Personalizing large-scale diffusion models poses serious privacy risks, especially when adapting to small, sensitive datasets. A common approach is to fine-tune the model using differentially private stochastic gradient descent (DP-SGD), but this suffers from severe utility degradation due to the high noise needed for privacy, particularly in the small data regime. We propose an alternative that leverages Textual Inversion (TI), which learns an embedding vector for an image or set of images, to enable adaptation under differential privacy (DP) constraints. Our approach, Differentially Private Aggregation via Textual Inversion (DPAgg-TI), adds calibrated noise to the aggregation of per-image embeddings to ensure formal DP guarantees while preserving high output fidelity. We show that DPAgg-TI outperforms DP-SGD finetuning in both utility and robustness under the same privacy budget, achieving results closely matching the non-private baseline on style adaptation tasks using private artwork from a single artist and Paris 2024 Olympic pictograms. In contrast, DP-SGD fails to generate meaningful outputs in this setting.","author":[{"family":"Peetathawatchai","given":"Pura"},{"family":"Chen","given":"Wei"},{"family":"Isik","given":"Berivan"},{"family":"Koyejo","given":"Sanmi"},{"family":"No","given":"Albert"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2411.14639","URL":"https://doi.org/10.48550/arxiv.2411.14639","source":"datacite"},{"id":"oa:W4387891747","type":"manuscript","title":"pFedLoRA: Model-Heterogeneous Personalized Federated Learning with LoRA Tuning","abstract":"Federated learning (FL) is an emerging machine learning paradigm in which a central server coordinates multiple participants (clients) collaboratively to train on decentralized data. In practice, FL often faces statistical, system, and model heterogeneities, which inspires the field of Model-Heterogeneous Personalized Federated Learning (MHPFL). With the increased interest in adopting large language models (LLMs) in FL, the existing MHPFL methods cannot achieve acceptable computational and communication costs, while maintaining satisfactory model performance. To bridge this gap, we propose a novel and efficient model-heterogeneous personalized Federated learning framework based on LoRA tuning (pFedLoRA). Inspired by the popular LoRA method for fine-tuning pre-trained LLMs with a low-rank model (a.k.a., an adapter), we design a homogeneous small adapter to facilitate federated client's heterogeneous local model training with our proposed iterative training for global-local knowledge exchange. The homogeneous small local adapters are aggregated on the FL server to generate a global adapter. We theoretically prove the convergence of pFedLoRA. Extensive experiments on two benchmark datasets demonstrate that pFedLoRA outperforms six state-of-the-art baselines, beating the best method by 1.35% in test accuracy, 11.81 times computation overhead reduction and 7.41 times communication cost saving.","author":[{"family":"Yi","given":"Liping"},{"family":"Yu","given":"Han"},{"family":"Wang","given":"Gang"},{"family":"Liu","given":"Xiaoguang"},{"family":"Li","given":"Xiaoxiao"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2310.13283","URL":"https://doi.org/10.48550/arxiv.2310.13283","source":"openalex"},{"id":"oa:W4401918924","type":"article-journal","title":"Automated machine learning with interpretation: A systematic review of methodologies and applications in healthcare","abstract":"Abstract Machine learning (ML) has achieved substantial success in performing healthcare tasks in which the configuration of every part of the ML pipeline relies heavily on technical knowledge. To help professionals with borderline expertise to better use ML techniques, Automated ML (AutoML) has emerged as a prospective solution. However, most models generated by AutoML are black boxes that are challenging to comprehend and deploy in healthcare settings. We conducted a systematic review to examine AutoML with interpretation systems for healthcare. We searched four databases (MEDLINE, EMBASE, Web of Science, and Scopus) complemented with seven prestigious ML conferences (AAAI, ACL, ICLR, ICML, IJCAI, KDD, and NeurIPS) that reported AutoML with interpretation for healthcare before September 1, 2023. We included 118 articles related to AutoML with interpretation in healthcare. First, we illustrated AutoML techniques used in the included publications, including automated data preparation, automated feature engineering, and automated model development, accompanied by a real‐world case study to demonstrate the advantages of AutoML over classic ML. Then, we summarized interpretation methods: feature interaction and importance, data dimensionality reduction, intrinsically interpretable models, and knowledge distillation and rule extraction. Finally, we detailed how AutoML with interpretation has been used for six major data types: image, free text, tabular data, signal, genomic sequences, and multi‐modality. To some extent, AutoML with interpretation provides effortless development and improves users' trust in ML in healthcare settings. In future studies, researchers should explore automated data preparation, seamless integration of automation and interpretation, compatibility with multi‐modality, and utilization of foundation models.","author":[{"family":"Yuan","given":"Han"},{"family":"Yu","given":"Kunyu"},{"family":"Xie","given":"Feng"},{"family":"Liu","given":"M"},{"family":"Sun","given":"Shenghuan"}],"issued":{"date-parts":[[2024]]},"DOI":"10.1002/med4.75","URL":"https://doi.org/10.1002/med4.75","source":"openalex"},{"id":"oa:W4405358103","type":"article-journal","title":"Federated Learning‐Based Intrusion Detection Systems for Massive IoT","abstract":"Despite the significant benefits that the 6G-enabled massive Internet of Things (IoT) applications will bring to the economy and society in the coming years, it is expected that the exponential increase in the number and the diversity of IoT devices and connections in the massive IoT ecosystem will raise a plethora of known and unknown security and privacy challenges for these applications in the 6G era. Consequently, novel security solutions effectively protecting massive IoT applications from future adversaries, while taking into consideration the resource-constrained characteristics of the massive IoT ecosystem and the stringent network performance and privacy-preserving requirements of these applications, are critical for their acceptance and wide adoption in the upcoming 6G era. In this context, Federated Learning (FL)-based intrusion detection is a promising solution for effective intrusion detection in the massive IoT ecosystem, as it scales well with the massive growth of resource-constrained IoT devices and the wide geographical spread of generated IoT data across wide-area IoT networks. In addition, the FL approach enables local model training at each client based on its locally available training dataset instead of sending it to a remote central server (i.e. centralized intrusion detection), which may bring single-point failure risks and compromise the privacy of the dataset. Furthermore, the aggregation of locally trained models, supported by FL, allows the quick development of accurate models even for devices generating only few training data. Thus, this chapter is focused on the investigation of existing FL-based Intrusion Detection Systems (FL-based IDSs) in order to provide a roadmap to support the design and development of effective, efficient, and privacy-preserving IDSs for protecting the emerging disruptive massive IoT applications in the 6G era.","author":[{"family":"Pelekoudasoikonomou","given":"Filippos"},{"family":"Mirzaee","given":"Parya"},{"family":"Hathal","given":"Waleed"},{"family":"Μαντάς","given":"Γεώργιος"},{"family":"Rodrıguez","given":"Jonathan"},{"family":"Cruickshank","given":"Haitham"},{"family":"Sun","given":"Zhili"}],"issued":{"date-parts":[[2024]]},"DOI":"10.1002/9781119988007.ch4","URL":"https://doi.org/10.1002/9781119988007.ch4","source":"openalex"},{"id":"oa:W4388517677","type":"article-journal","title":"The Sugeno Integral Used for Federated Learning with Uncertainty for Unbalanced Data","abstract":"Data is crucial in the digital economy. Many businesses collect and use their data to enhance their performance. However, limited data or low data quality can hinder model development, particularly in dynamic environments. To overcome this, companies collecting similar data may opt to exchange knowledge without sharing their data, due to privacy or legal issues. This is where federated learning comes in. In horizontal federated learning, each client (organization) iteratively improves its model, so that it can be regularly aggregated and shared with all clients participating in the federation for further improvements. In federated averaging, the aggregation mechanism is based on the weighted average and the weights depend on the amount of data available to each client. In this paper, we propose to use a more advanced aggregation mechanism, namely the Sugeno integral. The initial results are promising.","author":[{"family":"Wilbik","given":"Anna"},{"family":"Pȩkala","given":"Barbara"},{"family":"Szkoła","given":"Jarosław"},{"family":"Dyczkowski","given":"Krzysztof"}],"issued":{"date-parts":[[2023]]},"DOI":"10.1109/fuzz52849.2023.10309680","URL":"https://doi.org/10.1109/fuzz52849.2023.10309680","source":"openalex"},{"id":"oa:W4401780339","type":"article-journal","title":"The blockchain‐based privacy‐preserving searchable attribute‐based encryption scheme for federated learning model in IoMT","abstract":"Abstract Federated learning enables training healthcare diagnostic models across multiple decentralized devices containing local private health data samples, without transferring data to a central server, providing privacy‐preserving services for healthcare professionals. However, for a model of a specific field, some medical data from non‐target participants may be included in model training, compromising model accuracy. Moreover, diagnostic queries for healthcare models stored in cloud servers may result in the leakage of the privacy of healthcare participants and the parameters of models. Furthermore, the records of model searching and usage could be tracked causing privacy disclosure risk. To address these issues, we propose a blockchain‐based privacy‐preserving searchable attribute‐based encryption scheme for the diagnostic model federated learning in the Internet of Medical Things (BSAEM‐FL). We first adopt fine‐grained model trainer participation policies for federated learning, using the attribute‐based encryption (ABE) mechanism, to realize model accuracy and local data privacy. Then, We employ searchable encryption technology for model training and usage to protect the security of models stored in the cloud server. Blockchain is utilized to implement distributed healthcare models' keyword‐based search and model users' attribute‐based authentication. Lastly, we transfer most of the computational overhead of user terminals in model searching and decryption to edge nodes, achieving lightweight computation of IoMT terminals. The security analysis proves the security of the proposed healthcare scheme. The performance evaluation indicates our scheme is of better feasibility, efficiency, and decentralization.","author":[{"family":"Zhou","given":"Ziyu"},{"family":"Wang","given":"Na"},{"family":"Liu","given":"Jianwei"},{"family":"Fu","given":"Junsong"},{"family":"Deng","given":"Lunzhi"}],"issued":{"date-parts":[[2024]]},"DOI":"10.1002/cpe.8257","URL":"https://doi.org/10.1002/cpe.8257","source":"openalex"},{"id":"oa:W4380481519","type":"article-journal","title":"FedCache: A Knowledge Cache-driven Federated Learning Architecture for Personalized Edge Intelligence","abstract":"Edge Intelligence (EI) enables Artificial Intelligence (AI) applications to run at the edge, where data analysis and decision-making can be performed in real-time and close to data sources. To protect data privacy and unify data silos distributed among end devices in EI, Federated Learning (FL) is proposed for collaborative training shared AI models across multiple devices without compromising data security. However, the prevailing FL approaches cannot guarantee model generalization and adaptation on heterogeneous clients. Recently, Personalized Federated Learning (PFL) has drawn growing awareness in EI, as it enables striking a productive balance between local-specific training requirements inherent in devices and global-generalized optimization objectives for satisfactory performance. However, most existing PFL methods are based on the Parameters Interaction-based Architecture (PIA) represented by FedAvg, which causes unaffordable communication burdens due to large-scale parameters transmission between devices and the edge server. In contrast, Logits Interaction-based Architecture (LIA) enables to update model parameters with logits transfer, and gains the advantages of communication lightweight and heterogeneous on-device model allowance compared to PIA. Nevertheless, previous LIA methods attempt to achieve satisfactory performance either relying on unrealistic public datasets or increasing communication overhead for additional information transmission other than logits. To tackle this dilemma, we propose a knowledge cache-driven PFL architecture, named FedCache, which reserves a knowledge cache on the server for fetching personalized knowledge from the samples with similar hashes to each given on-device sample. During the training phase, ensemble distillation is applied to on-device models for constructive optimization with personalized knowledge transferred from the server-side knowledge cache. Empirical experiments on four datasets demonstrate the comparable performance of FedCache with state-of-art PFL approaches, with more than two orders of magnitude improvements in communication efficiency. Our code and DEMO are available at https://github.com/wuzhiyuan2000/FedCache.","author":[{"family":"Wu","given":"Zhiyuan"},{"family":"Sun","given":"Sheng"},{"family":"Wang","given":"Yuwei"},{"family":"Liu","given":"Min"},{"family":"Xu","given":"Ke"},{"family":"Wang","given":"Wen"},{"family":"Jiang","given":"Xuefeng"},{"family":"Gao","given":"Bo"},{"family":"Lu","given":"Jinda"}],"issued":{"date-parts":[[2023]]},"DOI":"10.36227/techrxiv.23255420.v2","URL":"https://doi.org/10.36227/techrxiv.23255420.v2","source":"openalex"},{"id":"oa:W4387130572","type":"article-journal","title":"A Survey on Collaborative Learning for Intelligent Autonomous Systems","abstract":"This survey examines approaches to promote Collaborative Learning in distributed systems for emergent Intelligent Autonomous Systems (IAS). The study involves a literature review of Intelligent Autonomous Systems based on Collaborative Learning, analyzing aspects in four dimensions: computing environment, performance concerns, system management, and privacy concerns, mapping the significant requirements of systems to the emerging Artificial intelligence models. Furthermore, the article explores Collaborative Learning Taxonomy for IAS to demonstrate the correlation between IoT, Big Data, and Human-in-the-Loop. Several technological open issues exist in the aforementioned domains (such as in applications of autonomous driving, robotics in healthcare, cyber security, and others) to effectively achieve the future deployment of Intelligent Autonomous Systems. This Survey aims to organize concepts around IAS, indicating the approaches used to extract knowledge from data in Collaborative Learning for IAS, and identifying open issues. Moreover, it presents a guide to overcoming the existing challenges in decision-making mechanisms with IAS, providing a holistic vision of Big Data and Human-in-the-Loop.","author":[{"family":"Anjos","given":"Julio"},{"family":"Matteussi","given":"Kassiano"},{"family":"Orlandi","given":"Fernanda"},{"family":"Barbosa","given":"Jorge"},{"family":"Silva","given":"Jorge"},{"family":"Bittencourt","given":"Luiz"},{"family":"Geyer","given":"Cláudio"}],"issued":{"date-parts":[[2023]]},"DOI":"10.1145/3625544","URL":"https://doi.org/10.1145/3625544","source":"openalex"},{"id":"oa:W4399401115","type":"manuscript","title":"Asynchronous Byzantine Federated Learning","abstract":"Federated learning (FL) enables a set of geographically distributed clients to collectively train a model through a server. Classically, the training process is synchronous, but can be made asynchronous to maintain its speed in presence of slow clients and in heterogeneous networks. The vast majority of Byzantine fault-tolerant FL systems however rely on a synchronous training process. Our solution is one of the first Byzantine-resilient and asynchronous FL algorithms that does not require an auxiliary server dataset and is not delayed by stragglers, which are shortcomings of previous works. Intuitively, the server in our solution waits to receive a minimum number of updates from clients on its latest model to safely update it, and is later able to safely leverage the updates that late clients might send. We compare the performance of our solution with state-of-the-art algorithms on both image and text datasets under gradient inversion, perturbation, and backdoor attacks. Our results indicate that our solution trains a model faster than previous synchronous FL solution, and maintains a higher accuracy, up to 1.54x and up to 1.75x for perturbation and gradient inversion attacks respectively, in the presence of Byzantine clients than previous asynchronous FL solutions.","author":[{"family":"Cox","given":"Bart"},{"family":"Mălan","given":"Abele"},{"family":"Chen","given":"Lydia"},{"family":"Decouchant","given":"Jérémie"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2406.01438","URL":"https://doi.org/10.48550/arxiv.2406.01438","source":"openalex"},{"id":"oa:W4375928946","type":"article-journal","title":"Machine Learning for Service Migration: A Survey","abstract":"Future communication networks are envisioned to satisfy increasingly granular and dynamic requirements to accommodate the application and user demands. Indeed, novel immersive and mission-critical services necessitate increased computing and network resources, reduced communication latency, and guaranteed reliability. Thus, efficient and adaptive resource management schemes are required to provide and maintain sufficient levels of Quality of Experience (QoE) during the service life-cycle. Service migration is considered a key enabler of dynamic service orchestration. Indeed, moving services on demand is an efficient mechanism for user mobility support, load balancing in case of fluctuations in service demands, and hardware failure mitigation. However, service migration requires planning, as multiple parameters must be optimized to reduce service disruption to a minimum. Recent breakthroughs in computational capabilities allowed the emergence of Machine Learning as a tool for decision making that is expected to enable seamless automation of network resource management by predicting events and learning optimal decision policies. This paper surveys contributions applying Machine Learning (ML) methods to optimize service migration, providing a detailed literature review on recent advances in the field and establishing a classification of current research efforts with an analysis of their strengths and limitations. Finally, the paper provides insights on the main directions for future research.","author":[{"family":"Toumi","given":"Nassima"},{"family":"Bagaa","given":"Miloud"},{"family":"Ksentini","given":"Adlen"}],"issued":{"date-parts":[[2023]]},"DOI":"10.1109/comst.2023.3273121","URL":"https://doi.org/10.1109/comst.2023.3273121","source":"openalex"},{"id":"oa:W4384938131","type":"article-journal","title":"Towards a robust, effective and resource efficient machine learning technique for IoT security monitoring","abstract":"The application of Deep Neural Networks (DNNs) for monitoring cyberattacks in Internet of Things (IoT) systems has gained significant attention in recent years. However, achieving optimal detection performance through DNN training has posed challenges due to computational intensity and vulnerability to adversarial samples. To address these issues, this paper introduces an optimization method that combines regularization and simulated micro-batching. This approach enables the training of DNNs in a robust, efficient, and resource-friendly manner for IoT security monitoring. Experimental results demonstrate that the proposed DNN model, including its performance in Federated Learning (FL) settings, exhibits improved attack detection and resistance to adversarial perturbations compared to benchmark baseline models and conventional Machine Learning (ML) methods typically employed in IoT security monitoring. Notably, the proposed method achieves significant reductions of 79.54% and 21.91% in memory and time usage, respectively, when compared to the benchmark baseline in simulated virtual worker environments. Moreover, in realistic testbed scenarios, the proposed method reduces memory footprint by 6.05% and execution time by 15.84%, while maintaining accuracy levels that are superior or comparable to state-of-the-art methods. These findings validate the feasibility and effectiveness of the proposed optimization method for enhancing the efficiency and robustness of DNN-based IoT security monitoring.","author":[{"family":"Zakariyya","given":"Idris"},{"family":"Kalutarage","given":"Harsha"},{"family":"Al-Kadri","given":"MO"}],"issued":{"date-parts":[[2023]]},"DOI":"10.1016/j.cose.2023.103388","URL":"https://doi.org/10.1016/j.cose.2023.103388","source":"openalex"},{"id":"oa:W4395482942","type":"manuscript","title":"Model Poisoning Attacks to Federated Learning via Multi-Round Consistency","abstract":"Model poisoning attacks are critical security threats to Federated Learning (FL). Existing model poisoning attacks suffer from two key limitations: 1) they achieve suboptimal effectiveness when defenses are deployed, and/or 2) they require knowledge of the model updates or local training data on genuine clients. In this work, we make a key observation that their suboptimal effectiveness arises from only leveraging model-update consistency among malicious clients within individual training rounds, making the attack effect self-cancel across training rounds. In light of this observation, we propose PoisonedFL, which enforces multi-round consistency among the malicious clients' model updates while not requiring any knowledge about the genuine clients. Our empirical evaluation on five benchmark datasets shows that PoisonedFL breaks eight state-of-the-art defenses and outperforms seven existing model poisoning attacks. Moreover, we also explore new defenses that are tailored to PoisonedFL, but our results show that we can still adapt PoisonedFL to break them. Our study shows that FL systems are considerably less robust than previously thought, underlining the urgency for the development of new defense mechanisms.","author":[{"family":"Xie","given":"Yueqi"},{"family":"Fang","given":"Minghong"},{"family":"Gong","given":"Neil"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2404.15611","URL":"https://doi.org/10.48550/arxiv.2404.15611","source":"openalex"},{"id":"oa:W4398186361","type":"article-journal","title":"RUL Prediction of Lithium-ion Batteries using a Federated and Homomorphically Encrypted Learning Method","abstract":"The increasing demand for lithium-ion batteries (LIB) across various industries has accentuated the importance of accurately predicting the Remaining Useful Life (RUL) of these energy storage devices. This article introduces a novel approach to RUL prediction by leveraging a federated learning (FL) and homomorphic encryption (HE) model, called FedHEONN. Traditional RUL prediction models often face challenges not only related to the accuracy and reliability of estimations but also to data privacy and security when dealing with sensitive information in Internet of Things (IoT) environments. In response, our approach employs FL, allowing multiple distributed nodes to collaboratively train a predictive model without sharing private data. This ensures data privacy and security while harnessing the collective knowledge from diverse edge computing devices. Furthermore, to address the issue of secure computation over encrypted data, FedHEONN has the capacity to incorporate HE into the learning process. This enables the model to operate directly on encrypted data, providing an additional layer of protection to that of the federated model itself.","author":[{"family":"López","given":"Víctor"},{"family":"Fontenla-Romero","given":"Óscar"},{"family":"Hernández-Pereira","given":"Elena"},{"family":"Guijarroberdiñas","given":"Bertha"},{"family":"Blanco-Seijo","given":"Carlos"},{"family":"Fernández-Paz","given":"Samuel"}],"issued":{"date-parts":[[2024]]},"DOI":"10.1145/3605098.3636045","URL":"https://doi.org/10.1145/3605098.3636045","source":"openalex"},{"id":"oa:W4394750287","type":"article-journal","title":"Federated Learning with Pareto Optimality for Resource Efficiency and Fast Model Convergence in Mobile Environments","abstract":"Federated learning (FL) is an emerging distributed learning technique through which models can be trained using the data collected by user devices in resource-constrained situations while protecting user privacy. However, FL has three main limitations: First, the parameter server (PS), which aggregates the local models that are trained using local user data, is typically far from users. The large distance may burden the path links between the PS and local nodes, thereby increasing the consumption of the network and computing resources. Second, user device resources are limited, but this aspect is not considered in the training of the local model and transmission of the model parameters. Third, the PS-side links tend to become highly loaded as the number of participating clients increases. The links become congested owing to the large size of model parameters. In this study, we propose a resource-efficient FL scheme. We follow the Pareto optimality concept with the biased client selection to limit client participation, thereby ensuring efficient resource consumption and rapid model convergence. In addition, we propose a hierarchical structure with location-based clustering for device-to-device communication using k-means clustering. Simulation results show that with prate at 0.75, the proposed scheme effectively reduced transmitted and received network traffic by 75.89% and 78.77%, respectively, compared to the FedAvg method. It also achieves faster model convergence compared to other FL mechanisms, such as FedAvg and D2D-FedAvg.","author":[{"family":"Jung","given":"June"},{"family":"Ko","given":"Young‐bae"},{"family":"Lim","given":"Sung‐hwa"}],"issued":{"date-parts":[[2024]]},"DOI":"10.3390/s24082476","URL":"https://doi.org/10.3390/s24082476","source":"openalex"},{"id":"oa:W4391534311","type":"article-journal","title":"Pediatric ECG-Based Deep Learning to Predict Left Ventricular Dysfunction and Remodeling","abstract":"BACKGROUND: Artificial intelligence-enhanced ECG analysis shows promise to detect ventricular dysfunction and remodeling in adult populations. However, its application to pediatric populations remains underexplored. METHODS: A convolutional neural network was trained on paired ECG-echocardiograms (≤2 days apart) from patients ≤18 years of age without major congenital heart disease to detect human expert-classified greater than mild left ventricular (LV) dysfunction, hypertrophy, and dilation (individually and as a composite outcome). Model performance was evaluated on single ECG-echocardiogram pairs per patient at Boston Children's Hospital and externally at Mount Sinai Hospital using area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC). RESULTS: The training cohort comprised 92 377 ECG-echocardiogram pairs (46 261 patients; median age, 8.2 years). Test groups included internal testing (12 631 patients; median age, 8.8 years; 4.6% composite outcomes), emergency department (2830 patients; median age, 7.7 years; 10.0% composite outcomes), and external validation (5088 patients; median age, 4.3 years; 6.1% composite outcomes) cohorts. Model performance was similar on internal test and emergency department cohorts, with model predictions of LV hypertrophy outperforming the pediatric cardiologist expert benchmark. Adding age and sex to the model added no benefit to model performance. When using quantitative outcome cutoffs, model performance was similar between internal testing (composite outcome: AUROC, 0.88, AUPRC, 0.43; LV dysfunction: AUROC, 0.92, AUPRC, 0.23; LV hypertrophy: AUROC, 0.88, AUPRC, 0.28; LV dilation: AUROC, 0.91, AUPRC, 0.47) and external validation (composite outcome: AUROC, 0.86, AUPRC, 0.39; LV dysfunction: AUROC, 0.94, AUPRC, 0.32; LV hypertrophy: AUROC, 0.84, AUPRC, 0.25; LV dilation: AUROC, 0.87, AUPRC, 0.33), with composite outcome negative predictive values of 99.0% and 99.2%, respectively. Saliency mapping highlighted ECG components that influenced model predictions (precordial QRS complexes for all outcomes; T waves for LV dysfunction). High-risk ECG features include lateral T-wave inversion (LV dysfunction), deep S waves in V1 and V2 and tall R waves in V6 (LV hypertrophy), and tall R waves in V4 through V6 (LV dilation). CONCLUSIONS: This externally validated algorithm shows promise to inexpensively screen for LV dysfunction and remodeling in children, which may facilitate improved access to care by democratizing the expertise of pediatric cardiologists.","author":[{"family":"Mayourian","given":"Joshua"},{"family":"Cava","given":"William"},{"family":"Vaid","given":"Akhil"},{"family":"Nadkarni","given":"Girish"},{"family":"Ghelani","given":"Sunil"},{"family":"Mannix","given":"Rebekah"},{"family":"Geva","given":"Tal"},{"family":"Dionne","given":"Audrey"},{"family":"Alexander","given":"Mark"},{"family":"Duong","given":"Son"},{"family":"Triedman","given":"John"}],"issued":{"date-parts":[[2024]]},"DOI":"10.1161/circulationaha.123.067750","URL":"https://doi.org/10.1161/circulationaha.123.067750","source":"openalex"},{"id":"oa:W4401509080","type":"article-journal","title":"Federated Learning-Based Intrusion Detection Framework for Internet of Things and Edge Computing Backed Critical Infrastructure","abstract":"Modern Critical Infrastructure (CI) sectors including Smart Girds operate on the Internet of Things and edge computing paradigm. With the enormous growth of these sectors, there are emerging and escalating cyber threats. Traditional Machine Mearning (ML) approaches strive to provide a certain level of resilience against cyber threats but at the cost of privacy leading towards a single point of vulnerability. Following an exhaustive analysis of traditional ML algorithms and cyber threats to the CI, this work introduces a privacy-preserving Federated Learning (FL) driven intrusion detection framework to identify cyber threats focusing on the use case of Smart Grids within the CI. This paper firstly implements and compares various traditional ML algorithms such as Support Vector Machine, Random Forest, and Logistic Regression which works on a centeralised dataset. Secondly, an analysis has been carried out using the proposed FL-based approach to further improve security and privacy along with minimising the need for centralised dataset. Experimental results highlight that our traditional RF-based approach and FL-based approach achieve high intrusion detection accuracy. However, FL has more significant advantages in distributed and privacy-sensitive environments, protecting privacy and reducing the need for data centralisation.","author":[{"family":"Meng","given":"Ruofei"},{"family":"Shah","given":"Awais"},{"family":"Jamshed","given":"Muhammad"},{"family":"Pezaros","given":"Dimitrios"}],"issued":{"date-parts":[[2024]]},"DOI":"10.1109/iccworkshops59551.2024.10615814","URL":"https://doi.org/10.1109/iccworkshops59551.2024.10615814","source":"openalex"},{"id":"doi:10.48550/arxiv.2607.04189","type":"manuscript","title":"SpecGradFilter: A Spectral Gradient Filtering Framework for Taming Federated Heterogeneity","abstract":"Federated Learning (FL) is fundamentally challenged by statistical heterogeneity, where non-identically distributed (non-IID) data induces client drift that severely hampers global convergence. While existing approaches attempt to mitigate this drift through spatial-domain gradient correction or regularization, they overlook the intrinsic spectral structure of optimization signals. In this work, we revisit client drift from a novel frequency-domain perspective and uncover a critical Spectral Bias of Drift: inter-client gradient divergence is predominantly concentrated in low-frequency components which encode client-specific distributional shifts, while high-frequency components representing fine-grained features remain relatively consistent. Motivated by this, we propose SpecGradFilter, a unified Spectral Gradient Filtering Framework that tames heterogeneity by suppressing discordant low-frequency signals. Crucially, we demonstrate that SpecGradFilter is a generalizable principle, effective not only via precise FFT-based truncation but also through spatial approximations like Gaussian detrending. Extensive experiments on benchmarks such as CIFAR-10/100 and Tiny-ImageNet demonstrate that SpecGradFilter significantly performs better performance in highly Non-IID settings with negligible communication overhead, establishing a new paradigm for robust federated optimization.","author":[{"family":"Yuan","given":"Liyang"},{"family":"Yang","given":"Yibo"},{"family":"Guo","given":"Dandan"},{"family":"Richtarik","given":"Peter"},{"family":"Lin","given":"Zhouchen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.04189","URL":"https://doi.org/10.48550/arxiv.2607.04189","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.07130","type":"manuscript","title":"Ampere: Communication-Efficient and High-Accuracy Split Federated Learning","abstract":"A Federated Learning (FL) system collaboratively trains neural networks across devices and a server but is limited by significant on-device computation costs. Split Federated Learning (SFL) systems mitigate this by offloading a block of layers of the network from the device to a server. However, in doing so, it introduces large communication overheads due to frequent exchanges of intermediate activations and gradients between devices and the server and reduces model accuracy for non-IID data. We propose Ampere, a novel collaborative training system that simultaneously minimizes on-device computation and device-server communication while improving model accuracy. Unlike SFL, which uses a global loss by iterative end-to-end training, Ampere develops unidirectional inter-block training to sequentially train the device and server blocks with a local loss, eliminating the transfer of gradients. A lightweight auxiliary network generation method decouples training between the device and server, reducing frequent intermediate exchanges to a single transfer, which significantly reduces the communication overhead. Ampere mitigates the impact of data heterogeneity by consolidating activations generated by the trained device block to train the server block, in contrast to SFL, which trains on device-specific, non-IID activations. Extensive experiments on multiple CNNs and Transformers show that, compared to state-of-the-art SFL baseline systems, Ampere (i) improves model accuracy by up to 11.70 percentage points while training up to 18.6x faster, (ii) incurs up to 911x lower device-server communication overhead and up to 14.5x lower on-device computation, and (iii) reduces standard deviation of accuracy by 71.13% for various non-IID degrees highlighting superior performance when faced with heterogeneous data. Ampere is available from https://github.com/blessonvar/Ampere.","author":[{"family":"Zhang","given":"Zihan"},{"family":"Wong","given":"Leon"},{"family":"Varghese","given":"Blesson"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.07130","URL":"https://doi.org/10.48550/arxiv.2507.07130","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.27715","type":"manuscript","title":"Beyond Non-IID: Learner--Client Distribution Mismatch in Federated Learning","abstract":"Federated learning systems are increasingly deployed to facilitate collaborative model training across a heterogeneous client population. Existing practice mostly implicitly assumes that the aggregated client data distribution is representative of the learner's target distribution or that learning from all available clients is uniformly beneficial for the learner distribution. However, such an assumption often does not hold in reality. Traditional client selection strategies in FL literature largely overlook such misalignment, while most existing work on multi-source transfer learning either requires direct access to local data or uses one-shot model/feature aggregation. In this paper, we take the initiative to understand and mitigate the impacts of such learner-client population misalignment. In particular, we consider the practical setting where the learner keeps a small proxy dataset. We observe that client contributions vary significantly across training rounds, and traditional technology is insufficient to identify beneficial sources under multi-source transfer diversity. Then, we propose a dynamic, influence-aware client selection framework that estimates each client's potential utility to the learner's optimization objective using proxy influence signals on a learner-specific proxy set. Via using leave-one-out evaluations, we prioritize the most informative sources of knowledge while controlling the negative impacts of statistical noise and data heterogeneity. Experiments on CIFAR-10 under heterogeneous data partitions demonstrate that our approach consistently outperforms static and dynamic baselines, achieving faster convergence and higher accuracy.","author":[{"family":"Xie","given":"Yiming"},{"family":"Su","given":"Lili"},{"family":"Mi","given":"Ningfang"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.27715","URL":"https://doi.org/10.48550/arxiv.2608.27715","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.27713","type":"manuscript","title":"DART-FL: Burst-Aware Multitask Federated Learning under Dynamic Inference Demand at the Edge","abstract":"Edge intelligence systems increasingly require model training and online inference to coexist on resource-constrained devices, while inference demand can vary substantially across tasks over time. This creates two coupled challenges: sufficient computation must be reserved for inference to maintain service-level objectives (SLOs), while the remaining training capacity should adapt to task-specific demand so that frequently requested tasks can improve earlier during training. We propose an SLO-aware, demand-driven multitask federated learning framework (DART-FL) that jointly adapts the inference-training resource split and task-level training emphasis. At each scheduling interval, DART-FL uses the inference backlog and profiled service capacity to determine the minimum resource allocation required for inference. The remaining training capacity is then distributed across tasks using a queue-aware DPP-inspired scheduler, and the resulting task allocations are mapped to dynamic loss weights. This allows tasks experiencing higher inference demand to receive greater training emphasis in earlier communication rounds. Clients train a shared backbone with task-specific heads, and the complete multitask model is aggregated through FedAvg. We evaluate DART-FL using Stanford Cars and Oxford Flowers 102 under both synthetic and real Alibaba trace-derived workloads. Results show that DART-FL dynamically adapts the inference-training resource split to time-varying inference demand and shifts the learning progress of high-demand tasks toward their burst periods, improving model accuracy when those tasks are frequently requested while maintaining comparable long-term multitask performance.","author":[{"family":"Xie","given":"Yiming"},{"family":"Yu","given":"Pinrui"},{"family":"Yuan","given":"Geng"},{"family":"Lin","given":"Xue"},{"family":"Mi","given":"Ningfang"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.27713","URL":"https://doi.org/10.48550/arxiv.2608.27713","source":"datacite"},{"id":"doi:10.25532/opara-1534","type":"article-journal","title":"FedSurg EndoVis 2024: Challenge Subset of Appendix300","abstract":"This deposit contains the supplementary data for the FedSurg EndoVis 2024 Challenge, the first federated learning challenge in Surgical Data Science, held at MICCAI 2024. The challenge used a preliminary subset of the Appendix300 dataset, a multi-institutional collection of laparoscopic appendectomy recordings from German university and community hospitals, annotated with intraoperative appendicitis severity grades 0 to 5 following Gomes et al. The challenge cohort comprises 223 recordings across four centers, split into 153 training and 70 test cases. The deposit includes a CSV file specifying which Appendix300 samples were used in the challenge and their assignment to centers and to the training and test splits, together with two recordings that were used in the challenge but excluded from the final Appendix300 dataset because the available footage was shorter than the nominal 100 second window or the appendix was not clearly visible at the annotated timestamp. This deposit does not contain the appendectomy video recordings themselves, with the exception of the two excluded cases named above. The video data and accompanying clinical metadata are available separately as the Appendix300 dataset at https://doi.org/10.25532/OPARA-1173. The challenge evaluation and ranking code is available at https://gitlab.com/nct_tso_public/challenges/miccai2024/snippet, and the example federated learning setup at https://gitlab.com/nct_tso_public/challenges/miccai2024/FedSurg24. Use of this material requires citation of both the FedSurg challenge paper and the Appendix300 data descriptor. Released under CC BY.","author":[{"family":"Kirchner","given":"Max"},{"family":"Kolbinger","given":"Fiona"},{"family":"Jenke","given":"Alexander"},{"family":"Saldanha","given":"Oliver"},{"family":"Pfeiffer","given":"Kevin"},{"family":"Kanjo","given":"Weam"},{"family":"Kather","given":"Jakob"},{"family":"Bodenstedt","given":"Sebastian"},{"family":"Speidel","given":"Stefanie"}],"issued":{"date-parts":[[2026]]},"DOI":"10.25532/opara-1534","URL":"https://doi.org/10.25532/opara-1534","source":"datacite"},{"id":"doi:10.5281/zenodo.21550062","type":"article-journal","title":"Machine Learning Approaches for Autism Spectrum Disorder Detection: A Systematic Review of Age-Specific Applications and Performance Metrics","abstract":"Autism Spectrum Disorder is one of the biggest concerns in the healthcare sector, and it's crucial to diagnose it at an early stage for patients with Autism Spectrum Disorder. This review focuses on the use of machine learning in diagnosing Autism Spectrum Disorder, drawing data from 100 papers between 2015 and 2024. We touched every possible method starting from the classic ones like Support Vector Machines (SVMs) to the new ones like federated learning. Proving the federated learning is actually great since it is very precise (up to 98%) while keeping people's information personal, which is a crucial matter in the healthcare industry. But one cannot write-off the basic framework where people use standard machine learning models such as SVMs, which at this point achieve around 92% accuracy. Also, they are more convenient to be implemented in small clinics that do not possess many great computers, and etcetera. This review suggests that the most suitable ML approaches for Autism Spectrum Disorder detection need to consider accuracy, privacy and availability of resources. Lately, more developed technologies provide even better outcomes; nevertheless, conventional techniques provide terrific options for clinics without much complicated systems available. Thus, the study offers meaningful suggestions to facilitate the choice of the most suitable methods based on the comparison between these approaches. In sum, this review spans the existing gap between research advancements in state-of-art machine learning techniques and practical healthcare settings and provides important recommendations for enhancing Autism Spectrum Disorder screening across various contexts.","author":[{"family":"Patil","given":"Pooja"},{"family":"Patil","given":"Jaydeep"},{"family":"Patil","given":"Sangram"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21550062","URL":"https://doi.org/10.5281/zenodo.21550062","source":"datacite"},{"id":"doi:10.5281/zenodo.21550063","type":"article-journal","title":"Machine Learning Approaches for Autism Spectrum Disorder Detection: A Systematic Review of Age-Specific Applications and Performance Metrics","abstract":"Autism Spectrum Disorder is one of the biggest concerns in the healthcare sector, and it's crucial to diagnose it at an early stage for patients with Autism Spectrum Disorder. This review focuses on the use of machine learning in diagnosing Autism Spectrum Disorder, drawing data from 100 papers between 2015 and 2024. We touched every possible method starting from the classic ones like Support Vector Machines (SVMs) to the new ones like federated learning. Proving the federated learning is actually great since it is very precise (up to 98%) while keeping people's information personal, which is a crucial matter in the healthcare industry. But one cannot write-off the basic framework where people use standard machine learning models such as SVMs, which at this point achieve around 92% accuracy. Also, they are more convenient to be implemented in small clinics that do not possess many great computers, and etcetera. This review suggests that the most suitable ML approaches for Autism Spectrum Disorder detection need to consider accuracy, privacy and availability of resources. Lately, more developed technologies provide even better outcomes; nevertheless, conventional techniques provide terrific options for clinics without much complicated systems available. Thus, the study offers meaningful suggestions to facilitate the choice of the most suitable methods based on the comparison between these approaches. In sum, this review spans the existing gap between research advancements in state-of-art machine learning techniques and practical healthcare settings and provides important recommendations for enhancing Autism Spectrum Disorder screening across various contexts.","author":[{"family":"Patil","given":"Pooja"},{"family":"Patil","given":"Jaydeep"},{"family":"Patil","given":"Sangram"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21550063","URL":"https://doi.org/10.5281/zenodo.21550063","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.08014","type":"manuscript","title":"FedTR: Federated Learning Framework with Transfer Learning for Industrial Visual Inspection","abstract":"Federated learning (FL) is a collaborative learning scheme to train deep learning models, where collaborating parties can consolidate their models without sharing local data with other parties, hence preserving data privacy. Nevertheless, when implementing FL in Industrial visual inspection (IVI), the constraints posed by limited data availability and the intricate nature of the inspection tasks significantly impact the performance of the resulting model. This paper introduces FedTR, a novel FL framework incorporating transfer learning designed for Autonomous IVI, focusing on the challenging task of identifying label defects through end-to-end text recognition. Transfer learning is a method that leverages the knowledge of a pre-trained model to adapt to a different dataset. FedTR initially trains the model using a publicly available dataset, after which performs the essential federated learning process with model fine-tuning on the distributed and limited private data. Extensive experiment results demonstrate the effectiveness and feasibility of FedTR on private ink cartridge datasets for label defect identification. FedTR achieves an end-to-end text recognition word-level accuracy of 95.5% and 94.2% on homogeneous and heterogeneous data respectively. Additionally, it attains performance levels that are on par with those achieved through centralized training.","author":[{"family":"Sathiamoorthy","given":"Vikash"},{"family":"Huai","given":"Shuo"},{"family":"Kong","given":"Hao"},{"family":"Liu","given":"Di"},{"family":"Loy","given":"Wendy"},{"family":"Makaya","given":"Christian"},{"family":"Ho","given":"Daren"},{"family":"Subramaniam","given":"Ravi"},{"family":"Lin","given":"Qian"},{"family":"Liu","given":"Weichen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.08014","URL":"https://doi.org/10.48550/arxiv.2607.08014","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.03714","type":"manuscript","title":"Don't Trust Us: A privacy-by-design android malware detection pipeline","abstract":"Android malware detection increasingly relies on collecting and processing sensitive user data, including device identifiers, network artifacts, and runtime traces, while privacy is too often treated as a secondary concern. Existing privacy-aware approaches typically enforce privacy after data collection, for example, through anonymization, encryption, or federated learning, yet still require access to user information and therefore demand a high level of user trust in systems that already operate with privileged access to device activity. We argue that this requirement should be removed rather than managed. Android malware detection should be privacy-aware by design, so that effective analysis does not depend on sensitive data being accessed in the first place. To this end, we first formalize a set of design requirements for privacy-by-design detection and then implement each requirement in a comprehensive pipeline. First, static analysis is performed to extract relevant data from each APK, following the Drebin representation, which is then submitted to an SVM after vectorization. The model is equipped with a dual-reject threshold rule that either commits to a confident decision or defers uncertain samples to a dynamic analysis stage within a sandboxed environment, so that genuine user information never enters the analysis loop. Results confirm that, on a temporally split dataset spanning from 2024 to 2025, the pipeline achieves an F1 score of 0.87 with the first static analysis stage, deferring only 6.7% of test samples to secondary dynamic analysis. Additionally, dynamic sandboxing helps recognize applications' maliciousness with high confidence without extracting any sensitive data. These results demonstrate that strong detection performance is achievable without sacrificing user privacy.","author":[{"family":"Massidda","given":"Emmanuele"},{"family":"Soi","given":"Diego"},{"family":"Giacinto","given":"Giorgio"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.03714","URL":"https://doi.org/10.48550/arxiv.2606.03714","source":"datacite"},{"id":"doi:10.5281/zenodo.20094300","type":"article-journal","title":"Dynamic Latency Optimization for Edge-Based Machine Learning Models in 6G-Enabled Industrial Internet of Things (IIoT)","abstract":"Abstract The integration of 6G technology into the Industrial Internet of Things (IIoT) promises to redefine manufacturing through \"Hyper-Reliable Low-Latency Communication\" (HRLLC). However, the deployment of complex Machine Learning (ML) models at the edge remains constrained by the heterogeneous nature of industrial data and the limited computational resources of edge nodes. This article proposes a novel framework for Dynamic Latency Optimization (DLO) that leverages Deep Reinforcement Learning (DRL) for intelligent task offloading and resource allocation. By utilizing 6G's Terahertz (THz) spectrum and AI-native Network Slicing, the proposed framework dynamically adapts to fluctuating network conditions to maintain sub-millisecond latency. Our simulation results demonstrate a 42% reduction in end-to-end delay and a 30% improvement in energy efficiency compared to traditional 5G-MEC architectures. Furthermore, we explore the integration of Reconfigurable Intelligent Surfaces (RIS), Semantic Communication, and Zero-Trust Edge Security to further optimize the data-intelligence pipeline for Industry 5.0 applications, focusing on the critical synergy between human operators and autonomous systems within a resilient, sustainable, and cognitively aware industrial fabric. Keywords: 6G Networks, Industrial IoT (IIoT), Edge Intelligence, Deep Reinforcement Learning, Latency Optimization 1. Introduction: From Automation to Human-Centric Intelligence The transition from Industry 4.0 to Industry 5.0 marks a profound shift toward human-centric, resilient, and sustainable manufacturing systems. While Industry 4.0 was characterized by the digitalization of physical assets and the rise of cyber-physical systems, Industry 5.0 emphasizes the \"Tactile Internet\" and \"Human-Robot Co-evolution.\" In this new paradigm, the focus shifts from pure efficiency to the seamless collaboration between humans and increasingly autonomous machines. The \"Tactile Internet\" concept is particularly revolutionary, as it requires a \"haptic control loop\"—the ability to transmit touch and feel sensations over the network with such low latency that the human brain perceives no delay. This necessitates an end-to-end latency below 1ms, encompassing both the transmission and the computational processing of sensory feedback. This evolution necessitates a communication infrastructure capable of supporting advanced applications such as ultra-responsive autonomous mobile robots (AMRs), synchronized multi-robot assembly lines, and high-fidelity haptic feedback for remote maintenance in hazardous environments. For example, a specialist surgeon operating a robotic arm in a factory cleanup of toxic waste requires instantaneous haptic feedback to \"feel\" the resistance of the materials being handled. If the feedback loop exceeds 10ms, the mismatch between visual and tactile input can lead to \"operator sickness\" or mechanical errors that jeopardize safety. Furthermore, we must consider proprioceptive alignment—the sense of self-movement and body position. In 6G-enabled IIoT, the network must act as an extension of the human nervous system, where the delay jitter is so minimal that the robotic actuator feels like a literal extension of the operator's limb. This requires not just low latency, but Isochronous Communication, where packets arrive at precisely regular intervals to maintain the temporal rhythm of human motor-sensory systems. This synchronization is critical for Tele-Operation in nanomanufacturing, where even a micro-stutter in the feedback loop can cause the robotic probe to crush a microscopic wafer. The biological threshold for \"instantaneous\" feedback in human motor control is roughly 1-10ms for tactile sensations and less than 1ms for the suppression of \"visual-vestibular conflict.\" In 6G, we move into the regime of \"Sub-Perceptual Jitter,\" where the network variance is lower than the biological noise of the human nervous system. This enables \"Neuromorphic Manufacturi","author":[{"family":"Patil","given":"Seema"},{"family":"Doddamani","given":"Harshavardhana"},{"family":"Rivers","given":"Julianne"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20094300","URL":"https://doi.org/10.5281/zenodo.20094300","source":"datacite"},{"id":"doi:10.5281/zenodo.20094301","type":"article-journal","title":"Dynamic Latency Optimization for Edge-Based Machine Learning Models in 6G-Enabled Industrial Internet of Things (IIoT)","abstract":"Abstract The integration of 6G technology into the Industrial Internet of Things (IIoT) promises to redefine manufacturing through \"Hyper-Reliable Low-Latency Communication\" (HRLLC). However, the deployment of complex Machine Learning (ML) models at the edge remains constrained by the heterogeneous nature of industrial data and the limited computational resources of edge nodes. This article proposes a novel framework for Dynamic Latency Optimization (DLO) that leverages Deep Reinforcement Learning (DRL) for intelligent task offloading and resource allocation. By utilizing 6G's Terahertz (THz) spectrum and AI-native Network Slicing, the proposed framework dynamically adapts to fluctuating network conditions to maintain sub-millisecond latency. Our simulation results demonstrate a 42% reduction in end-to-end delay and a 30% improvement in energy efficiency compared to traditional 5G-MEC architectures. Furthermore, we explore the integration of Reconfigurable Intelligent Surfaces (RIS), Semantic Communication, and Zero-Trust Edge Security to further optimize the data-intelligence pipeline for Industry 5.0 applications, focusing on the critical synergy between human operators and autonomous systems within a resilient, sustainable, and cognitively aware industrial fabric. Keywords: 6G Networks, Industrial IoT (IIoT), Edge Intelligence, Deep Reinforcement Learning, Latency Optimization 1. Introduction: From Automation to Human-Centric Intelligence The transition from Industry 4.0 to Industry 5.0 marks a profound shift toward human-centric, resilient, and sustainable manufacturing systems. While Industry 4.0 was characterized by the digitalization of physical assets and the rise of cyber-physical systems, Industry 5.0 emphasizes the \"Tactile Internet\" and \"Human-Robot Co-evolution.\" In this new paradigm, the focus shifts from pure efficiency to the seamless collaboration between humans and increasingly autonomous machines. The \"Tactile Internet\" concept is particularly revolutionary, as it requires a \"haptic control loop\"—the ability to transmit touch and feel sensations over the network with such low latency that the human brain perceives no delay. This necessitates an end-to-end latency below 1ms, encompassing both the transmission and the computational processing of sensory feedback. This evolution necessitates a communication infrastructure capable of supporting advanced applications such as ultra-responsive autonomous mobile robots (AMRs), synchronized multi-robot assembly lines, and high-fidelity haptic feedback for remote maintenance in hazardous environments. For example, a specialist surgeon operating a robotic arm in a factory cleanup of toxic waste requires instantaneous haptic feedback to \"feel\" the resistance of the materials being handled. If the feedback loop exceeds 10ms, the mismatch between visual and tactile input can lead to \"operator sickness\" or mechanical errors that jeopardize safety. Furthermore, we must consider proprioceptive alignment—the sense of self-movement and body position. In 6G-enabled IIoT, the network must act as an extension of the human nervous system, where the delay jitter is so minimal that the robotic actuator feels like a literal extension of the operator's limb. This requires not just low latency, but Isochronous Communication, where packets arrive at precisely regular intervals to maintain the temporal rhythm of human motor-sensory systems. This synchronization is critical for Tele-Operation in nanomanufacturing, where even a micro-stutter in the feedback loop can cause the robotic probe to crush a microscopic wafer. The biological threshold for \"instantaneous\" feedback in human motor control is roughly 1-10ms for tactile sensations and less than 1ms for the suppression of \"visual-vestibular conflict.\" In 6G, we move into the regime of \"Sub-Perceptual Jitter,\" where the network variance is lower than the biological noise of the human nervous system. This enables \"Neuromorphic Manufacturi","author":[{"family":"Patil","given":"Seema"},{"family":"Doddamani","given":"Harshavardhana"},{"family":"Rivers","given":"Julianne"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20094301","URL":"https://doi.org/10.5281/zenodo.20094301","source":"datacite"},{"id":"doi:10.5281/zenodo.19430884","type":"article-journal","title":"Federated Learning: A Systematic Review of Architecture, Challenges and Research Directions","abstract":"Federated Learning (FL) has emerged as a distributed machine learning paradigm that enables collaborative model training while preserving data privacy. Unlike traditional centralized learning frameworks which require collecting raw data at a single server, FL allows multiple clients to train models locally and share only model updates for global aggregation. In this review we examine twelve peer-reviewed surveys and research papers published between 2017 and 2025 that analyze federated learning architectures, communication mechanisms, privacy-preserving techniques, security threats, types of FL and real-world deployment scenarios. Drawing substantially from the comprehensive IEEE Access survey by Aledhari et al., our analysis shows that FL still faces major technical challenges including non-IID data distributions, high communication costs, scalability constraints and adversarial threats. We also highlight emerging research directions such as lightweight optimization, fairness-aware aggregation, blockchain-based trust mechanisms and personalized FL. This review consolidates existing work, presents a full 12-paper literature summary table, and outlines key open problems to guide future research on federated learning systems.","author":[{"family":"Rathod","given":"Dr"},{"family":"Pete","given":"Shravani"},{"family":"Dhoke","given":"Divya"},{"family":"Kurhekar","given":"Ishika"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19430884","URL":"https://doi.org/10.5281/zenodo.19430884","source":"datacite"},{"id":"doi:10.5281/zenodo.19430885","type":"article-journal","title":"Federated Learning: A Systematic Review of Architecture, Challenges and Research Directions","abstract":"Federated Learning (FL) has emerged as a distributed machine learning paradigm that enables collaborative model training while preserving data privacy. Unlike traditional centralized learning frameworks which require collecting raw data at a single server, FL allows multiple clients to train models locally and share only model updates for global aggregation. In this review we examine twelve peer-reviewed surveys and research papers published between 2017 and 2025 that analyze federated learning architectures, communication mechanisms, privacy-preserving techniques, security threats, types of FL and real-world deployment scenarios. Drawing substantially from the comprehensive IEEE Access survey by Aledhari et al., our analysis shows that FL still faces major technical challenges including non-IID data distributions, high communication costs, scalability constraints and adversarial threats. We also highlight emerging research directions such as lightweight optimization, fairness-aware aggregation, blockchain-based trust mechanisms and personalized FL. This review consolidates existing work, presents a full 12-paper literature summary table, and outlines key open problems to guide future research on federated learning systems.","author":[{"family":"Rathod","given":"Dr"},{"family":"Pete","given":"Shravani"},{"family":"Dhoke","given":"Divya"},{"family":"Kurhekar","given":"Ishika"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19430885","URL":"https://doi.org/10.5281/zenodo.19430885","source":"datacite"},{"id":"doi:10.5281/zenodo.21760942","type":"article-journal","title":"Advanced IoT-Integrated Systems for Intelligent Water Quality Monitoring in Aquaculture: Emerging Technologies, AI-Driven Analytics, and Future Perspectives","abstract":"Between 2020 and 2025, the integration of Internet of Things (IoT) sensor networks into aquaculture water quality management has seen a transformational acceleration due to the convergence of advanced sensor miniaturization, edge computing, artificial intelligence (AI), and next-generation wireless connectivity. This paper expands on the basic bibliometric and systematic analysis of IoT sensor applications in aquaculture (Flores-Iwasaki et al., 2025), taking the knowledge frontier a step further by critically reviewing recent technologies such as 5G-enabled real-time monitoring, AI-driven digital twins, flexible nano sensors, federated learning frameworks, and autonomous unmanned aerial/aquatic vehicle (UAV/AUV) platforms. The parameters most monitored were still pH (100%), temperature (96.7%) and dissolved oxygen (79.4%) and attention increased to total ammonia nitrogen (TAN), nitrite, nitrate and chlorophyll-a in recirculating aquaculture systems (RAS) and integrated aquaponics. The key findings suggest that AI-aided predictive models, particularly Long Short-Term Memory (LSTM), Transformers, and ensemble methods, demonstrated water quality prediction accuracies exceeding 95% in controlled settings. Microcontrollers (TinyML) Edge AI deployment reduced cloud latency up to 87%, enabling near-instant anomaly detection. Identified key gaps are lack of automated TAN sensing in most commercial deployments, limited self-cleaning sensor maintenance mechanisms, and inequitable technology access in developing-nation aquaculture. Future research directions include AI-digital twin co-simulation, bio-inspired nano sensor arrays, and 5G-LPWAN hybrid architectures. Addressing these challenges is critical for the deployment of fully automated, sustainable and economically viable aquaculture ecosystems","author":[{"family":"Nandankar","given":"Praful"},{"family":"Dhawas","given":"Prashantkumar"},{"family":"Kalambe","given":"Shilpa"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21760942","URL":"https://doi.org/10.5281/zenodo.21760942","source":"datacite"},{"id":"doi:10.5281/zenodo.21760943","type":"article-journal","title":"Advanced IoT-Integrated Systems for Intelligent Water Quality Monitoring in Aquaculture: Emerging Technologies, AI-Driven Analytics, and Future Perspectives","abstract":"Between 2020 and 2025, the integration of Internet of Things (IoT) sensor networks into aquaculture water quality management has seen a transformational acceleration due to the convergence of advanced sensor miniaturization, edge computing, artificial intelligence (AI), and next-generation wireless connectivity. This paper expands on the basic bibliometric and systematic analysis of IoT sensor applications in aquaculture (Flores-Iwasaki et al., 2025), taking the knowledge frontier a step further by critically reviewing recent technologies such as 5G-enabled real-time monitoring, AI-driven digital twins, flexible nano sensors, federated learning frameworks, and autonomous unmanned aerial/aquatic vehicle (UAV/AUV) platforms. The parameters most monitored were still pH (100%), temperature (96.7%) and dissolved oxygen (79.4%) and attention increased to total ammonia nitrogen (TAN), nitrite, nitrate and chlorophyll-a in recirculating aquaculture systems (RAS) and integrated aquaponics. The key findings suggest that AI-aided predictive models, particularly Long Short-Term Memory (LSTM), Transformers, and ensemble methods, demonstrated water quality prediction accuracies exceeding 95% in controlled settings. Microcontrollers (TinyML) Edge AI deployment reduced cloud latency up to 87%, enabling near-instant anomaly detection. Identified key gaps are lack of automated TAN sensing in most commercial deployments, limited self-cleaning sensor maintenance mechanisms, and inequitable technology access in developing-nation aquaculture. Future research directions include AI-digital twin co-simulation, bio-inspired nano sensor arrays, and 5G-LPWAN hybrid architectures. Addressing these challenges is critical for the deployment of fully automated, sustainable and economically viable aquaculture ecosystems","author":[{"family":"Nandankar","given":"Praful"},{"family":"Dhawas","given":"Prashantkumar"},{"family":"Kalambe","given":"Shilpa"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21760943","URL":"https://doi.org/10.5281/zenodo.21760943","source":"datacite"},{"id":"doi:10.5281/zenodo.19614178","type":"article-journal","title":"Future Directions in Clinical and Epidemiological Research","abstract":"Clinical and epidemiological research is undergoing a substantial transition shaped by multiple converging trends:generative AI and large language models transforming literature synthesis, protocol design, and clinical decision support;European Health Data Space implementation enabling cross-border research at unprecedented scale; learning healthsystems integrating research with routine care; federated analytics addressing data sovereignty while enabling multi-sitestudies; patient advocacy integration reshaping research priorities; and climate-health research emerging ascross-cutting priority. The research methodological implications span every phase of the research cycle and createopportunities and challenges that require strategic response from European research communities. We evaluated fivefuture-oriented clinical and epidemiological research infrastructure frameworks applied to 124 research programmedesigns across 22 European research institutions in Vienna, Stockholm, Rome, Amsterdam, and London developed forimplementation between 2025 and 2030. Frameworks ranged from conventional research programme designs throughintegrated future-adapted research infrastructure incorporating AI-enabled methodology, EHDS compliance, learninghealth system integration, and climate-health capability. Performance was benchmarked using our Future ResearchInfrastructure Effectiveness Score (FRIES), integrating methodological innovation, cross-border capability, patientengagement, climate-health capability, and sustainability. The integrated future-adapted framework achieved the highestFRIES (0.912), with 3.8-fold research productivity advantage and substantial patient engagement improvement overconventional research infrastructure.","author":[{"family":"Jensen","given":"Hugo"},{"family":"Dubois","given":"Jonas"},{"family":"Lindberg","given":"Eva"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.19614178","URL":"https://doi.org/10.5281/zenodo.19614178","source":"datacite"},{"id":"doi:10.5281/zenodo.19614179","type":"article-journal","title":"Future Directions in Clinical and Epidemiological Research","abstract":"Clinical and epidemiological research is undergoing a substantial transition shaped by multiple converging trends:generative AI and large language models transforming literature synthesis, protocol design, and clinical decision support;European Health Data Space implementation enabling cross-border research at unprecedented scale; learning healthsystems integrating research with routine care; federated analytics addressing data sovereignty while enabling multi-sitestudies; patient advocacy integration reshaping research priorities; and climate-health research emerging ascross-cutting priority. The research methodological implications span every phase of the research cycle and createopportunities and challenges that require strategic response from European research communities. We evaluated fivefuture-oriented clinical and epidemiological research infrastructure frameworks applied to 124 research programmedesigns across 22 European research institutions in Vienna, Stockholm, Rome, Amsterdam, and London developed forimplementation between 2025 and 2030. Frameworks ranged from conventional research programme designs throughintegrated future-adapted research infrastructure incorporating AI-enabled methodology, EHDS compliance, learninghealth system integration, and climate-health capability. Performance was benchmarked using our Future ResearchInfrastructure Effectiveness Score (FRIES), integrating methodological innovation, cross-border capability, patientengagement, climate-health capability, and sustainability. The integrated future-adapted framework achieved the highestFRIES (0.912), with 3.8-fold research productivity advantage and substantial patient engagement improvement overconventional research infrastructure.","author":[{"family":"Jensen","given":"Hugo"},{"family":"Dubois","given":"Jonas"},{"family":"Lindberg","given":"Eva"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.19614179","URL":"https://doi.org/10.5281/zenodo.19614179","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.08906","type":"manuscript","title":"Federated Attention Autoencoders with a Stochastic Aggregation Scheme for Anomaly Detection","abstract":"Outlier detection in decentralized data environments is a challenging task for many machine learning implementations, particularly in settings where data cannot be shared. Recently, there have been advances in federated outlier detection, some of which are based on the use of autoencoder networks. The introduction of attention mechanisms to autoencoders boosts their efficiency. However, the application of attention-based models in federated learning remains underdeveloped due to the absence of proper aggregation functions for these types of networks. In our work, we propose two novel aggregation functions tailored for attention-based autoencoders, which better preserve the learned information stored within the memory modules of these networks. We evaluated our approach on the KDDCUP10 dataset, and we showed that the proposed methods achieve up to 2.9\\% and 5.1\\% better results for F1 score and AUC ROC respectively when compared to traditional autoencoders.","author":[{"family":"Ilić","given":"Mihailo"},{"family":"Savić","given":"Miloš"},{"family":"Kurbalija","given":"Vladimir"},{"family":"Ivanović","given":"Mirjana"},{"family":"Fortino","given":"Giancarlo"},{"family":"Jakovetić","given":"Dušan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.08906","URL":"https://doi.org/10.48550/arxiv.2608.08906","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.05386","type":"manuscript","title":"Sparse Principal Component Analysis via Wavelets for Distributed Data","abstract":"The large volume of data and concerns about data privacy have motivated the development of techniques for distributed data, a problem also known as federated learning. In this scenario, sub-samples of the data are divided across different machines, and statistics must be computed over that data without direct access to the full sample. Johnstone &amp; Lu (2009, JASA) show that principal component analysis (PCA) is statistically inconsistent in the high-dimensional regime, and propose a way to recover consistency through wavelet-based sparsification and variable selection. Fan et al. (2019, AoS) show a way to perform this same estimation -- specifically, to estimate the eigenspace that would be obtained if all the data were pooled together, even though it remains effectively distributed -- without addressing the high-dimensional regime. This work incorporates the wavelet-based sparsification of Johnstone &amp; Lu (2009) into the distributed PCA framework of Fan et al. (2019), aiming to reduce communication cost without compromising the quality of the eigenspace estimation. Simulations across $d \\in [52, 5000]$ show that the proposed method overtakes Fan et al. (2019) in estimation error beyond a clear dimensional threshold ($d \\geq 152$ for $λ=25$, $d \\geq 252$ for $λ=50$), while transmitting systematically fewer coefficients throughout the entire range studied. This study was financed by the Sao Paulo Research Foundation (FAPESP), Brazil. Process Number #2023/02538-0 and Number #2025/21250-2.","author":[{"family":"Herrero","given":"Giovanni"},{"family":"Fonseca","given":"Rodney"},{"family":"Pinheiro","given":"Aluísio"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.05386","URL":"https://doi.org/10.48550/arxiv.2608.05386","source":"datacite"},{"id":"doi:10.5281/zenodo.19434964","type":"article-journal","title":"Creaboost: A Blockchain-Enabled Federated Learning Platform for Privacy-Preserving Ad Preference Insights and Creator Rewards","abstract":"Privacy preserving ad targeting is a major challenge in today's digital environment. It is such an environment where the centralized platforms expose user data and provide creators with low lying compensation. Although current federated learning (FL) models help reduce data leakage but do not have on chain reward systems. Blockchain based reward systems often face issues with slow speeds and high costs on public networks like Ethereum. This paper introduces Creaboost, a decentralized application on the XDC Network. It combines FL with smart contract rewards to allow secure, low-lying cost, and fair ad preference collection to its users. Here the users are going to submit their engagement metrics locally. After five unparalleled submissions, we aggregate model parameters off chain using federated learning. The resulting preference (0 or 1) then triggers an on-chain reward (1 or 2 ether) through a user signed transaction. Creaboost extends through five layers: Client, Frontend (HTML/JS), Backend (Flask), FL Server (PyTorch), and Blockchain. This setup guarantees that the raw data exposure leaves the Ethereum device. We ensure security with SSL/TLS, reCAPTCHA v2, email verification, and XDC signing. Creaboost exceeds both centralized and XDC alternatives by minimizing data exposure, execution costs, and XDC creating a scalable framework for transaction advertising.","author":[{"family":"Shaw","given":"Ritesh"},{"family":"Ahmed","given":"Safiya"},{"family":"Sinha","given":"Stuti"},{"family":"Chatterjee","given":"Subhangi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19434964","URL":"https://doi.org/10.5281/zenodo.19434964","source":"datacite"},{"id":"doi:10.5281/zenodo.19434963","type":"article-journal","title":"Creaboost: A Blockchain-Enabled Federated Learning Platform for Privacy-Preserving Ad Preference Insights and Creator Rewards","abstract":"Privacy preserving ad targeting is a major challenge in today's digital environment. It is such an environment where the centralized platforms expose user data and provide creators with low lying compensation. Although current federated learning (FL) models help reduce data leakage but do not have on chain reward systems. Blockchain based reward systems often face issues with slow speeds and high costs on public networks like Ethereum. This paper introduces Creaboost, a decentralized application on the XDC Network. It combines FL with smart contract rewards to allow secure, low-lying cost, and fair ad preference collection to its users. Here the users are going to submit their engagement metrics locally. After five unparalleled submissions, we aggregate model parameters off chain using federated learning. The resulting preference (0 or 1) then triggers an on-chain reward (1 or 2 ether) through a user signed transaction. Creaboost extends through five layers: Client, Frontend (HTML/JS), Backend (Flask), FL Server (PyTorch), and Blockchain. This setup guarantees that the raw data exposure leaves the Ethereum device. We ensure security with SSL/TLS, reCAPTCHA v2, email verification, and XDC signing. Creaboost exceeds both centralized and XDC alternatives by minimizing data exposure, execution costs, and XDC creating a scalable framework for transaction advertising.","author":[{"family":"Shaw","given":"Ritesh"},{"family":"Ahmed","given":"Safiya"},{"family":"Sinha","given":"Stuti"},{"family":"Chatterjee","given":"Subhangi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19434963","URL":"https://doi.org/10.5281/zenodo.19434963","source":"datacite"},{"id":"doi:10.5167/uzh-276716","type":"article-journal","title":"Artificial intelligence for dental implant classification and peri-implant pathology identification in 2D radiographs: A systematic review","abstract":"OBJECTIVE This systematic review aimed to summarize and evaluate the available information regarding the performance of artificial intelligence on dental implant classification and peri-implant pathology identification in 2D radiographs. DATA SOURCES Electronic databases (Medline, Embase, and Cochrane) were searched up to September 2024 for relevant observational studies and both randomized and controlled clinical trials. The search was limited to studies published in English from the last 7 years. Two reviewers independently conducted both study selection and data extraction. Risk of bias assessment was also performed individually by both operators using the Quality Assessment Diagnostic Tool (QUADAS-2). STUDY SELECTION Of the 1,465 records identified, 29 references were selected to perform qualitative analysis. The study characteristics were tabulated in a self-designed table. QUADAS-2 tool identified 10 and 15 studies to respectively have a high and an unclear risk of bias, while only four were categorized as low risk of bias. Overall, accuracy rates for dental implant classification ranged from 67 % to 99 %. Peri-implant pathology identification showed results with accuracy detection rates over 78,6 %. CONCLUSIONS While AI-based models, particularly convolutional neural networks, have shown high accuracy in dental implant classification and peri-implant pathology detection, several limitations must be addressed before widespread clinical application. More advanced AI techniques, such as Federated Learning should be explored to improve the generalizability and efficiency of these models in clinical practice. CLINICAL SIGNIFICANCE AI-based models offer can and clinicians to accurately classify unknown dental implants and enable early detection of peri-implantitis, improving patient outcomes and streamline treatment planning.","author":[{"family":"Bonfanti-Gris","given":"M"},{"family":"Ruales","given":"E"},{"family":"Salido","given":"MP"},{"family":"Martinez-Rus","given":"F"},{"family":"Özcan","given":"Mutlu"},{"family":"Pradies","given":"G"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5167/uzh-276716","URL":"https://doi.org/10.5167/uzh-276716","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.22037","type":"manuscript","title":"A Critical Look into Threshold Homomorphic Encryption for Private Average Aggregation","abstract":"Threshold Homomorphic Encryption (Threshold HE) is a good fit for implementing private federated average aggregation, a key operation in Federated Learning (FL). Despite its potential, recent studies have shown that threshold schemes available in mainstream HE libraries can introduce unexpected security vulnerabilities if an adversary has access to a restricted decryption oracle. This oracle reflects the FL clients' capacity to collaboratively decrypt the aggregated result without knowing the secret key. This work surveys the use of threshold RLWE-based HE for federated average aggregation and examines the performance impact of using smudging noise with a large variance as a countermeasure. We provide a detailed comparison of threshold variants of BFV and CKKS, finding that CKKS-based aggregations perform comparably to BFV-based solutions.","author":[{"family":"Morona-Mínguez","given":"Miguel"},{"family":"Pedrouzo-Ulloa","given":"Alberto"},{"family":"Pérez-González","given":"Fernando"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.22037","URL":"https://doi.org/10.48550/arxiv.2602.22037","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.08314","type":"manuscript","title":"Minimizing Layerwise Activation Norm Improves Generalization in Federated Learning","abstract":"Federated Learning (FL) is an emerging machine learning framework that enables multiple clients (coordinated by a server) to collaboratively train a global model by aggregating the locally trained models without sharing any client's training data. It has been observed in recent works that learning in a federated manner may lead the aggregated global model to converge to a 'sharp minimum' thereby adversely affecting the generalizability of this FL-trained model. Therefore, in this work, we aim to improve the generalization performance of models trained in a federated setup by introducing a 'flatness' constrained FL optimization problem. This flatness constraint is imposed on the top eigenvalue of the Hessian computed from the training loss. As each client trains a model on its local data, we further re-formulate this complex problem utilizing the client loss functions and propose a new computationally efficient regularization technique, dubbed 'MAN,' which Minimizes Activation's Norm of each layer on client-side models. We also theoretically show that minimizing the activation norm reduces the top eigenvalue of the layer-wise Hessian of the client's loss, which in turn decreases the overall Hessian's top eigenvalue, ensuring convergence to a flat minimum. We apply our proposed flatness-constrained optimization to the existing FL techniques and obtain significant improvements, thereby establishing new state-of-the-art.","author":[{"family":"Yashwanth","given":"M"},{"family":"Nayak","given":"Gaurav"},{"family":"Rangwani","given":"Harsh"},{"family":"Singh","given":"Arya"},{"family":"Babu","given":"RV"},{"family":"Chakraborty","given":"Anirban"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.08314","URL":"https://doi.org/10.48550/arxiv.2512.08314","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.25277","type":"manuscript","title":"A Privacy-Preserving Ecosystem for Developing Machine Learning Algorithms Using Patient Data: Insights from the TUM.ai Makeathon","abstract":"The integration of clinical data offers significant potential for the development of personalized medicine. However, its use is severely restricted by the General Data Protection Regulation (GDPR), especially for small cohorts with rare diseases. High-quality, structured data is essential for the development of predictive medical AI. In this case study, we propose a novel, multi-stage approach to secure AI training: (1) The model is designed on a simulated clinical knowledge graph (cKG). This graph is used exclusively to represent the structural characteristics of the real cKG without revealing any sensitive content. (2) The model is then integrated into the FeatureCloud (FC) federated learning framework, where it is prepared in a single-client configuration within a protected execution environment. (3) Training then takes place within the hospital environment on the real cKG, either under the direct supervision of hospital staff or via a fully automated pipeline controlled by the hospital. (4) Finally, verified evaluation scripts are executed, which only return aggregated performance metrics. This enables immediate performance feedback without sensitive patient data or individual predictions, leaving the clinic. A fundamental element of this approach involves the incorporation of a cKG, which serves to organize multi-omics and patient data within the context of real-world hospital environments. This approach was successfully validated during the TUM.ai Makeathon 2024 (TUMaiM24) challenge set by the Dr. von Hauner Children's Hospital (HCH-LMU): 50 students developed models for patient classification and diagnosis without access to real data. Deploying secure algorithms via federated frameworks, such as the FC framework, could be a practical way of achieving privacy-preserving AI in healthcare.","author":[{"family":"Süwer","given":"Simon"},{"family":"Mai","given":"Mai"},{"family":"Klein","given":"Christoph"},{"family":"Götzenberger","given":"Nicola"},{"family":"Dalić","given":"Denis"},{"family":"Maier","given":"Andreas"},{"family":"Baumbach","given":"Jan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.25277","URL":"https://doi.org/10.48550/arxiv.2510.25277","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.11400","type":"manuscript","title":"FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management","abstract":"Federated Learning (FL) emerges as a new learning paradigm that enables multiple devices to collaboratively train a shared model while preserving data privacy. However, one fundamental and prevailing challenge that hinders the deployment of FL on mobile devices is the memory limitation. This paper proposes \\textit{FedHybrid}, a novel framework that effectively reduces the memory footprint during the training process while guaranteeing the model accuracy and the overall training progress. Specifically, \\textit{FedHybrid} first selects the participating devices for each training round by jointly evaluating their memory budget, computing capability, and data diversity. After that, it judiciously analyzes the computational graph and generates an execution plan for each selected client in order to meet the corresponding memory budget while minimizing the training delay through employing a hybrid of recomputation and compression techniques according to the characteristic of each tensor. During the local training process, \\textit{FedHybrid} carries out the execution plan with a well-designed activation compression technique to effectively achieve memory reduction with minimum accuracy loss. We conduct extensive experiments to evaluate \\textit{FedHybrid} on both simulation and off-the-shelf mobile devices. The experiment results demonstrate that \\textit{FedHybrid} achieves up to a 39.1\\% increase in model accuracy and a 15.5$\\times$ reduction in wall clock time under various memory budgets compared with the baselines.","author":[{"family":"Tam","given":"Kahou"},{"family":"Tian","given":"Chunlin"},{"family":"Li","given":"Li"},{"family":"Zhao","given":"Haikai"},{"family":"Xu","given":"Chengzhong"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.11400","URL":"https://doi.org/10.48550/arxiv.2510.11400","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.21596","type":"manuscript","title":"In-network Attack Detection with Federated Deep Learning in IoT Networks: Real Implementation and Analysis","abstract":"The rapid expansion of the Internet of Things (IoT) and its integration with backbone networks have heightened the risk of security breaches. Traditional centralized approaches to anomaly detection, which require transferring large volumes of data to central servers, suffer from privacy, scalability, and latency limitations. This paper proposes a lightweight autoencoder-based anomaly detection framework designed for deployment on resource-constrained edge devices, enabling real-time detection while minimizing data transfer and preserving privacy. Federated learning is employed to train models collaboratively across distributed devices, where local training occurs on edge nodes and only model weights are aggregated at a central server. A real-world IoT testbed using Raspberry Pi sensor nodes was developed to collect normal and attack traffic data. The proposed federated anomaly detection system, implemented and evaluated on the testbed, demonstrates its effectiveness in accurately identifying network attacks. The communication overhead was reduced significantly while achieving comparable performance to the centralized method.","author":[{"family":"Chaudhary","given":"Devashish"},{"family":"Rajasegarar","given":"Sutharshan"},{"family":"Pokhrel","given":"Shiva"},{"family":"Pan","given":"Lei"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.21596","URL":"https://doi.org/10.48550/arxiv.2603.21596","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.08797","type":"manuscript","title":"Wireless Decentralized Federated Learning via Device Clustering and Inter-Cluster Link Enhancement","abstract":"Decentralized federated learning (DFL) dispenses with the central server of classical FL by utilizing peer-to-peer model exchanges among edge devices. This server-free architecture enables ad-hoc, flexible distributed learning in large device-to-device (D2D) networks. However, wireless DFL converges slowly because peer-to-peer model aggregation incurs high delays and errors. Each DFL training round involves many-to-many gradient sharing over wireless channels, resulting in uncoordinated channel access, large communication errors from stragglers, and slow model consensus, especially in large-scale D2D networks with pronounced clustering structures. We address these aggregation bottlenecks by provisioning a few reliable backhaul links at straggling nodes to enhance network connectivity. Building on this idea, our budget-aware, cluster-centric DFL framework first partitions the network into densely connected clusters, and then allocates the limited backhaul budget to selected cluster heads. The resulting two-tier protocol executes fast, parallel model aggregation within clusters and infrequent inter-cluster exchanges among the heads, yielding an O(1/t) convergence rate in t iterations. Numerical experiments on image-classification tasks confirm that our approach accelerates convergence compared to state-of-the-art DFL baselines with only a few strategically placed backhaul links.","author":[{"family":"Zheng","given":"William"},{"family":"Liu","given":"Hang"},{"family":"Zhang","given":"Ying"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.08797","URL":"https://doi.org/10.48550/arxiv.2607.08797","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.01750","type":"manuscript","title":"Communication-Aware Knowledge Distillation for Federated LLM Fine-Tuning over Wireless Networks","abstract":"Federated learning (FL) for large language models (LLMs) offers a privacy-preserving scheme, enabling clients to collaboratively fine-tune locally deployed LLMs or smaller language models (SLMs) without exchanging raw data. While parameter-sharing methods in traditional FL models solves number of technical challenges, they still incur high communication overhead and struggle with adapting to heterogeneous model architectures. Federated distillation, a framework for mutual knowledge transfer via shared logits, typically offers lower communication overhead than parameter-sharing methods. However, transmitting logits from LLMs remains challenging for bandwidth-limited clients due to their high dimensionality. In this work, we focus on a federated LLM distillation with efficient communication overhead. To achieve this, we first propose an adaptive Top-k logit selection mechanism, dynamically sparsifying logits according to real-time communication conditions. Then to tackle the dimensional inconsistency introduced by the adaptive sparsification, we design an adaptive logits aggregation scheme, effectively alleviating the artificial and uninformative inputs introduced by conventional zero-padding methods. Finally, to enhance the distillation effect, we incorporate LoRA-adapted hidden-layer projection from LLM into the distillation loss, reducing the communication overhead further while providing richer representation. Experimental results demonstrate that our scheme achieves superior performance compared to baseline methods while effectively reducing communication overhead by approximately 50%.","author":[{"family":"Zhang","given":"Xinlu"},{"family":"Yan","given":"Na"},{"family":"Su","given":"Yang"},{"family":"Deng","given":"Yansha"},{"family":"Mahmoodi","given":"Toktam"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.01750","URL":"https://doi.org/10.48550/arxiv.2509.01750","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.21474","type":"manuscript","title":"Towards Transparent Mental Health Insights: An Explainable AI Model for Career-Related Depression and Anxiety Among University Students Using Structured Data","abstract":"Career anxiety and depression among university students present a growing challenge to mental health and academic achievement. This study proposes an Explainable AI (XAI) framework using multimodal data and Federated Learning (FL) to identify early indicators of career-related mental health problems in a privacy-preserving and culturally responsive manner. The framework combines structured behavioral data and facial emotion features from interview videos via an intermediate fusion neural network with attention mechanisms. Label smoothing was applied to improve model generalizability. FL was used across institutions to enable collaborative training without raw data sharing. Evaluation was conducted using the Student Mental Health Survey dataset from university students across Pakistan. Our model attained an F1-score of 89.12%, recall of 86.54%, accuracy of 92.08%, and precision of 91.88%. Using Integrated Gradients and SHAP, the model identified key behavioral markers of depression including avoidance of direct gaze, lower facial expressiveness, and social withdrawal, consistent with psychological theory. This research presents an interpretable, scalable, and context-sensitive AI system for mental health pre-diagnosis with potential integration into student support services globally.","author":[{"family":"Azam","given":"Arsham"},{"family":"Ali","given":"Rasikh"},{"family":"Farhat","given":"Tayyaba"},{"family":"Akram","given":"Sheeraz"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.21474","URL":"https://doi.org/10.48550/arxiv.2606.21474","source":"datacite"},{"id":"doi:10.5281/zenodo.20529861","type":"article-journal","title":"FEDERATED LEARNING FOR PRIVACY-PRESERVING THREAT INTELLIGENCE SHARING IN DISTRIBUTED CYBERSECURITY ECOSYSTEMS","abstract":"Effective cybersecurity threat intelligence depends fundamentally on the breadth and timeliness of threat data — yet the organizations most capable of generating actionable intelligence are simultaneously most constrained in sharing it due to privacy regulations (GDPR, HIPAA, PDPA), competitive concerns, legal liability, and national security classifications. This tension between intelligence sharing and data privacy represents one of the most consequential unsolved challenges in cybersecurity: organizations that share threat intelligence detect attacks 2.4 times faster and suffer 47.3% lower breach costs, yet fewer than 23% of enterprises engage in structured threat intelligence sharing due to these barriers. This paper presents FedThreat-AI, a novel federated learning framework enabling privacy-preserving threat intelligence sharing across distributed cybersecurity ecosystems without requiring any organization to expose its raw security data, proprietary detection rules, or sensitive network topology. FedThreat-AI integrates four privacy-enhancing technologies — differential privacy (DP), homomorphic encryption (HE), secure multiparty computation (SMPC), and Byzantine-robust gradient aggregation — into a unified federated learning pipeline trained on distributed threat telemetry across participating organizations. The framework produces a continuously improving global threat detection model incorporating the collective intelligence of all participants, distributed back to each organization as model updates rather than data. Evaluated across a consortium of 24 organizations spanning financial services, healthcare, government, and technology sectors over 18 months (2023–2025), FedThreat-AI achieves global threat detection accuracy of 96.8% — only 1.4 percentage points below a centralized baseline that requires full data sharing — while providing mathematically provable privacy guarantees (ε = 0.8, δ = 10⁻⁵ per training round). The framework further demonstrates resilience against Byzantine poisoning attacks from up to 30% malicious participants and reduces mean time to detect novel threat campaigns by 67.4% compared to organization-siloed detection. FedThreat-AI is fully compatible with STIX 2.1 and TAXII 2.1 standards, enabling integration with existing threat intelligence platforms and ISACs","author":[{"family":"Dr Angira A","given":"Patel"},{"family":"Nilam","given":"Joshi"},{"family":"Vaidehi","given":"Patel"},{"family":"Avani","given":"Vagadiya"},{"family":"Dhruvi","given":"Pandya"},{"family":"Dr Kamalesh","given":"VN"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20529861","URL":"https://doi.org/10.5281/zenodo.20529861","source":"datacite"},{"id":"doi:10.5281/zenodo.20529862","type":"article-journal","title":"FEDERATED LEARNING FOR PRIVACY-PRESERVING THREAT INTELLIGENCE SHARING IN DISTRIBUTED CYBERSECURITY ECOSYSTEMS","abstract":"Effective cybersecurity threat intelligence depends fundamentally on the breadth and timeliness of threat data — yet the organizations most capable of generating actionable intelligence are simultaneously most constrained in sharing it due to privacy regulations (GDPR, HIPAA, PDPA), competitive concerns, legal liability, and national security classifications. This tension between intelligence sharing and data privacy represents one of the most consequential unsolved challenges in cybersecurity: organizations that share threat intelligence detect attacks 2.4 times faster and suffer 47.3% lower breach costs, yet fewer than 23% of enterprises engage in structured threat intelligence sharing due to these barriers. This paper presents FedThreat-AI, a novel federated learning framework enabling privacy-preserving threat intelligence sharing across distributed cybersecurity ecosystems without requiring any organization to expose its raw security data, proprietary detection rules, or sensitive network topology. FedThreat-AI integrates four privacy-enhancing technologies — differential privacy (DP), homomorphic encryption (HE), secure multiparty computation (SMPC), and Byzantine-robust gradient aggregation — into a unified federated learning pipeline trained on distributed threat telemetry across participating organizations. The framework produces a continuously improving global threat detection model incorporating the collective intelligence of all participants, distributed back to each organization as model updates rather than data. Evaluated across a consortium of 24 organizations spanning financial services, healthcare, government, and technology sectors over 18 months (2023–2025), FedThreat-AI achieves global threat detection accuracy of 96.8% — only 1.4 percentage points below a centralized baseline that requires full data sharing — while providing mathematically provable privacy guarantees (ε = 0.8, δ = 10⁻⁵ per training round). The framework further demonstrates resilience against Byzantine poisoning attacks from up to 30% malicious participants and reduces mean time to detect novel threat campaigns by 67.4% compared to organization-siloed detection. FedThreat-AI is fully compatible with STIX 2.1 and TAXII 2.1 standards, enabling integration with existing threat intelligence platforms and ISACs","author":[{"family":"Dr Angira A","given":"Patel"},{"family":"Nilam","given":"Joshi"},{"family":"Vaidehi","given":"Patel"},{"family":"Avani","given":"Vagadiya"},{"family":"Dhruvi","given":"Pandya"},{"family":"Dr Kamalesh","given":"VN"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20529862","URL":"https://doi.org/10.5281/zenodo.20529862","source":"datacite"},{"id":"doi:10.5281/zenodo.20443410","type":"article-journal","title":"BioHackathon Europe 2025 Report","abstract":"This report documents BioHackathon Europe 2025, ELIXIR's annual flagship collaborative hacking event, held 3 to 7 November 2025 at the Hotel Esplanade Resort & Spa, Bad Saarow, Germany. Organised by the ELIXIR Hub, the event brought together an international community of life scientists, developers, data stewards and infrastructure experts to accelerate the development of open, interoperable solutions for the life sciences. The 2025 event featured 31 hacking projects across five days of intensive collaboration. Projects spanned a broad range of ELIXIR strategic priority areas including research data management and FAIR implementation, AI and machine learning readiness, biodiversity genomics, secure and federated data access, cloud and workflow infrastructures, and sustainable computing. The event programme included opening flash presentations, daily uninterrupted hacking sessions, a new interactive mid-week poster session (replacing the traditional reporting format), and final presentations. Shared social activities fostered community building and cross-project collaboration. A post-event participant survey indicated strong satisfaction with the format, with dedicated hacking time rated as the most valuable component. The majority of respondents reported improved technical skills as a result of participation. Related resources Event photos: Flickr album Project outputs: GitHub repository Preprints: BioHackrXiv","author":[{"family":"Van Wyk","given":"Deborah"},{"family":"Anton","given":"Mihail"},{"family":"Heil","given":"Katharina"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20443410","URL":"https://doi.org/10.5281/zenodo.20443410","source":"datacite"},{"id":"doi:10.5281/zenodo.20446860","type":"article-journal","title":"BioHackathon Europe 2025 Report","abstract":"This report documents BioHackathon Europe 2025, ELIXIR's annual flagship collaborative hacking event, held 3 to 7 November 2025 at the Hotel Esplanade Resort & Spa, Bad Saarow, Germany. Organised by the ELIXIR Hub, the event brought together an international community of life scientists, developers, data stewards and infrastructure experts to accelerate the development of open, interoperable solutions for the life sciences. The 2025 event featured 31 hacking projects across five days of intensive collaboration. Projects spanned a broad range of ELIXIR strategic priority areas including research data management and FAIR implementation, AI and machine learning readiness, biodiversity genomics, secure and federated data access, cloud and workflow infrastructures, and sustainable computing. The event programme included opening flash presentations, daily uninterrupted hacking sessions, a new interactive mid-week poster session (replacing the traditional reporting format), and final presentations. Shared social activities fostered community building and cross-project collaboration. A post-event participant survey indicated strong satisfaction with the format, with dedicated hacking time rated as the most valuable component. The majority of respondents reported improved technical skills as a result of participation. Related resources Event photos: Flickr album Project outputs: GitHub repository Preprints: BioHackrXiv","author":[{"family":"Van Wyk","given":"Deborah"},{"family":"Anton","given":"Mihail"},{"family":"Heil","given":"Katharina"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20446860","URL":"https://doi.org/10.5281/zenodo.20446860","source":"datacite"},{"id":"doi:10.5281/zenodo.20446063","type":"article-journal","title":"BioHackathon Europe 2025 Report","abstract":"This report documents BioHackathon Europe 2025, ELIXIR's annual flagship collaborative hacking event, held 3 to 7 November 2025 at the Hotel Esplanade Resort & Spa, Bad Saarow, Germany. Organised by the ELIXIR Hub, the event brought together an international community of life scientists, developers, data stewards and infrastructure experts to accelerate the development of open, interoperable solutions for the life sciences. The 2025 event featured 31 hacking projects across five days of intensive collaboration. Projects spanned a broad range of ELIXIR strategic priority areas including research data management and FAIR implementation, AI and machine learning readiness, biodiversity genomics, secure and federated data access, cloud and workflow infrastructures, and sustainable computing. The event programme included opening flash presentations, daily uninterrupted hacking sessions, a new interactive mid-week poster session (replacing the traditional reporting format), and final presentations. Shared social activities fostered community building and cross-project collaboration. A post-event participant survey indicated strong satisfaction with the format, with dedicated hacking time rated as the most valuable component. The majority of respondents reported improved technical skills as a result of participation. Related resources Event photos: Flickr album Project outputs: GitHub repository Preprints: BioHackrXiv","author":[{"family":"Van Wyk","given":"Deborah"},{"family":"Anton","given":"Mihail"},{"family":"Heil","given":"Katharina"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20446063","URL":"https://doi.org/10.5281/zenodo.20446063","source":"datacite"},{"id":"doi:10.5281/zenodo.20443411","type":"article-journal","title":"BioHackathon Europe 2025 Report","abstract":"This report documents BioHackathon Europe 2025, ELIXIR's annual flagship collaborative hacking event, held 3 to 7 November 2025 at the Hotel Esplanade Resort & Spa, Bad Saarow, Germany. Organised by the ELIXIR Hub, the event brought together an international community of life scientists, developers, data stewards and infrastructure experts to accelerate the development of open, interoperable solutions for the life sciences. The 2025 event featured 31 hacking projects across five days of intensive collaboration. Projects spanned a broad range of ELIXIR strategic priority areas including research data management and FAIR implementation, AI and machine learning readiness, biodiversity genomics, secure and federated data access, cloud and workflow infrastructures, and sustainable computing. The event programme included opening flash presentations, daily uninterrupted hacking sessions, a new interactive mid-week poster session (replacing the traditional reporting format), and final presentations. Shared social activities fostered community building and cross-project collaboration. A post-event participant survey indicated strong satisfaction with the format, with dedicated hacking time rated as the most valuable component. The majority of respondents reported improved technical skills as a result of participation. Hybrid participation was supported throughout, with around three quarters of project teams including at least one remote participant. Following the event, 11 preprints were published on BioHackrXiv. Related resources Event photos: https://www.flickr.com/photos/elixir-europe/albums/72177720330241026/with/54916490416/ Project outputs: BioHackathon Europe 2025 GitHub repository Preprints: BioHackrXiv (indexed in Europe PMC since 2021) Event coordination: Slack channel #BioHackEU25","author":[{"family":"Van Wyk","given":"Deborah"},{"family":"Anton","given":"Mihail"},{"family":"Heil","given":"Katharina"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20443411","URL":"https://doi.org/10.5281/zenodo.20443411","source":"datacite"},{"id":"doi:10.6084/m9.figshare.29637053","type":"article-journal","title":"2025-CyberTraining PIMeeting-Poster-T3-CIDERS.pdf","abstract":"This poster introduces the T3-CIDERS project, a Train-the-Trainer initiative designed to build a national community of practice around cyberinfrastructure (CI)- and data-enabled cybersecurity research and education. The project equips faculty-student teams (Future Trainers) with hands-on experience in CI technologies, such as HPC, big data, ML, and cryptography, and pedagogical strategies to teach these concepts through cybersecurity applications. Highlights include the successful completion of the first cohort, six outreach events reaching nearly 100 students nationwide, and the development of a new module on Federated Learning Security. Upcoming activities include a Winter Institute in January 2026 and recruitment for the second cohort in Fall 2025.","author":[{"family":"Jiang","given":"Peng"},{"family":"Sosonkina","given":"Masha"},{"family":"Wu","given":"Hongyi"},{"family":"Purwanto","given":"Wirawan"},{"family":"Yang","given":"Mohan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.6084/m9.figshare.29637053","URL":"https://doi.org/10.6084/m9.figshare.29637053","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.18416","type":"manuscript","title":"Exploring Potential Prompt Injection Attacks in Federated Military LLMs and Their Mitigation","abstract":"Federated Learning (FL) is increasingly being adopted in military collaborations to develop Large Language Models (LLMs) while preserving data sovereignty. However, prompt injection attacks-malicious manipulations of input prompts-pose new threats that may undermine operational security, disrupt decision-making, and erode trust among allies. This perspective paper highlights four vulnerabilities in federated military LLMs: secret data leakage, free-rider exploitation, system disruption, and misinformation spread. To address these risks, we propose a human-AI collaborative framework with both technical and policy countermeasures. On the technical side, our framework uses red/blue team wargaming and quality assurance to detect and mitigate adversarial behaviors of shared LLM weights. On the policy side, it promotes joint AI-human policy development and verification of security protocols.","author":[{"family":"Lee","given":"Youngjoon"},{"family":"Park","given":"Taehyun"},{"family":"Lee","given":"Yunho"},{"family":"Gong","given":"Jinu"},{"family":"Kang","given":"Joonhyuk"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.18416","URL":"https://doi.org/10.48550/arxiv.2501.18416","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.12661","type":"manuscript","title":"An Efficient and Adaptive Framework for Achieving Underwater High-performance Maintenance Networks","abstract":"With the development of space-air-ground-aqua integrated networks (SAGAIN), high-speed and reliable network services are accessible at any time and any location. However, the long propagation delay and limited network capacity of underwater communication networks (UCN) negatively impact the service quality of SAGAIN. To address this issue, this paper presents U-HPNF, a hierarchical framework designed to achieve a high-performance network with self-management, self-configuration, and self-optimization capabilities. U-HPNF leverages the sensing and decision-making capabilities of deep reinforcement learning (DRL) to manage limited resources in UCNs, including communication bandwidth, computational resources, and energy supplies. Additionally, we incorporate federated learning (FL) to iteratively optimize the decision-making model, thereby reducing communication overhead and protecting the privacy of node observation information. By deploying digital twins (DT) at both the intelligent sink layer and aggregation layer, U-HPNF can mimic numerous network scenarios and adapt to varying network QoS requirements. Through a three-tier network design with two-levels DT, U-HPNF provides an AI-native high-performance underwater network. Numerical results demonstrate that the proposed U-HPNF framework can effectively optimize network performance across various situations and adapt to changing QoS requirements.","author":[{"family":"Gou","given":"Yu"},{"family":"Zhang","given":"Tong"},{"family":"Liu","given":"Jun"},{"family":"Qi","given":"Zhongyang"},{"family":"Zheng","given":"Dezhi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.12661","URL":"https://doi.org/10.48550/arxiv.2508.12661","source":"datacite"},{"id":"doi:10.48550/arxiv.2502.05547","type":"manuscript","title":"Dual Defense: Enhancing Privacy and Mitigating Poisoning Attacks in Federated Learning","abstract":"Federated learning (FL) is inherently susceptible to privacy breaches and poisoning attacks. To tackle these challenges, researchers have separately devised secure aggregation mechanisms to protect data privacy and robust aggregation methods that withstand poisoning attacks. However, simultaneously addressing both concerns is challenging; secure aggregation facilitates poisoning attacks as most anomaly detection techniques require access to unencrypted local model updates, which are obscured by secure aggregation. Few recent efforts to simultaneously tackle both challenges offen depend on impractical assumption of non-colluding two-server setups that disrupt FL's topology, or three-party computation which introduces scalability issues, complicating deployment and application. To overcome this dilemma, this paper introduce a Dual Defense Federated learning (DDFed) framework. DDFed simultaneously boosts privacy protection and mitigates poisoning attacks, without introducing new participant roles or disrupting the existing FL topology. DDFed initially leverages cutting-edge fully homomorphic encryption (FHE) to securely aggregate model updates, without the impractical requirement for non-colluding two-server setups and ensures strong privacy protection. Additionally, we proposes a unique two-phase anomaly detection mechanism for encrypted model updates, featuring secure similarity computation and feedback-driven collaborative selection, with additional measures to prevent potential privacy breaches from Byzantine clients incorporated into the detection process. We conducted extensive experiments on various model poisoning attacks and FL scenarios, including both cross-device and cross-silo FL. Experiments on publicly available datasets demonstrate that DDFed successfully protects model privacy and effectively defends against model poisoning threats.","author":[{"family":"Xu","given":"Runhua"},{"family":"Gao","given":"Shiqi"},{"family":"Li","given":"Chao"},{"family":"Joshi","given":"James"},{"family":"Li","given":"Jianxin"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2502.05547","URL":"https://doi.org/10.48550/arxiv.2502.05547","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.13213","type":"manuscript","title":"Distributed Intrusion Detection in Dynamic Networks of UAVs using Few-Shot Federated Learning","abstract":"Flying Ad Hoc Networks (FANETs), which primarily interconnect Unmanned Aerial Vehicles (UAVs), present distinctive security challenges due to their distributed and dynamic characteristics, necessitating tailored security solutions. Intrusion detection in FANETs is particularly challenging due to communication costs, and privacy concerns. While Federated Learning (FL) holds promise for intrusion detection in FANETs with its cooperative and decentralized model training, it also faces drawbacks such as large data requirements, power consumption, and time constraints. Moreover, the high speeds of nodes in dynamic networks like FANETs may disrupt communication among Intrusion Detection Systems (IDS). In response, our study explores the use of few-shot learning (FSL) to effectively reduce the data required for intrusion detection in FANETs. The proposed approach called Few-shot Federated Learning-based IDS (FSFL-IDS) merges FL and FSL to tackle intrusion detection challenges such as privacy, power constraints, communication costs, and lossy links, demonstrating its effectiveness in identifying routing attacks in dynamic FANETs.This approach reduces both the local models and the global model's training time and sample size, offering insights into reduced computation and communication costs and extended battery life. Furthermore, by employing FSL, which requires less data for training, IDS could be less affected by lossy links in FANETs.","author":[{"family":"Ceviz","given":"Ozlem"},{"family":"Sen","given":"Sevil"},{"family":"Sadioglu","given":"Pinar"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.13213","URL":"https://doi.org/10.48550/arxiv.2501.13213","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.00732","type":"manuscript","title":"Gradient Compression and Correlation Driven Federated Learning for Wireless Traffic Prediction","abstract":"Wireless traffic prediction plays an indispensable role in cellular networks to achieve proactive adaptation for communication systems. Along this line, Federated Learning (FL)-based wireless traffic prediction at the edge attracts enormous attention because of the exemption from raw data transmission and enhanced privacy protection. However FL-based wireless traffic prediction methods still rely on heavy data transmissions between local clients and the server for local model updates. Besides, how to model the spatial dependencies of local clients under the framework of FL remains uncertain. To tackle this, we propose an innovative FL algorithm that employs gradient compression and correlation-driven techniques, effectively minimizing data transmission load while preserving prediction accuracy. Our approach begins with the introduction of gradient sparsification in wireless traffic prediction, allowing for significant data compression during model training. We then implement error feedback and gradient tracking methods to mitigate any performance degradation resulting from this compression. Moreover, we develop three tailored model aggregation strategies anchored in gradient correlation, enabling the capture of spatial dependencies across diverse clients. Experiments have been done with two real-world datasets and the results demonstrate that by capturing the spatio-temporal characteristics and correlation among local clients, the proposed algorithm outperforms the state-of-the-art algorithms and can increase the communication efficiency by up to two orders of magnitude without losing prediction accuracy. Code is available at https://github.com/chuanting/FedGCC.","author":[{"family":"Zhang","given":"Chuanting"},{"family":"Zhang","given":"Haixia"},{"family":"Dang","given":"Shuping"},{"family":"Shihada","given":"Basem"},{"family":"Alouini","given":"Mohamed"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.00732","URL":"https://doi.org/10.48550/arxiv.2501.00732","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.16881","type":"manuscript","title":"Federated Multi Agent Deep Learning and Neural Networks for Advanced Distributed Sensing in Wireless Networks","abstract":"Multi-agent deep learning (MADL), including multi-agent deep reinforcement learning (MADRL), distributed/federated training, and graph-structured neural networks, is becoming a unifying framework for decision-making and inference in wireless systems where sensing, communication, and computing are tightly coupled. Recent 5G-Advanced and 6G visions strengthen this coupling through integrated sensing and communication, edge intelligence, open programmable RAN, and non-terrestrial/UAV networking, which create decentralized, partially observed, time-varying, and resource-constrained control problems. This survey synthesizes the state of the art, with emphasis on 2021-2025 research, on MADL for distributed sensing and wireless communications. We present a task-driven taxonomy across (i) learning formulations (Markov games, Dec-POMDPs, CTDE), (ii) neural architectures (GNN-based radio resource management, attention-based policies, hierarchical learning, and over-the-air aggregation), (iii) advanced techniques (federated reinforcement learning, communication-efficient federated deep RL, and serverless edge learning orchestration), and (iv) application domains (MEC offloading with slicing, UAV-enabled heterogeneous networks with power-domain NOMA, intrusion detection in sensor networks, and ISAC-driven perceptive mobile networks). We also provide comparative tables of algorithms, training topologies, and system-level trade-offs in latency, spectral efficiency, energy, privacy, and robustness. Finally, we identify open issues including scalability, non-stationarity, security against poisoning and backdoors, communication overhead, and real-time safety, and outline research directions toward 6G-native sense-communicate-compute-learn systems.","author":[{"family":"Muller","given":"Nadine"},{"family":"Derosa","given":"Stefano"},{"family":"Zhang","given":"Su"},{"family":"Huan","given":"Chun"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.16881","URL":"https://doi.org/10.48550/arxiv.2603.16881","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.08372","type":"manuscript","title":"Rethinking the Backbone in Class Imbalanced Federated Source Free Domain Adaptation: The Utility of Vision Foundation Models","abstract":"Federated Learning (FL) offers a framework for training models collaboratively while preserving data privacy of each client. Recently, research has focused on Federated Source-Free Domain Adaptation (FFREEDA), a more realistic scenario wherein client-held target domain data remains unlabeled, and the server can access source domain data only during pre-training. We extend this framework to a more complex and realistic setting: Class Imbalanced FFREEDA (CI-FFREEDA), which takes into account class imbalances in both the source and target domains, as well as label shifts between source and target and among target clients. The replication of existing methods in our experimental setup lead us to rethink the focus from enhancing aggregation and domain adaptation methods to improving the feature extractors within the network itself. We propose replacing the FFREEDA backbone with a frozen vision foundation model (VFM), thereby improving overall accuracy without extensive parameter tuning and reducing computational and communication costs in federated learning. Our experimental results demonstrate that VFMs effectively mitigate the effects of domain gaps, class imbalances, and even non-IID-ness among target clients, suggesting that strong feature extractors, not complex adaptation or FL methods, are key to success in the real-world FL.","author":[{"family":"Kihara","given":"Kosuke"},{"family":"Mori","given":"Junki"},{"family":"Miyagawa","given":"Taiki"},{"family":"Ebihara","given":"Akinori"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.08372","URL":"https://doi.org/10.48550/arxiv.2509.08372","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.04887","type":"manuscript","title":"Federated Modality-specific Encoders and Partially Personalized Fusion Decoder for Multimodal Brain Tumor Segmentation","abstract":"Most existing federated learning (FL) methods for medical image analysis only considered intramodal heterogeneity, limiting their applicability to multimodal imaging applications. In practice, some FL participants may possess only a subset of the complete imaging modalities, posing intermodal heterogeneity as a challenge to effectively training a global model on all participants' data. Meanwhile, each participant expects a personalized model tailored to its local data characteristics in FL. This work proposes a new FL framework with federated modality-specific encoders and partially personalized multimodal fusion decoders (FedMEPD) to address the two concurrent issues. Specifically, FedMEPD employs an exclusive encoder for each modality to account for the intermodal heterogeneity. While these encoders are fully federated, the decoders are partially personalized to meet individual needs -- using the discrepancy between global and local parameter updates to dynamically determine which decoder filters are personalized. Implementation-wise, a server with full-modal data employs a fusion decoder to fuse representations from all modality-specific encoders, thus bridging the modalities to optimize the encoders via backpropagation. Moreover, multiple anchors are extracted from the fused multimodal representations and distributed to the clients in addition to the model parameters. Conversely, the clients with incomplete modalities calibrate their missing-modal representations toward the global full-modal anchors via scaled dot-product cross-attention, making up for the information loss due to absent modalities. FedMEPD is validated on the BraTS 2018 and 2020 multimodal brain tumor segmentation benchmarks. Results show that it outperforms various up-to-date methods for multimodal and personalized FL, and its novel designs are effective.","author":[{"family":"Liu","given":"Hong"},{"family":"Wei","given":"Dong"},{"family":"Dai","given":"Qian"},{"family":"Wu","given":"Xian"},{"family":"Zheng","given":"Yefeng"},{"family":"Wang","given":"Liansheng"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.04887","URL":"https://doi.org/10.48550/arxiv.2603.04887","source":"datacite"},{"id":"doi:10.5281/zenodo.18766337","type":"article-journal","title":"A Comparative Review of Machine Learning Approaches for Manufacturing Applications in Industry 4.0","abstract":"Industry 4.0 has redefined modern manufacturing by integrating cyber–physical systems, Industrial Internet of Things (IIoT), cloud–edge computing, and data-driven intelligence. Among these enablers, machine learning (ML) has emerged as a foundational technology for extracting actionable insights from heterogeneous manufacturing data. This paper presents an extended and comparative review of ML and deep learning (DL) techniques—including supervised, unsupervised, semi-supervised, reinforcement learning, and hybrid models—applied across core manufacturing domains such as predictive maintenance, quality inspection and defect detection, process optimization, production planning, and supply chain management. Based on a systematic analysis of literature published between 2015 and 2025, the review compares algorithmic performance, computational complexity, interpretability, and deployment feasibility. Mathematical formulations of commonly used models, including regression, support vector machines, convolutional neural networks (CNNs), and long short-term memory (LSTM) networks, are presented to enhance methodological clarity. Emerging trends such as transfer learning, federated learning, edge AI, and explainable artificial intelligence (XAI) are discussed in the context of industrial scalability and reliability. The study concludes that context-aware model selection, combined with hybrid and explainable frameworks, is critical for bridging the gap between laboratory-scale ML models and real-world smart manufacturing systems.","author":[{"family":"Paswan","given":"Veeru"},{"family":"Gupta","given":"Shalu"},{"family":"Gurleen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18766337","URL":"https://doi.org/10.5281/zenodo.18766337","source":"datacite"},{"id":"doi:10.5281/zenodo.18766338","type":"article-journal","title":"A Comparative Review of Machine Learning Approaches for Manufacturing Applications in Industry 4.0","abstract":"Industry 4.0 has redefined modern manufacturing by integrating cyber–physical systems, Industrial Internet of Things (IIoT), cloud–edge computing, and data-driven intelligence. Among these enablers, machine learning (ML) has emerged as a foundational technology for extracting actionable insights from heterogeneous manufacturing data. This paper presents an extended and comparative review of ML and deep learning (DL) techniques—including supervised, unsupervised, semi-supervised, reinforcement learning, and hybrid models—applied across core manufacturing domains such as predictive maintenance, quality inspection and defect detection, process optimization, production planning, and supply chain management. Based on a systematic analysis of literature published between 2015 and 2025, the review compares algorithmic performance, computational complexity, interpretability, and deployment feasibility. Mathematical formulations of commonly used models, including regression, support vector machines, convolutional neural networks (CNNs), and long short-term memory (LSTM) networks, are presented to enhance methodological clarity. Emerging trends such as transfer learning, federated learning, edge AI, and explainable artificial intelligence (XAI) are discussed in the context of industrial scalability and reliability. The study concludes that context-aware model selection, combined with hybrid and explainable frameworks, is critical for bridging the gap between laboratory-scale ML models and real-world smart manufacturing systems.","author":[{"family":"Paswan","given":"Veeru"},{"family":"Gupta","given":"Shalu"},{"family":"Gurleen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18766338","URL":"https://doi.org/10.5281/zenodo.18766338","source":"datacite"},{"id":"doi:10.5281/zenodo.18625147","type":"article-journal","title":"White paper #1: HPC-Cloud and Quantum Computing: State of the Art and Innovation Roadmap","abstract":"This white paper presents an overview of the current state of the art in high-performance computing (HPC) and its convergence with cloud technologies, with a strategic focus on innovation management and exploitation. It outlines recent advances in cloud-based HPC services, hybrid architectures, and federated systems that integrate edge, cloud, and HPC resources. Emerging paradigms such as federated learning, AIdriven optimization, and sustainable computing are analyzed for their transformative potential. Special emphasis is placed on the NOUS project, which exemplifies a holistic approach to federated HPC-cloud services, combining technical innovation with robust exploitation strategies. NOUS goes one step beyond and explores the integration of quantum computing in handling data existing in the cloud. As of late 2025, the synergy between Cloud Computing and Quantum Computing (QC) has matured into a functional \"Quantum-as-a-Service\" (QaaS) model. While physical quantum hardware remains too fragile for on-premise deployment, cloud providers have democratized access to the \"Quantum Stack.\" This report highlights this progress, the issues and real future applications. NOUS addresses European priorities for digital sovereignty and data interoperability, supporting scalable, privacy-preserving, and AI-enabled HPC workflows. The paper concludes with a roadmap that positions NOUS as a reference architecture and innovation catalyst for Europe's distributed computing ecosystem.","author":[{"family":"Krokidas","given":"Panagiotis"},{"family":"Rekatsinas","given":"Christoforos"},{"family":"Terlixidis","given":"Periklis"},{"family":"Giannopoulos","given":"Georgios"},{"family":"Rallis","given":"Konstantinos"},{"family":"Dimitrakis","given":"Panagiotis"},{"family":"Melissourgos","given":"Nikolaos"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.18625147","URL":"https://doi.org/10.5281/zenodo.18625147","source":"datacite"},{"id":"doi:10.5281/zenodo.18625148","type":"article-journal","title":"White paper #1: HPC-Cloud and Quantum Computing: State of the Art and Innovation Roadmap","abstract":"This white paper presents an overview of the current state of the art in high-performance computing (HPC) and its convergence with cloud technologies, with a strategic focus on innovation management and exploitation. It outlines recent advances in cloud-based HPC services, hybrid architectures, and federated systems that integrate edge, cloud, and HPC resources. Emerging paradigms such as federated learning, AIdriven optimization, and sustainable computing are analyzed for their transformative potential. Special emphasis is placed on the NOUS project, which exemplifies a holistic approach to federated HPC-cloud services, combining technical innovation with robust exploitation strategies. NOUS goes one step beyond and explores the integration of quantum computing in handling data existing in the cloud. As of late 2025, the synergy between Cloud Computing and Quantum Computing (QC) has matured into a functional \"Quantum-as-a-Service\" (QaaS) model. While physical quantum hardware remains too fragile for on-premise deployment, cloud providers have democratized access to the \"Quantum Stack.\" This report highlights this progress, the issues and real future applications. NOUS addresses European priorities for digital sovereignty and data interoperability, supporting scalable, privacy-preserving, and AI-enabled HPC workflows. The paper concludes with a roadmap that positions NOUS as a reference architecture and innovation catalyst for Europe's distributed computing ecosystem.","author":[{"family":"Krokidas","given":"Panagiotis"},{"family":"Rekatsinas","given":"Christoforos"},{"family":"Terlixidis","given":"Periklis"},{"family":"Giannopoulos","given":"Georgios"},{"family":"Rallis","given":"Konstantinos"},{"family":"Dimitrakis","given":"Panagiotis"},{"family":"Melissourgos","given":"Nikolaos"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.18625148","URL":"https://doi.org/10.5281/zenodo.18625148","source":"datacite"},{"id":"doi:10.48550/arxiv.2504.15717","type":"manuscript","title":"Trusted Compute Units: A Framework for Chained Verifiable Computations","abstract":"Blockchain and distributed ledger technologies (DLTs) facilitate decentralized computations across trust boundaries. However, ensuring complex computations with low gas fees and confidentiality remains challenging. Recent advances in Confidential Computing -- leveraging hardware-based Trusted Execution Environments (TEEs) -- and Proof-carrying Data -- employing cryptographic Zero-Knowledge Virtual Machines (zkVMs) -- hold promise for secure, privacy-preserving off-chain and layer-2 computations. On the other side, a homogeneous reliance on a single technology, such as TEEs or zkVMs, is impractical for decentralized environments with heterogeneous computational requirements. This paper introduces the Trusted Compute Unit (TCU), a unifying framework that enables composable and interoperable verifiable computations across heterogeneous technologies. Our approach allows decentralized applications (dApps) to flexibly offload complex computations to TCUs, obtaining proof of correctness. These proofs can be anchored on-chain for automated dApp interactions, while ensuring confidentiality of input data, and integrity of output data. We demonstrate how TCUs can support a prominent blockchain use case, such as federated learning. By enabling secure off-chain interactions without incurring on-chain confirmation delays or gas fees, TCUs significantly improve system performance and scalability. Experimental insights and performance evaluations confirm the feasibility and practicality of this unified approach, advancing the state of the art in verifiable off-chain services for the blockchain ecosystem.","author":[{"family":"Castillo","given":"Fernando"},{"family":"Heiss","given":"Jonathan"},{"family":"Werner","given":"Sebastian"},{"family":"Tai","given":"Stefan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2504.15717","URL":"https://doi.org/10.48550/arxiv.2504.15717","source":"datacite"},{"id":"doi:10.48550/arxiv.2503.11146","type":"manuscript","title":"Layer-wise Update Aggregation with Recycling for Communication-Efficient Federated Learning","abstract":"Expensive communication cost is a common performance bottleneck in Federated Learning (FL), which makes it less appealing in real-world applications. Many communication-efficient FL methods focus on discarding a part of model updates mostly based on gradient magnitude. In this study, we find that recycling previous updates, rather than simply dropping them, more effectively reduces the communication cost while maintaining FL performance. We propose FedLUAR, a Layer-wise Update Aggregation with Recycling scheme for communication-efficient FL. We first define a useful metric that quantifies the extent to which the aggregated gradients influences the model parameter values in each layer. FedLUAR selects a few layers based on the metric and recycles their previous updates on the server side. Our extensive empirical study demonstrates that the update recycling scheme significantly reduces the communication cost while maintaining model accuracy. For example, our method achieves nearly the same AG News accuracy as FedAvg, while reducing the communication cost to just 17%.","author":[{"family":"Kim","given":"Jisoo"},{"family":"Kang","given":"Sungmin"},{"family":"Lee","given":"Sunwoo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2503.11146","URL":"https://doi.org/10.48550/arxiv.2503.11146","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.00718","type":"manuscript","title":"Federated Learning at the Forefront of Fairness: A Multifaceted Perspective","abstract":"Fairness in Federated Learning (FL) is emerging as a critical factor driven by heterogeneous clients' constraints and balanced model performance across various scenarios. In this survey, we delineate a comprehensive classification of the state-of-the-art fairness-aware approaches from a multifaceted perspective, i.e., model performance-oriented and capability-oriented. Moreover, we provide a framework to categorize and address various fairness concerns and associated technical aspects, examining their effectiveness in balancing equity and performance within FL frameworks. We further examine several significant evaluation metrics leveraged to measure fairness quantitatively. Finally, we explore exciting open research directions and propose prospective solutions that could drive future advancements in this important area, laying a solid foundation for researchers working toward fairness in FL.","author":[{"family":"Mukhtiar","given":"Noorain"},{"family":"Mahmood","given":"Adnan"},{"family":"Zhou","given":"Yipeng"},{"family":"Yang","given":"Jian"},{"family":"Teng","given":"Jing"},{"family":"Sheng","given":"Quan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.00718","URL":"https://doi.org/10.48550/arxiv.2602.00718","source":"datacite"},{"id":"doi:10.48550/arxiv.2502.08488","type":"manuscript","title":"One-Shot Federated Learning with Classifier-Free Diffusion Models","abstract":"Federated learning (FL) enables collaborative learning without data centralization but introduces significant communication costs due to multiple communication rounds between clients and the server. One-shot federated learning (OSFL) addresses this by forming a global model with a single communication round, often relying on the server's model distillation or auxiliary dataset generation - mostly through pre-trained diffusion models (DMs). Existing DM-assisted OSFL methods, however, typically employ classifier-guided DMs, which require training auxiliary classifier models at each client, introducing additional computation overhead. This work introduces OSCAR (One-Shot Federated Learning with Classifier-Free Diffusion Models), a novel OSFL approach that eliminates the need for auxiliary models. OSCAR uses foundation models to devise category-specific data representations at each client which are integrated into a classifier-free diffusion model pipeline for server-side data generation. In our experiments, OSCAR outperforms the state-of-the-art on four benchmark datasets while reducing the communication load by at least 99%.","author":[{"family":"Zaland","given":"Obaidullah"},{"family":"Jin","given":"Shutong"},{"family":"Pokorny","given":"Florian"},{"family":"Bhuyan","given":"Monowar"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2502.08488","URL":"https://doi.org/10.48550/arxiv.2502.08488","source":"datacite"},{"id":"doi:10.48550/arxiv.2601.19345","type":"manuscript","title":"AI-driven Intrusion Detection for UAV in Smart Urban Ecosystems: A Comprehensive Survey","abstract":"UAVs have the potential to revolutionize urban management and provide valuable services to citizens. They can be deployed across diverse applications, including traffic monitoring, disaster response, environmental monitoring, and numerous other domains. However, this integration introduces novel security challenges that must be addressed to ensure safe and trustworthy urban operations. This paper provides a structured, evidence-based synthesis of UAV applications in smart cities and their associated security challenges as reported in the literature over the last decade, with particular emphasis on developments from 2019 to 2025. We categorize these challenges into two primary classes: 1) cyber-attacks targeting the communication infrastructure of UAVs and 2) unwanted or unauthorized physical intrusions by UAVs themselves. We examine the potential of Artificial Intelligence (AI) techniques in developing intrusion detection mechanisms to mitigate these security threats. We analyze how AI-based methods, such as machine/deep learning for anomaly detection and computer vision for object recognition, can play a pivotal role in enhancing UAV security through unified detection systems that address both cyber and physical threats. Furthermore, we consolidate publicly available UAV datasets across network traffic and vision modalities suitable for Intrusion Detection Systems (IDS) development and evaluation. The paper concludes by identifying ten key research directions, including scalability, robustness, explainability, data scarcity, automation, hybrid detection, large language models, multimodal approaches, federated learning, and privacy preservation. Finally, we discuss the practical challenges of implementing UAV IDS solutions in real-world smart city environments.","author":[{"family":"Khanfor","given":"Abdullah"},{"family":"Hamadi","given":"Raby"},{"family":"Lasla","given":"Noureddine"},{"family":"Ghazzai","given":"Hakim"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2601.19345","URL":"https://doi.org/10.48550/arxiv.2601.19345","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.10654","type":"manuscript","title":"Efficient Transmission of Radiomaps via Physics-Enhanced Semantic Communications","abstract":"Enriching information of spectrum coverage, radiomap plays an important role in many wireless communication applications, such as resource allocation and network optimization. To enable real-time, distributed spectrum management, particularly in the scenarios with unstable and dynamic environments, the efficient transmission of spectrum coverage information for radiomaps from edge devices to the central server emerges as a critical problem. In this work, we propose an innovative physics-enhanced semantic communication framework tailored for efficient radiomap transmission based on generative learning models. Specifically, instead of bit-wise message passing, we only transmit the key \"semantics\" in radiomaps characterized by the radio propagation behavior and surrounding environments, where semantic compression schemes are utilized to reduce the communication overhead. Incorporating the novel concepts of Radio Depth Maps, the radiomaps are reconstructed from the delivered semantic information backboned on the conditional generative adversarial networks. Our framework is further extended to facilitate its implementation in the scenarios of multi-user edge computing, by integrating with federated learning for collaborative model training while preserving the data privacy. Experimental results show that our approach achieves high accuracy in radio coverage information recovery at ultra-high bandwidth efficiency, which has great potentials in many wireless-generated data transmission applications.","author":[{"family":"Zhou","given":"Yueling"},{"family":"Wijesinghe","given":"Achintha"},{"family":"Wang","given":"Yue"},{"family":"Zhang","given":"Songyang"},{"family":"Cai","given":"Zhipeng"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.10654","URL":"https://doi.org/10.48550/arxiv.2501.10654","source":"datacite"},{"id":"doi:10.48550/arxiv.2503.11151","type":"manuscript","title":"Enabling Weak Client Participation via On-device Knowledge Distillation in Heterogeneous Federated Learning","abstract":"Online Knowledge Distillation (KD) is recently highlighted to train large models in Federated Learning (FL) environments. Many existing studies adopt the logit ensemble method to perform KD on the server side. However, they often assume that unlabeled data collected at the edge is centralized on the server. Moreover, the logit ensemble method personalizes local models, which can degrade the quality of soft targets, especially when data is highly non-IID. To address these critical limitations,we propose a novel on-device KD-based heterogeneous FL method. Our approach leverages a small auxiliary model to learn from labeled local data. Subsequently, a subset of clients with strong system resources transfers knowledge to a large model through on-device KD using their unlabeled data. Our extensive experiments demonstrate that our on-device KD-based heterogeneous FL method effectively utilizes the system resources of all edge devices as well as the unlabeled data, resulting in higher accuracy compared to SOTA KD-based FL methods.","author":[{"family":"Lim","given":"Jihyun"},{"family":"Jo","given":"Junhyuk"},{"family":"Zhang","given":"Tuo"},{"family":"Lee","given":"Sunwoo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2503.11151","URL":"https://doi.org/10.48550/arxiv.2503.11151","source":"datacite"},{"id":"doi:10.5281/zenodo.21803858","type":"article-journal","title":"Federated Foundation Models and Multi-Agent Orchestration: Enabling Autonomous Decision Intelligence for Enterprise AI Systems and Adaptive Governance","abstract":"The rapid maturation of large language models (LLMs), retrieval-augmented generation (RAG), and multi-agent orchestration frameworks has catalyzed a new generation of autonomous decision-support systems capable of reasoning over heterogeneous enterprise data, invoking external tools, and coordinating multi-step workflows with minimal human supervision. This paper investigates the architecture, performance characteristics, and deployment challenges of federated, multi-agent foundation-model systems designed to support autonomous decision intelligence in financial, customer-service, and clinical-support environments. We examine how the combination of lightweight distilled small language models (SLMs), retrieval-grounded reasoning, and cross-organizational federated fine-tuning enables enterprises to deploy capable AI agents without centralizing sensitive data or incurring unsustainable inference cost. A novel three-tier reference architecture — comprising a Model Tier, an Agent Tier, and a Governance Tier — is proposed to unify context acquisition, distributed reasoning, multi-agent coordination, and cross-organizational policy compliance within a single framework, designated the Federated Orchestration for Reasoning, Governance and Execution (FORGE) framework. The Model Tier employs distilled small language models and semantic context-compression pipelines to achieve sub-300 ms local inference latency on commodity inference hardware. The Agent Tier hosts a multi-agent orchestration engine coordinated by a cost-aware task scheduler that dynamically routes sub-tasks between local SLMs and larger upstream foundation models based on task complexity and real-time budget constraints. The Governance Tier provides federated fine-tuning, policy synchronization, and cross-domain compliance auditing consistent with emerging AI-governance regulation. Experimental evaluations conducted on representative workloads — spanning financial risk analysis, customer-service automation, and clinical decision support — demonstrate that the proposed architecture achieves up to 52% reduction in end-to-end decision latency, 38% improvement in agent resource utilization, and 29% reduction in inference token cost compared to monolithic, single-model deployments. The framework sustains near-linear horizontal scalability up to 5,000 concurrent agent sessions, and mean time to service restoration following node failure is reduced to 4.3 seconds through integrated state replication and rapid failover protocols. A federated fine-tuning extension enables privacy-preserving model adaptation across heterogeneous organizational datasets, achieving accuracy within 4% of centralized training baselines while satisfying differential-privacy guarantees. Our findings indicate that a unified framework co-designing model efficiency, agent coordination, and governance constraints is essential for the next generation of trustworthy, autonomous enterprise AI systems.","author":[{"family":"Yuanyuan","given":"Wu"},{"family":"Wang","given":"Ruxing"},{"family":"Guo","given":"Lijuan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21803858","URL":"https://doi.org/10.5281/zenodo.21803858","source":"datacite"},{"id":"doi:10.5281/zenodo.21803859","type":"article-journal","title":"Federated Foundation Models and Multi-Agent Orchestration: Enabling Autonomous Decision Intelligence for Enterprise AI Systems and Adaptive Governance","abstract":"The rapid maturation of large language models (LLMs), retrieval-augmented generation (RAG), and multi-agent orchestration frameworks has catalyzed a new generation of autonomous decision-support systems capable of reasoning over heterogeneous enterprise data, invoking external tools, and coordinating multi-step workflows with minimal human supervision. This paper investigates the architecture, performance characteristics, and deployment challenges of federated, multi-agent foundation-model systems designed to support autonomous decision intelligence in financial, customer-service, and clinical-support environments. We examine how the combination of lightweight distilled small language models (SLMs), retrieval-grounded reasoning, and cross-organizational federated fine-tuning enables enterprises to deploy capable AI agents without centralizing sensitive data or incurring unsustainable inference cost. A novel three-tier reference architecture — comprising a Model Tier, an Agent Tier, and a Governance Tier — is proposed to unify context acquisition, distributed reasoning, multi-agent coordination, and cross-organizational policy compliance within a single framework, designated the Federated Orchestration for Reasoning, Governance and Execution (FORGE) framework. The Model Tier employs distilled small language models and semantic context-compression pipelines to achieve sub-300 ms local inference latency on commodity inference hardware. The Agent Tier hosts a multi-agent orchestration engine coordinated by a cost-aware task scheduler that dynamically routes sub-tasks between local SLMs and larger upstream foundation models based on task complexity and real-time budget constraints. The Governance Tier provides federated fine-tuning, policy synchronization, and cross-domain compliance auditing consistent with emerging AI-governance regulation. Experimental evaluations conducted on representative workloads — spanning financial risk analysis, customer-service automation, and clinical decision support — demonstrate that the proposed architecture achieves up to 52% reduction in end-to-end decision latency, 38% improvement in agent resource utilization, and 29% reduction in inference token cost compared to monolithic, single-model deployments. The framework sustains near-linear horizontal scalability up to 5,000 concurrent agent sessions, and mean time to service restoration following node failure is reduced to 4.3 seconds through integrated state replication and rapid failover protocols. A federated fine-tuning extension enables privacy-preserving model adaptation across heterogeneous organizational datasets, achieving accuracy within 4% of centralized training baselines while satisfying differential-privacy guarantees. Our findings indicate that a unified framework co-designing model efficiency, agent coordination, and governance constraints is essential for the next generation of trustworthy, autonomous enterprise AI systems.","author":[{"family":"Yuanyuan","given":"Wu"},{"family":"Wang","given":"Ruxing"},{"family":"Guo","given":"Lijuan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21803859","URL":"https://doi.org/10.5281/zenodo.21803859","source":"datacite"},{"id":"doi:10.5281/zenodo.20130383","type":"article-journal","title":"Edge Intelligence and 5G Networks: Enabling Smart IoT Systems for Real-Time Automation and Adaptive Control","abstract":"The rapid convergence of edge computing, fifth-generation (5G) wireless networks, and the Internet of Things (IoT) has catalyzed a new generation of intelligent, distributed automation systems capable of processing vast volumes of sensor data at or near the source. This paper investigates the architecture, performance characteristics, and deployment challenges of edge-native intelligence frameworks designed to support real-time IoT workloads in industrial, urban, and healthcare environments. We explore how the ultra-low latency and massive device connectivity provided by 5G networks enable edge computing nodes to execute time-critical inference and control tasks that were previously confined to centralized cloud data centers. A novel three-tier reference architecture — comprising a Device Tier, an Edge Tier, and an Orchestration Tier — is proposed to unify data acquisition, on-device inference, hierarchical coordination, and cloud-based analytics within a single framework, designated the Edge-Intelligence and Smart Automation (EISA) framework. The Device Tier employs lightweight TinyML models and event-driven data pipelines to achieve sub-10 ms local inference latency on resource-constrained microcontrollers. The Edge Tier hosts multi-model serving backends coordinated by an adaptive task scheduler that dynamically balances computation between edge nodes and upstream cloud resources based on real-time load estimates. The Orchestration Tier provides federated model training, global policy synchronization, and cross-domain resource governance compliant with emerging IoT data-sovereignty regulations. Experimental evaluations conducted on representative workloads — spanning predictive maintenance in industrial IoT, intelligent traffic management in smart cities, and patient vital-sign monitoring in connected healthcare — demonstrate that the proposed architecture achieves up to 47% reduction in end-to-end response latency, 41% improvement in edge node compute utilization, and 34% reduction in cellular backhaul bandwidth consumption compared to cloud-centric deployments. The framework sustains near-linear horizontal scalability up to 10,000 concurrent IoT endpoints, and mean time to service restoration following node failure is reduced to 6.1 seconds through integrated state replication and rapid failover protocols. A federated learning extension enables privacy-preserving model refinement across heterogeneous device populations, achieving convergence performance within 5% of centralized training baselines while satisfying differential-privacy guarantees. Our findings indicate that a unified framework co-designing edge hardware constraints, 5G network capabilities, and IoT application semantics is essential for the next generation of intelligent, autonomous cyber-physical systems.","author":[{"family":"Yuanyuan","given":"Wu"},{"family":"Wang","given":"Ruxing"},{"family":"Guo","given":"Lijuan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.20130383","URL":"https://doi.org/10.5281/zenodo.20130383","source":"datacite"},{"id":"doi:10.5281/zenodo.20130384","type":"article-journal","title":"Edge Intelligence and 5G Networks: Enabling Smart IoT Systems for Real-Time Automation and Adaptive Control","abstract":"The rapid convergence of edge computing, fifth-generation (5G) wireless networks, and the Internet of Things (IoT) has catalyzed a new generation of intelligent, distributed automation systems capable of processing vast volumes of sensor data at or near the source. This paper investigates the architecture, performance characteristics, and deployment challenges of edge-native intelligence frameworks designed to support real-time IoT workloads in industrial, urban, and healthcare environments. We explore how the ultra-low latency and massive device connectivity provided by 5G networks enable edge computing nodes to execute time-critical inference and control tasks that were previously confined to centralized cloud data centers. A novel three-tier reference architecture — comprising a Device Tier, an Edge Tier, and an Orchestration Tier — is proposed to unify data acquisition, on-device inference, hierarchical coordination, and cloud-based analytics within a single framework, designated the Edge-Intelligence and Smart Automation (EISA) framework. The Device Tier employs lightweight TinyML models and event-driven data pipelines to achieve sub-10 ms local inference latency on resource-constrained microcontrollers. The Edge Tier hosts multi-model serving backends coordinated by an adaptive task scheduler that dynamically balances computation between edge nodes and upstream cloud resources based on real-time load estimates. The Orchestration Tier provides federated model training, global policy synchronization, and cross-domain resource governance compliant with emerging IoT data-sovereignty regulations. Experimental evaluations conducted on representative workloads — spanning predictive maintenance in industrial IoT, intelligent traffic management in smart cities, and patient vital-sign monitoring in connected healthcare — demonstrate that the proposed architecture achieves up to 47% reduction in end-to-end response latency, 41% improvement in edge node compute utilization, and 34% reduction in cellular backhaul bandwidth consumption compared to cloud-centric deployments. The framework sustains near-linear horizontal scalability up to 10,000 concurrent IoT endpoints, and mean time to service restoration following node failure is reduced to 6.1 seconds through integrated state replication and rapid failover protocols. A federated learning extension enables privacy-preserving model refinement across heterogeneous device populations, achieving convergence performance within 5% of centralized training baselines while satisfying differential-privacy guarantees. Our findings indicate that a unified framework co-designing edge hardware constraints, 5G network capabilities, and IoT application semantics is essential for the next generation of intelligent, autonomous cyber-physical systems.","author":[{"family":"Yuanyuan","given":"Wu"},{"family":"Wang","given":"Ruxing"},{"family":"Guo","given":"Lijuan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.20130384","URL":"https://doi.org/10.5281/zenodo.20130384","source":"datacite"},{"id":"doi:10.5281/zenodo.20760280","type":"article-journal","title":"A DECENTRALIZED MACHINE LEARNING APPROACH FOR FRAUD DETECTION WITH BLOCKCHAIN-DRIVEN PRIVACY PROTECTION","abstract":"Modern digital ecosystems have the significant difficulty of detecting fraud while managing large, real-time data transfer and privacy preservation. This study presents an innovative architecture that combines blockchain technology with machine learning to provide safe, transparent, and privacy-conscious fraud detection. The solution utilizes federated learning and differential privacy methods to train machine learning models without revealing raw user data, while using blockchain's decentralized framework to guarantee data immutability and reliability. A dynamic incentive framework using smart contracts further motivates users to provide detection-ready, high-quality data. The suggested method promotes cooperation among entities, protects user data security, and achieves enhanced fraud detection accuracy via the integration of privacy-preserving computing and decentralized trust. Experimental assessments on both simulated and actual financial datasets illustrate the system's precision, robustness, and scalability in detecting intricate and evolving fraud patterns.","author":[{"family":"Koppula","given":"Anitha"},{"family":"Dr","given":"Chava"},{"family":"Dr","given":"Vunnava"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20760280","URL":"https://doi.org/10.5281/zenodo.20760280","source":"datacite"},{"id":"doi:10.5281/zenodo.20760281","type":"article-journal","title":"A DECENTRALIZED MACHINE LEARNING APPROACH FOR FRAUD DETECTION WITH BLOCKCHAIN-DRIVEN PRIVACY PROTECTION","abstract":"Modern digital ecosystems have the significant difficulty of detecting fraud while managing large, real-time data transfer and privacy preservation. This study presents an innovative architecture that combines blockchain technology with machine learning to provide safe, transparent, and privacy-conscious fraud detection. The solution utilizes federated learning and differential privacy methods to train machine learning models without revealing raw user data, while using blockchain's decentralized framework to guarantee data immutability and reliability. A dynamic incentive framework using smart contracts further motivates users to provide detection-ready, high-quality data. The suggested method promotes cooperation among entities, protects user data security, and achieves enhanced fraud detection accuracy via the integration of privacy-preserving computing and decentralized trust. Experimental assessments on both simulated and actual financial datasets illustrate the system's precision, robustness, and scalability in detecting intricate and evolving fraud patterns.","author":[{"family":"Koppula","given":"Anitha"},{"family":"Dr","given":"Chava"},{"family":"Dr","given":"Vunnava"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20760281","URL":"https://doi.org/10.5281/zenodo.20760281","source":"datacite"},{"id":"doi:10.5281/zenodo.20548461","type":"article-journal","title":"Code and trained models for \"SNR-Conditioned Residual CNN with Feature-wise Linear Modulation for OFDM Channel Estimation: Ablation Analysis and Federated Deployment","abstract":"Reproducibility code, trained model weights, and result files for the paper \"SNR-Conditioned Residual CNN with Feature-wise Linear Modulation for OFDM Channel Estimation: Ablation Analysis and Federated Deployment.\" Includes a self-contained PyTorch implementation of the SNR-conditioned residual 1-D CNN with FiLM for pilot-aided OFDM channel estimation, the Ye et al. fully-connected DNN baseline, the undersampled (L>P) experiment, trained model weights, result CSV files, and figure-generation scripts. The OFDM simulator (N=64, P=8 comb pilots, L=8 uniform PDP, QPSK) generates all channel realizations on the fly; no external dataset is required. See README.md for usage.","author":[{"family":"Abdulhameed","given":"Zainab"},{"family":"Hatem","given":"Haraa"},{"family":"Shehab","given":"Jinan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20548461","URL":"https://doi.org/10.5281/zenodo.20548461","source":"datacite"},{"id":"doi:10.5281/zenodo.20548462","type":"article-journal","title":"Code and trained models for \"SNR-Conditioned Residual CNN with Feature-wise Linear Modulation for OFDM Channel Estimation: Ablation Analysis and Federated Deployment","abstract":"Reproducibility code, trained model weights, and result files for the paper \"SNR-Conditioned Residual CNN with Feature-wise Linear Modulation for OFDM Channel Estimation: Ablation Analysis and Federated Deployment.\" Includes a self-contained PyTorch implementation of the SNR-conditioned residual 1-D CNN with FiLM for pilot-aided OFDM channel estimation, the Ye et al. fully-connected DNN baseline, the undersampled (L>P) experiment, trained model weights, result CSV files, and figure-generation scripts. The OFDM simulator (N=64, P=8 comb pilots, L=8 uniform PDP, QPSK) generates all channel realizations on the fly; no external dataset is required. See README.md for usage.","author":[{"family":"Abdulhameed","given":"Zainab"},{"family":"Hatem","given":"Haraa"},{"family":"Shehab","given":"Jinan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20548462","URL":"https://doi.org/10.5281/zenodo.20548462","source":"datacite"},{"id":"doi:10.5281/zenodo.21531664","type":"article-journal","title":"A Robust Self-Supervised Contrastive Learning Framework  for Secure and Reliable Financial Transaction Mining","abstract":"ABSTRACT The rapid expansion of digital payment ecosystems, online banking platforms, and fintech services has led to an unprecedented growth in financial transaction data. While this digital transformation enhances accessibility and efficiency, it also increases vulnerability to sophisticated fraud schemes, money laundering activities, and cyber-financial crimes. Conventional supervised machine learning models for financial transaction mining rely heavily on large volumes of labeled data, which are often scarce, imbalanced, costly to annotate, and subject to strict privacy regulations. These limitations significantly hinder the scalability, adaptability, and reliability of traditional fraud detection systems. To address these challenges, this paper proposes a robust self-supervised contrastive learning framework for secure and reliable financial transaction mining. The proposed approach leverages vast amounts of unlabeled transaction data to learn meaningful and discriminative latent representations without requiring manual annotation. By employing contrastive learning principles, the framework maximizes agreement between augmented views of similar transactions while simultaneously minimizing similarity between dissimilar transaction pairs. Domain-specific augmentation strategies such as feature masking, temporal perturbation, and controlled noise injection are introduced to ensure robustness against transaction variability and adversarial manipulation. The learned representations are subsequently finetuned using a lightweight supervised classifier with a limited labeled dataset, significantly reducing dependency on annotated data while improving fraud detection accuracy. The framework enhances generalization capability, reduces false positive rates, and demonstrates resilience against class imbalance and evolving fraud patterns (concept drift). Furthermore, the architecture supports privacy-preserving extensions such as federated learning, enabling secure distributed training across financial institutions without sharing sensitive raw data. Experimental evaluation on benchmark financial transaction datasets demonstrates that the proposed selfsupervised contrastive framework outperforms conventional supervised and semi-supervised models in terms of ROC-AUC, F1-score, recall, and robustness under noisy conditions. The results confirm that contrastive selfsupervised learning provides a scalable, reliable, and secure artificial intelligence solution for next-generation financial transaction analytics. The proposed methodology contributes to advancing intelligent financial systems by combining representation learning, security awareness, and data efficiency, thereby paving the way for adaptive and trustworthy financial transaction mining in dynamic real-world environments.","author":[{"family":"Nagapriya","given":"Mrs"},{"family":"Princerani","given":"Dr"},{"family":"Haq","given":"Dr"},{"family":"Chowdhury","given":"Mohammad"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21531664","URL":"https://doi.org/10.5281/zenodo.21531664","source":"datacite"},{"id":"doi:10.5281/zenodo.21531665","type":"article-journal","title":"A Robust Self-Supervised Contrastive Learning Framework  for Secure and Reliable Financial Transaction Mining","abstract":"ABSTRACT The rapid expansion of digital payment ecosystems, online banking platforms, and fintech services has led to an unprecedented growth in financial transaction data. While this digital transformation enhances accessibility and efficiency, it also increases vulnerability to sophisticated fraud schemes, money laundering activities, and cyber-financial crimes. Conventional supervised machine learning models for financial transaction mining rely heavily on large volumes of labeled data, which are often scarce, imbalanced, costly to annotate, and subject to strict privacy regulations. These limitations significantly hinder the scalability, adaptability, and reliability of traditional fraud detection systems. To address these challenges, this paper proposes a robust self-supervised contrastive learning framework for secure and reliable financial transaction mining. The proposed approach leverages vast amounts of unlabeled transaction data to learn meaningful and discriminative latent representations without requiring manual annotation. By employing contrastive learning principles, the framework maximizes agreement between augmented views of similar transactions while simultaneously minimizing similarity between dissimilar transaction pairs. Domain-specific augmentation strategies such as feature masking, temporal perturbation, and controlled noise injection are introduced to ensure robustness against transaction variability and adversarial manipulation. The learned representations are subsequently finetuned using a lightweight supervised classifier with a limited labeled dataset, significantly reducing dependency on annotated data while improving fraud detection accuracy. The framework enhances generalization capability, reduces false positive rates, and demonstrates resilience against class imbalance and evolving fraud patterns (concept drift). Furthermore, the architecture supports privacy-preserving extensions such as federated learning, enabling secure distributed training across financial institutions without sharing sensitive raw data. Experimental evaluation on benchmark financial transaction datasets demonstrates that the proposed selfsupervised contrastive framework outperforms conventional supervised and semi-supervised models in terms of ROC-AUC, F1-score, recall, and robustness under noisy conditions. The results confirm that contrastive selfsupervised learning provides a scalable, reliable, and secure artificial intelligence solution for next-generation financial transaction analytics. The proposed methodology contributes to advancing intelligent financial systems by combining representation learning, security awareness, and data efficiency, thereby paving the way for adaptive and trustworthy financial transaction mining in dynamic real-world environments.","author":[{"family":"Nagapriya","given":"Mrs"},{"family":"Princerani","given":"Dr"},{"family":"Haq","given":"Dr"},{"family":"Chowdhury","given":"Mohammad"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21531665","URL":"https://doi.org/10.5281/zenodo.21531665","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.27191","type":"manuscript","title":"Personalized Federated Learning for Tensor Regression","abstract":"The growing availability of tensor-valued data across multiple institutions creates opportunities for collaborative analysis, but also raises challenges related to data privacy, high dimensionality, and client heterogeneity. This paper introduces a personalized federated tensor regression framework that addresses all three simultaneously. Each client's coefficient tensor is decomposed into a globally shared low-Tucker-rank component and a locally sparse deviation, estimated via a two-stage privacy-preserving procedure. We establish finite-sample upper bounds and minimax lower bounds that quantify the privacy-accuracy trade-off, and prove the consistency of the supporting initialization and rank-selection steps. Simulation studies confirm that the federated approach improves estimation and prediction over purely local methods, especially when per-client data are scarce, and an MRI-based ADHD study illustrates its strong performance under real privacy constraints.","author":[{"family":"Chen","given":"Kejun"},{"family":"Wei","given":"Xianqi"},{"family":"Zhu","given":"Qianqian"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.27191","URL":"https://doi.org/10.48550/arxiv.2608.27191","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.27108","type":"manuscript","title":"SecureDrive-FL: Joint Differential Privacy and Gradient-Aware Selective Homomorphic Encryption for Federated Driver Monitoring","abstract":"Federated Learning (FL) enables privacy-aware distributed training, yet gradient updates remain exploitable: Man-in-the-Middle (MitM) interception exposes updates in transit, while model poisoning corrupts global convergence. We first introduce GASHE (Gradient-Aware Selective Homomorphic Encryption), a novel selective encryption strategy that dynamically identifies and encrypts only the gradient components exceeding a DP-calibrated sensitivity threshold, rather than encrypting all parameters uniformly as in static layer-based or full-parameter CKKS schemes. Building on GASHE, we introduce SecureDrive-FL, a federated driver monitoring framework that couples DP-SGD with GASHE to create the first closed-loop DP+HE privacy pipeline: DP-SGD calibration parameters directly derive the GASHE encryption mask, unifying training-time privacy and communication-time confidentiality. Evaluated on a ten-class distracted driver classification task under non-IID federated splits, SecureDrive-FL matches DP-SGD alone's poisoning resistance (73.6% vs. 74.0% accuracy, 3.9% Attack Success Rate for both) while additionally withstanding MitM interception, where DP-SGD alone collapses to near-random accuracy (78.2% vs. 10.4%), all under only approx. 8--10% additional runtime overhead relative to DP-SGD alone---under DP-SGD noise injection with per-round privacy parameter epsilon_0=4.","author":[{"family":"Gül","given":"Baran"},{"family":"Tunuguntla","given":"Hanuma"},{"family":"Naik","given":"Anjana"},{"family":"Potekar","given":"Abhishek"},{"family":"Jazdi","given":"Nasser"},{"family":"Weyrich","given":"Michael"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.27108","URL":"https://doi.org/10.48550/arxiv.2608.27108","source":"datacite"},{"id":"doi:10.5281/zenodo.20829078","type":"article-journal","title":"The Future of Artificial Intelligence in Breast Cancer Diagnosis and Prognosis: Trends, Innovations, and Breakthroughs","abstract":"Breast cancer remains one of the most significant health challenges worldwide, requiring continuous advancements in early detection, diagnosis, and prognosis. With the rapid evolution of artificial intelligence (AI), the landscape of breast cancer management has been transformed, offering unprecedented improvements in accuracy, efficiency, and personalized care. The Future of Artificial Intelligence in Breast Cancer Diagnosis and Prognosis: Trends, Innovations, and Breakthroughs explores the cutting-edge role of AI in revolutionizing breast cancer detection, treatment, and patient outcomes.This book delves into the fundamental principles of AI, machine learning, and deep learning as applied to oncology. It examines the integration of AI with medical imaging techniques such as mammography, ultrasound, and MRI, along with its applications in histopathological analysis, liquid biopsy, and biomarker discovery. Special emphasis is placed on AI-driven prognostic models, radiomics, and precision medicine, highlighting how AI enhances risk stratification, treatment response prediction, and personalized therapy selection.A crucial aspect of this book is its discussion on the ethical, regulatory, and practical challenges of AI adoption in healthcare. Issues such as bias in AI algorithms, data privacy, security concerns, and compliance with global regulatory frameworks are explored in depth. Additionally, the book sheds light on the future trajectory of AI in breast cancer research, including emerging technologies such as federated learning, quantum computing, and AI-driven clinical trials.Written by Dr. Ashok Kumar Dogra, PhD in Biochemistry, this book serves as a comprehensive resource for oncologists, radiologists, medical researchers, data scientists, and healthcare policymakers. By bridging the gap between AI innovation and clinical application, it aims to equip professionals with the knowledge to leverage AI for improved breast cancer care.As we stand at the forefront of a technological revolution in oncology, this book seeks to inspire further advancements, ensuring AI continues to drive transformative breakthroughs in breast cancer diagnosis and prognosis. I hope that this work will contribute to the ongoing efforts to enhance patient outcomes and shape the future of AI-driven healthcare.","author":[{"family":"Dogra","given":"Ashok"},{"family":"Prakash","given":"Archana"},{"family":"Gupta","given":"Meenu"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.20829078","URL":"https://doi.org/10.5281/zenodo.20829078","source":"datacite"},{"id":"doi:10.5281/zenodo.20829079","type":"article-journal","title":"The Future of Artificial Intelligence in Breast Cancer Diagnosis and Prognosis: Trends, Innovations, and Breakthroughs","abstract":"Breast cancer remains one of the most significant health challenges worldwide, requiring continuous advancements in early detection, diagnosis, and prognosis. With the rapid evolution of artificial intelligence (AI), the landscape of breast cancer management has been transformed, offering unprecedented improvements in accuracy, efficiency, and personalized care. The Future of Artificial Intelligence in Breast Cancer Diagnosis and Prognosis: Trends, Innovations, and Breakthroughs explores the cutting-edge role of AI in revolutionizing breast cancer detection, treatment, and patient outcomes.This book delves into the fundamental principles of AI, machine learning, and deep learning as applied to oncology. It examines the integration of AI with medical imaging techniques such as mammography, ultrasound, and MRI, along with its applications in histopathological analysis, liquid biopsy, and biomarker discovery. Special emphasis is placed on AI-driven prognostic models, radiomics, and precision medicine, highlighting how AI enhances risk stratification, treatment response prediction, and personalized therapy selection.A crucial aspect of this book is its discussion on the ethical, regulatory, and practical challenges of AI adoption in healthcare. Issues such as bias in AI algorithms, data privacy, security concerns, and compliance with global regulatory frameworks are explored in depth. Additionally, the book sheds light on the future trajectory of AI in breast cancer research, including emerging technologies such as federated learning, quantum computing, and AI-driven clinical trials.Written by Dr. Ashok Kumar Dogra, PhD in Biochemistry, this book serves as a comprehensive resource for oncologists, radiologists, medical researchers, data scientists, and healthcare policymakers. By bridging the gap between AI innovation and clinical application, it aims to equip professionals with the knowledge to leverage AI for improved breast cancer care.As we stand at the forefront of a technological revolution in oncology, this book seeks to inspire further advancements, ensuring AI continues to drive transformative breakthroughs in breast cancer diagnosis and prognosis. I hope that this work will contribute to the ongoing efforts to enhance patient outcomes and shape the future of AI-driven healthcare.","author":[{"family":"Dogra","given":"Ashok"},{"family":"Prakash","given":"Archana"},{"family":"Gupta","given":"Meenu"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.20829079","URL":"https://doi.org/10.5281/zenodo.20829079","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.26433","type":"manuscript","title":"FedCMAPSS: A Benchmark for Federated Learning in Remaining Useful Life Estimation","abstract":"Data-driven prognostics and health management has emerged as a key enabler for Industry 4.0, yet the development of robust remaining useful life (RUL) estimation models is often limited by the scarcity of run-to-failure data. While federated learning offers a promising paradigm to collaboratively train predictive models without sharing sensor data, research efforts have operated so far in the absence of a common evaluation framework. To address this gap, this paper introduces FedCMAPSS, a benchmark for federated RUL estimation based on the commonly-used NASA C-MAPSS dataset. We define a set of five standardized tasks designed to simulate real-world industrial challenges, ranging from ideal IID settings to extreme statistical heterogeneity, and conduct a systematic evaluation of state-of-the-art federated optimization algorithms across multiple neural architectures. By establishing reproducible baselines and making the source code and data splits publicly available, this work aims to provide a standard foundation for developing and comparing federated predictive maintenance solutions.","author":[{"family":"Sorrenti","given":"Amelia"},{"family":"Pennisi","given":"Matteo"},{"family":"Spampinato","given":"Concetto"},{"family":"Palazzo","given":"Simone"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.26433","URL":"https://doi.org/10.48550/arxiv.2608.26433","source":"datacite"},{"id":"doi:10.5281/zenodo.21604094","type":"article-journal","title":"Challenges and Future Directions in Breast Cancer Segmentation: A Research Perspective","abstract":"Breast cancer is one of the most aggressive and widespread illnesses afflicting women across the globe. Fast and precise segmentation techniques are essential for early detection, diagnosis, and treatment planning. This paper reviews comprehensively segmentation methods used in breast cancer detection operations, including traditional methods of thresholding and edge detection and eliciting advanced deep learning techniques such as Convolutional Neural Networks (CNN), U-Net, Generative Adversarial Networks (GANs), and Transformer-based models. The review stresses the merits of hybrid approaches combining many segmentation paradigms for better accuracy and robustness. This new round of research underlines the recent progress in segmentation with the help of attention mechanisms, precise mapping, and multimodal imaging integration. Yet, problems such as dataset-level issues, generalization issues, computational complexity, and lack of explainability still remain. Future research will design lightweight architectures, explainable AI, federated learning, and advanced multimodal data fusion techniques. This paper highlights the dynamic nature of breast cancer segmentation and marks that without continued innovation, achieving clinically relevant and accurate automated segmentation systems will remain a challenge.","author":[{"family":"Iyer","given":"Swathi"},{"family":"Tudilkar","given":"Salwa"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21604094","URL":"https://doi.org/10.5281/zenodo.21604094","source":"datacite"},{"id":"doi:10.5281/zenodo.21604095","type":"article-journal","title":"Challenges and Future Directions in Breast Cancer Segmentation: A Research Perspective","abstract":"Breast cancer is one of the most aggressive and widespread illnesses afflicting women across the globe. Fast and precise segmentation techniques are essential for early detection, diagnosis, and treatment planning. This paper reviews comprehensively segmentation methods used in breast cancer detection operations, including traditional methods of thresholding and edge detection and eliciting advanced deep learning techniques such as Convolutional Neural Networks (CNN), U-Net, Generative Adversarial Networks (GANs), and Transformer-based models. The review stresses the merits of hybrid approaches combining many segmentation paradigms for better accuracy and robustness. This new round of research underlines the recent progress in segmentation with the help of attention mechanisms, precise mapping, and multimodal imaging integration. Yet, problems such as dataset-level issues, generalization issues, computational complexity, and lack of explainability still remain. Future research will design lightweight architectures, explainable AI, federated learning, and advanced multimodal data fusion techniques. This paper highlights the dynamic nature of breast cancer segmentation and marks that without continued innovation, achieving clinically relevant and accurate automated segmentation systems will remain a challenge.","author":[{"family":"Iyer","given":"Swathi"},{"family":"Tudilkar","given":"Salwa"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21604095","URL":"https://doi.org/10.5281/zenodo.21604095","source":"datacite"},{"id":"doi:10.5281/zenodo.21608193","type":"article-journal","title":"Interpreting Federated Learning (FL) Models on Edge Devices by Enhancing Model Explainability with Computational Geometry and Advanced Database Architectures","abstract":"Federated learning (FL) on edge devices has emerged as a promising approach for decentralized model training, enabling data privacy and efficiency in distributed networks. However, the complexity of these models presents significant challenges in terms of transparency and interpretability, which are critical for trust and accountability in real-world applications. This paper explores the integration of explainable AI techniques to enhance model interpretability within federated learning systems. By incorporating computational geometry, we aim to optimize model structure and decision-making processes, providing clearer insights into how models generate predictions. Additionally, we examine the role of advanced database architectures in managing the complexity of federated learning models on edge devices, ensuring efficient data handling and storage. Together, these approaches contribute to a more transparent, efficient, and scalable framework for federated learning on edge networks, addressing key challenges in both model explainability and performance optimization. This review highlights recent advancements and suggests future directions for research at the intersection of federated learning (FL), edge computing, explainability, and computational techniques.","author":[{"family":"Enyejo","given":"Lawrence"},{"family":"Adewoye","given":"Michael"},{"family":"Ugochukwu","given":"Uchenna"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21608193","URL":"https://doi.org/10.5281/zenodo.21608193","source":"datacite"},{"id":"doi:10.5281/zenodo.21608194","type":"article-journal","title":"Interpreting Federated Learning (FL) Models on Edge Devices by Enhancing Model Explainability with Computational Geometry and Advanced Database Architectures","abstract":"Federated learning (FL) on edge devices has emerged as a promising approach for decentralized model training, enabling data privacy and efficiency in distributed networks. However, the complexity of these models presents significant challenges in terms of transparency and interpretability, which are critical for trust and accountability in real-world applications. This paper explores the integration of explainable AI techniques to enhance model interpretability within federated learning systems. By incorporating computational geometry, we aim to optimize model structure and decision-making processes, providing clearer insights into how models generate predictions. Additionally, we examine the role of advanced database architectures in managing the complexity of federated learning models on edge devices, ensuring efficient data handling and storage. Together, these approaches contribute to a more transparent, efficient, and scalable framework for federated learning on edge networks, addressing key challenges in both model explainability and performance optimization. This review highlights recent advancements and suggests future directions for research at the intersection of federated learning (FL), edge computing, explainability, and computational techniques.","author":[{"family":"Enyejo","given":"Lawrence"},{"family":"Adewoye","given":"Michael"},{"family":"Ugochukwu","given":"Uchenna"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21608194","URL":"https://doi.org/10.5281/zenodo.21608194","source":"datacite"},{"id":"doi:10.5281/zenodo.21579767","type":"article-journal","title":"Deep Learning Models for Predicting and Mitigating Environmental Impact of Industrial Processes in Real-Time","abstract":"Industrial processes contribute significantly to environmental degradation through emissions, waste, and resource depletion. The need for real-time monitoring and mitigation strategies has led to the adoption of deep learning (DL) models for predictive analytics and automated decision-making. This study explores the application of deep learning techniques in predicting and mitigating the environmental impact of industrial activities. We review state-of-the-art deep learning architectures, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory (LSTM) networks, and transformers, in processing large-scale environmental data. These models analyze real-time sensor data, satellite imagery, and industrial parameters to forecast pollution levels, detect anomalies, and optimize industrial operations for sustainability. Key advancements in deep learning, such as hybrid architectures integrating deep reinforcement learning (DRL) and generative adversarial networks (GANs), enhance predictive accuracy and robustness in environmental monitoring systems. Transfer learning and federated learning approaches facilitate scalable and adaptive solutions across diverse industrial sectors. The study highlights the role of DL in early detection of air and water pollution, energy consumption optimization, and emission control through predictive maintenance and process adjustments. Moreover, integrating explainable artificial intelligence (XAI) ensures model interpretability, fostering trust among policymakers and industry stakeholders. Challenges in deploying deep learning models include data heterogeneity, computational complexity, and model interpretability. To address these issues, we discuss techniques such as data augmentation, adversarial training, and edge AI implementation for real-time processing. Ethical and regulatory considerations surrounding AI-driven environmental monitoring are also examined to ensure compliance with sustainability standards. This research underscores the transformative potential of deep learning in industrial sustainability, emphasizing its role in real-time decision support systems. Future directions involve integrating quantum computing and neuromorphic computing for enhanced model efficiency and expanding interdisciplinary collaborations for AI-driven environmental governance. By leveraging deep learning for predictive environmental impact assessment, industries can transition toward greener and more efficient operational frameworks.","author":[{"family":"Ojadi","given":"Jessica"},{"family":"Owulade","given":"Olumide"},{"family":"Odionu","given":"Chinekwu"},{"family":"Onukwulu","given":"Ekene"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21579767","URL":"https://doi.org/10.5281/zenodo.21579767","source":"datacite"},{"id":"doi:10.5281/zenodo.21579768","type":"article-journal","title":"Deep Learning Models for Predicting and Mitigating Environmental Impact of Industrial Processes in Real-Time","abstract":"Industrial processes contribute significantly to environmental degradation through emissions, waste, and resource depletion. The need for real-time monitoring and mitigation strategies has led to the adoption of deep learning (DL) models for predictive analytics and automated decision-making. This study explores the application of deep learning techniques in predicting and mitigating the environmental impact of industrial activities. We review state-of-the-art deep learning architectures, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory (LSTM) networks, and transformers, in processing large-scale environmental data. These models analyze real-time sensor data, satellite imagery, and industrial parameters to forecast pollution levels, detect anomalies, and optimize industrial operations for sustainability. Key advancements in deep learning, such as hybrid architectures integrating deep reinforcement learning (DRL) and generative adversarial networks (GANs), enhance predictive accuracy and robustness in environmental monitoring systems. Transfer learning and federated learning approaches facilitate scalable and adaptive solutions across diverse industrial sectors. The study highlights the role of DL in early detection of air and water pollution, energy consumption optimization, and emission control through predictive maintenance and process adjustments. Moreover, integrating explainable artificial intelligence (XAI) ensures model interpretability, fostering trust among policymakers and industry stakeholders. Challenges in deploying deep learning models include data heterogeneity, computational complexity, and model interpretability. To address these issues, we discuss techniques such as data augmentation, adversarial training, and edge AI implementation for real-time processing. Ethical and regulatory considerations surrounding AI-driven environmental monitoring are also examined to ensure compliance with sustainability standards. This research underscores the transformative potential of deep learning in industrial sustainability, emphasizing its role in real-time decision support systems. Future directions involve integrating quantum computing and neuromorphic computing for enhanced model efficiency and expanding interdisciplinary collaborations for AI-driven environmental governance. By leveraging deep learning for predictive environmental impact assessment, industries can transition toward greener and more efficient operational frameworks.","author":[{"family":"Ojadi","given":"Jessica"},{"family":"Owulade","given":"Olumide"},{"family":"Odionu","given":"Chinekwu"},{"family":"Onukwulu","given":"Ekene"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21579768","URL":"https://doi.org/10.5281/zenodo.21579768","source":"datacite"},{"id":"doi:10.5281/zenodo.21578631","type":"article-journal","title":"Federated Meta-Learning for Few-Shot Image Classification with Personalized Model Adaptation: A Comprehensive Review","abstract":"Federated meta-learning represents a paradigm shift in machine learning that combines the privacy-preserving benefits of federated learning with the rapid adaptation capabilities of meta-learning for few-shot image classification tasks. This comprehensive literature review critically evaluates the latest advancements in federated meta-learning approaches, with particular emphasis on personalized model adaptation strategies for image classification scenarios with limited labelled data examining classical federated learning methods, advanced meta-learning techniques, and their integration for few-shot learning applications. The review highlights significant challenges including non-independent and identically distributed data, communication efficiency, privacy preservation, and model personalization across diverse client populations. Advanced techniques utilizing deep learning architectures, optimization-based meta-learning, and adaptive aggregation mechanisms have demonstrated promising results in enhancing classification accuracy while maintaining privacy constraints. This review serves as a comprehensive guide for researchers and practitioners, providing thorough understanding of state-of-the-art federated meta-learning techniques, their implications for few-shot image classification, and potential avenues for further development.","author":[{"family":"Thenmozhi","given":"R"},{"family":"Santhalakshmi","given":"M"},{"family":"Shanthakumar","given":"M"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21578631","URL":"https://doi.org/10.5281/zenodo.21578631","source":"datacite"},{"id":"doi:10.5281/zenodo.21578632","type":"article-journal","title":"Federated Meta-Learning for Few-Shot Image Classification with Personalized Model Adaptation: A Comprehensive Review","abstract":"Federated meta-learning represents a paradigm shift in machine learning that combines the privacy-preserving benefits of federated learning with the rapid adaptation capabilities of meta-learning for few-shot image classification tasks. This comprehensive literature review critically evaluates the latest advancements in federated meta-learning approaches, with particular emphasis on personalized model adaptation strategies for image classification scenarios with limited labelled data examining classical federated learning methods, advanced meta-learning techniques, and their integration for few-shot learning applications. The review highlights significant challenges including non-independent and identically distributed data, communication efficiency, privacy preservation, and model personalization across diverse client populations. Advanced techniques utilizing deep learning architectures, optimization-based meta-learning, and adaptive aggregation mechanisms have demonstrated promising results in enhancing classification accuracy while maintaining privacy constraints. This review serves as a comprehensive guide for researchers and practitioners, providing thorough understanding of state-of-the-art federated meta-learning techniques, their implications for few-shot image classification, and potential avenues for further development.","author":[{"family":"Thenmozhi","given":"R"},{"family":"Santhalakshmi","given":"M"},{"family":"Shanthakumar","given":"M"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21578632","URL":"https://doi.org/10.5281/zenodo.21578632","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.17529","type":"manuscript","title":"CryptDough: A Unified Analytics Engine for Secure Multiparty Computation","abstract":"We present CryptDough, a unified analytics engine for secure multiparty computation (MPC). CryptDough enables multiple distrusting parties to jointly execute a data analysis pipeline on their private inputs and learn nothing beyond the result (e.g., aggregate statistics). Unlike existing MPC solutions that support a single threat model or workload type, CryptDough provides built-in support for cross-domain analytics (relational, time series, ML inference) under various threat models, all within the same system runtime. CryptDough contributes (i) a hierarchical system design that facilitates modularity and extensibility through progressive lowering of abstractions, and (ii) the concept of virtual vectors that enable users to write single-threaded code across all layers of the software stack, while pushing the complexity of communication, parallelization, and memory management down to the execution engine. We show that CryptDough generalizes the functionality of state-of-the-art MPC systems and remains competitive on the analytics they support, often outperforming them by more than $2\\times$.","author":[{"family":"Faisal","given":"Muhammad"},{"family":"Lanz","given":"Alessandra"},{"family":"Buxbaum","given":"Sam"},{"family":"Godel","given":"Adam"},{"family":"Kalavri","given":"Vasiliki"},{"family":"Varia","given":"Mayank"},{"family":"Liagouris","given":"John"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.17529","URL":"https://doi.org/10.48550/arxiv.2608.17529","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.17131","type":"manuscript","title":"Blind Transpiler: An open-source library for universally blind and homomorphic quantum computations","abstract":"Blind quantum computation is a cryptographic primitive that allows a limited-capability client to delegate its complex computation to a remote server without revealing its data and/or computation. This branch of quantum cryptography has been bifurcated into two distinct primitives, quantum homomorphic encryption (concerning the security of only data) and universal blind quantum computation (concerning the security of data and the computing algorithm). These primitives have immense applicability in problems like secure cloud computing, secure quantum variational algorithms, quantum federated learning, and secure multiparty computation. However, no software tools exist for the rapid prototyping of such protocols, hindering the academic interrogation for potential applications. In this paper, we describe the development of the first such library for transpiling circuits written in Qiskit to its blind counterpart, which can then be delegated in a client-server architecture without revealing the client's data and/or computation. The proposed library is designed in modular and reusable component layers, enabling easier scalability to newer BQC primitives and robustness against changes in underlying primitives. We show the implementation of these primitives to a blind variational quantum classifier for the IRIS dataset.","author":[{"family":"Joshi","given":"Mohit"},{"family":"Mishra","given":"Manoj"},{"family":"Karthikeyan","given":"S"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.17131","URL":"https://doi.org/10.48550/arxiv.2607.17131","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.12557","type":"manuscript","title":"New bounds on private simultaneous quantum message passing","abstract":"In the private simultaneous message (PSM) setting, $k$ players obtain inputs $x_i\\in\\{0,1\\}^n$ and then each send messages to a referee, who should learn $f(x_1,...,x_k)$ but no other information about $(x_1,...,x_k)$. The PSM setting was introduced as a minimal model for secure multiparty computation and has connections to Boolean function complexity. In the quantum setting, PSM has been related to non-local quantum computation (NLQC). The communication and correlation cost of implementing PSM remains poorly understood. Here, we give new upper and lower bounds on the (quantum) PSM model. For lower bounds, we show: 1) Nečiporuk's measure lower bounds the entanglement required for $k$-player quantum PSM with perfect correctness. This leads to quadratic lower bounds for explicit functions. 2) The rank of the communication matrix of $f(x_1,x_2)$ lower bounds 2-player quantum PSM with perfect privacy but imperfect correctness. This implies a previously unknown lower bound on classical PSM with imperfect correctness. When allowing quantum communication and shared entanglement, these are the first lower bounds on quantum PSM that make use of the privacy condition. For upper bounds, we show: 1) Letting $s$ be the size of a quantum circuit computing $f$, $d_f$ be the circuit depth, $k$ the number of players, $n$ the number of bits received by each player, and $ε$ a correctness parameter, we obtain $\\mathsf{PSM}_k^*(f) \\leq (kn +s) \\cdot \\log^{O(d_f)}(s/ε)$. 2) The square of the Fourier 1 norm of $f$, $\\Vert \\hat{f}\\Vert_1^2$, upper bounds the classical PSM complexity, $\\mathsf{PSM}(f)\\leq O(\\Vert \\hat{f} \\Vert^2_1)$. In proving the first upper bound, we generalize existing $T$-depth based techniques for NLQC from $2$ to $k\\geq 2$ parties, and consider cases where the Clifford layers are restricted to having small light cones.","author":[{"family":"Girish","given":"Uma"},{"family":"May","given":"Alex"},{"family":"Parham","given":"Natalie"},{"family":"Yuen","given":"Henry"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.12557","URL":"https://doi.org/10.48550/arxiv.2606.12557","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.27456","type":"manuscript","title":"Federated Generation of Synthetic RNA-seq Data","abstract":"Access to genomic data is highly regulated due to its sensitive nature. While safeguards are essential, cumbersome data access processes pose a significant barrier to the development of AI methods for genomics. Synthetic data generation can mitigate this tension by enabling broader data sharing without exposing sensitive information. Synthetic genomic data are produced by training generative models on real data and subsequently sampling artificial data that preserves relevant statistics while limiting disclosures about the underlying individuals. In some settings, a single data holder may have sufficient data to train such generative models; however, in many applications data must be combined across multiple sites to achieve adequate scale. This need arises, e.g., in rare disease studies, where individual hospitals typically hold data for only a small number of patients. The solution we present in this paper enables multiple data holders to jointly train a synthetic data generator without revealing their raw data. Our approach combines secure multiparty computation (MPC) to ensure input privacy, so that no party ever discloses its data in unencrypted form, with differential privacy (DP) to provide output privacy by mitigating information leakage from the released synthetic data. We empirically demonstrate the effectiveness of the proposed method by generating high-utility synthetic datasets from multiple real RNA-seq cohorts in federated settings, showing that our approach enables privacy-preserving data synthesis even when data are distributed across institutions.","author":[{"family":"Filienko","given":"Daniil"},{"family":"De Cock","given":"Martine"},{"family":"Pentyala","given":"Sikha"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.27456","URL":"https://doi.org/10.48550/arxiv.2604.27456","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.07009","type":"manuscript","title":"Fast Bounded-Independence Functions and Their Duals","abstract":"We continue the study of {\\em fast} functions, computable by linear-size circuits, that share useful properties of random functions. Motivated by cryptographic applications, we generalize and improve on previous results in this area, obtaining the following results: - For any constant $t$, we construct a fast $t$-wise independent hash function with algebraic degree $\\log_2 t$ (over $\\mathbb F_2$), simultaneously optimizing both asymptotic circuit size and degree. - We simplify and improve a recent construction (ITCS 2026) of a family of fast codes with fast duals, both meeting the Gilbert-Varshamov bound. Unlike the previous construction, our construction has negligible failure probability, can accommodate general fields and rates, supports a systematic encoding, and admits fast universal encoders. - We strengthen the above to support stronger random-like properties, such as optimal combinatorial list-decoding. This is achieved by constructing, for any constant $t$, a family of fast linear functions that map any $t$ linearly independent inputs to uniform and statistically independent outputs. Prior to our work, this was only known for $t=1$. We demonstrate the usefulness of the above results to cryptography. This includes the first nontrivial protocols for perfectly secure multiparty computation whose circuit complexity scales linearly with the number of parties, as well as protocols for computing encrypted matrix-vector products with optimal asymptotic circuit complexity.","author":[{"family":"Brehm","given":"Martijn"},{"family":"Ishai","given":"Yuval"},{"family":"Resch","given":"Nicolas"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.07009","URL":"https://doi.org/10.48550/arxiv.2606.07009","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.00169","type":"manuscript","title":"Beyond Latency: A System-Level Characterization of MPC and FHE for PPML","abstract":"Privacy protection has become an increasing concern in modern machine learning applications. Privacy-preserving machine learning (PPML) has attracted growing research attention, with approaches such as secure multiparty computation (MPC) and fully homomorphic encryption (FHE) being actively explored. However, existing evaluations of these approaches have frequently been done on a narrow, fragmented setup and only focused on a specific performance metric, such as the online inference latency of a specific batch size. From the existing reports, it is hard to compare different approaches, especially when considering other metrics like energy/cost or broader system setups (various hyperparameters, offline overheads, future hardware/network configurations, etc.). We present a unified characterization of three popular approaches -- two variants of MPC based on arithmetic/binary sharing conversion and function secret sharing, and FHE -- on their performance and cost in performing privacy-preserving inference on multiple CNN and Transformer models. We study a range of LAN and WAN environments, model sizes, batch sizes, and input sequence lengths. We evaluate not only the performance but also the energy consumption and monetary cost of deploying under a realistic scenario, taking into account their offline and online computation/communication overheads. We provide empirical guidance for selecting, optimizing, and deploying these privacy-preserving compute paradigms, and outline how evolving hardware and network trends are likely to shift trade-offs between the two MPC schemes and FHE. This work provides system-level insights for researchers and practitioners who seek to understand or accelerate PPML workloads.","author":[{"family":"Huang","given":"Pengzhi"},{"family":"Maeng","given":"Kiwan"},{"family":"Suh","given":"GE"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.00169","URL":"https://doi.org/10.48550/arxiv.2604.00169","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.23031","type":"manuscript","title":"FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning","abstract":"Federated Learning (FL) enables distributed clients to collaboratively train models without sharing raw data, making it promising for leveraging massive devices in communication networks. In distillation-based FL, each client applies its local model on an unlabeled public dataset, and shares only prediction results with the server. While heterogeneous local data introduces label distribution skew, thus biasing client models toward majority classes and leading to potentially inaccurate predictions. The lack of ground-truth labels in the public dataset hampers the server's ability to calibrate predictions, which ultimately degrades overall performance. To address this, we propose FedCC, a simple and effective algorithm for mitigating client misclassification. Instead of being forced to classify and risking error propagation, clients are allowed to tag ambiguous samples as 'unknown'. This additional class, together with calibrated pseudo-labels on the public data, balances confidence in majority classes against uncertainty in under-represented ones. Extensive experiments demonstrate that FedCC significantly outperforms existing methods, especially under severe label skew. In the extreme scenario where each client holds samples from only one of ten classes, FedCC achieves 67.3% accuracy, while baselines collapse to near-random results.","author":[{"family":"Ye","given":"Wenxuan"},{"family":"Ayan","given":"Onur"},{"family":"An","given":"Xueli"},{"family":"Carle","given":"Georg"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.23031","URL":"https://doi.org/10.48550/arxiv.2608.23031","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.25496","type":"manuscript","title":"FedQoS: Federated QoS-Risk Learning for Heterogeneous Indoor-Outdoor Access Selection","abstract":"Reliable access selection in dynamic and heterogeneous indoor-outdoor environments is challenging because instantaneous radio measurements alone cannot capture future QoS degradation caused by mobility, blockage, traffic load, and resource competition. This paper proposes FedQoS, a federated QoS-risk learning framework for predicting the future reliability of candidate access links and supporting access-node selection without centralizing user-level network data. In FedQoS, each access node locally learns from its observed network logs, including radio, traffic, load, and service-context features, while a global QoS-risk predictor is trained through federated aggregation. The learned model estimates the probability of QoS failure for each candidate link, and the controller uses these risk scores to select reliable access nodes under dynamic network conditions. To evaluate the framework, we construct physics-based synthetic indoor-outdoor wireless datasets using the Sionna framework, covering normal traffic, mobility, event-driven congestion, and non-IID client observations. Simulation results show that learning-based access selection substantially reduces the QoS-failure rate compared with signal-based and historical-QoS heuristic methods. FedQoS achieves near-centralized predictive performance and provides clear reliability gains under mild non-IID data while remaining competitive under the more challenging severe non-IID condition. These results demonstrate the potential of federated QoS-risk learning for reliable, data-local access selection in dynamic wireless environments.","author":[{"family":"Van Thieu","given":"Nguyen"},{"family":"Nguyen","given":"Ti"},{"family":"Aouedi","given":"Ons"},{"family":"Huruy","given":"Zerihun"},{"family":"Ha","given":"Vu"},{"family":"Chatzinotas","given":"Symeon"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.25496","URL":"https://doi.org/10.48550/arxiv.2608.25496","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.25133","type":"manuscript","title":"Rethinking the Transferable Adversarial Attacks and Robust Defense in Federated Learning","abstract":"The development of federated learning (FL) techniques has helped improve the privacy preservation of users' data and extended the applications of machine learning models. However, the involvement of a large number of users in FL also creates open opportunities for different adversaries, such as poisoning attacks, Byzantine attacks, and adversarial example attacks. Yet, recent research has disclosed that existing poisoning attacks and Byzantine attacks can not achieve satisfactory penetration in realistic FL scenarios caused by strong assumptions, \\textit{e.g.,} client selection rate, and the ratio of malicious attackers. In this paper, the transferability of adversarial examples among different client models is analyzed to understand the relation between adversarial examples and clients' data distribution. Moreover, to mitigate the attacks of transferable adversarial examples, we design a defense mechanism stemming from the transferability of model robustness by adversarial training. As a result, through theoretical analysis of transferability, we gain insights into adversarial examples and the vulnerability of federated learning systems. Our proposed adversarial attack and defense methods are evaluated via real-life datasets in various settings to show their performance over the existing state-of-the-art methods.","author":[{"family":"Xiong","given":"Zuobin"},{"family":"Mukherjee","given":"Deval"},{"family":"Cho","given":"Homook"},{"family":"Li","given":"Wei"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.25133","URL":"https://doi.org/10.48550/arxiv.2608.25133","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.06637","type":"manuscript","title":"Bypassing Krum: Selection-Aware Backdoor Attacks in Federated Learning","abstract":"Robust aggregation methods are widely used in federated learning to mitigate the impact of adversarial client behavior. Distance-based aggregation rules, such as Krum and Multi-Krum, select updates that are closest to the majority under the assumption that benign updates form a compact cluster. However, these methods rely on geometric properties that can be exploited by adaptive adversaries. We introduce the Krum-Proxy attack, a selection-aware backdoor injection strategy that consistently bypasses Byzantine-robust aggregation. Rather than relying on naive scaling or constraining, our method actively optimizes malicious updates to infiltrate the dense core of the benign distribution. The proposed method constructs adversarial updates that are not only similar to benign updates but are also optimized to lie in regions of the update space that are favored during aggregation. This is achieved through a two-stage optimization procedure that separates task-specific attack objectives from geometry-aware refinement, using a nearest-neighbor proxy, stochastic reference modeling, and anchor-guided alignment. To maintain stealth, we introduce a projection mechanism that constrains adversarial updates within realistic norm and variance bounds. Experiments on standard federated learning benchmarks show that Krum-Proxy achieves higher attack success while preserving clean accuracy, highlighting the vulnerability of distance-based aggregation to selection-aware adversaries.","author":[{"family":"Subramanian","given":"Srinivasan"},{"family":"Khan","given":"Md"},{"family":"Islam","given":"Kazi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.06637","URL":"https://doi.org/10.48550/arxiv.2608.06637","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.00364","type":"manuscript","title":"Understanding Federated Learning Through the Lens of Mechanism Design: The Role of Data Heterogeneity","abstract":"Federated learning (FL) requires effective incentive mechanisms to motivate data sharing and prevent strategic free-riding. Recent FL mechanisms such as the Shapley value mechanism M^Shap guarantee reciprocal fairness for agents. However, a complete analysis of how such mechanisms impact social optimality and individual rationality under realistic, standalone outside options remains unknown. In this paper, we address this gap by adapting the classical Externality mechanism M^E to the federated learning setting. We conduct a comparison of M^Shap and M^E across three dimensions: social optimality, individual rationality, and fairness/reciprocity. First, we establish that M^Shap generally does not maximize social welfare because its marginal incentives drive agents to over-contribute resources, while M^E maximizes social welfare by design. Second, we evaluate participation incentives through the individual rationality gap when considering agents' outside options as standalone training on their own data. We find that both mechanisms ensure individual rationality in homogeneous settings. We further show that under mild conditions, M^E maintains this guarantee under agent heterogeneity, whereas M^Shap does not. Third, we demonstrate that while M^Shap maintains perfect reciprocity by design, M^E generally does not, and only ensures that individual benefits match Shapley contributions at symmetric equilibria under homogeneity, as it sacrifices individual fairness to maximize collective welfare under heterogeneity. Empirical simulations validate our theoretical findings and illustrate a tradeoff between reciprocal fairness and social efficiency.","author":[{"family":"Alkarmi","given":"Lina"},{"family":"Chen","given":"Po"},{"family":"Liu","given":"Mingyan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.00364","URL":"https://doi.org/10.48550/arxiv.2608.00364","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.07209","type":"manuscript","title":"Continual Learning With Participation Privacy: An Auditable Buffering-Aggregation Recipe","abstract":"Modern federated and streaming learning systems often release intermediate models, so privacy must hold for the full trajectory under adaptive interaction. Motivated by participation privacy, we study single-edit neighboring user streams, where one insertion/deletion shifts all subsequent updates and defeats standard Hamming-neighbor continual-release analyses. We give an auditable modular recipe. A randomized buffering wrapper emits bins of size $[U,2U]$, reducing single-edit streams to a Hamming-style per-bin update stream with explicit backlog/delay guarantees, where $U$ is calibrated by the privacy parameters $(\\varepsilon,δ)$. We then prove a certification theorem identifying when a non-adaptive Hamming-neighbor DP proof for a continual primitive lifts to adaptive inputs: the primitive must use fresh per-round randomness and have a stable one-round privacy profile under common adaptive context. Together, these ingredients yield trajectory-level $(\\varepsilon,δ)$-DP for single-edit streams using standard primitives (e.g., tree prefix sums), with an explicit privacy--latency link via $U$.","author":[{"family":"Chan","given":"THH"},{"family":"Shi","given":"Elaine"},{"family":"Zhao","given":"Mengshi"},{"family":"Zhou","given":"Mingxun"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.07209","URL":"https://doi.org/10.48550/arxiv.2607.07209","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.21539","type":"manuscript","title":"Federated Continual Learning as a Distributed Drift-Plus-Penalty Control Problem","abstract":"Federated Continual Learning (FCL) is fundamental to real-world distributed learning systems, requiring models to adapt to sequential, non-IID data across clients while mitigating catastrophic forgetting and client drift. Existing approaches formulate continual learning (CL) as a sequence of per-task optimization problems, applied locally at each client and coupled through aggregation, using heuristic mechanisms such as replay, regularization, or projection-based constraints. However, forgetting in FCL is inherently a long-term, distributed phenomenon, arising from the interaction of temporal task evolution and cross-client heterogeneity, which is not explicitly regulated. In this work, we cast FCL as a stochastic control problem and propose Federated Queue-regulated Continual Learning (FedQCL), a framework based on Lyapunov drift-plus-penalty (DPP) optimization. FedQCL introduces virtual queues to track the accumulation of forgetting across tasks and clients, enabling explicit control of the stability-plasticity trade-off. By optimizing a DPP objective, the method jointly improves current-task performance while the queue-based formulation provides an interpretable and tunable mechanism to balance adaptation and retention through a single parameter, without requiring gradient projection or additional communication overhead. Empirical evaluations on standard benchmarks, including Split-CIFAR-10, Split-CIFAR-100, and Split-TinyImageNet, demonstrate that FedQCL outperforms state-of-the-art baselines with respect to accuracy while significantly reducing forgetting under heterogeneous data distributions.","author":[{"family":"Shah","given":"Nazreen"},{"family":"Somireddy","given":"Naveen"},{"family":"Shaban","given":"Zubair"},{"family":"Prasad","given":"Ranjitha"},{"family":"Bharath","given":"BN"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.21539","URL":"https://doi.org/10.48550/arxiv.2608.21539","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.21096","type":"manuscript","title":"FlatLand: Personalized Graph Federated Learning via Tailored Lorentz Space","abstract":"Federated learning enables privacy-preserving collaborative training, but highly heterogeneous client data remain challenging, especially in graph federated learning where clients possess structurally diverse graphs. Existing personalized federated learning (PFL) methods ignore the intrinsic geometric properties of diverse graph structures. We propose FlatLand, a novel personalized federated learning method that embeds different clients' data in tailored Lorentz space of hyperbolic geometry. Our key insight is that hyperbolic geometry naturally accommodates the intrinsic negative curvature prevalent in real-world graphs, while the time-like dimension in Lorentz space provides a principled way to encode client-specific heterogeneity. We develop a parameter decoupling strategy that separates heterogeneous information (captured in time-like parameters) from common knowledge (preserved in space-like parameters), enabling direct aggregation without requiring client similarity estimation and extra calculation modules. Empirical results on diverse federated graph learning tasks demonstrate that FlatLand achieves superior performance, particularly in low-dimensional settings.","author":[{"family":"Liu","given":"Jiahong"},{"family":"Fu","given":"Xinyu"},{"family":"Yang","given":"Menglin"},{"family":"Zhang","given":"Weixi"},{"family":"Ying","given":"Rex"},{"family":"King","given":"Irwin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.21096","URL":"https://doi.org/10.48550/arxiv.2608.21096","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.20518","type":"manuscript","title":"FL-MAESTRO: Multi-Agent LLM Orchestration for Resource-Constrained Federated Learning","abstract":"In Federated Learning (FL), the communication topology is a runtime variable rather than a fixed design choice, since links and edge devices drop in and out during training. Each round, the server must commit three coupled decisions, namely the communication topology, per-client resource allocation, and the aggregation rule for combining local updates. Recent agentic systems have begun bringing large language models (LLM) into FL, but the existing line of work either operates at setup time or handles a single runtime dimension such as client selection. We propose FL-MAESTRO, a multi-agent orchestrator that makes the joint runtime FL decision directly through three specialist LLM agents, one per decision dimension. A coordinator combines their analyses into a single decision, and a non-LLM feasibility check confirms it before the round executes. Because the orchestrator consumes the server's predicted-failure list, it withholds clients whose updates would never be aggregated, which removes the dominant source of wasted round energy in classical FL on volatile edge networks. Because client state is read as natural-text profiles, the same orchestrator extends to heterogeneous device classes without per-class energy models. On a non-IID CIFAR-10 benchmark, FL-MAESTRO matches the accuracy of the strongest energy-aware baseline while cutting wasted round energy from over a third to near zero. Code is available at https://github.com/denoslab/FL-MAESTRO.","author":[{"family":"Wu","given":"Jiajun"},{"family":"Wang","given":"Zirui"},{"family":"Zhou","given":"Jiayu"},{"family":"Ye","given":"Qiang"},{"family":"Drew","given":"Steve"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.20518","URL":"https://doi.org/10.48550/arxiv.2608.20518","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.19650","type":"manuscript","title":"Enhancing Privacy in Federated Learning via Dual Obfuscation of Gradients and Training Images","abstract":"Federated learning enables collaborative model training while keeping data locally at each client; however, recent studies have shown that training data can be reconstructed from shared model updates. To address this issue, this paper proposes a dual obfuscation method that enhances robustness against image restoration attacks by jointly obfuscating updated information and training images. The proposed method combines a robustness enhancement technique based on random binary weights, which randomly sets a portion of gradient elements to zero, with an image encryption technique. These techniques provide complementary protection by reducing the amount of original gradient information available to an attacker and the visual interpretability of reconstructed images, respectively. Furthermore, the image encryption technique allows independent keys to be used for each client and each image, avoiding explicit key sharing. Experimental results on an image classification task using a Vision Transformer (ViT) show that the proposed method reduces the visual information recovered by Attention Privacy Leakage (APRIL) under the evaluated settings without causing additional degradation in classification performance beyond that caused by image encryption. Although the proposed combination does not provide an absolute security guarantee, the results demonstrate the potential benefit of combining gradient modification and image encryption for privacy-enhanced federated learning.","author":[{"family":"Itabashi","given":"Yuki"},{"family":"Sawada","given":"Hiroto"},{"family":"Hirose","given":"Mare"},{"family":"Imaizumi","given":"Shoko"},{"family":"Kiya","given":"Hitoshi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.19650","URL":"https://doi.org/10.48550/arxiv.2608.19650","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.19649","type":"manuscript","title":"Differential Privacy in Feature Reconstruction Aided Federated Learning for Agent's Semantic Communication Model Update","abstract":"This paper proposes a differentially private federated learning (FL) framework built upon an FL algorithm with semantic feature reconstruction (FedSFR) for training semantic communication modules for image transmission. By allowing clients with unfavorable uplink capacity to transmit low-dimensional semantic feature vectors extracted from locally trained joint source-channel coding (JSCC) encoders, FedSFR enhances communication efficiency and training stability under heterogeneous wireless conditions. To protect client privacy, we incorporate the oneshot Laplace mechanism and theoretically demonstrate that feature-based transmission achieves strictly stronger differential privacy (DP) guarantees than gradient-based transmission under an identical communication budget. In addition, a model selection mechanism is introduced to alleviate performance degradation caused by privacy-preserving perturbations. Experimental results on multiple datasets show that the proposed DP-aided FedSFR outperforms DP-enabled FedAvg in training stability and image reconstruction quality in heterogeneous wireless systems.","author":[{"family":"Huh","given":"Yoon"},{"family":"Kim","given":"Bumjun"},{"family":"Choi","given":"Wan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.19649","URL":"https://doi.org/10.48550/arxiv.2608.19649","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.06692","type":"manuscript","title":"HeteRo-Select: Informativeness as the Participation Driver in Heterogeneous Federated Learning","abstract":"Federated learning systems typically allocate gradient compression by link speed. This is sensible when bandwidth and data informativeness align. However, under non-IID data, these signals often decorrelate or invert. A bandwidth-driven allocator then risks compressing the most informative gradients hardest. We propose HeteRo-Select, a framework that replaces bandwidth with a per-client informativeness score as the primary driver of compression. The score jointly governs three decisions per round: client selection, compression ratio, and server aggregation weight, with bandwidth retained only as a hard ceiling. Score-proportional selection provably reduces the effective heterogeneity of the chosen subset; score-proportional compression provably lowers aggregate top-$k$ error at fixed traffic. Under the exact FedCG simulation protocol, HeteRo-Select delivers a $1.78\\times$ speedup and an $18.2\\%$ reduction in traffic on CIFAR-10. The same configuration, unchanged, scales from a $7{,}850$-parameter logistic regression to an $11.27$M-parameter ResNet-18, hitting the accuracy target on three of four benchmarks. When bandwidth and informativeness are deliberately anti-correlated, the method still achieves the target accuracy with less traffic than the normal-bandwidth run.","author":[{"family":"Masud","given":"Md"},{"family":"Jahin","given":"Md"},{"family":"Hasan","given":"Mahmud"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.06692","URL":"https://doi.org/10.48550/arxiv.2508.06692","source":"datacite"},{"id":"doi:10.5281/zenodo.18643524","type":"article-journal","title":"Privacy-Preserving Machine Learning Technique Using One-Way Hashing for Enhancing  Data Security in the Nigerian Healthcare Ecosystem","abstract":"Safeguarding healthcare data in Nigeria remains a pressing challenge, complicated by fragmented infrastructure, limited resources, and evolving regulatory frameworks. This paper presents a conceptual analysis of one-way hashing as a lightweight cryptographic technique for pseudonymizing patient identifiers within federated learning pipelines. By situating hashing in contrast to heavier cryptographic methods such as homomorphic encryption and secure multiparty computation, the study highlights its relative efficiency, scalability, and compliance with the Nigeria Data Protection Regulation (NDPR, 2019) and the Nigeria Data Protection Act (NDPA, 2023). Through analytical benchmarking and illustrative scenarios, hashing is shown to offer a pragmatic balance between privacy preservation and operational feasibility in resource-constrained healthcare environments. The paper concludes that one-way hashing provides a viable conceptual pathway for operationalizing privacy-preserving machine learning in Nigerian healthcare systems, while laying the foundation for future empirical validation.","author":[{"family":"Damang"},{"family":"Fs"},{"family":"Aimufua","given":"And"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18643524","URL":"https://doi.org/10.5281/zenodo.18643524","source":"datacite"},{"id":"doi:10.5281/zenodo.18643523","type":"article-journal","title":"Privacy-Preserving Machine Learning Technique Using One-Way Hashing for Enhancing  Data Security in the Nigerian Healthcare Ecosystem","abstract":"Safeguarding healthcare data in Nigeria remains a pressing challenge, complicated by fragmented infrastructure, limited resources, and evolving regulatory frameworks. This paper presents a conceptual analysis of one-way hashing as a lightweight cryptographic technique for pseudonymizing patient identifiers within federated learning pipelines. By situating hashing in contrast to heavier cryptographic methods such as homomorphic encryption and secure multiparty computation, the study highlights its relative efficiency, scalability, and compliance with the Nigeria Data Protection Regulation (NDPR, 2019) and the Nigeria Data Protection Act (NDPA, 2023). Through analytical benchmarking and illustrative scenarios, hashing is shown to offer a pragmatic balance between privacy preservation and operational feasibility in resource-constrained healthcare environments. The paper concludes that one-way hashing provides a viable conceptual pathway for operationalizing privacy-preserving machine learning in Nigerian healthcare systems, while laying the foundation for future empirical validation.","author":[{"family":"Damang"},{"family":"Fs"},{"family":"Aimufua","given":"And"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18643523","URL":"https://doi.org/10.5281/zenodo.18643523","source":"datacite"},{"id":"doi:10.5281/zenodo.18643521","type":"article-journal","title":"Privacy-Preserving Machine Learning Technique Using One-Way Hashing for Enhancing  Data Security in the Nigerian Healthcare Ecosystem","abstract":"Safeguarding healthcare data in Nigeria remains a pressing challenge, complicated by fragmented infrastructure, limited resources, and evolving regulatory frameworks. This paper presents a conceptual analysis of one-way hashing as a lightweight cryptographic technique for pseudonymizing patient identifiers within federated learning pipelines. By situating hashing in contrast to heavier cryptographic methods such as homomorphic encryption and secure multiparty computation, the study highlights its relative efficiency, scalability, and compliance with the Nigeria Data Protection Regulation (NDPR, 2019) and the Nigeria Data Protection Act (NDPA, 2023). Through analytical benchmarking and illustrative scenarios, hashing is shown to offer a pragmatic balance between privacy preservation and operational feasibility in resource-constrained healthcare environments. The paper concludes that one-way hashing provides a viable conceptual pathway for operationalizing privacy-preserving machine learning in Nigerian healthcare systems, while laying the foundation for future empirical validation.","author":[{"family":"Damang"},{"family":"Fs"},{"family":"Aimufua","given":"And"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18643521","URL":"https://doi.org/10.5281/zenodo.18643521","source":"datacite"},{"id":"doi:10.5281/zenodo.18643522","type":"article-journal","title":"Privacy-Preserving Machine Learning Technique Using One-Way Hashing for Enhancing  Data Security in the Nigerian Healthcare Ecosystem","abstract":"Safeguarding healthcare data in Nigeria remains a pressing challenge, complicated by fragmented infrastructure, limited resources, and evolving regulatory frameworks. This paper presents a conceptual analysis of one-way hashing as a lightweight cryptographic technique for pseudonymizing patient identifiers within federated learning pipelines. By situating hashing in contrast to heavier cryptographic methods such as homomorphic encryption and secure multiparty computation, the study highlights its relative efficiency, scalability, and compliance with the Nigeria Data Protection Regulation (NDPR, 2019) and the Nigeria Data Protection Act (NDPA, 2023). Through analytical benchmarking and illustrative scenarios, hashing is shown to offer a pragmatic balance between privacy preservation and operational feasibility in resource-constrained healthcare environments. The paper concludes that one-way hashing provides a viable conceptual pathway for operationalizing privacy-preserving machine learning in Nigerian healthcare systems, while laying the foundation for future empirical validation.","author":[{"family":"Damang"},{"family":"Fs"},{"family":"Aimufua","given":"And"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18643522","URL":"https://doi.org/10.5281/zenodo.18643522","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.23483","type":"manuscript","title":"Towards a Functionally Complete and Parameterizable TFHE Processor","abstract":"Fully homomorphic encryption allows the evaluation of arbitrary functions on encrypted data. It can be leveraged to secure outsourced and multiparty computation. TFHE is a fast torus-based fully homomorphic encryption scheme that allows both linear operations, as well as the evaluation of arbitrary non-linear functions. It currently provides the fastest bootstrapping operation performance of any other FHE scheme. Despite its fast performance, TFHE suffers from a considerably higher computational overhead for the evaluation of homomorphic circuits. Computations in the encrypted domain are orders of magnitude slower than their unencrypted equivalents. This bottleneck hinders the widespread adoption of (T)FHE for the protection of sensitive data. While state-of-the-art implementations focused on accelerating and outsourcing single operations, their scalability and practicality are constrained by high memory bandwidth costs. In order to overcome this, we propose an FPGA-based hardware accelerator for the evaluation of homomorphic circuits. Specifically, we design a functionally complete TFHE processor for FPGA hardware capable of processing instructions on the data completely on the FPGA. In order to achieve a higher throughput from our TFHE processor, we implement an improved programmable bootstrapping module, which outperforms the current state-of-the-art by 240% to 480% more bootstrappings per second. Our efficient, compact, and scalable design lays the foundation for implementing complete FPGA-based TFHE processor architectures.","author":[{"family":"Häusler","given":"Valentin"},{"family":"Ott","given":"Gabriel"},{"family":"Jayasena","given":"Aruna"},{"family":"Peter","given":"Andreas"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.23483","URL":"https://doi.org/10.48550/arxiv.2510.23483","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.21525","type":"manuscript","title":"Enhancing Distributed Authorization With Lagrange Interpolation And Attribute-Based Encryption","abstract":"In todays security landscape, every user wants to access large amounts of data with confidentiality and authorization. To maintain confidentiality, various researchers have proposed several techniques. However, to access secure data, researchers use access control lists to grant authentication and provide authorization. The above several steps will increase the server's computation overhead and response time. To cope with these two problems, we proposed multiparty execution on the server. In this paper, we introduce two different approaches. The first approach is encryption, utilizing the Involution Function Based Stream Cipher to encrypt the file data. The second approach is key distribution, using the Shamir secret sharing scheme to divide and distribute the symmetric key to every user. The decryption process required key reconstruction, which used second order Lagrange interpolation to reconstruct the secret keys from the hidden points. The process will reduce the server's computational overhead. The results are evaluated based on the encryption and decryption time, throughput, computational overhead, and security analysis. In the future, the proposed mechanism will be used to share large-scale, secure data within the organization.","author":[{"family":"Sinha","given":"Keshav"},{"family":"Sumitra"},{"family":"Kumari","given":"Richa"},{"family":"Bhardwaj","given":"Akashdeep"},{"family":"Rahman","given":"Shawon"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.21525","URL":"https://doi.org/10.48550/arxiv.2512.21525","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.20688","type":"manuscript","title":"Guard-GBDT: Efficient Privacy-Preserving Approximated GBDT Training on Vertical Dataset","abstract":"In light of increasing privacy concerns and stringent legal regulations, using secure multiparty computation (MPC) to enable collaborative GBDT model training among multiple data owners has garnered significant attention. Despite this, existing MPC-based GBDT frameworks face efficiency challenges due to high communication costs and the computation burden of non-linear operations, such as division and sigmoid calculations. In this work, we introduce Guard-GBDT, an innovative framework tailored for efficient and privacy-preserving GBDT training on vertical datasets. Guard-GBDT bypasses MPC-unfriendly division and sigmoid functions by using more streamlined approximations and reduces communication overhead by compressing the messages exchanged during gradient aggregation. We implement a prototype of Guard-GBDT and extensively evaluate its performance and accuracy on various real-world datasets. The results show that Guard-GBDT outperforms state-of-the-art HEP-XGB (CIKM'21) and SiGBDT (ASIA CCS'24) by up to $2.71\\times$ and $12.21 \\times$ on LAN network and up to $2.7\\times$ and $8.2\\times$ on WAN network. Guard-GBDT also achieves comparable accuracy with SiGBDT and plaintext XGBoost (better than HEP-XGB ), which exhibits a deviation of $\\pm1\\%$ to $\\pm2\\%$ only. Our implementation code is provided at https://github.com/XidianNSS/Guard-GBDT.git.","author":[{"family":"Song","given":"Anxiao"},{"family":"Cui","given":"Shujie"},{"family":"Bai","given":"Jianli"},{"family":"Cheng","given":"Ke"},{"family":"Shen","given":"Yulong"},{"family":"Russello","given":"Giovanni"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.20688","URL":"https://doi.org/10.48550/arxiv.2507.20688","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.03973","type":"manuscript","title":"Performance of Practical Quantum Oblivious Key Distribution","abstract":"Motivated by the applications of secure multiparty computation as a privacy-protecting data analysis tool, and identifying oblivious transfer as one of its main practical enablers, we propose a practical realization of randomized quantum oblivious transfer. By using only symmetric cryptography primitives to implement commitments, we construct computationally-secure randomized oblivious transfer without the need for public-key cryptography or assumptions imposing limitations on the adversarial devices. We show that the protocol is secure under an indistinguishability-based notion of security and demonstrate an experimental implementation to test its real-world performance. Its security and performance are then compared to both quantum and classical alternatives, showing potential advantages over existing solutions based on the noisy storage model and public-key cryptography.","author":[{"family":"Lemus","given":"Mariano"},{"family":"Schiansky","given":"Peter"},{"family":"Goulão","given":"Manuel"},{"family":"Bozzio","given":"Mathieu"},{"family":"Elkouss","given":"David"},{"family":"Paunković","given":"Nikola"},{"family":"Mateus","given":"Paulo"},{"family":"Walther","given":"Philip"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.03973","URL":"https://doi.org/10.48550/arxiv.2501.03973","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.20101","type":"manuscript","title":"Secure Multi-Key Homomorphic Encryption with Application to Privacy-Preserving Federated Learning","abstract":"Multi-Key Homomorphic Encryption (MKHE), proposed by Lopez-Alt et al. (STOC 2012), allows for performing arithmetic computations directly on ciphertexts encrypted under distinct keys. Subsequent works by Chen and Dai et al. (CCS 2019) and Kim and Song et al. (CCS 2023) extended this concept by proposing multi-key BFV/CKKS variants, referred to as the CDKS scheme. These variants incorporate asymptotically optimal techniques to facilitate secure computation across multiple data providers. In this paper, we identify a critical security vulnerability in the CDKS scheme when applied to multiparty secure computation tasks, such as privacy-preserving federated learning (PPFL). In particular, we show that CDKS may inadvertently leak plaintext information from one party to others. To mitigate this issue, we propose a new scheme, SMHE (Secure Multi-Key Homomorphic Encryption), which incorporates a novel masking mechanism into the multi-key BFV and CKKS frameworks to ensure that plaintexts remain confidential throughout the computation. We implement a PPFL application using SMHE and demonstrate that it provides significantly improved security with only a modest overhead in homomorphic evaluation. For instance, our PPFL model based on multi-key CKKS incurs less than a 2\\times runtime and communication traffic increase compared to the CDKS-based PPFL model. The code is publicly available at https://github.com/JiahuiWu2022/SMHE.git.","author":[{"family":"Wu","given":"Jiahui"},{"family":"Sun","given":"Tiecheng"},{"family":"Luo","given":"Fucai"},{"family":"Wang","given":"Haiyan"},{"family":"Zhang","given":"Weizhe"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.20101","URL":"https://doi.org/10.48550/arxiv.2506.20101","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.07814","type":"manuscript","title":"Utilizing Model-Free Reinforcement Learning for Optimizing Secure Multi-Party Computation Protocols","abstract":"In this manuscript, we explore the application of model-free reinforcement learning in optimizing secure multiparty computation (SMPC) protocols. SMPC is a crucial tool for performing computations on private data without the need to disclose it, holding significant importance in various domains, including information security and privacy. However, the efficiency of current protocols is often suboptimal due to computational and communicational complexities. Our proposed approach leverages model-free reinforcement learning algorithms to enhance the performance of these protocols. We have designed a reinforcement learning model capable of dynamically learning and adapting optimal strategies for secure computations. Our experimental results demonstrate that employing this method leads to a substantial reduction in execution time and communication costs of the protocols. These achievements highlight the high potential of reinforcement learning in improving the efficiency of secure multiparty computation protocols, providing an effective solution to the existing challenges in this field.","author":[{"family":"Sayyadi","given":"Javad"},{"family":"Nangir","given":"Mahdi"},{"family":"Feghhi","given":"Mahmood"},{"family":"Sayyadi","given":"Hamid"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.07814","URL":"https://doi.org/10.48550/arxiv.2510.07814","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.24761","type":"manuscript","title":"Federated Sharing and Continuous Improvement of Medical Device Knowledge Artifacts: A Conceptual Model","abstract":"Healthcare organisations use digital systems to exchange information from clinical cases. Medical centres with digital production facilities create device designs during care. These designs and production records often remain at the site that made them. Other sites may struggle to find a suitable design or learn what happened when staff used it. Mobile medical centres may also lose access when they work away from hospital systems. This paper proposes an artifact-centred model for a federated exchange infrastructure that lets hospitals and mobile medical centres share and improve medical knowledge while controlling their own records and decisions. An integrative literature review screened 910 records and mapped 240 publications across six questions. We read 72 publications in detail to trace the path from local use to a decision about shared knowledge. The review found no common process that links a record of local use to a decision about changing the knowledge shared with later users. The model keeps case data at each site and records which artifact version informed each use. Federation lets sites share reviewed versions and return records from use for review. Mobile units can receive artifacts before deployment and record their use while offline. After reconnecting, they can exchange these records with other sites. Point-of-care manufacturing shows how a medical device knowledge artifact connects a design to production records while the care site controls product release. Federation could form a governed learning network where one site's experience improves medical knowledge artifacts used elsewhere while authority remains local.","author":[{"family":"Mariscal-Melgar","given":"JC"},{"family":"Buxbaum-Conradi","given":"Sonja"},{"family":"Wenzelmann","given":"Victoria"},{"family":"Redlich","given":"Tobias"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.24761","URL":"https://doi.org/10.48550/arxiv.2608.24761","source":"datacite"},{"id":"doi:10.5281/zenodo.20962002","type":"article-journal","title":"Intelligent Power Infrastructure: A Structured Examination of Artificial Intelligence Techniques, Operational Challenges, and Evolutionary Pathways in Next-Generation Smart Grid Systems","abstract":"Abstract Global momentum toward intelligent power infrastructure is accelerating as electricity demand grows, renewable energy penetration deepens, and sustainability imperatives intensify. Artificial intelligence (AI) technologies—encompassing machine learning, deep learning, reinforcement learning, expert systems, fuzzy logic, and hybrid frameworks—have become indispensable enablers of next-generation grid operations. These approaches support intelligent monitoring, accurate load and renewable-energy forecasting, autonomous fault detection, demand-response orchestration, and optimal distributed energy resource integration. Notwithstanding these advantages, practical deployment faces persistent barriers including cyber security vulnerabilities, data-privacy constraints, infrastructure investment costs, device interoperability deficits, and insufficient regulatory frameworks. This study presents a structured original review of AI-enabled smart grid architectures and operational paradigms, systematically evaluates capabilities and trade-offs of principal AI methods, analyzes prevailing challenges alongside viable mitigation strategies, quantifies documented operational benefits, and maps prospective technological trajectories—including edge AI, digital twins, explainable AI, block chain, and federated learning—toward autonomous grid operation. Findings indicate that strategic AI adoption is critical to achieving resilient, efficient, and low-carbon power systems for the future.","author":[{"family":"Gopinath","given":"K"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20962002","URL":"https://doi.org/10.5281/zenodo.20962002","source":"datacite"},{"id":"doi:10.5281/zenodo.20962003","type":"article-journal","title":"Intelligent Power Infrastructure: A Structured Examination of Artificial Intelligence Techniques, Operational Challenges, and Evolutionary Pathways in Next-Generation Smart Grid Systems","abstract":"Abstract Global momentum toward intelligent power infrastructure is accelerating as electricity demand grows, renewable energy penetration deepens, and sustainability imperatives intensify. Artificial intelligence (AI) technologies—encompassing machine learning, deep learning, reinforcement learning, expert systems, fuzzy logic, and hybrid frameworks—have become indispensable enablers of next-generation grid operations. These approaches support intelligent monitoring, accurate load and renewable-energy forecasting, autonomous fault detection, demand-response orchestration, and optimal distributed energy resource integration. Notwithstanding these advantages, practical deployment faces persistent barriers including cyber security vulnerabilities, data-privacy constraints, infrastructure investment costs, device interoperability deficits, and insufficient regulatory frameworks. This study presents a structured original review of AI-enabled smart grid architectures and operational paradigms, systematically evaluates capabilities and trade-offs of principal AI methods, analyzes prevailing challenges alongside viable mitigation strategies, quantifies documented operational benefits, and maps prospective technological trajectories—including edge AI, digital twins, explainable AI, block chain, and federated learning—toward autonomous grid operation. Findings indicate that strategic AI adoption is critical to achieving resilient, efficient, and low-carbon power systems for the future.","author":[{"family":"Gopinath","given":"K"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20962003","URL":"https://doi.org/10.5281/zenodo.20962003","source":"datacite"},{"id":"doi:10.6084/m9.figshare.c.7942699","type":"article-journal","title":"FedscGen: privacy-preserving federated batch effect correction of single-cell RNA sequencing data","abstract":"Abstract Single-cell RNA-seq data from clinical samples often suffer from batch effects, but data sharing is limited due to genomic privacy concerns. We present FedscGen, a privacy-preserving communication-efficient federated method built upon the scGen model, enhanced with secure multiparty computation. FedscGen supports federated training and batch effect correction workflows, including the integration of new studies. We benchmark FedscGen across diverse datasets, showing competitive performance—matching scGen on key metrics like NMI, GC, ILF1, ASW_C, kBET, and EBM on the Human Pancreas dataset. Published as a FeatureCloud app, FedscGen enables secure, real-world collaboration for scRNA-seq batch effect correction.","author":[{"family":"Bakhtiari","given":"Mohammad"},{"family":"Bonn","given":"Stefan"},{"family":"Theis","given":"Fabian"},{"family":"Zolotareva","given":"Olga"},{"family":"Baumbach","given":"Jan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.6084/m9.figshare.c.7942699","URL":"https://doi.org/10.6084/m9.figshare.c.7942699","source":"datacite"},{"id":"doi:10.6084/m9.figshare.c.7942699.v1","type":"article-journal","title":"FedscGen: privacy-preserving federated batch effect correction of single-cell RNA sequencing data","abstract":"Abstract Single-cell RNA-seq data from clinical samples often suffer from batch effects, but data sharing is limited due to genomic privacy concerns. We present FedscGen, a privacy-preserving communication-efficient federated method built upon the scGen model, enhanced with secure multiparty computation. FedscGen supports federated training and batch effect correction workflows, including the integration of new studies. We benchmark FedscGen across diverse datasets, showing competitive performance—matching scGen on key metrics like NMI, GC, ILF1, ASW_C, kBET, and EBM on the Human Pancreas dataset. Published as a FeatureCloud app, FedscGen enables secure, real-world collaboration for scRNA-seq batch effect correction.","author":[{"family":"Bakhtiari","given":"Mohammad"},{"family":"Bonn","given":"Stefan"},{"family":"Theis","given":"Fabian"},{"family":"Zolotareva","given":"Olga"},{"family":"Baumbach","given":"Jan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.6084/m9.figshare.c.7942699.v1","URL":"https://doi.org/10.6084/m9.figshare.c.7942699.v1","source":"datacite"},{"id":"doi:10.48550/arxiv.2601.06404","type":"manuscript","title":"Stitch the Fragments: One-Shot Hierarchical Federated Clustering","abstract":"Federated Clustering (FC) faces a critical bottleneck in real-world scenarios, i.e., global clusters are rarely intact, often fragmenting into incomplete, multi-granular unlabeled ``clusterlets'' distributed across Non-IID clients. Although hierarchical clustering is theoretically well-suited to model such nested distributions, its recursive nature strictly relies on multi-round communication, introducing prohibitive computational overhead and severe privacy vulnerabilities. This paper, therefore, proposes a novel one-shot hierarchical federated clustering framework designed to seamlessly ``stitch'' the fragmented local clusterlets into a holistic global distribution. Our approach enables clients to perform autonomous fine-grained distribution exploration, uploading prototype-level knowledge via a dynamic parameter-interleaving mechanism to scramble transmission trajectories, which effectively prevents the server from tracing individual client data distributions. Subsequently, a multi-granular learning mechanism at the server fuses these granularly inconsistent local clusterlets, reconstructing a coherent global hierarchy for ultimate clustering. Extensive experiments on real benchmark datasets illustrate the superiority of the proposed approach, which effectively bridges the granularity gap among heterogeneous clients while minimizing privacy exposure risks via anonymized informative one-shot communication.","author":[{"family":"Cai","given":"Shenghong"},{"family":"Yang","given":"Zihua"},{"family":"Lu","given":"Yang"},{"family":"Li","given":"Mengke"},{"family":"Ji","given":"Yuzhu"},{"family":"Zhang","given":"Yiqun"},{"family":"Cheung","given":"Yiu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2601.06404","URL":"https://doi.org/10.48550/arxiv.2601.06404","source":"datacite"},{"id":"doi:10.48550/arxiv.2505.04535","type":"manuscript","title":"FDA-Opt: Federated Fine-Tuning via Dynamic Update Schedules","abstract":"Federated Learning (FL) enables the utilization of vast, previously inaccessible data sources. At the same time, pre-trained Language Models (LMs) have taken the world by storm and for good reason. They exhibit remarkable emergent abilities and are readily adapted to downstream tasks. This opens one of the most exciting frontiers in FL: fine-tuning LMs. Yet, a persistent challenge in FL is the frequent, rigid communication of parameters -- a problem magnified by the sheer size of these contemporary models. The FedOpt family of algorithms has become the go-to approach for FL, relying on fixed but arbitrary intervals for model exchanges. Recently, the FDA algorithm prescribed a dynamic approach by monitoring the training progress. However, it introduced a hard-to-calibrate parameter and imposed a rigid synchronization scheme. In this work, we address these limitations by proposing the FDA-Opt family of algorithms -- a unified generalization of both FDA and FedOpt. Our experimental evaluation focuses on fine-tuning LMs on downstream NLP tasks and demonstrates that FDA-Opt outperforms FedOpt even when it is configured with hyper-parameters specifically optimized for the latter. In other words, we show that FDA-Opt is a practical, drop-in replacement for FedOpt in modern FL libraries and systems: it requires no additional configuration and delivers superior performance out of the box.","author":[{"family":"Theologitis","given":"Michael"},{"family":"Samoladas","given":"Vasilis"},{"family":"Deligiannakis","given":"Antonios"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2505.04535","URL":"https://doi.org/10.48550/arxiv.2505.04535","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.14242","type":"manuscript","title":"Could Model Partitioning Make Federated Learning More Sustainable?","abstract":"As federated learning (FL) extends from distributed machine learning between low-power devices to cross-silo scenarios involving edge servers and data centres, its carbon footprint has become a growing concern. Addressing this, methods for sustainable FL align training with low-carbon energy availability or low grid demand and reduce the energy consumption of clients powered by high-carbon sources by decreasing the size of their models. We propose applying model partitioning, which can shift energy consumption by offloading parts of a model to another participant, in response to carbon- or grid-aware signals. Our preliminary findings show that for some partition points, model partitioning can reduce a participant's energy consumption by up to 76% without any significant time or energy consumption overhead compared to non-partitioned training.","author":[{"family":"Frohlich","given":"Tobias"},{"family":"Vlaar","given":"Tiffany"},{"family":"Thamsen","given":"Lauritz"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.14242","URL":"https://doi.org/10.48550/arxiv.2608.14242","source":"datacite"},{"id":"doi:10.5281/zenodo.21937534","type":"article-journal","title":"T3-CIDERS: Train-the-trainer for CI Upskilling in Cybersecurity Disciplines: Insights from Fostering the First-Year Cohort","abstract":"We present the T3-CIDERS project, a “Train-the-Trainer” (T3) initiative to foster a community of practice in cyberinfrastructure (CI)- and data-enabled cybersecurity research and education. Building on the NSF-funded DeapSECURE project, T3-CIDERS prepares Future Trainers (FTs)—faculty-student teams—with strong technical foundations in CI (HPC, big data, machine learning, cryptography) and instructional methods for teaching CI-related contents in the context of cybersecurity research and education. Since its launch in 2023, the project has supported the first cohort through pre-training, a summer institute, and ongoing learning engagements. In the first cohort, six local CI training events led by FTs were conducted across multiple U.S. states, introducing CI and cybersecurity concepts to nearly 100 students, from K-12 to graduate levels. A new “Module X” on the Security of Federated Learning is close to completion and will be featured in the upcoming winter institute in January 2026 in Arizona. We are currently recruiting for the winter institute second cohort. The project’s scholarly outcomes to date include three conferences and three journal papers spanning both HPC- and education-focused conferences.","author":[{"family":"Sosonkina","given":"Masha"},{"family":"Wu","given":"Hongyi"},{"family":"Jiang","given":"Peng"},{"family":"Purwanto","given":"Wirawan"},{"family":"Yang","given":"Mohan"},{"family":"Parry","given":"Dorothy"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21937534","URL":"https://doi.org/10.5281/zenodo.21937534","source":"datacite"},{"id":"doi:10.5281/zenodo.21937535","type":"article-journal","title":"T3-CIDERS: Train-the-trainer for CI Upskilling in Cybersecurity Disciplines: Insights from Fostering the First-Year Cohort","abstract":"We present the T3-CIDERS project, a “Train-the-Trainer” (T3) initiative to foster a community of practice in cyberinfrastructure (CI)- and data-enabled cybersecurity research and education. Building on the NSF-funded DeapSECURE project, T3-CIDERS prepares Future Trainers (FTs)—faculty-student teams—with strong technical foundations in CI (HPC, big data, machine learning, cryptography) and instructional methods for teaching CI-related contents in the context of cybersecurity research and education. Since its launch in 2023, the project has supported the first cohort through pre-training, a summer institute, and ongoing learning engagements. In the first cohort, six local CI training events led by FTs were conducted across multiple U.S. states, introducing CI and cybersecurity concepts to nearly 100 students, from K-12 to graduate levels. A new “Module X” on the Security of Federated Learning is close to completion and will be featured in the upcoming winter institute in January 2026 in Arizona. We are currently recruiting for the winter institute second cohort. The project’s scholarly outcomes to date include three conferences and three journal papers spanning both HPC- and education-focused conferences.","author":[{"family":"Sosonkina","given":"Masha"},{"family":"Wu","given":"Hongyi"},{"family":"Jiang","given":"Peng"},{"family":"Purwanto","given":"Wirawan"},{"family":"Yang","given":"Mohan"},{"family":"Parry","given":"Dorothy"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21937535","URL":"https://doi.org/10.5281/zenodo.21937535","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.04791","type":"manuscript","title":"On the Effectiveness of Adaptation Strategies for VLM-Based Federated Learning in Remote Sensing","abstract":"Federated learning (FL) enables collaborative training of deep learning models across decentralized image archives without requiring data centralization. This paradigm is particularly relevant in remote sensing (RS), where legal regulations, privacy concerns, and bandwidth constraints restrict data sharing. However, the presence of training data heterogeneity across clients (known as non-IID data) can impede convergence and limit the generalization capability of the aggregated global model. To mitigate the adverse effects of training data heterogeneity, vision-language models (VLMs) can be leveraged in FL due to their transferable representations, which have demonstrated robustness under distribution shifts. However, their large parameter size may substantially increase communication overhead and local computational complexity in federated settings. Therefore, it is crucial to select an appropriate VLM adaptation strategy that balances the generalization ability with the communication and computational constraints. To address this issue, in this paper, we present the first comparative study of VLM adaptation strategies for FL in the context of RS image classification. We investigate full fine-tuning, encoder-specific fine-tuning, prompt learning, and low-rank adaptation (LoRA) tuning, and analyze them with respect to three criteria: 1) generalization capability under non-IID data, 2) communication overhead, and 3) local computational complexity. Experiments on BigEarthNet-S2, EuroSAT, RESISC45, and ImageNet reveal distinct trade-offs between task specialization, cross-domain generalization, and efficiency. Based on our findings, we derive a guideline for the selection of an appropriate VLM adaptation strategy in FL for RS image classification under different operational constraints. The code of this work is publicly available at https://git.tu-berlin.de/rsim/FL-RS-VLM.","author":[{"family":"Lösche","given":"Simon"},{"family":"Büyüktaş","given":"Barış"},{"family":"Adler","given":"Mathis"},{"family":"Zavras","given":"Angelos"},{"family":"Papoutsis","given":"Ioannis"},{"family":"Demir","given":"Begüm"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.04791","URL":"https://doi.org/10.48550/arxiv.2608.04791","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.04753","type":"manuscript","title":"Attention, Anomalies! Handling Attention Layers in Unsupervised Federated Outlier Detection","abstract":"Attention layers are the backbone of today's most powerful and impactful models. Models with multi-million and billion parameters rely on contextual knowledge provided by attention layers. However, their use goes well beyond just being the core component of large language models. One particularly interesting application is in Memory Augmented Autoencoders (MemAE), specifically for unsupervised representation learning in outlier detection tasks. It was shown that attention helps these models be more effective in centralized learning scenarios. Our work aims to address the lack of specialized aggregation techniques in Federated Learning (FL) when it comes to MemAE models. In this paper we analyze the intricacies of the architecture behind Memory Augmented Autoencoders, and propose novel, guided approaches to effectively aggregate these models in federated scenarios. We demonstrate our approach on non-IID datasets and show that these novel aggregation schemes are more robust when dealing with numerous edge nodes in environments with unbalanced datasets, specifically for unsupervised anomaly detection scenarios. This approach improves the performance of even very shallow autoencoders, allowing them to be used in resource constrained environments.","author":[{"family":"Ilić","given":"Mihailo"},{"family":"Savić","given":"Miloš"},{"family":"Kurbalija","given":"Vladimir"},{"family":"Ivanović","given":"Mirjana"},{"family":"Fortino","given":"Giancarlo"},{"family":"Jakovetić","given":"Dušan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.04753","URL":"https://doi.org/10.48550/arxiv.2608.04753","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.01521","type":"manuscript","title":"MineGrad: Gradient Inversion Attacks on LoRA Fine-Tuning","abstract":"Parameter-efficient fine-tuning (PEFT), such as low-rank adaptation (LoRA), has recently been adopted in federated learning to reduce communication and computation costs. In this setup, users download a pretrained model from the server prior to fine-tuning, and then fine-tune lightweight LoRA modules locally while keeping the pretrained model frozen, sharing only the gradients of the fine-tuning parameters with the server. Despite its growing popularity, robustness of federated fine-tuning against an adversarial server remains underexplored, where the server maliciously tampers with the training protocol to breach the privacy of users' data. In this work, we investigate gradient inversion attacks on LoRA fine-tuning. We propose an analytical attack that enables a malicious server to recover private user data by leveraging a poisoned pretrained model and fine-tuning parameters. Our design embeds fine-tuning data within the shared gradients, to allow the server to analytically reconstruct user data. Unlike prior works, our attack is applicable to both language and vision tasks, does not rely on computationally expensive (adversarial) pretraining with public datasets or require the number of training tokens to be less than the rank of LoRA modules. Experimental results on both language and vision tasks demonstrate high-fidelity data recovery across multiple baselines, revealing several critical vulnerabilities.","author":[{"family":"Sami","given":"Hasin"},{"family":"Sen","given":"Swapneel"},{"family":"Guler","given":"Basak"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.01521","URL":"https://doi.org/10.48550/arxiv.2608.01521","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.01129","type":"manuscript","title":"FeDepth: Federated Learning for Depth Estimation under Robot Heterogeneity","abstract":"Although recent robot perception research emphasizes training on data from diverse environments to improve generalization, most existing methods still rely on centralized learning, which is inefficient and difficult to scale across heterogeneous robot platforms. Federated learning (FL) offers an alternative by enabling distributed training without raw data transfer, but it suffers from severe performance degradation under domain shifts caused by heterogeneity across clients. In real robotic deployments, data distributions often overlap across platforms, environments, and sensing conditions, making it difficult to partition clients into clearly separated domains. However, this characteristic breaks the assumption of clearly separable client domains commonly used in clustered FL. To address this gap in robot perception, particularly in depth estimation, we introduce two realistic and unexplored non-IID scenarios that reflect heterogeneity in terms of platform, environment, and depth distribution. We then propose FeDepth, a descriptor-based clustered FL framework that models client relationships through soft clustering. Unlike hard clustering methods that assume clearly separated clusters, FeDepth allows clients to participate in multiple clusters, capturing continuous and ambiguous domain transitions commonly observed in robotic environments. Extensive experiments demonstrate that FeDepth consistently improves robustness over standard FL and clustered FL baselines across multiple depth estimation architectures, providing a practical and effective solution for federated robot perception. Our project page is available at https://vision3d-lab.github.io/fedepth/.","author":[{"family":"Lee","given":"Ganghyeon"},{"family":"Lee","given":"Inha"},{"family":"Lee","given":"Junhee"},{"family":"Lee","given":"Jeongeon"},{"family":"Yoon","given":"Sung"},{"family":"Joo","given":"Kyungdon"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.01129","URL":"https://doi.org/10.48550/arxiv.2608.01129","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.00855","type":"manuscript","title":"Partially-Observable Transmission Control for UAV-Enabled Federated Learning in IoT Networks","abstract":"Uncrewed aerial vehicle (UAV)-enabled federated learning (FL) can provide flexible, on-demand edge intelligence for large-scale IoT deployments, but operating in shared unlicensed bands makes uplink update delivery interference-coupled and unreliable. In this paper, we develop a packet-level transmission framework that captures buffer overflow, delay violations, and transmission errors, and uses the resulting packet delivery ratio (PDR) to represent partial-update reception through a packetized, Bernoulli-masked FL aggregation process. We then formulate a fairness-consensus bilevel (FCB) optimization that jointly controls (i) transmission thresholds to maximize the average PDR while reaching consensus under partial observability and (ii) transmission powers to improve the worst PDR and enforce fairness across IoT learners. To solve this problem, we propose an alternating FCB optimizer composed of a consensus-based threshold controller (CTC), which drives the IoT learners toward a PDR-efficient consensus on transmission thresholds, and a fairness-based power controller (FPC), which updates transmission powers to improve the worst PDR and ensure fairness under the resulting consensus thresholds. Numerical results on CNN-based FL tasks show that the FCB optimizer improves FL aggregation and training performance by enhancing packet-level update delivery, consistently outperforming baseline transmission policies.","author":[{"family":"Ghazikor","given":"Masoud"},{"family":"Ni","given":"Zhou"},{"family":"Hashemi","given":"Morteza"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.00855","URL":"https://doi.org/10.48550/arxiv.2608.00855","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.23017","type":"manuscript","title":"Nautilus: A Verifiable Hierarchical Federated Learning Framework for Vehicular-Edge-Cloud Systems","abstract":"Federated Learning (FL) enables privacy-preserving collaborative learning for Internet of Vehicles (IoV) scenarios, but extreme heterogeneity of vehicular-edge-cloud resources severely limits system efficiency. Dynamic scheduling strategies mitigate this issue but introduce new trust concerns: verifying fair scheduling decisions and faithful client execution of compression instructions without privacy leakage remains an open challenge. We propose Nautilus, a verifiable efficient federated learning framework. First, a multi-dimensional resource-aware scheduling algorithm dynamically allocates compression ratios and training tasks based on vehicle bandwidth, latency and computing power, improving training efficiency. Second, a Zero-Knowledge Proof (ZKP) mechanism ensures scheduling fairness and execution compliance while preserving privacy. Experiments show the framework reduces communication overhead and accelerates convergence with guaranteed system integrity.","author":[{"family":"Wu","given":"Linyang"},{"family":"Jia","given":"Linpeng"},{"family":"Zhang","given":"Hanwen"},{"family":"Duan","given":"Tiantian"},{"family":"Sun","given":"Yi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.23017","URL":"https://doi.org/10.48550/arxiv.2606.23017","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.23245","type":"manuscript","title":"FedTaste: Topology-Aware Structural Transfer for Multimodal Federated Learning with Missing Modalities","abstract":"Multimodal Federated Learning is often challenged by arbitrary modality missingness and Non-IID data distributions, which lead to severe representation drift and hinder effective collaboration across clients. Existing methods typically rely on generative imputation, external auxiliary data, or isolated unimodal training to bridge modality gaps, often incurring substantial communication and computational costs as well as potential privacy risks. To address these limitations, we propose FedTaste, a parameter-efficient framework for topology-aware structural transfer in Multimodal Federated Learning with missing modalities. Instead of aligning fragile first-order features, FedTaste focuses on more stable group-level semantic relations. Specifically, FedTaste leverages frozen foundation models to extract a joint multimodal topology from full-modality clients, which is then consolidated by the server into a global structural blueprint. To adapt clients with missing modalities, we introduce Modality-Adaptive Structural Prompts together with spectral consistency regularization, enabling lightweight branch-specific adaptation that aligns local partial representations with the shared blueprint. In this way, FedTaste avoids explicit modality imputation while preserving shared semantic structure across clients. Extensive experiments demonstrate that FedTaste consistently achieves superior performance across multiple datasets and challenging Non-IID settings, while substantially reducing communication overhead compared with existing methods.","author":[{"family":"Liang","given":"Haochen"},{"family":"Zhang","given":"Jie"},{"family":"Ochiai","given":"Hideya"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.23245","URL":"https://doi.org/10.48550/arxiv.2607.23245","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.21449","type":"manuscript","title":"QuantumChain: Blockchain-Backed Quantum Federated Learning for Financial Fraud Detection","abstract":"Financial fraud detection is challenged by decentralized data, severe class imbalance, and privacy constraints. This paper presents QuantumChain, a secure Quantum Federated Learning (QFL) framework that combines hybrid quantum-classical neural networks, encrypted federated aggregation, blockchain-based auditability, and quantum-secure communication. Each client trains a local hybrid model in which a variational quantum circuit is embedded between classical neural layers, while model updates are protected through homomorphic encryption, threshold secret sharing, and QKD-based keying. A permissioned blockchain records aggregation events and supports reputation-weighted trust among participants. We evaluate QuantumChain on financial transaction data using a compact, size-matched classical baseline to isolate the effect of the quantum layer. Results show that the HQNN achieves comparable accuracy while improving fraud-class recall in most settings, reaching 94.6% recall compared with 93.2% for the classical model. The Deep QLayer improves performance in full-data settings, suggesting that added circuit depth helps recover representational capacity when the shallow circuit becomes limited. Mixed-state simulations further show that the recall trend persists under non-ideal quantum evolution. In federated deployment with 10 heterogeneous clients, global accuracy increases from 97.7% to 98.8% over five rounds before stabilizing. These results show that QuantumChain can integrate depth-aware hybrid quantum models into a secure federated fraud-detection pipeline while maintaining stable global convergence.","author":[{"family":"Douros","given":"Epameinondas"},{"family":"Dalampekis","given":"Konstantinos"},{"family":"Innan","given":"Nouhaila"},{"family":"Theodonis","given":"Ioannis"},{"family":"Shafique","given":"Muhammad"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.21449","URL":"https://doi.org/10.48550/arxiv.2607.21449","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.05157","type":"manuscript","title":"Onboarding Without Forgetting: Hypernetwork Personalization with Data-Free Replay for Personalized Federated Learning","abstract":"Federated Learning (FL) enables collaborative training across distributed clients without sharing raw data, offering strong privacy benefits. However, most methods assume all clients remain available throughout training, which is unrealistic as new clients often join over time. We study this setting, where the task and label space stay fixed but clients arrive in batches. Our analysis reveals two key challenges: updating the shared model only with new clients harms existing clients, while freezing it protects them but blocks gains from new knowledge. To capture these trade-offs, we introduce Proactive Adaptation (PA) for onboarding gains and Retroactive Improvement (RI) for changes in earlier clients without retraining. We then propose pFedDSH, which combines a central hypernetwork for personalized initialization, batch-specific binary masks for capacity preservation and allocation, and server-side data-free replay to propagate improvements without exposing client data. Experiments show that pFedDSH preserves stability for existing clients while keeping communication and adaptation costs unchanged for new clients.","author":[{"family":"Nguyen","given":"Thinh"},{"family":"Khiem","given":"Le"},{"family":"Tran","given":"Van"},{"family":"Doan","given":"Khoa"},{"family":"Chawla","given":"Nitesh"},{"family":"Wong","given":"Kok"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.05157","URL":"https://doi.org/10.48550/arxiv.2508.05157","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.19881","type":"manuscript","title":"FedS2R: One-Shot Federated Domain Generalization for Synthetic-to-Real Semantic Segmentation in Autonomous Driving","abstract":"Federated domain generalization has shown promising progress in image classification by enabling collaborative training across multiple clients without sharing raw data. However, its potential in the semantic segmentation of autonomous driving remains underexplored. In this paper, we propose FedS2R, the first one-shot federated domain generalization framework for synthetic-to-real semantic segmentation in autonomous driving. FedS2R comprises two components: an inconsistency-driven data augmentation strategy that generates images for unstable classes, and a multi-client knowledge distillation scheme with feature fusion that distills a global model from multiple client models. Experiments on five real-world datasets, Cityscapes, BDD100K, Mapillary, IDD, and ACDC, show that the global model significantly outperforms individual client models and is only 2 mIoU points behind the model trained with simultaneous access to all client data. These results demonstrate the effectiveness of FedS2R in synthetic-to-real semantic segmentation for autonomous driving under federated learning","author":[{"family":"Lian","given":"Tao"},{"family":"Gómez","given":"Jose"},{"family":"López","given":"Antonio"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.19881","URL":"https://doi.org/10.48550/arxiv.2507.19881","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.00216","type":"manuscript","title":"FedCod: An Efficient Communication Protocol for Cross-Silo Federated Learning with Coding","abstract":"Federated Learning (FL) is an innovative distributed machine learning paradigm that enables multiple parties to collaboratively train a model without sharing their raw data, thereby preserving data privacy. Communication efficiency concerns arise in cross-silo FL, particularly due to the network heterogeneity and fluctuations associated with geo-distributed silos. Most existing solutions to these problems focus on algorithmic improvements that alter the FL algorithm but sacrificing the training performance. How to address these problems from a network perspective that is decoupled from the FL algorithm remains an open challenge. In this paper, we propose FedCod, a new application layer communication protocol designed for cross-silo FL. FedCod transparently utilizes a coding mechanism to enhance the efficient use of idle bandwidth through client-to-client communication, and dynamically adjusts coding redundancy to mitigate network bottlenecks and fluctuations, thereby improving the communication efficiency and accelerating the training process. In our real-world experiments, FedCod demonstrates a significant reduction in average communication time by up to 62% compared to the baseline, while maintaining FL training performance and optimizing inter-client communication traffic.","author":[{"family":"Yan","given":"Peishen"},{"family":"Li","given":"Jun"},{"family":"Wang","given":"Hao"},{"family":"Song","given":"Tao"},{"family":"Hua","given":"Yang"},{"family":"Peng","given":"Lu"},{"family":"Guan","given":"Haibing"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.00216","URL":"https://doi.org/10.48550/arxiv.2501.00216","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.17913","type":"manuscript","title":"AutoEncoder-Compressed Parallel Split Learning for Pre-trained Model Fine-Tuning","abstract":"Distributed Fine-Tuning (DFT) of large-scale Foundation Models (FMs) on resource-constrained edge devices is limited by local compute constraints and communication overhead. Parallel Split Learning (PSL) reduces client-side computation by keeping few model layers on each client and offloading the remaining computation to the server; however, clients must exchange intermediate activations and gradients with the server at every training step. Existing SL communication-compression methods mainly rely on task-agnostic heuristics, such as sparsification and quantization. While learnable SL compressors can better adapt to intermediate representations, they require co-training with the target model. Therefore, directly inserting them into off-the-shelf FMs introduces feature-distribution misalignment and degrades DFT performance. To address this, we propose AE-PSL, a communication-efficient PSL framework that compresses intermediate activations and gradients using a lightweight AutoEncoder (AE) placed at the split layer. To ensure compatibility of AE compression with pre-trained FMs, AE-PSL introduces a novel two-stage alignment mechanism, which adapts the AE to the pre-trained model's feature manifold and client-specific feature distributions before DFT.","author":[{"family":"Meuwissen","given":"Bas"},{"family":"Tsouvalas","given":"Vasileios"},{"family":"Meratnia","given":"Nirvana"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.17913","URL":"https://doi.org/10.48550/arxiv.2607.17913","source":"datacite"},{"id":"doi:10.5281/zenodo.20344083","type":"article-journal","title":"Deep Learning Strategies for Predicting Drug–Target Interactions: Advances, Challenges, and Future Perspectives","abstract":"Drug–target interaction (DTI) prediction lies at the heart of modern drug discovery, determining whether a candidate small molecule will bind to and modulate a biological macromolecule of therapeutic relevance. Traditional experimental high-throughput screening is expensive, time-consuming, and constrained by library size, while classical computational approaches—docking, pharmacophore modelling, and quantitative structure–activity relationship (QSAR) modelling suffer from limitations in scalability and generalizability. The emergence of deep learning (DL) has fundamentally transformed the field, enabling end-to-end learning of molecular representations and interaction patterns from heterogeneous, large-scale biomedical data. This review provides a comprehensive synthesis of DL-based DTI prediction methodologies, covering convolutional neural networks (CNNs), recurrent neural networks (RNNs), graph neural networks (GNNs), transformer-based architectures, autoencoders, and multi-modal fusion frameworks. We discuss the critical role of molecular representation—from one-dimensional SMILES strings and circular fingerprints to three-dimensional molecular graphs and protein contact maps. Benchmark datasets including Davis, KIBA, BindingDB, ChEMBL, and PDBbind are reviewed with respect to their composition, metric conventions, and appropriate use. We critically examine key challenges: data scarcity and imbalance, negative-sample bias, interpretability deficits, cold-start generalization, and the limited availability of experimental three-dimensional protein structures. Emerging solutions—pre-trained chemical language models, AlphaFold3 integration, federated learning, knowledge graph-augmented GNNs, and causal interpretability methods—are discussed as future directions. This review aims to serve as an authoritative reference for computational chemists, bioinformaticians, and medicinal chemists seeking to leverage DL for accelerated, cost-effective drug discovery.","author":[{"family":"Bajpai","given":"Avinash"},{"family":"Sharma","given":"Sachin"},{"family":"Singh","given":"Birender"},{"family":"Nisha","given":"Km"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20344083","URL":"https://doi.org/10.5281/zenodo.20344083","source":"datacite"},{"id":"doi:10.5281/zenodo.20344084","type":"article-journal","title":"Deep Learning Strategies for Predicting Drug–Target Interactions: Advances, Challenges, and Future Perspectives","abstract":"Drug–target interaction (DTI) prediction lies at the heart of modern drug discovery, determining whether a candidate small molecule will bind to and modulate a biological macromolecule of therapeutic relevance. Traditional experimental high-throughput screening is expensive, time-consuming, and constrained by library size, while classical computational approaches—docking, pharmacophore modelling, and quantitative structure–activity relationship (QSAR) modelling suffer from limitations in scalability and generalizability. The emergence of deep learning (DL) has fundamentally transformed the field, enabling end-to-end learning of molecular representations and interaction patterns from heterogeneous, large-scale biomedical data. This review provides a comprehensive synthesis of DL-based DTI prediction methodologies, covering convolutional neural networks (CNNs), recurrent neural networks (RNNs), graph neural networks (GNNs), transformer-based architectures, autoencoders, and multi-modal fusion frameworks. We discuss the critical role of molecular representation—from one-dimensional SMILES strings and circular fingerprints to three-dimensional molecular graphs and protein contact maps. Benchmark datasets including Davis, KIBA, BindingDB, ChEMBL, and PDBbind are reviewed with respect to their composition, metric conventions, and appropriate use. We critically examine key challenges: data scarcity and imbalance, negative-sample bias, interpretability deficits, cold-start generalization, and the limited availability of experimental three-dimensional protein structures. Emerging solutions—pre-trained chemical language models, AlphaFold3 integration, federated learning, knowledge graph-augmented GNNs, and causal interpretability methods—are discussed as future directions. This review aims to serve as an authoritative reference for computational chemists, bioinformaticians, and medicinal chemists seeking to leverage DL for accelerated, cost-effective drug discovery.","author":[{"family":"Bajpai","given":"Avinash"},{"family":"Sharma","given":"Sachin"},{"family":"Singh","given":"Birender"},{"family":"Nisha","given":"Km"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20344084","URL":"https://doi.org/10.5281/zenodo.20344084","source":"datacite"},{"id":"doi:10.5281/zenodo.19610735","type":"article-journal","title":"AI-Assisted Clinical Trial Recruitment for Devices","abstract":"Clinical trial recruitment is the most frequent cause of device trial delay and the most preventable. More than 80% ofdevice trials fail to meet original enrolment targets on schedule, extending development timelines by a median of 8-14months and generating costs that fall disproportionately on smaller device developers without the buffer capital to absorbslippage. AI-assisted recruitment addresses the root causes of this problem: slow patient identification, high screenfailure rates from manual eligibility review, and inadequate representation of the real-world device use population in trialcohorts. This study presents the AI-Assisted Device Trial Recruitment Framework (AADTRF), evaluating five AIrecruitment approaches -- NLP-based electronic health record eligibility screening (NLP-EHR), machine learning patientmatching from device registries (ML-Reg), digital outreach and social media recruitment (DO-SM), federated multi-sitecohort identification (Fed-CI), and predictive retention and dropout risk modelling (PR-DRM) -- across four device trialcontexts: cardiac implant trials, orthopaedic device trials, neurostimulation device trials, and diagnostic imaging trials.Performance was scored using the Recruitment Effectiveness Score (RES), a weighted composite of enrolment rateimprovement (0.30), screen failure reduction (0.20), time to enrolment completion (0.20), population representativeness(0.15), and cost efficiency (0.15). NLP-based EHR screening achieved the highest RES (0.890), improving enrolmentrates by a mean 47% and reducing screen failure from 38.4% to 16.2% across four trial contexts. ML registry matchingranked second (0.888) with the best screen failure reduction (0.920) in device-specific contexts. Federated cohortidentification (0.877) demonstrated unique value for rare device trial indications, generating eligible patient pools 3.4times larger than single-site approaches. Digital outreach led on cost efficiency (0.940) and populationrepresentativeness (0.920) but produced the highest screen failure rates (0.800) due to self-selection bias in volunteerpopulations.","author":[{"family":"Mulle","given":"Marta"},{"family":"Klein","given":"Nina"},{"family":"Bianchi","given":"Andreas"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.19610735","URL":"https://doi.org/10.5281/zenodo.19610735","source":"datacite"},{"id":"doi:10.5281/zenodo.19610736","type":"article-journal","title":"AI-Assisted Clinical Trial Recruitment for Devices","abstract":"Clinical trial recruitment is the most frequent cause of device trial delay and the most preventable. More than 80% ofdevice trials fail to meet original enrolment targets on schedule, extending development timelines by a median of 8-14months and generating costs that fall disproportionately on smaller device developers without the buffer capital to absorbslippage. AI-assisted recruitment addresses the root causes of this problem: slow patient identification, high screenfailure rates from manual eligibility review, and inadequate representation of the real-world device use population in trialcohorts. This study presents the AI-Assisted Device Trial Recruitment Framework (AADTRF), evaluating five AIrecruitment approaches -- NLP-based electronic health record eligibility screening (NLP-EHR), machine learning patientmatching from device registries (ML-Reg), digital outreach and social media recruitment (DO-SM), federated multi-sitecohort identification (Fed-CI), and predictive retention and dropout risk modelling (PR-DRM) -- across four device trialcontexts: cardiac implant trials, orthopaedic device trials, neurostimulation device trials, and diagnostic imaging trials.Performance was scored using the Recruitment Effectiveness Score (RES), a weighted composite of enrolment rateimprovement (0.30), screen failure reduction (0.20), time to enrolment completion (0.20), population representativeness(0.15), and cost efficiency (0.15). NLP-based EHR screening achieved the highest RES (0.890), improving enrolmentrates by a mean 47% and reducing screen failure from 38.4% to 16.2% across four trial contexts. ML registry matchingranked second (0.888) with the best screen failure reduction (0.920) in device-specific contexts. Federated cohortidentification (0.877) demonstrated unique value for rare device trial indications, generating eligible patient pools 3.4times larger than single-site approaches. Digital outreach led on cost efficiency (0.940) and populationrepresentativeness (0.920) but produced the highest screen failure rates (0.800) due to self-selection bias in volunteerpopulations.","author":[{"family":"Mulle","given":"Marta"},{"family":"Klein","given":"Nina"},{"family":"Bianchi","given":"Andreas"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.19610736","URL":"https://doi.org/10.5281/zenodo.19610736","source":"datacite"},{"id":"doi:10.5281/zenodo.21605297","type":"article-journal","title":"Utilizing Federated Health Databases and AI-Enhanced Neurodevelopmental Trajectory Mapping for Early Diagnosis of Autism Spectrum Disorder: A Review of Scalable Computational Models","abstract":"Early diagnosis of Autism Spectrum Disorder (ASD) remains a critical challenge in pediatric neurodevelopmental care, with significant implications for intervention effectiveness and lifelong outcomes. Traditional diagnostic methods often rely on subjective behavioral assessments, limiting early detection, especially in resource-limited settings. This review explores the convergence of federated health databases and AI-driven neurodevelopmental trajectory mapping as scalable computational solutions for timely ASD diagnosis. Federated learning frameworks enable collaborative model training across decentralized data silos while preserving patient privacy, thereby overcoming the limitations of fragmented healthcare data ecosystems. Simultaneously, advanced machine learning techniques—including temporal graph networks, multimodal deep learning, and probabilistic modeling—facilitate individualized developmental path prediction. The paper systematically analyzes current architectures, datasets, and computational methods used to infer ASD risk from longitudinal health and behavioral data. It also discusses interpretability, scalability, and ethical considerations surrounding data governance and model transparency. The review concludes with recommendations for improving early ASD diagnosis through integrated AI-federated frameworks in global pediatric healthcare systems.","author":[{"family":"Omolayo","given":"Olasehinde"},{"family":"Okare","given":"Babawale"},{"family":"Taiwo","given":"Ajao"},{"family":"Aduloju","given":"Tope"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21605297","URL":"https://doi.org/10.5281/zenodo.21605297","source":"datacite"},{"id":"doi:10.5281/zenodo.21605298","type":"article-journal","title":"Utilizing Federated Health Databases and AI-Enhanced Neurodevelopmental Trajectory Mapping for Early Diagnosis of Autism Spectrum Disorder: A Review of Scalable Computational Models","abstract":"Early diagnosis of Autism Spectrum Disorder (ASD) remains a critical challenge in pediatric neurodevelopmental care, with significant implications for intervention effectiveness and lifelong outcomes. Traditional diagnostic methods often rely on subjective behavioral assessments, limiting early detection, especially in resource-limited settings. This review explores the convergence of federated health databases and AI-driven neurodevelopmental trajectory mapping as scalable computational solutions for timely ASD diagnosis. Federated learning frameworks enable collaborative model training across decentralized data silos while preserving patient privacy, thereby overcoming the limitations of fragmented healthcare data ecosystems. Simultaneously, advanced machine learning techniques—including temporal graph networks, multimodal deep learning, and probabilistic modeling—facilitate individualized developmental path prediction. The paper systematically analyzes current architectures, datasets, and computational methods used to infer ASD risk from longitudinal health and behavioral data. It also discusses interpretability, scalability, and ethical considerations surrounding data governance and model transparency. The review concludes with recommendations for improving early ASD diagnosis through integrated AI-federated frameworks in global pediatric healthcare systems.","author":[{"family":"Omolayo","given":"Olasehinde"},{"family":"Okare","given":"Babawale"},{"family":"Taiwo","given":"Ajao"},{"family":"Aduloju","given":"Tope"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21605298","URL":"https://doi.org/10.5281/zenodo.21605298","source":"datacite"},{"id":"doi:10.5281/zenodo.21604736","type":"article-journal","title":"Deep Learning–Based Soybean Leaf Disease Classification : A Comprehensive Review","abstract":"Soybean is one of the most economically important crops worldwide, yet its productivity is severely affected by a wide range of foliar diseases. Traditional disease diagnosis relies on expert visual inspection, which is time-consuming, subjective, and impractical for large-scale monitoring. In recent years, computer vision and artificial intelligence have emerged as promising tools for automated soybean leaf disease classification. This review presents a comprehensive analysis of state-of-the-art image-based soybean leaf disease classification techniques, with particular emphasis on deep learning, transfer learning, federated learning, and transformer-based models. The paper systematically examines preprocessing strategies, feature extraction methods, classification architectures, and evaluation protocols used in recent studies. Furthermore, comparative insights are drawn across convolutional neural networks, lightweight models, ensemble approaches, and vision transformers. The review also highlights key research findings, existing challenges, and unresolved limitations such as dataset imbalance, real-field variability, explainability, and deployment constraints. By synthesizing recent advances and identifying open research directions, this review aims to guide researchers toward more robust, scalable, and intelligent soybean disease diagnosis systems for sustainable agriculture.","author":[{"family":"Rathod","given":"Drashti"},{"family":"Patel","given":"Ketan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21604736","URL":"https://doi.org/10.5281/zenodo.21604736","source":"datacite"},{"id":"doi:10.5281/zenodo.21604737","type":"article-journal","title":"Deep Learning–Based Soybean Leaf Disease Classification : A Comprehensive Review","abstract":"Soybean is one of the most economically important crops worldwide, yet its productivity is severely affected by a wide range of foliar diseases. Traditional disease diagnosis relies on expert visual inspection, which is time-consuming, subjective, and impractical for large-scale monitoring. In recent years, computer vision and artificial intelligence have emerged as promising tools for automated soybean leaf disease classification. This review presents a comprehensive analysis of state-of-the-art image-based soybean leaf disease classification techniques, with particular emphasis on deep learning, transfer learning, federated learning, and transformer-based models. The paper systematically examines preprocessing strategies, feature extraction methods, classification architectures, and evaluation protocols used in recent studies. Furthermore, comparative insights are drawn across convolutional neural networks, lightweight models, ensemble approaches, and vision transformers. The review also highlights key research findings, existing challenges, and unresolved limitations such as dataset imbalance, real-field variability, explainability, and deployment constraints. By synthesizing recent advances and identifying open research directions, this review aims to guide researchers toward more robust, scalable, and intelligent soybean disease diagnosis systems for sustainable agriculture.","author":[{"family":"Rathod","given":"Drashti"},{"family":"Patel","given":"Ketan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21604737","URL":"https://doi.org/10.5281/zenodo.21604737","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.14489","type":"manuscript","title":"EdgeFaaS: A Function-based Framework for Edge Computing","abstract":"Edge computing brings unique challenges as the resources on the edge are highly diverse in capabilities and capacities, and highly distributed across many users and the physical world. Existing distributed computing frameworks cannot adequately handle this level of heterogeneity and distribution. This paper proposes EdgeFaaS, a novel function-based edge computing framework to enable edge applications to effectively utilize heterogeneous resources distributed across the Internet of Things (IoT), edge, and cloud for computing. It proposes function virtualization and storage virtualization to abstract distributed and heterogeneous physical resources and provides consistent virtual interfaces for deploying and executing functions and storing and accessing data. EdgeFaaS provides comprehensive support to diverse edge computing workflows, and at the same time allows users to flexibly adjust the configurations and explore various important tradeoffs. To demonstrate its usability, the paper also presents the implementation and evaluation of three representative workflows on EdgeFaaS for video analytics, federated learning, and audio classification, on a real testbed of 100+ geographically distributed IoT devices, edge servers, and cloud services. EdgeFaaS allows users to flexibly explore the deployment configurations of these workflows over distributed and heterogeneous resources. For example, users can easily vary the function placement of the video processing pipeline across IoT, edge, and cloud resources and study the tradeoff between computation and communication costs; users can also flexibly adjust the cluster count and size in the hierarchical federated learning system and explore the tradeoff between training accuracy and speed.","author":[{"family":"Vadnere","given":"Neha"},{"family":"Wang","given":"Yu"},{"family":"Chen","given":"Yitao"},{"family":"Sadesh","given":"Sreehari"},{"family":"Zhao","given":"Ming"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.14489","URL":"https://doi.org/10.48550/arxiv.2607.14489","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.13754","type":"manuscript","title":"PriEval-Protect: A Unified Framework for Privacy Evaluation and Protection in Healthcare Systems","abstract":"Safeguarding patient privacy while enabling meaningful healthcare data use remains critical under GDPR and HIPAA. Existing compliance methods are manual, error-prone, and separate policy audits from data-level assessments. This paper presents PriEval-Protect, a two-phase framework for unified privacy risk evaluation and mitigation. The evaluation phase combines regulatory compliance scoring using a fine-tuned legal LLM with RAG, and technical analysis via encryption type, data architecture, and metrics including similarity, uncertainty, adversary success, and information gain/loss. A composite risk score uses weighted aggregation via Analytic Hierarchy Process. The protection phase recommends countermeasures including federated learning and differential privacy based on assessed risk. Results on hospital documents and datasets demonstrate regulation-aligned, explainable assessments, bridging legal conformance and data-level risk analysis.","author":[{"family":"Chebil","given":"Ilef"},{"family":"Hadj","given":"Asma"},{"family":"Yousfi","given":"Souheib"},{"family":"Hedhili","given":"Aroua"},{"family":"Sliman","given":"Layth"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.13754","URL":"https://doi.org/10.48550/arxiv.2607.13754","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.05774","type":"manuscript","title":"First-Order Softmax Weighted Switching Gradient Method for Distributed Stochastic Minimax Optimization with Stochastic Constraints","abstract":"This paper addresses the distributed stochastic minimax optimization problem subject to stochastic constraints. We propose a novel first-order Softmax-Weighted Switching Gradient method tailored for federated learning. Under full client participation, our algorithm achieves the standard $\\tilde{\\mathcal{O}}(ε^{-4})$ oracle complexity to satisfy a unified bound $ε$ for both the optimality gap and feasibility tolerance. We extend our theoretical analysis to the practical partial participation regime by quantifying client sampling noise through a stochastic superiority assumption. Furthermore, by relaxing standard boundedness assumptions on the objective functions, we establish a strictly tighter lower bound for the softmax hyperparameter. We provide a unified error decomposition and establish a sharp $\\mathcal{O}(\\log\\frac{1}δ)$ high-probability convergence guarantee. Ultimately, our framework demonstrates that a single-loop primal-only switching mechanism provides a stable alternative for optimizing worst-case client performance, effectively bypassing the hyperparameter sensitivity and convergence oscillations often encountered in traditional primal-dual or penalty-based approaches. We verify the efficacy of our algorithm via experiment on the Neyman-Pearson (NP) classification, fair classification, and federated safe reinforcement learning tasks.","author":[{"family":"Luo","given":"Zhankun"},{"family":"Upadhyay","given":"Antesh"},{"family":"Moon","given":"Sang"},{"family":"Hashemi","given":"Abolfazl"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.05774","URL":"https://doi.org/10.48550/arxiv.2603.05774","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.08368","type":"manuscript","title":"FedOPAL: One-Shot Federated Learning via Analytic Visual Prompt Tuning","abstract":"With the widespread deployment of basic models in edge intelligence, communication bandwidth has become a core bottleneck restricting the scalability of federated learning. Although one-shot federated learning alleviates this problem by minimizing communication rounds, existing iterative fine-tuning or knowledge distillation methods still face challenges such as high server-side computational costs and hyperparameter sensitivity. Analytical federated learning achieves efficient gradientfree aggregation using least-squares closed-form solutions, but in environments with non-independent and identically distributed data, its static feature assumptions fail, leading to feature manifold misalignment and severely impairing model performance. To address this contradiction, this paper proposes the FedOPAL framework. This framework adapts the visual prompts as feature rectifiers, actively correcting the feature distribution of heterogeneous data to a linearly separable space by applying local proximal constraints, thereby satisfying the theoretical assumptions of analytical federated learning. Experimental results show that FedOPAL not only significantly outperforms the original analytical methods on several benchmarks, but also achieves accuracy comparable to state-of-the-art iterative methods while maintaining zero server-side training costs, providing a new engineering paradigm for efficient collaboration of large models on the edge.","author":[{"family":"Qiu","given":"Lingyu"},{"family":"Annunziata","given":"Daniela"},{"family":"Izzo","given":"Stefano"},{"family":"Giampaolo","given":"Fabio"},{"family":"Piccialli","given":"Francesco"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.08368","URL":"https://doi.org/10.48550/arxiv.2607.08368","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.06838","type":"manuscript","title":"An Adaptive Differentially Private Federated Learning Framework","abstract":"Federated learning enables collaborative model training across distributed clients while preserving data privacy. However, in practical deployments, device heterogeneity and non-independent and identically distributed (Non-IID) data often lead to unstable and biased gradient. When differential privacy is enforced, conventional fixed gradient clipping and Gaussian noise injection may further amplify gradient perturbations, resulting in training oscillation and degraded model performance. To address these challenges, we propose an adaptive differentially private federated learning framework that explicitly targets model efficiency under heterogeneous and privacy-constrained settings. On the client side, a lightweight local dimensionality reduction module is introduced to learn reduced-dimensional intermediate representations and produce more structured gradients during backpropagation, thereby mitigating noise amplification during local optimization. On the server side, an adaptive gradient clipping strategy dynamically adjusts clipping thresholds based on historical update statistics to avoid over-clipping and noise domination. Furthermore, a constraint-aware robust aggregation mechanism is designed to suppress unreliable or noise-dominated client updates and stabilize global optimization. Extensive experiments on CIFAR-10, SVHN, and STL-10 demonstrate that the proposed method consistently improves convergence stability and classification performance under differential privacy.","author":[{"family":"Wang","given":"Jin"},{"family":"Ma","given":"Hui"},{"family":"Zhang","given":"Yajun"},{"family":"Pei","given":"Xinjun"},{"family":"Yan","given":"Ming"},{"family":"Xing","given":"Fei"},{"family":"Chen","given":"Yikun"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.06838","URL":"https://doi.org/10.48550/arxiv.2602.06838","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.05720","type":"manuscript","title":"Inertia-Informed Federated Learning Control Framework for Distributed Smart Grid Resilience","abstract":"Resilient-by-design smart grid control demands frameworks capable of maintaining stability under physical disturbances and communication failures, without reliance on centralized coordination. While Centralized Training Decentralized Execution (CTDE) enables a learning-based control paradigm at the grid edge, individually trained models fail to generalize across unseen fault contingencies and fall short of fully decentralized deployment. Federated learning (FL) restores generalization through collaborative training; however, standard aggregation strategies remain agnostic to the physical heterogeneity of synchronous generators. This work proposes Inertia-Informed Weighted FedAvg (IIWFedAvg), a physics-informed aggregation strategy that embeds generator inertia directly into global model fusion for transient stability control in transmission networks. The proposed framework further integrates interpretable Chebyshev Kolmogorov-Arnold Network (ChebyKAN)-based controllers, augmented with Rate-of-Change-of-Frequency (RoCoF) features to enhance dynamic response awareness. Evaluated on the IEEE 39-bus benchmark under full decentralized deployment, IIWFedAvg achieves a 75% generalization success rate across unseen fault contingencies. It also surpasses the centralized baseline in two out of three stabilized faults, while delivering a 3x improvement in stabilization speed at zero centralized coordination overhead.","author":[{"family":"Shahbaz","given":"Ibrahim"},{"family":"Al-Refai","given":"Omar"},{"family":"Hammad","given":"Eman"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.05720","URL":"https://doi.org/10.48550/arxiv.2607.05720","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.14401","type":"manuscript","title":"pFedNavi: Structure-Aware Personalized Federated Vision-Language Navigation for Embodied AI","abstract":"Vision-Language Navigation VLN requires large-scale trajectory instruction data from private indoor environments, raising significant privacy concerns. Federated Learning FL mitigates this by keeping data on-device, but vanilla FL struggles under VLNs' extreme cross-client heterogeneity in environments and instruction styles, making a single global model suboptimal. This paper proposes pFedNavi, a structure-aware and dynamically adaptive personalized federated learning framework tailored for VLN. Our key idea is to personalize where it matters: pFedNavi adaptively identifies client-specific layers via layer-wise mixing coefficients, and performs fine-grained parameter fusion on the selected components (e.g., the encoder-decoder projection and environment-sensitive decoder layers) to balance global knowledge sharing with local specialization. We evaluate pFedNavi on two standard VLN benchmarks, R2R and RxR, using both ResNet and CLIP visual representations. Across all metrics, pFedNavi consistently outperforms the FedAvg-based VLN baseline, achieving up to 7.5% improvement in navigation success rate and up to 7.8% gain in trajectory fidelity, while converging 1.38x faster under non-IID conditions.","author":[{"family":"Yang","given":"Qingqian"},{"family":"Wang","given":"Hao"},{"family":"Zhang","given":"Sai"},{"family":"Li","given":"Jian"},{"family":"Hua","given":"Yang"},{"family":"Pan","given":"Miao"},{"family":"Song","given":"Tao"},{"family":"Qi","given":"Zhengwei"},{"family":"Guan","given":"Haibing"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.14401","URL":"https://doi.org/10.48550/arxiv.2602.14401","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.30161","type":"manuscript","title":"Federated Learning with Energy-Based Structured Probabilistic Inference","abstract":"Federated learning typically aggregates client updates using fixed or heuristic weighting rules, which can be suboptimal when clients have heterogeneous data and varying contributions to the global model. We propose a framework that refines client aggregation weights using Conditional Random Fields (CRFs). Our method defines unary potentials for individual clients and pairwise potentials for all client pairs, allowing the server to model both client-specific reliability and interactions between clients. The resulting CRF inference produces aggregation weights that enable better convergence of the global training objective. Experiments show that, under non-IID heterogeneity, our approach consistently improves performance over well-established federated learning baselines.","author":[{"family":"Fenoglio","given":"Dario"},{"family":"Kirilenko","given":"Daniil"},{"family":"Gjoreski","given":"Martin"},{"family":"Langheinrich","given":"Marc"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.30161","URL":"https://doi.org/10.48550/arxiv.2606.30161","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.28493","type":"manuscript","title":"The Role of Artificial Intelligence in the SKA Era","abstract":"The Square Kilometre Array Observatory (SKAO) will usher in an era of unprecedented data complexity and scientific opportunity in radio astronomy, producing petabyte-scale datasets and terabit-per-second streams that challenge traditional analysis paradigms. Artificial Intelligence (AI) stands at the forefront of this transformation, offering scalable, adaptive solutions to the most pressing problems in radio astronomy and astrophysics. This chapter explores the pivotal role of AI in the SKA era, from real-time operations to scientific discovery. We examine how deep learning models enable automated source detection, radio-frequency interference mitigation, anomaly detection, and parameter inference, while generative approaches accelerate sky simulations, calibration, and imaging. Reinforcement learning promises dynamic scheduling and autonomous system control, and federated learning could address the distributed nature of SKA data. Beyond performance, we emphasize the necessity of explainability, uncertainty quantification, and physics-informed inductive biases to ensure scientific integrity. By mapping SKAO's core challenges - data volume, complexity, and interpretability - onto modern AI methodologies, we review how deep learning, self-supervised frameworks, and probabilistic models can unlock new frontiers in cosmology, galaxy evolution, and time-domain astrophysics. AI is not merely an automation tool for coping with scale. It is a catalyst for discovery, redefining how we observe, model, and understand the Universe.","author":[{"family":"Denzel","given":"Philipp"},{"family":"Schilling","given":"Frank"},{"family":"Gavagnin","given":"Elena"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.28493","URL":"https://doi.org/10.48550/arxiv.2606.28493","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.27511","type":"manuscript","title":"When the Aggregator Cheats: Data-Free Backdoors in Federated LLM-based QA Systems","abstract":"Large Language Model (LLM)-based question-answering (QA) systems are increasingly deployed in sensitive domains such as healthcare, mental health counseling, and legal consultation. Federated learning (FL) enables collaborative training without sharing raw client data, for which locally trained models are aggregated at a central server (i.e., a cloud service provider) to obtain a global model. In this paper, we explore the potential vulnerability where a malicious aggregator, who may collude with a third-party vendor, stealthily implants advertisement-type backdoors into federated QA models, without ever accessing client data. The attacker's goals are twofold: (1) preserve clean QA fidelity (i.e., the poisoned model behaves like a clean model on non-triggered queries); and (2) generate highly natural, contextually relevant responses with target advertisements when a trigger appears. Achieving these two goals simultaneously is highly challenging, as naive backdoor injection without knowledge about private data may degrade model's clean performance or fail to inject the target. Motivated by this, we propose to leverage clients' uploaded gradients during training, and develop a two-stage framework for data-free and stealthy poisoning: (1) recover representative training samples from client gradients, and (2) construct poisoning datasets utilizing recovered samples and trigger phrases to inject backdoors into the global model. Experiments across representative QA datasets and LLM families under full fine-tuning and LoRA settings demonstrate that, our method achieves nearly 100% Attack Success Rate (ASR) while incurring negligible degradation on clean tasks. Crucially, reconstructing only 5-20% of gradients suffices to mount a reliable attack, exposing a practical blind spot in the pipeline of federated training of QA LLMs.","author":[{"family":"Zhu","given":"Chenqing"},{"family":"Dai","given":"Yanbo"},{"family":"Tian","given":"Yulong"},{"family":"Li","given":"Qingming"},{"family":"Li","given":"Songze"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.27511","URL":"https://doi.org/10.48550/arxiv.2606.27511","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.22875","type":"manuscript","title":"FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs","abstract":"Training Latent Diffusion Models (LDMs) within Federated Learning (FL) has attracted increasing attention due to its ability to combine the powerful generative capacity of LDMs with the privacy-preserving properties of FL. However, FL requires sharing the global model with multiple participants, which risks unauthorized model distribution or resale by malicious clients. While an intuitive approach is to adopt existing VAE-based watermarking techniques for LDMs in FL, this strategy falls short in addressing such threats due to two fundamental challenges: (1) Existing methods support ownership verification but lack the ability to trace model leakage to a specific malicious client; (2) VAE-based watermarks are vulnerable, as they can be removed simply by replacing the decoder with a clean counterpart. In this paper, we propose FedOT, the first framework for ownership verification and leakage tracing in federated LDMs. Specifically, to address the first challenge, we design a chunked watermark, where the first part is for ownership verification, and the second part is used for client identification. Furthermore, to overcome the second challenge and secure the model against VAE replacement attack, we introduce Latent Vector Transformation (LVT), which strengthens the connection between the VAE and U-Net latent spaces by modifying the original latent distribution of the VAE. Consequently, any attempt to replace the VAE for watermark removal leads to significant image quality degradation, making the LDM model unusable. Extensive experiments demonstrate that FedOT achieves superior performance in both ownership verification and traceability. Project page: https://spyzixuan.github.io/FedOT/.","author":[{"family":"Cheng","given":"Wenlong"},{"family":"Gan","given":"Yuan"},{"family":"Xu","given":"Yunqiu"},{"family":"Miao","given":"Jiaxu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.22875","URL":"https://doi.org/10.48550/arxiv.2606.22875","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.19734","type":"manuscript","title":"Federated Bilevel Performative Prediction","abstract":"Federated bilevel optimization is widely used for nested learning problems across distributed clients, such as federated hyperparameter tuning and meta-learning under privacy and communication constraints. Most existing formulations assume fixed client data distributions, which can be violated by performativity, where deployed decisions reshape client behavior and data collection, inducing client-specific, decision-dependent distribution shift. We study federated bilevel performative prediction, where both upper-level (UL) and lower-level (LL) objectives are evaluated under client-dependent, decision-dependent distributions. We formalize the federated bilevel performatively stable (FBPS) point under a decoupled-risk perspective and provide sufficient conditions for its existence and uniqueness. We then develop two federated methods to compute the FBPS solution: FBi-RRM, which converges linearly under a contraction condition, and FBi-SGD, a communication-efficient stochastic method based on federated hypergradient estimation with convergence guarantees under diminishing step sizes when sensitivities are sufficiently small. Experiments on strategic regression and meta strategic classification validate the predicted stability thresholds and demonstrate improved meta-generalization over non-performative baselines, and CNN-based classification further demonstrates the practical effectiveness of the proposed methods in nonconvex neural network settings.","author":[{"family":"Qian","given":"Liangxin"},{"family":"Liu","given":"Chang"},{"family":"Cao","given":"Xuanyu"},{"family":"Zhao","given":"Jun"},{"family":"Lam","given":"Kwok"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.19734","URL":"https://doi.org/10.48550/arxiv.2606.19734","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.16655","type":"manuscript","title":"Distribution Alignment for One-Shot Federated Learning via Optimal Transport","abstract":"One-Shot Federated Learning (OSFL) addresses extreme communication regimes in which clients interact with the server only once, amplifying the impact of heterogeneous client data distributions. In particular, the interaction of domain shift and label shift across clients induces misaligned feature representations that cannot be corrected through iterative optimization. Existing OSFL methods rely on distillation, server-side generation or ensemble-based aggregation, but assume aligned representations or address domain and label shift separately. We introduce SLOT-Align (Single-round, Learning-free Optimal Transport Alignment), a geometry-aware feature harmonization framework for OSFL. SLOT-Align uses a shared frozen encoder to extract compact feature statistics, constructs a global reference via Bures-Wasserstein barycenters, and aligns local representations using closed-form geodesic optimal transport maps. The method is computationally efficient and can be combined with existing OSFL pipelines relying on frozen encoders without modifying their training procedures. Extensive experiments across multiple benchmarks, pretrained backbones, and OSFL methods show that SLOT-Align consistently improves accuracy and robustness under joint domain and label shift.","author":[{"family":"Berardini","given":"Daniele"},{"family":"Pastore","given":"Vito"},{"family":"Murino","given":"Vittorio"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.16655","URL":"https://doi.org/10.48550/arxiv.2606.16655","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.13748","type":"manuscript","title":"FedSPC: Shared Parameter Correction for Personalized Federated Learning","abstract":"Personalized federated learning (PFL) is one of the important approaches in federated learning for addressing statistical heterogeneity while enabling client-specific adaptation. Many PFL methods split the model into shared and personalized parameters, which are jointly trained on each client. However, this creates an optimization issue: shared parameters are updated by clients optimizing different local objectives, which can lead to inconsistent shared updates and weaken the shared representation. To address this problem, we propose Federated Shared Parameter Correction (FedSPC), a modular correction method for PFL. FedSPC applies control-variate correction only to the shared parameters of a given PFL method, while leaving personalized parameters unchanged. It can be integrated into three common PFL settings: shared feature extractors, shared classifiers, and fully shared models with local regularization. Experiments on CIFAR-100 and Tiny-ImageNet with ViT, ResNet-34, and VGG-11 show that FedSPC improves performance across representative PFL methods, including FedPer, FedRep, FedBABU, LG-FedAvg, and Ditto.","author":[{"family":"Menon","given":"Kannanthodath"},{"family":"Prehofer","given":"Christian"},{"family":"Xu","given":"Yunfei"},{"family":"Hirano","given":"Toru"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.13748","URL":"https://doi.org/10.48550/arxiv.2606.13748","source":"datacite"},{"id":"doi:10.48550/arxiv.2601.01901","type":"manuscript","title":"FedBiCross: Personalized One-Shot Federated Learning on Medical Images","abstract":"Data-free knowledge distillation-based one-shot federated learning (OSFL) trains a model in a single communication round without sharing raw data, making OSFL attractive for privacy-sensitive medical applications. However, existing methods aggregate predictions from all clients to form a global teacher. Under non-IID data, conflicting predictions dilute each other during averaging, yielding less informative soft labels that weaken distillation. We propose FedBiCross, a personalized OSFL framework with three stages: (1) clustering clients by model output similarity to form coherent sub-ensembles, (2) bi-level cross-cluster optimization that learns adaptive weights to selectively leverage beneficial cross-cluster knowledge while suppressing negative transfer, and (3) personalized distillation for client-specific adaptation. Experiments on four medical image datasets demonstrate that FedBiCross consistently outperforms state-of-the-art baselines across different non-IID degrees.","author":[{"family":"Xia","given":"Yuexuan"},{"family":"Zhang","given":"Yinghao"},{"family":"Liu","given":"Yalin"},{"family":"Dai","given":"Hong"},{"family":"Xia","given":"Yong"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2601.01901","URL":"https://doi.org/10.48550/arxiv.2601.01901","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.29933","type":"manuscript","title":"GreenFLag: A Green Agentic Approach for Energy-Efficient Federated Learning","abstract":"Progressing toward a new generation of mobile networks, a clear focus on integrating distributed intelligence across the system is observed to drive performance, autonomy, and real-time adaptability. Federated learning (FL) stands out as a key emerging technique, enabling on-device model training while preserving data locality. However, its operation introduces substantial energy and resource demands. Energy needs are mostly met by grid power sources, while FL resource orchestration strategies remain limited. This work introduces GreenFLag, an agentic resource orchestration framework designed to minimize the energy consumption from the grid power to complete FL workflows, guarantee FL model performance, and reduce grid power reliance by incorporating renewable sources into the system. GreenFLag leverages a Soft-Actor Critic reinforcement learning approach to jointly optimize computational and communication resources, while accounting for communication contention and the dynamic availability of renewable energy. Evaluations using a real-world open dataset from Copernicus, demonstrate that GreenFLag significantly reduces grid energy consumption by 94.8% on average, compared to three state-of-the-art baselines, while primarily relying on green power.","author":[{"family":"Panagea","given":"Theodora"},{"family":"Koursioumpas","given":"Nikolaos"},{"family":"Magoula","given":"Lina"},{"family":"Khalili","given":"Ramin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.29933","URL":"https://doi.org/10.48550/arxiv.2603.29933","source":"datacite"},{"id":"doi:10.48550/arxiv.2505.15124","type":"manuscript","title":"A Survey On Secure Machine Learning","abstract":"In this survey, we will explore the interaction between secure multiparty computation and the area of machine learning. Recent advances in secure multiparty computation (MPC) have significantly improved its applicability in the realm of machine learning (ML), offering robust solutions for privacy-preserving collaborative learning. This review explores key contributions that leverage MPC to enable multiple parties to engage in ML tasks without compromising the privacy of their data. The integration of MPC with ML frameworks facilitates the training and evaluation of models on combined datasets from various sources, ensuring that sensitive information remains encrypted throughout the process. Innovations such as specialized software frameworks and domain-specific languages streamline the adoption of MPC in ML, optimizing performance and broadening its usage. These frameworks address both semi-honest and malicious threat models, incorporating features such as automated optimizations and cryptographic auditing to ensure compliance and data integrity. The collective insights from these studies highlight MPC's potential in fostering collaborative yet confidential data analysis, marking a significant stride towards the realization of secure and efficient computational solutions in privacy-sensitive industries. This paper investigates a spectrum of SecureML libraries that includes cryptographic protocols, federated learning frameworks, and privacy-preserving algorithms. By surveying the existing literature, this paper aims to examine the efficacy of these libraries in preserving data privacy, ensuring model confidentiality, and fortifying ML systems against adversarial attacks. Additionally, the study explores an innovative application domain for SecureML techniques: the integration of these methodologies in gaming environments utilizing ML.","author":[{"family":"Liao","given":"Taobo"},{"family":"Li","given":"Taoran"},{"family":"Nadkarni","given":"Prathamesh"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2505.15124","URL":"https://doi.org/10.48550/arxiv.2505.15124","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.21895","type":"manuscript","title":"PrivDNN: A Secure Multi-Party Computation Framework for Deep Learning using Partial DNN Encryption","abstract":"In the past decade, we have witnessed an exponential growth of deep learning models, platforms, and applications. While existing DL applications and Machine Learning as a service (MLaaS) frameworks assume fully trusted models, the need for privacy-preserving DNN evaluation arises. In a secure multi-party computation scenario, both the model and the data are considered proprietary, i.e., the model owner does not want to reveal the highly valuable DL model to the user, while the user does not wish to disclose their private data samples either. Conventional privacy-preserving deep learning solutions ask the users to send encrypted samples to the model owners, who must handle the heavy lifting of ciphertext-domain computation with homomorphic encryption. In this paper, we present a novel solution, namely, PrivDNN, which (1) offloads the computation to the user side by sharing an encrypted deep learning model with them, (2) significantly improves the efficiency of DNN evaluation using partial DNN encryption, (3) ensures model accuracy and model privacy using a core neuron selection and encryption scheme. Experimental results show that PrivDNN reduces privacy-preserving DNN inference time and memory requirement by up to 97% while maintaining model performance and privacy. Codes can be found at https://github.com/LiangqinRen/PrivDNN","author":[{"family":"Ren","given":"Liangqin"},{"family":"Liu","given":"Zeyan"},{"family":"Li","given":"Fengjun"},{"family":"Liang","given":"Kaitai"},{"family":"Li","given":"Zhu"},{"family":"Luo","given":"Bo"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.21895","URL":"https://doi.org/10.48550/arxiv.2607.21895","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.03191","type":"manuscript","title":"Private Embedding Lookup with Encrypted Compact Queries under Fully Homomorphic Encryption","abstract":"Many NLP or recommendation models begin by mapping discrete client inputs to embedding vectors. Since inputs can reveal sensitive information, the embedding step must be protected in privacy-preserving inference. Fully Homomorphic Encryption (FHE) enables inference over encrypted client data, but turns embedding lookup from simple table access into homomorphic computation. To keep the embedding table server-side and avoid transmitting encrypted embedding vectors from the client, we focus on server-side lookup: the client sends only a small encrypted index. Prior ICML 2024 work first builds a one-hot vector from the encrypted index before multiplying with the embedding table, and this one-hot generation is the dominant cost. One-hot-based methods are expensive in FHE: they construct a p-dimensional selection vector via an equality test for each coordinate, requiring $O(p \\log p)$ total homomorphic operations. Our key observation is that private embedding lookup only requires a linearly independent representation of the encrypted index, not the one-hot basis itself. Building on it, we propose Independent Vector Evaluation (IVE). Instead of constructing a one-hot vector, IVE evaluates a linearly independent vector built from successive powers of a single encrypted value, reducing vector-generation cost to $O(p)$. It then recovers the same embedding vector via a precomputed change of basis, instantiated with an orthogonal Discrete Cosine Transform to mitigate error amplification. Our implementation shows IVE improves amortized lookup time by up to 78.4x over prior method. We further evaluate its impact on end-to-end encrypted FastText inference, where embedding lookup is a major cost in the shallow model. On Enron-Spam dataset, replacing one-hot generation with IVE reduces the share of vector generation in encrypted inference time from 99.6% to 66.3%.","author":[{"family":"Cheon","given":"Jung"},{"family":"Jang","given":"Daehyun"},{"family":"Kang","given":"Jaehee"},{"family":"Rhee","given":"Hanee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.03191","URL":"https://doi.org/10.48550/arxiv.2606.03191","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.08400","type":"manuscript","title":"Classical Obfuscation of Quantum Circuits via Publicly-Verifiable QFHE","abstract":"A classical obfuscator for quantum circuits is a classical program that, given the classical description of a quantum circuit $Q$, outputs the classical description of a functionally equivalent quantum circuit $\\hat{Q}$ that hides as much as possible about $Q$. Previously, the only known feasibility result for classical obfuscation of quantum circuits (Bartusek and Malavolta, ITCS 2022) was limited to circuits that always reject. On the other hand, if the obfuscator is allowed to compile the quantum circuit $Q$ into a quantum state $|\\hat{Q}\\rangle$, there exist feasibility results for obfuscating all pseudo-deterministic quantum circuits (Bartusek, Kitagawa, Nishimaki and Yamakawa, STOC 2023, Bartusek, Brakerski and Vaikuntanathan, STOC 2024), and all unitaries (Huang and Tang, FOCS 2025). We show that (relative to a classical oracle) there exists a classical obfuscator for all pseudo-deterministic quantum circuits. We do this by giving the first construction of a compact quantum fully-homomorphic encryption (QFHE) scheme that supports public verification of (pseudo-deterministic) quantum evaluation, relative to a classical oracle. To construct our QFHE scheme, we improve on the approach of Bartusek, Kitagawa, Nishimaki and Yamakawa (STOC 2023), which required ciphertexts that are both quantum and non-compact due to the use of quantum coset states and their publicly-verifiable properties. We introduce new techniques for analyzing coset states that can be generated ''on the fly'', by proving new cryptographic properties of the one-shot signature scheme of Shmueli and Zhandry (CRYPTO 2025). Our techniques allow us to produce QFHE ciphertexts that are purely classical, compact, and publicly-verifiable. This also yields the first classical verification of quantum computation protocol for BQP that simultaneously satisfies blindness and public-verifiability.","author":[{"family":"Bartusek","given":"James"},{"family":"Gupte","given":"Aparna"},{"family":"Mutreja","given":"Saachi"},{"family":"Shmueli","given":"Omri"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.08400","URL":"https://doi.org/10.48550/arxiv.2510.08400","source":"datacite"},{"id":"doi:10.5281/zenodo.16901364","type":"article-journal","title":"The First IEEE World Technology Summit 2024: AI Infrastructure Importance and Challenges","abstract":"The inaugural IEEE World Technology Summit 2024 brought together global leaders in technology, academia, and industry to explore the transformative intersections of artificial intelligence (AI), energy systems, quantum computing, and governance. Central themes included the challenges of scaling AI infrastructure, achieving sustainable AI development, and fostering global standards for secure and ethical AI deployment. Highlights from the summit showcased advancements in industrial AI applications, digital twins, semiconductor design, and energy-efficient systems. Key insights included leveraging AI for predictive maintenance, optimizing chip design through reinforcement learning, and integrating digital twins with generative AI to revolutionize industrial workflows. Presentations underscored the growing energy demands of AI, with strategies such as liquid cooling, energy-efficient data centers, and carbon-neutral energy sources proposed to address sustainability challenges. Standards and security emerged as critical pillars, with discussions on post-quantum cryptography, fully homomorphic encryption, and AI governance frameworks. The summit’s multidisciplinary approach emphasized the convergence of innovation and ethical responsibility, setting the stage for AI’s transformative potential in reshaping industries and addressing global challenges. By fostering collaboration and advancing technical solutions, the IEEE World Technology Summit highlighted the pivotal role of AI in driving sustainable, inclusive, and secure technological progress.","author":[{"family":"Vyatkin","given":"Valeriy"},{"family":"Huang","given":"Victor"},{"family":"Condry","given":"Michael"},{"family":"Nihtianov","given":"Stoyan"},{"family":"Karnouskos","given":"Stamatis"},{"family":"Manic","given":"Milos"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16901364","URL":"https://doi.org/10.5281/zenodo.16901364","source":"datacite"},{"id":"doi:10.5281/zenodo.16901365","type":"article-journal","title":"The First IEEE World Technology Summit 2024: AI Infrastructure Importance and Challenges","abstract":"The inaugural IEEE World Technology Summit 2024 brought together global leaders in technology, academia, and industry to explore the transformative intersections of artificial intelligence (AI), energy systems, quantum computing, and governance. Central themes included the challenges of scaling AI infrastructure, achieving sustainable AI development, and fostering global standards for secure and ethical AI deployment. Highlights from the summit showcased advancements in industrial AI applications, digital twins, semiconductor design, and energy-efficient systems. Key insights included leveraging AI for predictive maintenance, optimizing chip design through reinforcement learning, and integrating digital twins with generative AI to revolutionize industrial workflows. Presentations underscored the growing energy demands of AI, with strategies such as liquid cooling, energy-efficient data centers, and carbon-neutral energy sources proposed to address sustainability challenges. Standards and security emerged as critical pillars, with discussions on post-quantum cryptography, fully homomorphic encryption, and AI governance frameworks. The summit’s multidisciplinary approach emphasized the convergence of innovation and ethical responsibility, setting the stage for AI’s transformative potential in reshaping industries and addressing global challenges. By fostering collaboration and advancing technical solutions, the IEEE World Technology Summit highlighted the pivotal role of AI in driving sustainable, inclusive, and secure technological progress.","author":[{"family":"Vyatkin","given":"Valeriy"},{"family":"Huang","given":"Victor"},{"family":"Condry","given":"Michael"},{"family":"Nihtianov","given":"Stoyan"},{"family":"Karnouskos","given":"Stamatis"},{"family":"Manic","given":"Milos"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.16901365","URL":"https://doi.org/10.5281/zenodo.16901365","source":"datacite"},{"id":"doi:10.5281/zenodo.21849330","type":"article-journal","title":"A Privacy-Preserving Explainable Artificial Intelligence Hybrid Framework for Abnormal Network Traffic Identification and Intelligent Threat Detection","abstract":"Abstract The rapid evolution of cyber threats, the widespread adoption of encrypted communications, and the increasing complexity of enterprise, cloud, edge, and Internet of Things (IoT) environments have exposed the limitations of conventional intrusion detection systems in accurately identifying abnormal network traffic and detecting sophisticated cyberattacks. Existing hybrid deep learning frameworks, including that of Wang [54], achieve high detection accuracy but lack privacy-preserving computation, explainable artificial intelligence, intelligent threat prioritization, and adaptive operational capabilities. This study therefore designed and developed a Hybrid Artificial Intelligence Framework for Abnormal Network Traffic Identification and Intelligent Threat Detection by integrating Long Short-Term Memory (LSTM), Transformer, Random Forest, CKKS Homomorphic Encryption, Explainable Artificial Intelligence (SHAP/LIME), weighted ensemble decision fusion, intelligent threat prioritization, and adaptive feedback learning. The study adopted the Design Science Research Methodology (DSRM), while the proposed framework was implemented using Python, TensorFlow/Keras, Scikit-learn, Microsoft SEAL/TenSEAL, and evaluated using the CICIDS2017 and UNSW-NB15 benchmark datasets. Experimental evaluation was performed using accuracy, precision, recall, F1-score, false positive rate, detection latency, zero-day detection rate, throughput, scalability, explainability, and encryption overhead as performance metrics. The proposed framework achieved detection accuracies of 99.12% and 98.76% on the CICIDS2017 and UNSW-NB15 datasets, respectively, with precision values of 98.87% and 98.42%, recall values of 99.05% and 98.61%, F1-scores of 98.96% and 98.51%, false positive rates of 0.84% and 1.12%, and zero-day detection rates of 94.30% and 92.75%. Comparative analysis demonstrated improved detection accuracy, lower false-positive rates, reduced detection latency, enhanced interpretability, and stronger privacy preservation compared with the hybrid CNN–LSTM–Transformer framework of Wang [54]. The study concludes that integrating hybrid artificial intelligence, privacy-preserving computation, explainable artificial intelligence, and intelligent threat prioritization provides a robust, scalable, and adaptive solution for modern cybersecurity. The proposed framework is recommended for deployment in enterprise networks, cloud computing, IoT, edge computing, and critical infrastructure environments to strengthen real-time cyber threat detection and response. Keywords: Hybrid Artificial Intelligence, Abnormal Network Traffic Identification, Intelligent Threat Detection, Intrusion Detection System, LSTM, Transformer, Random Forest, Explainable Artificial Intelligence, CKKS Homomorphic Encryption, Zero-Day Attack Detection","author":[{"family":"Onuma","given":"DA"},{"family":"Matthias","given":"D"},{"family":"Taylor","given":"OE"},{"family":"Nwiabu","given":"ND"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21849330","URL":"https://doi.org/10.5281/zenodo.21849330","source":"datacite"},{"id":"doi:10.5281/zenodo.21849331","type":"article-journal","title":"A Privacy-Preserving Explainable Artificial Intelligence Hybrid Framework for Abnormal Network Traffic Identification and Intelligent Threat Detection","abstract":"Abstract The rapid evolution of cyber threats, the widespread adoption of encrypted communications, and the increasing complexity of enterprise, cloud, edge, and Internet of Things (IoT) environments have exposed the limitations of conventional intrusion detection systems in accurately identifying abnormal network traffic and detecting sophisticated cyberattacks. Existing hybrid deep learning frameworks, including that of Wang [54], achieve high detection accuracy but lack privacy-preserving computation, explainable artificial intelligence, intelligent threat prioritization, and adaptive operational capabilities. This study therefore designed and developed a Hybrid Artificial Intelligence Framework for Abnormal Network Traffic Identification and Intelligent Threat Detection by integrating Long Short-Term Memory (LSTM), Transformer, Random Forest, CKKS Homomorphic Encryption, Explainable Artificial Intelligence (SHAP/LIME), weighted ensemble decision fusion, intelligent threat prioritization, and adaptive feedback learning. The study adopted the Design Science Research Methodology (DSRM), while the proposed framework was implemented using Python, TensorFlow/Keras, Scikit-learn, Microsoft SEAL/TenSEAL, and evaluated using the CICIDS2017 and UNSW-NB15 benchmark datasets. Experimental evaluation was performed using accuracy, precision, recall, F1-score, false positive rate, detection latency, zero-day detection rate, throughput, scalability, explainability, and encryption overhead as performance metrics. The proposed framework achieved detection accuracies of 99.12% and 98.76% on the CICIDS2017 and UNSW-NB15 datasets, respectively, with precision values of 98.87% and 98.42%, recall values of 99.05% and 98.61%, F1-scores of 98.96% and 98.51%, false positive rates of 0.84% and 1.12%, and zero-day detection rates of 94.30% and 92.75%. Comparative analysis demonstrated improved detection accuracy, lower false-positive rates, reduced detection latency, enhanced interpretability, and stronger privacy preservation compared with the hybrid CNN–LSTM–Transformer framework of Wang [54]. The study concludes that integrating hybrid artificial intelligence, privacy-preserving computation, explainable artificial intelligence, and intelligent threat prioritization provides a robust, scalable, and adaptive solution for modern cybersecurity. The proposed framework is recommended for deployment in enterprise networks, cloud computing, IoT, edge computing, and critical infrastructure environments to strengthen real-time cyber threat detection and response. Keywords: Hybrid Artificial Intelligence, Abnormal Network Traffic Identification, Intelligent Threat Detection, Intrusion Detection System, LSTM, Transformer, Random Forest, Explainable Artificial Intelligence, CKKS Homomorphic Encryption, Zero-Day Attack Detection","author":[{"family":"Anele","given":"Onuma"},{"family":"Matthias","given":"Prof"},{"family":"Taylor","given":"Dr"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21849331","URL":"https://doi.org/10.5281/zenodo.21849331","source":"datacite"},{"id":"doi:10.5281/zenodo.20608454","type":"article-journal","title":"Two-Way Confidential VMs (2cVM): Collaborative Confidential Computing for Mutually Distrustful Parties","abstract":"Collaborative computation across organizations is often constrained by the need to process sensitive data and proprietary code without exposing them to untrusted infrastructure or participants. Cryptographic approaches such as fully homomorphic encryption and secure multi-party computation provide strong confidentiality but remain impractical for general workloads due to their extreme com- putational cost. We present the Two-Way Confidential Virtual Machine (2cVM), a two-layer architecture that pairs a hardware trusted execution environment with an intra-workload isolation layer. Unlike regular Confidential Virtual Machines, 2cVM enforces mutual isolation between co-resident workloads, ensuring that participants retain control over their data and code. All computation in 2cVM is governed by a Commitment Manifest that enumerates participants, component composition, permitted data channels, and authorized outputs; the manifest is locked to the VM and incorporated into attestation evidence, making the policy immutable and independently verifiable throughout the VM’s lifetime. A proof-of-concept realization combines AMD SEV-SNP for hardware protection with the WebAssembly Component Model for fine-grained sandboxing of participant code. Evaluation on commodity hardware across four benchmark classes shows that the two isolation layers do not accumulate linearly: once a workload executes inside the WebAssembly sandbox, the marginal cost of enabling hardware memory protection is small. Overhead is workload-dependent, governed primarily by memory access pattern, ranging from negligible for sequential workloads to approximately 2×for irregular, pointer-chasing access patterns. These results indicate that 2cVM provides a practical and verifiable foundation for privacy-preserving collaborative computation.","author":[{"family":"Thijsman","given":"Jordi"},{"family":"Sebrechts","given":"Merlijn"},{"family":"Lefever","given":"Stefan"},{"family":"De Turck","given":"Filip"},{"family":"Volckaert","given":"Bruno"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20608454","URL":"https://doi.org/10.5281/zenodo.20608454","source":"datacite"},{"id":"doi:10.5281/zenodo.20610780","type":"article-journal","title":"Two-Way Confidential VMs (2cVM): Collaborative Confidential Computing for Mutually Distrustful Parties","abstract":"Collaborative computation across organizations is often constrained by the need to process sensitive data and proprietary code without exposing them to untrusted infrastructure or participants. Cryptographic approaches such as fully homomorphic encryption and secure multi-party computation provide strong confidentiality but remain impractical for general workloads due to their extreme com- putational cost. We present the Two-Way Confidential Virtual Machine (2cVM), a two-layer architecture that pairs a hardware trusted execution environment with an intra-workload isolation layer. Unlike regular Confidential Virtual Machines, 2cVM enforces mutual isolation between co-resident workloads, ensuring that participants retain control over their data and code. All computation in 2cVM is governed by a Commitment Manifest that enumerates participants, component composition, permitted data channels, and authorized outputs; the manifest is locked to the VM and incorporated into attestation evidence, making the policy immutable and independently verifiable throughout the VM’s lifetime. A proof-of-concept realization combines AMD SEV-SNP for hardware protection with the WebAssembly Component Model for fine-grained sandboxing of participant code. Evaluation on commodity hardware across four benchmark classes shows that the two isolation layers do not accumulate linearly: once a workload executes inside the WebAssembly sandbox, the marginal cost of enabling hardware memory protection is small. Overhead is workload-dependent, governed primarily by memory access pattern, ranging from negligible for sequential workloads to approximately 2×for irregular, pointer-chasing access patterns. These results indicate that 2cVM provides a practical and verifiable foundation for privacy-preserving collaborative computation.","author":[{"family":"Thijsman","given":"Jordi"},{"family":"Sebrechts","given":"Merlijn"},{"family":"Lefever","given":"Stefan"},{"family":"De Turck","given":"Filip"},{"family":"Volckaert","given":"Bruno"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20610780","URL":"https://doi.org/10.5281/zenodo.20610780","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.21131","type":"manuscript","title":"Billion-Scale Nearest-Neighbor Search under Fully Homomorphic Encryption on a Single GPU, Balancing Leakage and Cost","abstract":"We build a system that answers \"which database vectors are most similar to my query?\" without the server ever seeing the query. The query is encrypted with fully homomorphic en- cryption (FHE); the server does all its scoring on ciphertexts and returns encrypted results that only the client can read. The challenge is speed: at a billion vectors, scoring every row under encryption is far too slow, so we combine two ideas - rank reduction (shrink each vector's dimen- sion) and a hierarchy (route to a small candidate set instead of scanning everything) - executed under encryption on a single GPU. We evaluate on three corpora at very different scales: a face corpus of 222 049 centroids clustered from ~10 M face images (512-dim), DataComp-1B (1.39 x 10^9 vectors, 512-dim CLIP), and Deep1B (10^9 vectors, 96-dim). On DataComp-1B we reach a recall@10 of 0.90 against the single labeled answer, or 0.95 when a near-duplicate im- age in the top-10 also counts as correct (the data is web-scraped and full of duplicates), at ~6 s per encrypted query on a GPU; a lighter configuration reaches 0.78/0.83 at ~1.8 s. These are warm (deployable) server-side latencies - client decryption and network transfer are excluded. On Deep1B we reach recall@10 0.90 under all-levels FHE (0.9045 measured over 2000 FHE queries, matching the 0.906 plaintext routing - the 96 -&gt; 128 zero-pad is exact, correlation 1.0) at 2.3 s warm per query. We describe the full client-server protocol in enough detail to repro- duce it, and report accuracy and latency for every configuration. We also measure what this speed costs: the hierarchy's access pattern leaks the database geometry (an observer recovers 72% of the coarse-cell neighbor graph from access patterns alone), and we show that seeded (fixed-group) padding cuts this leak by ~35x (to ~2%), where naive padding is defeated by a repeated-query attack.","author":[{"family":"Isozaki","given":"Isamu"},{"family":"Bratina","given":"Madison"},{"family":"Kim","given":"Edward"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.21131","URL":"https://doi.org/10.48550/arxiv.2608.21131","source":"datacite"},{"id":"doi:10.5281/zenodo.20608453","type":"article-journal","title":"Two-Way Confidential VMs (2cVM): Collaborative Confidential Computing for Mutually Distrustful Parties","abstract":"Collaborative computation across organizations is often constrained by the need to process sensitive data and proprietary code without exposing them to untrusted infrastructure or participants. Cryptographic approaches such as fully homomorphic encryption and secure multi-party computation provide strong confidentiality but remain impractical for general workloads due to their extreme com- putational cost. We present the Two-Way Confidential Virtual Machine (2cVM), a two-layer architecture that pairs a hardware trusted execution environment with an intra-workload isolation layer. Unlike regular Confidential Virtual Machines, 2cVM enforces mutual isolation between co-resident workloads, ensuring that participants retain control over their data and code. All computation in 2cVM is governed by a Commitment Manifest that enumerates participants, component composition, permitted data channels, and authorized outputs; the manifest is locked to the VM and incorporated into attestation evidence, making the policy immutable and independently verifiable throughout the VM’s lifetime. A proof-of-concept realization combines AMD SEV-SNP for hardware protection with the WebAssembly Component Model for fine-grained sandboxing of participant code. Evaluation on commodity hardware across four benchmark classes shows that the two isolation layers do not accumulate linearly: once a workload executes inside the WebAssembly sandbox, the marginal cost of enabling hardware memory protection is small. Overhead is workload-dependent, governed primarily by memory access pattern, ranging from negligible for sequential workloads to approximately 2×for irregular, pointer-chasing access patterns. These results indicate that 2cVM provides a practical and verifiable foundation for privacy-preserving collaborative computation.","author":[{"family":"Thijsman","given":"Jordi"},{"family":"Sebrechts","given":"Merlijn"},{"family":"Lefever","given":"Stefan"},{"family":"De Turck","given":"Filip"},{"family":"Volckaert","given":"Bruno"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20608453","URL":"https://doi.org/10.5281/zenodo.20608453","source":"datacite"},{"id":"doi:10.5281/zenodo.21992945","type":"article-journal","title":"A Privacy-Preserving Explainable Artificial Intelligence Hybrid Framework for Abnormal Network Traffic Identification and Intelligent Threat Detection","abstract":"Abstract The rapid evolution of cyber threats, the widespread adoption of encrypted communications, and the increasing complexity of enterprise, cloud, edge, and Internet of Things (IoT) environments have exposed the limitations of conventional intrusion detection systems in accurately identifying abnormal network traffic and detecting sophisticated cyberattacks. Existing hybrid deep learning frameworks, including that of Wang [54], achieve high detection accuracy but lack privacy-preserving computation, explainable artificial intelligence, intelligent threat prioritization, and adaptive operational capabilities. This study therefore designed and developed a Hybrid Artificial Intelligence Framework for Abnormal Network Traffic Identification and Intelligent Threat Detection by integrating Long Short-Term Memory (LSTM), Transformer, Random Forest, CKKS Homomorphic Encryption, Explainable Artificial Intelligence (SHAP/LIME), weighted ensemble decision fusion, intelligent threat prioritization, and adaptive feedback learning. The study adopted the Design Science Research Methodology (DSRM), while the proposed framework was implemented using Python, TensorFlow/Keras, Scikit-learn, Microsoft SEAL/TenSEAL, and evaluated using the CICIDS2017 and UNSW-NB15 benchmark datasets. Experimental evaluation was performed using accuracy, precision, recall, F1-score, false positive rate, detection latency, zero-day detection rate, throughput, scalability, explainability, and encryption overhead as performance metrics. The proposed framework achieved detection accuracies of 99.12% and 98.76% on the CICIDS2017 and UNSW-NB15 datasets, respectively, with precision values of 98.87% and 98.42%, recall values of 99.05% and 98.61%, F1-scores of 98.96% and 98.51%, false positive rates of 0.84% and 1.12%, and zero-day detection rates of 94.30% and 92.75%. Comparative analysis demonstrated improved detection accuracy, lower false-positive rates, reduced detection latency, enhanced interpretability, and stronger privacy preservation compared with the hybrid CNN–LSTM–Transformer framework of Wang [54]. The study concludes that integrating hybrid artificial intelligence, privacy-preserving computation, explainable artificial intelligence, and intelligent threat prioritization provides a robust, scalable, and adaptive solution for modern cybersecurity. The proposed framework is recommended for deployment in enterprise networks, cloud computing, IoT, edge computing, and critical infrastructure environments to strengthen real-time cyber threat detection and response. Keywords: Hybrid Artificial Intelligence, Abnormal Network Traffic Identification, Intelligent Threat Detection, Intrusion Detection System, LSTM, Transformer, Random Forest, Explainable Artificial Intelligence, CKKS Homomorphic Encryption, Zero-Day Attack Detection","author":[{"family":"Onuma","given":"DA"},{"family":"Matthias","given":"D"},{"family":"Taylor","given":"OE"},{"family":"Nwiabu","given":"ND"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21992945","URL":"https://doi.org/10.5281/zenodo.21992945","source":"datacite"},{"id":"doi:10.5281/zenodo.19708361","type":"article-journal","title":"Federated Data Analytics in Smart Cities for Efficient Urban Intelligence","abstract":"The development of smart cities relies heavily on data-driven decision-making to improve urban services, enhance sustainability, and ensure efficient governance. With the rapid growth of IoT devices, sensors, and digital platforms, cities generate massive volumes of heterogeneous and sensitive data across domains such as healthcare, transportation, energy, and public safety. Centralized data analytics, while powerful, poses critical challenges related to privacy, security, bandwidth, and regulatory compliance. Federated Data Analytics (FDA) has emerged as a promising alternative, allowing decentralized model training and collaborative insights without the need to share raw data. This paper explores the role of FDA in the smart city ecosystem by reviewing its underlying methodologies, including federated learning frameworks, secure aggregation techniques, and privacy-preserving mechanisms such as homomorphic encryption and differential privacy. Key applications are examined in domains like traffic optimization, energy management, healthcare, and environmental monitoring. The study further highlights the technical, organizational, and ethical challenges hindering large-scale adoption, including data heterogeneity, communication overhead, governance issues, and legal constraints. Real-world use cases and pilot projects are analysed to demonstrate practical benefits and limitations. The findings suggest that FDA can balance innovation with privacy, enabling multi-stakeholder collaboration while safeguarding sensitive data. By integrating FDA with emerging technologies such as blockchain and edge computing, future smart cities can achieve secure, resilient, and citizen-centric urban development.","author":[{"family":"Dhome","given":"Shailesh"},{"family":"Nimbgoankar","given":"Prasad"},{"family":"Hakim","given":"Burhanoddin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19708361","URL":"https://doi.org/10.5281/zenodo.19708361","source":"datacite"},{"id":"doi:10.5281/zenodo.19708362","type":"article-journal","title":"Federated Data Analytics in Smart Cities for Efficient Urban Intelligence","abstract":"The development of smart cities relies heavily on data-driven decision-making to improve urban services, enhance sustainability, and ensure efficient governance. With the rapid growth of IoT devices, sensors, and digital platforms, cities generate massive volumes of heterogeneous and sensitive data across domains such as healthcare, transportation, energy, and public safety. Centralized data analytics, while powerful, poses critical challenges related to privacy, security, bandwidth, and regulatory compliance. Federated Data Analytics (FDA) has emerged as a promising alternative, allowing decentralized model training and collaborative insights without the need to share raw data. This paper explores the role of FDA in the smart city ecosystem by reviewing its underlying methodologies, including federated learning frameworks, secure aggregation techniques, and privacy-preserving mechanisms such as homomorphic encryption and differential privacy. Key applications are examined in domains like traffic optimization, energy management, healthcare, and environmental monitoring. The study further highlights the technical, organizational, and ethical challenges hindering large-scale adoption, including data heterogeneity, communication overhead, governance issues, and legal constraints. Real-world use cases and pilot projects are analysed to demonstrate practical benefits and limitations. The findings suggest that FDA can balance innovation with privacy, enabling multi-stakeholder collaboration while safeguarding sensitive data. By integrating FDA with emerging technologies such as blockchain and edge computing, future smart cities can achieve secure, resilient, and citizen-centric urban development.","author":[{"family":"Dhome","given":"Shailesh"},{"family":"Nimbgoankar","given":"Prasad"},{"family":"Hakim","given":"Burhanoddin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19708362","URL":"https://doi.org/10.5281/zenodo.19708362","source":"datacite"},{"id":"doi:10.5281/zenodo.19515793","type":"article-journal","title":"Cloud-Based e-Health Systems: A Comprehensive Review and Future Directions with Security and Privacy-Preserving Challenges","abstract":"Abstract Cloud computing has revolutionized e-health by enabling scalable storage and access to electronic health records (EHRs),By providing scalable storage and access to electronic health records (EHRs), cloud computing has transformed e-health. However, it also poses serious security and privacy threats that compromise patient confidence and legal compliance. In order to improve resilience against changing threats, this study carefully explores these issues, assesses current cryptographic and non-cryptographic solutions, and suggests a hybrid blockchain-integrated framework. We support proactive, patient-centric approaches to protect sensitive health data in multi-tenant cloud systems, arguing that present mechanisms are inadequate in tackling insider threats and data sovereignty challenges based on previous scholarly evaluations. The rapid adoption of cloud-based e-health systems promises enhanced interoperability, cost-efficiency, and data-driven diagnostics, but introduces formidable security and privacy-preserving challenges that threaten patient trust and regulatory adherence. This paper provides a comprehensive literature review of key threats—including data breaches, insider attacks, re-identification risks, and compliance conflicts under frameworks like GDPR and HIPAA—drawing from seminal works spanning 2017 to 2025. We critically evaluate cryptographic solutions (e.g., homomorphic encryption, attribute-based encryption), non-cryptographic approaches (e.g., differential privacy, role-based access controls), and emerging hybrids like blockchain-integrated architectures, highlighting their trade-offs in performance, scalability, and resilience. Keywords: Cloud-based e-health, Privacy-preserving techniques, Security challenges, Electronic health records (EHRs), Data confidentiality","author":[{"family":"Rkalaichelvan"},{"family":"Sgowthami"},{"family":"Kvanitha"},{"family":"Gsgeethamani"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19515793","URL":"https://doi.org/10.5281/zenodo.19515793","source":"datacite"},{"id":"doi:10.5281/zenodo.19515794","type":"article-journal","title":"Cloud-Based e-Health Systems: A Comprehensive Review and Future Directions with Security and Privacy-Preserving Challenges","abstract":"Abstract Cloud computing has revolutionized e-health by enabling scalable storage and access to electronic health records (EHRs),By providing scalable storage and access to electronic health records (EHRs), cloud computing has transformed e-health. However, it also poses serious security and privacy threats that compromise patient confidence and legal compliance. In order to improve resilience against changing threats, this study carefully explores these issues, assesses current cryptographic and non-cryptographic solutions, and suggests a hybrid blockchain-integrated framework. We support proactive, patient-centric approaches to protect sensitive health data in multi-tenant cloud systems, arguing that present mechanisms are inadequate in tackling insider threats and data sovereignty challenges based on previous scholarly evaluations. The rapid adoption of cloud-based e-health systems promises enhanced interoperability, cost-efficiency, and data-driven diagnostics, but introduces formidable security and privacy-preserving challenges that threaten patient trust and regulatory adherence. This paper provides a comprehensive literature review of key threats—including data breaches, insider attacks, re-identification risks, and compliance conflicts under frameworks like GDPR and HIPAA—drawing from seminal works spanning 2017 to 2025. We critically evaluate cryptographic solutions (e.g., homomorphic encryption, attribute-based encryption), non-cryptographic approaches (e.g., differential privacy, role-based access controls), and emerging hybrids like blockchain-integrated architectures, highlighting their trade-offs in performance, scalability, and resilience. Keywords: Cloud-based e-health, Privacy-preserving techniques, Security challenges, Electronic health records (EHRs), Data confidentiality","author":[{"family":"Rkalaichelvan"},{"family":"Sgowthami"},{"family":"Kvanitha"},{"family":"Gsgeethamani"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19515794","URL":"https://doi.org/10.5281/zenodo.19515794","source":"datacite"},{"id":"doi:10.5281/zenodo.21824885","type":"article-journal","title":"DEVELOPMENT OF AN ENHANCED BIT PLANE COMPLEXITY SEGMENTATION BASED HOMOMORPHIC ENCRYPTION FOR PRIVACY PRESERVATION IN DISTRIBUTED SYSTEM","abstract":"Privacy preservation in distributed systems faces persistent challenges in securing sensitive data while maintaining confidentiality and image quality. Conventional Bit Plane Complexity Segmentation (BPCS) techniques use fixed complexity thresholds, limiting embedding efficiency, and they do not protect embedded data if detected. This study developed an Enhanced BPCS based Homomorphic Encryption (EBPCS+HE) framework integrating an adaptive threshold-based region selection mechanism with the Paillier homomorphic encryption scheme to enhance data confidentiality and enable encrypted-domain operations. Implemented in MATLAB R2023a using thirty Chest X-ray images, the framework was evaluated using PSNR, payload ratio, encryption/decryption time, memory usage, and Shannon entropy. The developed framework achieved an average PSNR of 50.75 dB, payload ratio of 23.48%, embedding, encryption, and decryption times of 2.54 s, 1.80 s, and 1.41 s, respectively, and encryption memory usage of 2.92 MB. Comparative analysis showed improved performance over existing BPCS-based methods. Average Shannon entropy increased from 7.4168 in cover images to 7.8934 after encryption, indicating enhanced randomness and resistance to statistical attacks. The study concludes that the EBPCS+HE framework provides an effective solution for privacy preservation in distributed systems by improving embedding capacity, preserving image quality, strengthening data confidentiality, and enabling processing operations on encrypted data.","author":[{"family":"Kamaldeen","given":"Zubair"},{"family":"Olabiyisi","given":"Stephen"},{"family":"Alade","given":"Oluwaseun"},{"family":"Alabi","given":"Iretiolu"},{"family":"Olaitan","given":"Nurudeen"},{"family":"Seun","given":"Oyeranmi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21824885","URL":"https://doi.org/10.5281/zenodo.21824885","source":"datacite"},{"id":"doi:10.5281/zenodo.21824886","type":"article-journal","title":"DEVELOPMENT OF AN ENHANCED BIT PLANE COMPLEXITY SEGMENTATION BASED HOMOMORPHIC ENCRYPTION FOR PRIVACY PRESERVATION IN DISTRIBUTED SYSTEM","abstract":"Privacy preservation in distributed systems faces persistent challenges in securing sensitive data while maintaining confidentiality and image quality. Conventional Bit Plane Complexity Segmentation (BPCS) techniques use fixed complexity thresholds, limiting embedding efficiency, and they do not protect embedded data if detected. This study developed an Enhanced BPCS based Homomorphic Encryption (EBPCS+HE) framework integrating an adaptive threshold-based region selection mechanism with the Paillier homomorphic encryption scheme to enhance data confidentiality and enable encrypted-domain operations. Implemented in MATLAB R2023a using thirty Chest X-ray images, the framework was evaluated using PSNR, payload ratio, encryption/decryption time, memory usage, and Shannon entropy. The developed framework achieved an average PSNR of 50.75 dB, payload ratio of 23.48%, embedding, encryption, and decryption times of 2.54 s, 1.80 s, and 1.41 s, respectively, and encryption memory usage of 2.92 MB. Comparative analysis showed improved performance over existing BPCS-based methods. Average Shannon entropy increased from 7.4168 in cover images to 7.8934 after encryption, indicating enhanced randomness and resistance to statistical attacks. The study concludes that the EBPCS+HE framework provides an effective solution for privacy preservation in distributed systems by improving embedding capacity, preserving image quality, strengthening data confidentiality, and enabling processing operations on encrypted data.","author":[{"family":"Kamaldeen","given":"Zubair"},{"family":"Olabiyisi","given":"Stephen"},{"family":"Alade","given":"Oluwaseun"},{"family":"Alabi","given":"Iretiolu"},{"family":"Olaitan","given":"Nurudeen"},{"family":"Seun","given":"Oyeranmi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21824886","URL":"https://doi.org/10.5281/zenodo.21824886","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.13846","type":"manuscript","title":"Verified Pythagorean Composition for Adaptive Cryptographic Games: Noise Flooding in Homomorphic Encryption","abstract":"Noise flooding is a standard defense against decryption attacks on approximate homomorphic encryption, but its security proof is unusually sensitive to composition. Replacing each of $q$ adaptive decryption answers with a statistically close simulation and applying an ordinary hybrid argument loses linearly in $q$. The cryptographic proof instead accumulates conditional Kullback-Leibler (KL) costs and converts to statistical distance once, giving the parameter-critical square-root loss. We machine-check this argument using Rocq and SSProve. Given any fully homomorphic encryption scheme that is approximately correct and IND-CPA secure, we formalize a reduction for every $q$-query IND-CPAD adversary and prove \\[ \\Pr[\\mathsf{IND\\text{-}CPAD}_{\\mathsf{NF}}^{\\mathcal A}=1] \\leq β_{\\mathsf{CPA}}(\\mathcal B_{\\mathcal A,q}) + \\frac{\\sqrt{qn}}{2γ}. \\] where $n$ is the plaintext dimension and $γ$ is the flooding-width multiplier. Our proof constructs a new relational program logic over SSProve semantics. Its Pythagorean judgment composes conditional KL budgets without converting them to statistical distance, and a verified trace compiler lifts a local oracle rule to arbitrary adaptive programs with a single final conversion.","author":[{"family":"Lee","given":"Yi"},{"family":"Cojocaru","given":"Alexandru"},{"family":"Liu","given":"Junyi"},{"family":"Wu","given":"Xiaodi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.13846","URL":"https://doi.org/10.48550/arxiv.2608.13846","source":"datacite"},{"id":"doi:10.5281/zenodo.21608302","type":"article-journal","title":"Privacy-Preserving in Machine Learning: Bridging Security and Performance in Modern Applications","abstract":"Privacy-preserving machine learning (PPML) has emerged as a critical paradigm in the era of data-driven applications, addressing the fundamental tension between leveraging large-scale datasets and protecting individual privacy. This technical article examines recent advances in PPML techniques, focusing on three key approaches: federated learning, which enables distributed model training while keeping data localized; homomorphic encryption, allowing computation on encrypted data; and secure multi-party computation (MPC) for privacy-conscious collaborative learning. Through detailed architectural analysis and real-world case studies in mobile device personalization and healthcare analytics, this article demonstrates how these techniques can be effectively implemented while navigating computational overhead and implementation complexity. This article reveals current PPML approaches successfully preserve privacy in production environments, but they face significant challenges in computational efficiency and system integration. This article concludes by presenting optimization strategies and emerging research directions aimed at making PPML more practical for large-scale deployments.","author":[{"family":"Nalam","given":"Ramachandra"},{"family":"Nalam","given":"Pooja"},{"family":"Anuvalasetty","given":"Sruthi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21608302","URL":"https://doi.org/10.5281/zenodo.21608302","source":"datacite"},{"id":"doi:10.5281/zenodo.21608303","type":"article-journal","title":"Privacy-Preserving in Machine Learning: Bridging Security and Performance in Modern Applications","abstract":"Privacy-preserving machine learning (PPML) has emerged as a critical paradigm in the era of data-driven applications, addressing the fundamental tension between leveraging large-scale datasets and protecting individual privacy. This technical article examines recent advances in PPML techniques, focusing on three key approaches: federated learning, which enables distributed model training while keeping data localized; homomorphic encryption, allowing computation on encrypted data; and secure multi-party computation (MPC) for privacy-conscious collaborative learning. Through detailed architectural analysis and real-world case studies in mobile device personalization and healthcare analytics, this article demonstrates how these techniques can be effectively implemented while navigating computational overhead and implementation complexity. This article reveals current PPML approaches successfully preserve privacy in production environments, but they face significant challenges in computational efficiency and system integration. This article concludes by presenting optimization strategies and emerging research directions aimed at making PPML more practical for large-scale deployments.","author":[{"family":"Nalam","given":"Ramachandra"},{"family":"Nalam","given":"Pooja"},{"family":"Anuvalasetty","given":"Sruthi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.21608303","URL":"https://doi.org/10.5281/zenodo.21608303","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.00546","type":"manuscript","title":"Lightweight, Practical Encrypted Face Recognition with GPU Support","abstract":"Face recognition models operate in a client-server setting where a client extracts a compact face embedding and a server performs similarity search over a template database. This raises privacy concerns, as facial data is highly sensitive. To provide cryptographic privacy guarantees, one can use fully homomorphic encryption to perform end-to-end encrypted similarity search. However, existing FHE-based protocols are computationally costly and, impose high memory overhead. Building on prior work, HyDia (PoPETS 2025), we introduce algorithmic and system-level improvements targeting real-world deployment with resource-constrained clients. First, we propose BSGS-Diagonal, an algorithm delivering fast and memory-efficient similarity computation. BSGS-Diagonal substantially shrinks the rotation-key set, lowering both client and server memory requirements, and also improves practical server runtime. This yields a 91% reduction in the number of rotation keys, translating to approximately 14 GB less memory used on the client, and reducing overall CPU peak RAM from over 33 GB in the original HyDia to under 11 GB for databases up to size 1M. In addition, runtime is improved by up to 1.57x for the membership verification scenario and 1.43x for the identification scenario. Secondly, we introduce fully GPU-optimized similarity matrix computation kernels. The implementation is built upon FIDESlib, a CKKS-level GPU library based on OpenFHE. Rather than offloading individual CKKS primitives in isolation, the integrated kernels fuse operations to avoid repeated CPU-GPU ciphertext movement and costly FIDESlib/OpenFHE data-structure conversions. As a result, our GPU implementations of both HyDia and BSGS-Diagonal achieve up to 9x and 21x speedups, respectively, enabling sub-second encrypted face recognition for databases up to 32K entries while further reducing host memory usage.","author":[{"family":"De Micheli","given":"Gabrielle"},{"family":"Hafiz","given":"Syed"},{"family":"Pereira","given":"Geovandro"},{"family":"Cominetti","given":"Eduardo"},{"family":"Paiva","given":"Thales"},{"family":"Choi","given":"Jina"},{"family":"Simplicio","given":"Marcos"},{"family":"Yildiz","given":"Bahattin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.00546","URL":"https://doi.org/10.48550/arxiv.2604.00546","source":"datacite"},{"id":"doi:10.48550/arxiv.2601.17620","type":"manuscript","title":"Reconstructing Protected Biometric Templates from Binary Authentication Results","abstract":"Biometric data is considered to be very private and highly sensitive. As such, many methods for biometric template protection were considered over the years -- from biohashing and specially crafted feature extraction procedures, to the use of cryptographic solutions such as Fuzzy Commitments or the use of Fully Homomorphic Encryption (FHE). A key question that arises is how much protection these solutions can offer when the adversary can inject samples, and observe the outputs of the system. While for systems that return the similarity score, one can use attacks such as hill-climbing, for systems where the adversary can only learn whether the authentication attempt was successful, this question remained open. In this paper, we show that it is indeed possible to reconstruct the biometric template by just observing the success/failure of the authentication attempt (given the ability to inject a sufficient amount of templates). Our attack achieves negligible template reconstruction loss and enables full recovery of facial images through a generative inversion method, forming a pipeline from binary scores to high-resolution facial images that successfully pass the system more than 98\\% of the time. Our results, of course, are applicable for any protection mechanism that maintains the accuracy of the recognition.","author":[{"family":"Rahimi","given":"Eliron"},{"family":"Osadchy","given":"Margarita"},{"family":"Dunkelman","given":"Orr"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2601.17620","URL":"https://doi.org/10.48550/arxiv.2601.17620","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.24658","type":"manuscript","title":"Taking Advantage of Rational Canonical Form for Faster Ring-LWE based Encrypted Controller with Recursive Multiplication","abstract":"This paper aims to provide an efficient implementation of encrypted linear dynamic controllers that perform recursive multiplications on a Ring-Learning With Errors (Ring-LWE) based cryptosystem. By adopting a system-theoretical approach, we significantly reduce both time and space complexities, particularly the number of homomorphic operations required for recursive multiplications. Rather than encrypting the entire state matrix of a given controller, the state matrix is transformed into its rational canonical form, whose sparse and circulant structure enables that encryption and computation are required only on its nontrivial columns. Furthermore, we propose a novel method to ``pack'' each of the input and the output matrices into a single polynomial, thereby reducing the number of homomorphic operations. Simulation results demonstrate that the proposed design enables a remarkably fast implementation of encrypted controllers.","author":[{"family":"Song","given":"Donghyeon"},{"family":"Jang","given":"Yeongjun"},{"family":"Lee","given":"Joowon"},{"family":"Kim","given":"Junsoo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.24658","URL":"https://doi.org/10.48550/arxiv.2512.24658","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.04583","type":"manuscript","title":"Energy Consumption of TLS, Searchable Encryption and Fully Homomorphic Encryption","abstract":"Privacy-enhancing technologies (PETs) have attracted significant attention in response to privacy regulations, driving the development of applications that prioritize user data protection. At the same time, the information and communication technology (ICT) sector faces growing pressure to reduce its environmental footprint, particularly its energy consumption. While numerous studies have assessed the energy consumption of ICT applications, the environmental impact of cryptographic PETs remains largely unexplored. This work investigates this question by measuring the energy consumption increase induced by three PETs compared to their non-private counterparts: TLS, Searchable Encryption, and Fully Homomorphic Encryption (FHE). These technologies were chosen for two reasons. First, they cover different maturity levels -- from the widely deployed TLS protocol to the emerging FHE schemes -- allowing us to examine the influence of maturity on energy consumption. Second, they each have well-established applications in industry: web browsing, encrypted databases, and privacy-preserving machine learning. Our results reveal highly variable energy consumption increases, ranging from 2x for TLS to 10x for Searchable Encryption and 100,000x for FHE. Our experiments demonstrate a simple and reproducible methodology, based on existing open-source software, to quantify the energy costs of PETs. They also highlight the wide spectrum of energy demands across technologies, underscoring the importance of further research on sustainable PET design. Finally, we discuss orthogonal research directions, such as hardware acceleration, to outline promising directions toward sustainable PETs.","author":[{"family":"Damie","given":"Marc"},{"family":"Pop","given":"Mihai"},{"family":"Posthuma","given":"Merijn"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.04583","URL":"https://doi.org/10.48550/arxiv.2508.04583","source":"datacite"},{"id":"doi:10.4230/lipics.itc.2026.4","type":"article-journal","title":"Fiat-Shamir for Bounded-Depth Adversaries","abstract":"We study how to construct hash functions that can securely instantiate the Fiat-Shamir transformation against bounded-depth adversaries. The motivation is twofold. First, given the recent fruitful line of research of constructing cryptographic primitives against bounded-depth adversaries under worst-case complexity assumptions, and the rich applications of Fiat-Shamir, instantiating Fiat-Shamir hash functions against bounded-depth adversaries under worst-case complexity assumptions might lead to further applications (such as SNARG for P, showing the cryptographic hardness of PPAD, etc.) against bounded-depth adversaries. Second, we wonder whether it is possible to overcome the impossibility results of constructing Fiat-Shamir for arguments [Goldwasser, Kalai, FOCS '03] in the setting where the depth of the adversary is bounded, given that the known impossibility results (against p.p.t. adversaries) are contrived. Our main results give new insights for Fiat-Shamir against bounded-depth adversaries in both the positive and negative directions. On the positive side, for Fiat-Shamir for proofs with certain properties, we show that weak worst-case assumptions are enough for constructing explicit hash functions that give AC⁰[2]-soundness. In particular, we construct an AC⁰[2]-computable correlation-intractable hash family for constant-degree polynomials against AC⁰[2] adversaries, assuming ⊕L/poly ⊈ Sum̃_{n^{-c}}∘AC⁰[2] for some c > 0. This is incomparable to all currently-known constructions, which are typically useful for larger classes and against stronger adversaries, but based on arguably stronger assumptions. Our construction is inspired by the Fiat-Shamir hash function by Peikert and Shiehian [CRYPTO '19] and the fully-homomorphic encryption scheme against bounded-depth adversaries by Wang and Pan [EUROCRYPT '22]. On the negative side, we show Fiat-Shamir for arguments is still impossible to achieve against bounded-depth adversaries. In particular, - Assuming the existence of AC⁰[2]-computable CRHF against p.p.t. adversaries, for every poly-size hash function, there is a (p.p.t.-sound) interactive argument that is not AC⁰[2]-sound after applying Fiat-Shamir with this hash function. - Assuming the existence of AC⁰[2]-computable CRHF against AC⁰[2] adversaries, there is an AC⁰[2]-sound interactive argument such that for every hash function computable by AC⁰[2] circuits, the argument does not preserve AC⁰[2]-soundness when applying Fiat-Shamir with this hash function. This is a low-depth variant of Goldwasser and Kalai.","author":[{"family":"Chen","given":"Liyan"},{"family":"Chen","given":"Yilei"},{"family":"Huang","given":"Zikuan"},{"family":"Sun","given":"Nuozhou"},{"family":"Yang","given":"Tianqi"},{"family":"Zhang","given":"Yiding"}],"issued":{"date-parts":[[2026]]},"DOI":"10.4230/lipics.itc.2026.4","URL":"https://doi.org/10.4230/lipics.itc.2026.4","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.08010","type":"manuscript","title":"Sensor Attack Detection Method for Encrypted State Observers","abstract":"This paper proposes an encrypted state observer that is capable of detecting sensor attacks without decryption. We first design a state observer that operates over a finite field of integers with the modular arithmetic. The observer generates a residue signal that indicates the presence of attacks under sparse attack and sensing redundancy conditions. Then, we develop a homomorphic encryption scheme that enables the observer to operate over encrypted data while automatically disclosing the residue signal. Unlike our previous work restricted to single-input single-output systems, the proposed scheme is applicable to general multi-input multi-output systems. Given that the disclosed residue signal remains below a prescribed threshold, the full state can be recovered as an encrypted message.","author":[{"family":"Jang","given":"Yeongjun"},{"family":"Lee","given":"Sangwon"},{"family":"Kim","given":"Junsoo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.08010","URL":"https://doi.org/10.48550/arxiv.2512.08010","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.21381","type":"manuscript","title":"Privacy-Preserving Distributed Stochastic Optimization with Homomorphic Encryption and Heterogeneous Stepsizes","abstract":"Distributed stochastic optimization enables multi-agent collaboration in applications such as distributed learning and sensor networks, but also raises critical privacy concerns due to the involvement of sensitive data. While existing privacy-preserving approaches often face limitations in balancing accuracy with efficiency, we propose a novel distributed stochastic gradient descent algorithm that integrates Paillier homomorphic encryption with heterogeneous and time-varying random stepsizes. The proposed algorithm provides inherent privacy protection against both internal honest-but-curious agents and external eavesdroppers, without relying on any trusted neighbors. Furthermore, we incorporate an attenuation factor to effectively mitigate quantization error induced by the encryption process, ensuring almost sure convergence to the optimal solution while maintaining privacy preservation. Numerical simulations demonstrate the effectiveness and efficiency of the proposed approach.","author":[{"family":"Zhou","given":"Haoqiang"},{"family":"Chen","given":"Chi"},{"family":"Zhi","given":"Yongfeng"},{"family":"Gao","given":"Huan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.21381","URL":"https://doi.org/10.48550/arxiv.2604.21381","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.19890","type":"manuscript","title":"Efficient Arithmetic-and-Comparison Homomorphic Encryption with Space Switching","abstract":"Fully homomorphic encryption (FHE) enables computation on encrypted data without decryption, making it central to privacy-preserving applications. However, no existing scheme efficiently supports both arithmetic and comparison operations in a unified framework. Prior approaches such as scheme switching and polynomial approximation face serious limitations: switching incurs prohibitive overhead for large inputs, while approximation methods introduce errors near critical points, restricting use in accuracy-sensitive tasks. We propose space switching method to integrate arithmetic and comparison computation seamlessly within FV-style schemes. Our approach identifies that the two types of operations require different plaintext spaces and introduces two procedures: a reduction step to transition from the number space $\\mathbb{Z}_{p^r}$ to the digit space $\\mathbb{Z}_{p}$, and a modulus-raising step to map results back to $\\mathbb{Z}_{p^r}$. This design enables continuous evaluation of arithmetic and comparison within the same scheme. Experiments show that our method achieves up to $17\\times$ faster performance than scheme switching and $15\\times$ faster than direct comparison on database workloads, demonstrating its practicality for real-world privacy-preserving computation. Code and artifacts are available at https://github.com/UCF-Lou-Lab-PET/Universal-BGV.","author":[{"family":"Wahyudi","given":"Erwin"},{"family":"Solihin","given":"Yan"},{"family":"Lou","given":"Qian"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.19890","URL":"https://doi.org/10.48550/arxiv.2604.19890","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.26417","type":"manuscript","title":"Towards Privacy-Preserving Federated Learning using Hybrid Homomorphic Encryption","abstract":"Federated Learning (FL) enables collaborative training while keeping sensitive data on clients' devices, but local model updates can still leak private information. Hybrid Homomorphic Encryption (HHE) has recently been applied to FL to mitigate client overhead while preserving privacy. However, existing HHE-FL systems rely on a single homomorphic key pair shared across all clients, which forces them to assume an unrealistically weak threat model: if a client misbehaves or intercepts another's traffic, private updates can be exposed. We eliminate this weakness by integrating two alternative key protection mechanisms into the HHE-FL workflow. The first is masking, where client keys are blinded before homomorphic encryption and later unblinded homomorphically by the server. The second is RSA encapsulation, where homomorphically encrypted keys are additionally wrapped under the server's RSA public key. These countermeasures prevent key misuse by other clients and extend HHE-FL security to adversarial settings with malicious participants. We implement both approaches on top of the Flower framework using the PASTA/BFV HHE scheme and evaluate them on the MNIST dataset with 12 clients. Results show that both mechanisms preserve model accuracy while adding minimal overhead: masking incurs negligible cost, and RSA encapsulation introduces only modest runtime and communication overhead.","author":[{"family":"Costa","given":"Ivan"},{"family":"Correia","given":"Pedro"},{"family":"Amorim","given":"Ivone"},{"family":"Maia","given":"Eva"},{"family":"Praça","given":"Isabel"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.26417","URL":"https://doi.org/10.48550/arxiv.2603.26417","source":"datacite"},{"id":"doi:10.48550/arxiv.2601.05865","type":"manuscript","title":"Secure Change-Point Detection for Time Series under Homomorphic Encryption","abstract":"We introduce the first method for change-point detection on encrypted time series. Our approach employs the CKKS homomorphic encryption scheme to detect shifts in statistical properties (e.g., mean, variance, frequency) without ever decrypting the data. Unlike solutions based on differential privacy, which degrade accuracy through noise injection, our solution preserves utility comparable to plaintext baselines. We assess its performance through experiments on both synthetic datasets and real-world time series from healthcare and network monitoring. Notably, our approach can process one million points within 3 minutes.","author":[{"family":"Mazzone","given":"Federico"},{"family":"Micali","given":"Giorgio"},{"family":"Pronesti","given":"Massimiliano"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2601.05865","URL":"https://doi.org/10.48550/arxiv.2601.05865","source":"datacite"},{"id":"doi:10.48550/arxiv.2504.11604","type":"manuscript","title":"SoK: Can Fully Homomorphic Encryption Support General AI Computation? A Functional and Cost Analysis","abstract":"Artificial intelligence (AI) increasingly powers sensitive applications in domains such as healthcare and finance, relying on both linear operations (e.g., matrix multiplications in large language models) and non-linear operations (e.g., sorting in retrieval-augmented generation). Fully homomorphic encryption (FHE) has emerged as a promising tool for privacy-preserving computation, but it remains unclear whether existing methods can support the full spectrum of AI workloads that combine these operations. In this SoK, we ask: Can FHE support general AI computation? We provide both a functional analysis and a cost analysis. First, we categorize ten distinct FHE approaches and evaluate their ability to support general computation. We then identify three promising candidates and benchmark workloads that mix linear and non-linear operations across different bit lengths and SIMD parallelization settings. Finally, we evaluate five real-world, privacy-sensitive AI applications that instantiate these workloads. Our results quantify the costs of achieving general computation in FHE and offer practical guidance on selecting FHE methods that best fit specific AI application requirements. Our codes are available at https://github.com/UCF-ML-Research/FHE-AI-Generality.","author":[{"family":"Xue","given":"Jiaqi"},{"family":"Xin","given":"Xin"},{"family":"Zhang","given":"Wei"},{"family":"Zheng","given":"Mengxin"},{"family":"Song","given":"Qianqian"},{"family":"Zhou","given":"Minxuan"},{"family":"Dong","given":"Yushun"},{"family":"Wang","given":"Dongjie"},{"family":"Chen","given":"Xun"},{"family":"Xie","given":"Jiafeng"},{"family":"Wang","given":"Liqiang"},{"family":"Mohaisen","given":"David"},{"family":"Wu","given":"Hongyi"},{"family":"Lou","given":"Qian"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2504.11604","URL":"https://doi.org/10.48550/arxiv.2504.11604","source":"datacite"},{"id":"doi:10.6084/m9.figshare.32964164.v1","type":"article-journal","title":"<b>Security Challenges in Large Language Models Across the Data Life Cycle: </b><b>A Cryptography-Aware Comprehensive Review</b>","abstract":"Large Language Models (LLMs) are now embedded in security and privacy critical applications, yet they remain vulnerable to attacks that span their entire data life cycle. This survey provides a comprehensive, cryptography-aware review of these risks across three phases—training, inference, and deployment; while explicitly connecting them to classical security goals and primitives. We introduce a simple stage-wise risk scoring model inspired by NIST risk assessment that propagates vulnerabilities across the life cycle, and we instantiate it with a numeric example linking training time poisoning to inference time data extraction. We further propose a life cycle aligned evaluation framework that maps modern benchmarks (e.g., HarmBench, JailbreakBench, TrustLLM, DecodingTrust) to concrete threat classes and reports representative quantitative results, such as attack success rates under different defenses. Finally, we analyze the practicality of advanced defenses—including differential privacy, fully homomorphic encryption, secure multi-party computation, and zero knowledge proofs—in light of their computational overhead and deployment constraints, building on foundational cryptography and privacy works. Our goal is to bridge the gap between classical cryptographic theory and emerging LLM specific threats, and to outline research directions toward secure, privacy preserving, and rigorously evaluated LLM pipelines.","author":[{"family":"Chakoli","given":"Sepehr"},{"family":"Etrati","given":"Seyed"},{"family":"Etrati","given":"Seyed"},{"family":"Javadi","given":"Hamid"}],"issued":{"date-parts":[[2026]]},"DOI":"10.6084/m9.figshare.32964164.v1","URL":"https://doi.org/10.6084/m9.figshare.32964164.v1","source":"datacite"},{"id":"doi:10.6084/m9.figshare.32964164","type":"article-journal","title":"<b>Security Challenges in Large Language Models Across the Data Life Cycle: </b><b>A Cryptography-Aware Comprehensive Review</b>","abstract":"Large Language Models (LLMs) are now embedded in security and privacy critical applications, yet they remain vulnerable to attacks that span their entire data life cycle. This survey provides a comprehensive, cryptography-aware review of these risks across three phases—training, inference, and deployment; while explicitly connecting them to classical security goals and primitives. We introduce a simple stage-wise risk scoring model inspired by NIST risk assessment that propagates vulnerabilities across the life cycle, and we instantiate it with a numeric example linking training time poisoning to inference time data extraction. We further propose a life cycle aligned evaluation framework that maps modern benchmarks (e.g., HarmBench, JailbreakBench, TrustLLM, DecodingTrust) to concrete threat classes and reports representative quantitative results, such as attack success rates under different defenses. Finally, we analyze the practicality of advanced defenses—including differential privacy, fully homomorphic encryption, secure multi-party computation, and zero knowledge proofs—in light of their computational overhead and deployment constraints, building on foundational cryptography and privacy works. Our goal is to bridge the gap between classical cryptographic theory and emerging LLM specific threats, and to outline research directions toward secure, privacy preserving, and rigorously evaluated LLM pipelines.","author":[{"family":"Chakoli","given":"Sepehr"},{"family":"Etrati","given":"Seyed"},{"family":"Etrati","given":"Seyed"},{"family":"Javadi","given":"Hamid"}],"issued":{"date-parts":[[2026]]},"DOI":"10.6084/m9.figshare.32964164","URL":"https://doi.org/10.6084/m9.figshare.32964164","source":"datacite"},{"id":"doi:10.48550/arxiv.2605.29450","type":"manuscript","title":"Protecting On-Device AI Inference: A Systematic Review of Attacks and Defence Mechanisms","abstract":"The need for secure and private Artificial Intelligence (AI) and Machine Learning (ML) on edge and mobile devices has increased the necessity of protecting the architecture of these systems from threats to both security and privacy. With an ever-increasing number of pre-trained AI models being used on mobile platforms for client-side inference, there are rising concerns about the risks associated with the theft/extraction of AI models, adversarial attacks on AI models, and data breaches. As a result of this trend, a variety of defence mechanisms have been proposed to protect against these threats. These include Trusted Execution Environments (TEEs), homomorphic encryption, obfuscation, and differential privacy, among others. However, current surveys largely focus on edge intelligence, which includes distributed training, and thus overlook security and privacy issues that are specific to on-device AI inference. To the best of our knowledge, this paper presents the first comprehensive review of threats and corresponding defence mechanisms targeting on-device inference. Our results show that the attack and defence literature are unbalanced: approximately one quarter of the surveyed attack papers focus on Intellectual Property (IP) attacks, whereas half of the defence solutions tackle the same issue. More importantly, some attack categories have no defence paper associated to them, such as adversarial attacks that account for roughly one third of the attack literature. This asymmetry between known attacks and available mitigations highlights clear opportunities for future research on securing on-device AI inference.","author":[{"family":"Tsiatsikas","given":"Zisis"},{"family":"Fakis","given":"Alexandros"},{"family":"Karopoulos","given":"Georgios"},{"family":"Kouliaridis","given":"Vasileios"},{"family":"Anagnostopoulos","given":"Marios"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.29450","URL":"https://doi.org/10.48550/arxiv.2605.29450","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.11470","type":"manuscript","title":"Cachemir: Fully Homomorphic Encrypted Inference of Generative Large Language Model with KV Cache","abstract":"Generative large language models (LLMs) have revolutionized multiple domains. Modern LLMs predominantly rely on an autoregressive decoding strategy, which generates output tokens sequentially and employs a key-value cache (KV cache) to avoid redundant computation. However, the widespread deployment of LLMs has raised serious privacy concerns, as users are feeding all types of data into the model, motivating the development of secure inference frameworks based on fully homomorphic encryption (FHE). A major limitation of existing FHE-based frameworks is their inability to effectively integrate the KV cache, resulting in prohibitively high latency for autoregressive decoding. In this paper, we propose Cachemir, a KV Cache Accelerated Homomorphic Encrypted LLM Inference Regime to overcome this limitation. Cachemir comprises three key technical contributions: 1) a set of novel HE packing algorithms specifically designed to leverage the computational advantages of the KV cache; 2) an interleaved replicated packing algorithm to efficiently compute the vector-matrix multiplications that result from using the KV cache in Transformer linear layers; and 3) an augmented bootstrapping placement strategy that accounts for the KV cache to minimize bootstrapping cost. We demonstrate that Cachemir achieves $48.83\\times$ and $67.16\\times$ speedup over MOAI (ICML'25) and THOR (CCS'25) respectively on CPU and consumes less than 100 seconds on GPU to generate an output token for Llama-3-8B.","author":[{"family":"Yu","given":"Ye"},{"family":"Zhou","given":"Yifan"},{"family":"Chen","given":"Yi"},{"family":"Soto","given":"Pedro"},{"family":"Xiong","given":"Wenjie"},{"family":"Li","given":"Meng"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.11470","URL":"https://doi.org/10.48550/arxiv.2602.11470","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.02717","type":"manuscript","title":"On the Feasibility of Hybrid Homomorphic Encryption for Intelligent Transportation Systems","abstract":"Many Intelligent Transportation Systems (ITS) applications require strong privacy guarantees for both users and their data. Homomorphic encryption (HE) enables computation directly on encrypted messages and thus offers a compelling approach to privacy-preserving data processing in ITS. However, practical HE schemes incur substantial ciphertext expansion and communication overhead, which limits their suitability for time-critical transportation systems. Hybrid homomorphic encryption (HHE) addresses this challenge by combining a homomorphic encryption scheme with a symmetric cipher, enabling efficient encrypted computation while dramatically reducing communication cost. In this paper, we develop theoretical models of representative ITS applications that integrate HHE to protect sensitive vehicular data. We then perform a parameter-based evaluation of the HHE scheme Rubato to estimate ciphertext sizes and communication overhead under realistic ITS workloads. Our results show that HHE achieves orders-of-magnitude reductions in ciphertext size compared with conventional HE while maintaining cryptographic security, making it significantly more practical for latency-constrained ITS communication.","author":[{"family":"Yates","given":"Kyle"},{"family":"Mamun","given":"Abdullah"},{"family":"Chowdhury","given":"Mashrur"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.02717","URL":"https://doi.org/10.48550/arxiv.2602.02717","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.19525","type":"manuscript","title":"Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC","abstract":"This paper presents an efficient framework for private Transformer inference that combines Homomorphic Encryption (HE) and Secure Multi-party Computation (MPC) to protect data privacy. Existing methods often leverage HE for linear layers (e.g., matrix multiplications) and MPC for non-linear layers (e.g., Softmax activation functions), but the conversion between HE and MPC introduces significant communication costs. The proposed framework, dubbed BLB, overcomes this by breaking down layers into fine-grained operators and further fusing adjacent linear operators, reducing the need for HE/MPC conversions. To manage the increased ciphertext bit width from the fused linear operators, BLB proposes the first secure conversion protocol between CKKS and MPC and enables CKKS-based computation of the fused operators. Additionally, BLB proposes an efficient matrix multiplication protocol for fused computation in Transformers. Extensive evaluations on BERT-base, BERT-large, and GPT2-base show that BLB achieves a $21\\times$ reduction in communication overhead compared to BOLT (S\\&amp;P'24) and a $2\\times$ reduction compared to Bumblebee (NDSS'25), along with latency reductions of $13\\times$ and $1.8\\times$, respectively, when leveraging GPU acceleration.","author":[{"family":"Xu","given":"Tianshi"},{"family":"Lu","given":"Wen"},{"family":"Yu","given":"Jiangrui"},{"family":"Yi","given":"Chen"},{"family":"Lin","given":"Chenqi"},{"family":"Wang","given":"Runsheng"},{"family":"Li","given":"Meng"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.19525","URL":"https://doi.org/10.48550/arxiv.2508.19525","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.06086","type":"manuscript","title":"QuHE: Optimizing Utility-Cost in Quantum Key Distribution and Homomorphic Encryption Enabled Secure Edge Computing Networks","abstract":"Ensuring secure and efficient data processing in mobile edge computing (MEC) systems is a critical challenge. While quantum key distribution (QKD) offers unconditionally secure key exchange and homomorphic encryption (HE) enables privacy-preserving data processing, existing research fails to address the comprehensive trade-offs among QKD utility, HE security, and system costs. This paper proposes a novel framework integrating QKD, transciphering, and HE for secure and efficient MEC. QKD distributes symmetric keys, transciphering bridges symmetric encryption, and HE processes encrypted data at the server. We formulate an optimization problem balancing QKD utility, HE security, processing and wireless transmission costs. However, the formulated optimization is non-convex and NPhard. To solve it efficiently, we propose the Quantum-enhanced Homomorphic Encryption resource allocation (QuHE) algorithm. Theoretical analysis proves the proposed QuHE algorithm's convergence and optimality, and simulations demonstrate its effectiveness across multiple performance metrics.","author":[{"family":"Qian","given":"Liangxin"},{"family":"Li","given":"Yang"},{"family":"Zhao","given":"Jun"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.06086","URL":"https://doi.org/10.48550/arxiv.2507.06086","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.04775","type":"manuscript","title":"FIDESlib: A Fully-Fledged Open-Source FHE Library for Efficient CKKS on GPUs","abstract":"Word-wise Fully Homomorphic Encryption (FHE) schemes, such as CKKS, are gaining significant traction due to their ability to provide post-quantum-resistant, privacy-preserving approximate computing; an especially desirable feature in Machine-Learning-as-a-Service (MLaaS) cloud-computing paradigms. OpenFHE is a leading CPU-based FHE library with robust CKKS operations, but its server-side performance is not yet sufficient for practical cloud deployment. As GPU computing becomes more common in data centers, many FHE libraries are adding GPU support. However, integrating an efficient GPU backend into OpenFHE is challenging. While OpenFHE uses a Hardware Abstraction Layer (HAL), its flexible architecture sacrifices performance due to the abstraction layers required for multi-scheme and multi-backend compatibility. In this work, we introduce FIDESlib, the first open-source server-side CKKS GPU library that is fully interoperable with well-established client-side OpenFHE operations. Unlike other existing open-source GPU libraries, FIDESlib provides the first implementation featuring heavily optimized GPU kernels for all CKKS primitives, including bootstrapping. Our library also integrates robust benchmarking and testing, ensuring it remains adaptable to further optimization. Furthermore, its software architecture is designed to support extensions to a multi-GPU backend for enhanced acceleration. Our experiments across various GPU systems and the leading open-source CKKS library to date, Phantom, show that FIDESlib offers superior performance and scalability. For bootstrapping, FIDESlib achieves no less than 70x speedup over the AVX-optimized OpenFHE implementation.","author":[{"family":"Agulló-Domingo","given":"Carlos"},{"family":"Vera-López","given":"Óscar"},{"family":"Guzelhan","given":"Seyda"},{"family":"Daksha","given":"Lohit"},{"family":"Jerari","given":"Aymane"},{"family":"Shivdikar","given":"Kaustubh"},{"family":"Agrawal","given":"Rashmi"},{"family":"Kaeli","given":"David"},{"family":"Joshi","given":"Ajay"},{"family":"Abellán","given":"José"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.04775","URL":"https://doi.org/10.48550/arxiv.2507.04775","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.10399","type":"manuscript","title":"FicGCN: Unveiling the Homomorphic Encryption Efficiency from Irregular Graph Convolutional Networks","abstract":"Graph Convolutional Neural Networks (GCNs) have gained widespread popularity in various fields like personal healthcare and financial systems, due to their remarkable performance. Despite the growing demand for cloud-based GCN services, privacy concerns over sensitive graph data remain significant. Homomorphic Encryption (HE) facilitates Privacy-Preserving Machine Learning (PPML) by allowing computations to be performed on encrypted data. However, HE introduces substantial computational overhead, particularly for GCN operations that require rotations and multiplications in matrix products. The sparsity of GCNs offers significant performance potential, but their irregularity introduces additional operations that reduce practical gains. In this paper, we propose FicGCN, a HE-based framework specifically designed to harness the sparse characteristics of GCNs and strike a globally optimal balance between aggregation and combination operations. FicGCN employs a latency-aware packing scheme, a Sparse Intra-Ciphertext Aggregation (SpIntra-CA) method to minimize rotation overhead, and a region-based data reordering driven by local adjacency structure. We evaluated FicGCN on several popular datasets, and the results show that FicGCN achieved the best performance across all tested datasets, with up to a 4.10x improvement over the latest design.","author":[{"family":"Kan","given":"Zhaoxuan"},{"family":"Han","given":"Husheng"},{"family":"Shi","given":"Shangyi"},{"family":"Hua","given":"Tenghui"},{"family":"Lu","given":"Hang"},{"family":"Li","given":"Xiaowei"},{"family":"Mu","given":"Jianan"},{"family":"Hu","given":"Xing"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.10399","URL":"https://doi.org/10.48550/arxiv.2506.10399","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.09397","type":"manuscript","title":"Collision Risk Analysis for LEO Satellites with Confidential Orbital Data","abstract":"The growing number of satellites in low Earth orbit (LEO) has increased concerns about the risk of satellite collisions, which can ultimately result in the irretrievable loss of satellites and a growing amount of space debris. To mitigate this risk, accurate collision risk analysis is essential. However, this requires access to sensitive orbital data, which satellite operators are often unwilling to share due to privacy concerns. This contribution proposes a solution based on fully homomorphic encryption (FHE) and thus enables secure and private collision risk analysis. In contrast to existing methods, this approach ensures that collision risk analysis can be performed on sensitive orbital data without revealing it to other parties. To display the challenges and opportunities of FHE in this context, an implementation of the CKKS scheme is adapted and analyzed for its capacity to satisfy the theoretical requirements of precision and run time.","author":[{"family":"Lage","given":"Svenja"},{"family":"Hörmann","given":"Felicitas"},{"family":"Hanke","given":"Felix"},{"family":"Karl","given":"Michael"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.09397","URL":"https://doi.org/10.48550/arxiv.2501.09397","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.02461","type":"manuscript","title":"Experimental Evaluation of Post-Quantum Homomorphic Encryption for Privacy-Preserving I2I Communication in ITS","abstract":"This study experimentally evaluates the feasibility of post-quantum secure Homomorphic Encryption (HE) for privacy-preserving Infrastructure-to-Infrastructure (I2I) communication in Intelligent Transportation Systems (ITS). Unlike prior simulation-based efforts, this work implements three lattice-based HE schemes: Brakerski-Fan-Vercauteren (BFV), Brakerski-Gentry-Vaikuntanathan (BGV), and Cheon-Kim-Kim-Song (CKKS), within a real experimental pipeline representing roadside unit (RSU)-Cloud data exchange over Wi-Fi and Ethernet networks. The experiments benchmark encrypted addition and addition-plus-multiplication operations representing key analytical tasks, such as vehicle queue assessment and regional speed computation. Results show that while BFV achieves sub-5-second latency suitable for intersection-level analytics, BGV supports regional aggregation with 10 to 30-second updates. CKKS, though exhibiting higher latency (21-32 seconds), remains practical for minute-scale applications like eco-driving. These findings demonstrate that post-quantum HE can enable privacy-preserving ITS backhaul analytics when latency requirements align with application needs. The study also presents optimization pathways, including algorithmic tuning, network adaptation, and hardware acceleration, to reduce end-to-end delay.","author":[{"family":"Mamun","given":"Abdullah"},{"family":"Yates","given":"Kyle"},{"family":"Rakotondrafara","given":"Antsa"},{"family":"Chowdhury","given":"Mashrur"},{"family":"Cartor","given":"Ryann"},{"family":"Gao","given":"Shuhong"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.02461","URL":"https://doi.org/10.48550/arxiv.2508.02461","source":"datacite"},{"id":"doi:10.48550/arxiv.2511.09855","type":"manuscript","title":"Unlearning Imperative: Securing Trustworthy and Responsible LLMs through Engineered Forgetting","abstract":"The growing use of large language models in sensitive domains has exposed a critical weakness: the inability to ensure that private information can be permanently forgotten. Yet these systems still lack reliable mechanisms to guarantee that sensitive information can be permanently removed once it has been used. Retraining from the beginning is prohibitively costly, and existing unlearning methods remain fragmented, difficult to verify, and often vulnerable to recovery. This paper surveys recent research on machine unlearning for LLMs and considers how far current approaches can address these challenges. We review methods for evaluating whether forgetting has occurred, the resilience of unlearned models against adversarial attacks, and mechanisms that can support user trust when model complexity or proprietary limits restrict transparency. Technical solutions such as differential privacy, homomorphic encryption, federated learning, and ephemeral memory are examined alongside institutional safeguards including auditing practices and regulatory frameworks. The review finds steady progress, but robust and verifiable unlearning is still unresolved. Efficient techniques that avoid costly retraining, stronger defenses against adversarial recovery, and governance structures that reinforce accountability are needed if LLMs are to be deployed safely in sensitive applications. By integrating technical and organizational perspectives, this study outlines a pathway toward AI systems that can be required to forget, while maintaining both privacy and public trust.","author":[{"family":"Kang","given":"James"},{"family":"Bui","given":"Dang"},{"family":"Pham","given":"Thanh"},{"family":"Ling","given":"Huo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2511.09855","URL":"https://doi.org/10.48550/arxiv.2511.09855","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.01798","type":"manuscript","title":"A Survey on Privacy-Preserving Computing in the Automotive Domain","abstract":"As vehicles become increasingly connected and autonomous, they accumulate and manage various personal data, thereby presenting a key challenge in preserving privacy during data sharing and processing. This survey reviews applications of Secure Multi-Party Computation (MPC) and Homomorphic Encryption (HE) that address these privacy concerns in the automotive domain. First, we identify the scope of privacy-sensitive use cases for these technologies, by surveying existing works that address privacy issues in different automotive contexts, such as location-based services, mobility infrastructures, traffic management, etc. Then, we review recent works that employ MPC and HE as solutions for these use cases in detail. Our survey highlights the applicability of these privacy-preserving technologies in the automotive context, while also identifying challenges and gaps in the current research landscape. This work aims to provide a clear and comprehensive overview of this emerging field and to encourage further research in this domain.","author":[{"family":"Yuca","given":"Nergiz"},{"family":"Matyunin","given":"Nikolay"},{"family":"Arzoglou","given":"Ektor"},{"family":"Anagnostopoulos","given":"Nikolaos"},{"family":"Katzenbeisser","given":"Stefan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.01798","URL":"https://doi.org/10.48550/arxiv.2508.01798","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.27826","type":"manuscript","title":"Personalized and Multi-View Representation for Federated Cold-Start Recommendation","abstract":"Federated recommendation (FedRec) enables personalized modeling without centralizing users' interaction histories, but most existing methods assume a fixed item pool and thus overlook the practical cold-item setting where new items continuously arrive. Under the dual-sided constraint, where the server cannot access clients' interactions while clients cannot access the server's proprietary item attribute features, prior federated cold-start recommendation approaches suffer from three structural limitations: a lack of personalization, compositionality failure caused by forcing heterogeneous semantics into a single embedding space, and training- and communication-inefficiency arising from explicit alignment between separate collaborative and attribute representations. To address these challenges, we propose Personalized and Multi-view Representation for Federated Cold-Start Recommendation (PMFRec). PMFRec learns a personalized representation generator to produce user-specific item representations from attribute features, and introduces a global multi-view encoder with item-adaptive gating and an orthogonality objective to capture complementary semantic views while reducing cross-view redundancy. In addition, PMFRec fuses collaborative and attribute knowledge into a single exchanged item representation, eliminating the need for an explicit client-side regularizer and reducing communication overhead. Extensive experiments on real-world datasets show that PMFRec consistently outperforms strong baselines in cold-item recommendation and further improves user-level fairness, warm-scenario adaptability, and robustness under Local Differential Privacy (LDP).","author":[{"family":"Lim","given":"Jaehyung"},{"family":"Kweon","given":"Wonbin"},{"family":"Kim","given":"Woojoo"},{"family":"Kim","given":"Junyoung"},{"family":"Kim","given":"Dongha"},{"family":"Yu","given":"Hwanjo"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.27826","URL":"https://doi.org/10.48550/arxiv.2608.27826","source":"datacite"},{"id":"doi:10.17605/osf.io/xdwhs","type":"article-journal","title":"Large Language Models in the Acute Stroke Pathway: A Scoping Review of Applications, Evidence Maturity, and Implementation Readiness","abstract":"Background. Large language models (LLMs) have been rapidly adopted in medicine since late 2022, yet their role in the time-critical acute stroke pathway—from symptom recognition and prehospital triage to emergency diagnosis, imaging-related text tasks, reperfusion decision support, and acute-phase documentation and communication—has not been systematically mapped. Existing reviews cover the whole stroke-care continuum or mix LLMs with traditional NLP, leaving the acute phase under-characterized. Objective. To map the applications, evidence maturity, and implementation readiness of LLMs across the acute stroke pathway. Methods. This scoping review follows the PRISMA-ScR guideline. We search PubMed/MEDLINE, Europe PMC (including preprints), and Google Scholar for studies published from November 2022 onward. Eligible studies center on LLMs/generative AI applied to any stage of the acute stroke pathway. Two reviewers independently screen records and chart data using a piloted form. Evidence is synthesized along two dimensions: five pathway stages (prehospital recognition/dispatch; emergency triage and differential diagnosis; imaging-related text tasks; reperfusion decision support; acute documentation and communication) and three evidence-maturity tiers (simulation/benchmark; retrospective real-world data; prospective deployment). Implementation barriers (hallucination, bias, privacy, regulation, liability, integration, cost) are thematically summarized. Registration note. This review is registered on OSF; the full protocol is available in the attached files.","author":[{"family":"Luo","given":"Xianmu"},{"family":"Li","given":"Hongsong"},{"family":"Liao","given":"Xiaoli"}],"issued":{"date-parts":[[2026]]},"DOI":"10.17605/osf.io/xdwhs","URL":"https://doi.org/10.17605/osf.io/xdwhs","source":"datacite"},{"id":"oa:W4411700982","type":"article-journal","title":"LLM-Based Agents for Tool Learning: A Survey","abstract":"Abstract Human beings capable of making and using tools can accomplish tasks far beyond their innate abilities, and this paradigm of integration with tools may not be limited to humans themselves. Recently, the large language model (LLM) has demonstrated immense potential across various fields with its unique planning and reasoning abilities. However, there are still many challenges beyond its capabilities due to deficiencies in its training data and inherent illusions. Thus, integrating LLMs and tools into tool learning agents has become a new emerging research direction. To this end, we present a systematic investigation and comprehensive review of tool-learning agents in this paper. We start by introducing the definition of the tool learning task for Agents and then illustrating the typical architecture of the tool-learning models. Since these tools are all defined by users, LLM does not know what tools there are and what their functions are. Thus, LLMs should first find appropriate tools and split the tool retrieval methods into two categories: training-based and non-training-based. To accurately complete the user task, it is important to decompose the task into several sub-tasks and execute them in the correct order. Following that, we introduce the tool planning methods and organize these works by whether they rely on the model’s inherent reasoning capabilities for planning or utilize external reasoning tools. Due to the rapid development of this field, we also introduce an emerging frontier direction: using multimodal tools for LLM. In addition, we compile current open-source benchmarks and evaluation metrics, focusing on their scale, composition, calculation methods, and assessment dimensions. Next, we introduce several application scenarios for the LLM-based tool learning methods. Finally, we discuss the safety and ethical issues involved in tool learning.","author":[{"family":"Xu","given":"Weikai"},{"family":"Huang","given":"Chengrui"},{"family":"Gao","given":"Shen"},{"family":"Shang","given":"Shuo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s41019-025-00296-9","URL":"https://doi.org/10.1007/s41019-025-00296-9","source":"openalex"},{"id":"oa:W4409047908","type":"article-journal","title":"Practical Federated Learning without a Server","abstract":"Federated Learning (FL) enables end-user devices to collaboratively train ML models without sharing raw data, thereby preserving data privacy. In FL, a central parameter server coordinates the learning process by iteratively aggregating the trained models received from clients. Yet, deploying a central server is not always feasible due to hardware unavailability, infrastructure constraints, or operational costs. We present Plexus, a fully decentralized FL system for large networks that operates without the drawbacks originating from having a central server. Plexus distributes the responsibilities of model aggregation and sampling among participating nodes while avoiding network-wide coordination. We evaluate Plexus using realistic traces for compute speed, pairwise latency and network capacity. Our experiments on three common learning tasks and with up to 1000 nodes empirically show that Plexus reduces time-to-accuracy by 1.4--1.6×, communication volume by 15.8--292× and training resources needed for convergence by 30.5--77.9× compared to conventional decentralized learning algorithms.","author":[{"family":"Dhasade","given":"Akash"},{"family":"Kermarrec","given":"Anne"},{"family":"Lavoie","given":"Erick"},{"family":"Pouwelse","given":"Johan"},{"family":"Sharma","given":"Rishi"},{"family":"Vos","given":"Martijn"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1145/3721146.3721938","URL":"https://doi.org/10.1145/3721146.3721938","source":"openalex"},{"id":"oa:W4406193279","type":"manuscript","title":"A Survey on Federated Learning in Human Sensing","abstract":"Human Sensing, a field that leverages technology to monitor human activities, psycho-physiological states, and interactions with the environment, enhances our understanding of human behavior and drives the development of advanced services that improve overall quality of life. However, its reliance on detailed and often privacy-sensitive data as the basis for its machine learning (ML) models raises significant legal and ethical concerns. The recently proposed ML approach of Federated Learning (FL) promises to alleviate many of these concerns, as it is able to create accurate ML models without sending raw user data to a central server. While FL has demonstrated its usefulness across a variety of areas, such as text prediction and cyber security, its benefits in Human Sensing are under-explored, given the particular challenges in this domain. This survey conducts a comprehensive analysis of the current state-of-the-art studies on FL in Human Sensing, and proposes a taxonomy and an eight-dimensional assessment for FL approaches. Through the eight-dimensional assessment, we then evaluate whether the surveyed studies consider a specific FL-in-Human-Sensing challenge or not. Finally, based on the overall analysis, we discuss open challenges and highlight five research aspects related to FL in Human Sensing that require urgent research attention. Our work provides a comprehensive corpus of FL studies and aims to assist FL practitioners in developing and evaluating solutions that effectively address the real-world complexities of Human Sensing.","author":[{"family":"Li","given":"Mohan"},{"family":"Gjoreski","given":"Martin"},{"family":"Barbiero","given":"Pietro"},{"family":"Slapničar","given":"Gašper"},{"family":"Luštrek","given":"Mitja"},{"family":"Lane","given":"Nicholas"},{"family":"Langheinrich","given":"Marc"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.04000","URL":"https://doi.org/10.48550/arxiv.2501.04000","source":"openalex"},{"id":"doi:10.20944/preprints202506.0088.v1","type":"manuscript","title":"Federated Learning for Privacy-Preserving Defense in Power Cyber-Physical Systems: Frameworks, Techniques, and Challenges","abstract":"The increasing interconnection and digitalization of modern energy systems have intensified cybersecurity vulnerabilities in Power Cyber-Physical Systems (Power CPS). Traditional centralized defense approaches struggle to balance privacy preservation, scalability, and collaborative responsiveness across distributed infrastructures. Federated Learning (FL) emerges as a promising paradigm that enables distributed, privacy-preserving model training without sharing raw data. This review presents a comprehensive analysis of FL-based collaborative defense for Power CPS, spanning threat modeling, architectural taxonomies, privacy-preserving mechanisms, and real-world applications. We categorize FL techniques by learning structure, synchronization, and personalization, and examine privacy-enhancing technologies such as differential privacy, secure multiparty computation, homomorphic encryption, and trusted execution environments. Practical applications across substations, SCADA systems, WAMS, and EV infrastructures are reviewed alongside deployment challenges such as communication overhead, adversarial threats, and operational constraints. A roadmap is proposed for future research in cross-layer FL architectures, federated reinforcement learning, and regulatory standardization. The review concludes by advocating for cross-sector collaboration to operationalize federated defense as a cornerstone of resilient, secure, and privacy-compliant smart grids.","author":[{"family":"Wang","given":"Xiaokang"}],"issued":{"date-parts":[[2025]]},"DOI":"10.20944/preprints202506.0088.v1","URL":"https://doi.org/10.20944/preprints202506.0088.v1","source":"europepmc"},{"id":"doi:10.48550/arxiv.2605.28112","type":"manuscript","title":"A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG","abstract":"Federated Retrieval-Augmented Generation (FedRAG) is attractive for privacy-sensitive applications because full local corpora remain on clients. As a result, routing must rely on client-provided semantic profiles, creating a new opportunity for manipulation. We introduce Routing Hijacking, a routing-stage attack in which a malicious client forges its profile to attract target queries despite having irrelevant underlying data. We show that this vulnerability is severe. Across three representative FedRAG routing architectures, Routing Hijacking consistently misroutes target queries and leads to downstream disruptions and failures, including missing evidence, poisoning, incorrect answers, and hallucinations. In a controlled MedQA-USMLE stress test, we further show that poisoned retrieved evidence can mislead models across scales, leading to incorrect answers, hallucinations, and sycophantic failures. Existing defenses do not close this gap: encrypted routing preserves the exploited ranking, and Byzantine-robust Federated Learning (FL) rules transfer poorly to heterogeneous routing profiles. To address this gap, we propose a trust-aware post-routing framework that reweights clients using returned-evidence feedback, including retrieval relevance, profile consistency, and cross-client agreement; online experiments show that it suppresses persistent hijacking over recurring queries and transfers to a learned neural router. Our findings establish routing integrity as a security challenge in FedRAG and highlight the need for stronger defenses for secure federated retrieval.","author":[{"family":"Mu","given":"Junjie"},{"family":"Li","given":"Qiongxiu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2605.28112","URL":"https://doi.org/10.48550/arxiv.2605.28112","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.27836","type":"manuscript","title":"FISGuard: Defending Against Membership Inference via Fixed Input Subspaces","abstract":"As large language models are increasingly adopted in federated learning, protecting user privacy while performing parameter-efficient fine-tuning on distributed private data has become an important challenge. Although clients only share gradients instead of directly uploading raw data, the shared gradients may still leak membership information about training samples. ProjRes (S&amp;P, 2026) further increases this risk: with less information and without accessing model outputs, an attacker can effectively distinguish members from non-members solely based on the projection residual between a candidate representation and the subspace induced by server-observable gradients. Existing defenses against membership inference mostly rely on gradient perturbation or regularization, which can not only degrade model utility but also fail to effectively defend against the membership inference attack introduced by ProjRes, which exploits the geometric structure of gradients. To address this issue, we propose FISGuard, a lightweight defense. Its key idea is to construct and fix a low-dimensional representation subspace using independent public data, thereby restricting the space through which private representations are exposed via gradients while preserving the primary information required for downstream tasks. This substantially reduces the projection-residual discrepancy between members and non-members. We evaluate FISGuard against five representative defense methods across three NLP datasets, two LLMs, and two fine-tuning strategies, Adapter and LoRA. The results show that FISGuard reduces the ProjRes attack AUC to near the random-guessing level of 0.5 in most settings, while maintaining downstream task performance close to that of the undefended model and introducing only limited computational overhead, thereby achieving a favorable privacy--utility trade-off.","author":[{"family":"Jiang","given":"Haocheng"},{"family":"Shen","given":"Hua"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.27836","URL":"https://doi.org/10.48550/arxiv.2608.27836","source":"datacite"},{"id":"doi:10.5281/zenodo.22182979","type":"article-journal","title":"The Decentralized Classroom: A Narrative Review of Federated Learning from Google's Keyboards to the Privacy's Frontier","abstract":"Federated learning---the artificial intelligence whose subject is the decentralized classroom and whose lesson is the model's travel---moved from Dwork's 2006 differential privacy and Shokri and Shmatikov's 2015 gradients through Konečný's 2016 compression, McMahan's 2017 FedAvg, and Bonawitz's 2017 aggregation to Zhao's 2018 non-IID, Kairouz's 2021 survey, and Zhu's 2019 leakage. This article presents a narrative review of that arc's canonical line: Dwork's 2006 ICALP, Shokri and Shmatikov's 2015 CCS, Konečný and colleagues's 2016 strategies, McMahan, Moore, Ramage, Hampson, and Arcas's 2017 FedAvg, Bonawitz and colleagues's 2017 secure aggregation, Zhao and colleagues's 2018 non-IID, Hard and colleagues's 2018 keyboard, Zhu, Liu, and Han's 2019 gradients, Yang and colleagues's 2019 concept, Li and colleagues's 2020 convergence, Li and colleagues's 2020 challenges, and Kairouz and colleagues's 2021 advances. The review is organized around three themes: the privacy's premise and the communication's bottleneck, in which the Dwork's noise and the Shokri-Shmatikov's gradients founded the distributed's training; the algorithm's and the deployment's era, in which the FedAvg's averaging, the secure's aggregation, and the keyboard's deployment gave the federation its engine; and the heterogeneity's and the frontier's era, in which the non-IID's data, the gradient's leakage, the convergence's proofs, and the open's problems carried the field into the privacy's science. It is concluded that federated learning is the machine learning's decentralization---and that its arc is the classroom's reading from the centralized's server to the privacy's frontier.","author":[{"family":"Revista","given":"Zen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22182979","URL":"https://doi.org/10.5281/zenodo.22182979","source":"datacite"},{"id":"doi:10.5281/zenodo.22182978","type":"article-journal","title":"The Decentralized Classroom: A Narrative Review of Federated Learning from Google's Keyboards to the Privacy's Frontier","abstract":"Federated learning---the artificial intelligence whose subject is the decentralized classroom and whose lesson is the model's travel---moved from Dwork's 2006 differential privacy and Shokri and Shmatikov's 2015 gradients through Konečný's 2016 compression, McMahan's 2017 FedAvg, and Bonawitz's 2017 aggregation to Zhao's 2018 non-IID, Kairouz's 2021 survey, and Zhu's 2019 leakage. This article presents a narrative review of that arc's canonical line: Dwork's 2006 ICALP, Shokri and Shmatikov's 2015 CCS, Konečný and colleagues's 2016 strategies, McMahan, Moore, Ramage, Hampson, and Arcas's 2017 FedAvg, Bonawitz and colleagues's 2017 secure aggregation, Zhao and colleagues's 2018 non-IID, Hard and colleagues's 2018 keyboard, Zhu, Liu, and Han's 2019 gradients, Yang and colleagues's 2019 concept, Li and colleagues's 2020 convergence, Li and colleagues's 2020 challenges, and Kairouz and colleagues's 2021 advances. The review is organized around three themes: the privacy's premise and the communication's bottleneck, in which the Dwork's noise and the Shokri-Shmatikov's gradients founded the distributed's training; the algorithm's and the deployment's era, in which the FedAvg's averaging, the secure's aggregation, and the keyboard's deployment gave the federation its engine; and the heterogeneity's and the frontier's era, in which the non-IID's data, the gradient's leakage, the convergence's proofs, and the open's problems carried the field into the privacy's science. It is concluded that federated learning is the machine learning's decentralization---and that its arc is the classroom's reading from the centralized's server to the privacy's frontier.","author":[{"family":"Revista","given":"Zen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22182978","URL":"https://doi.org/10.5281/zenodo.22182978","source":"datacite"},{"id":"doi:10.5281/zenodo.20469968","type":"article-journal","title":"Federated Learning Aggregation Techniques and Multimodal Model Efficiency at the Edge","abstract":"This report synthesises findings from 13 peer-reviewed papers addressing the following research question: What is the impact of different federated learning aggregation techniques on the inference efficiency and latency of multimodal models deployed across heterogeneous edge devices. Abstract The rapid evolution of large language models (LLMs) has driven a transformative shift in artificial intelligence (AI), reshaping both research paradigms and practical applications. Distinguished from their predecessors by unprecedented scale and advanced capabilities. 9 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 7.8/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the impact of different federated learning aggregation techniques on the inference efficiency and latency of multimodal models deployed across heterogeneous edge devices? Autonomous literature synthesis. Automated review score: 7.8/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20469968","URL":"https://doi.org/10.5281/zenodo.20469968","source":"datacite"},{"id":"doi:10.5281/zenodo.20469969","type":"article-journal","title":"Federated Learning Aggregation Techniques and Multimodal Model Efficiency at the Edge","abstract":"This report synthesises findings from 13 peer-reviewed papers addressing the following research question: What is the impact of different federated learning aggregation techniques on the inference efficiency and latency of multimodal models deployed across heterogeneous edge devices. Abstract The rapid evolution of large language models (LLMs) has driven a transformative shift in artificial intelligence (AI), reshaping both research paradigms and practical applications. Distinguished from their predecessors by unprecedented scale and advanced capabilities. 9 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 7.8/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the impact of different federated learning aggregation techniques on the inference efficiency and latency of multimodal models deployed across heterogeneous edge devices? Autonomous literature synthesis. Automated review score: 7.8/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20469969","URL":"https://doi.org/10.5281/zenodo.20469969","source":"datacite"},{"id":"doi:10.5281/zenodo.22178214","type":"article-journal","title":"Computer-Assisted Resolution of the Direct Product Conjecture via Kullback-Leibler (KL) Divergence Information Tensorization and Pinsker-Bounded Simulation Operators (Cloud Security)","abstract":"Using the Lean 4 interactive theorem prover, this paper presents a machine-verified structural resolution of the Direct Product Conjecture in communication complexity. Extending Hilbert space orthogonal projections and ANOVA-Hoeffding decompositions beyond linear variance, we bridge non-linear information metrics through Kullback-Leibler divergence tensorization, Han's inequality, and a Pinsker-bounded simulation operator. By controlling cross-coordinate conditioning drift, our deductive pipeline formally proves the affirmative resolution: parallel execution cannot bypass single-instance information costs, establishing that: $$CC_\\epsilon(f^k) \\ge k \\cdot R_\\delta(f) - o(k)$$ This rigorously verified lower bound resolves the conjecture, providing essential, provable resource guarantees for secure multi-party computation and distributed cloud infrastructure.","author":[{"family":"Reed","given":"Jonathan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22178214","URL":"https://doi.org/10.5281/zenodo.22178214","source":"datacite"},{"id":"doi:10.5281/zenodo.22178215","type":"article-journal","title":"Computer-Assisted Resolution of the Direct Product Conjecture via Kullback-Leibler (KL) Divergence Information Tensorization and Pinsker-Bounded Simulation Operators (Cloud Security)","abstract":"Using the Lean 4 interactive theorem prover, this paper presents a machine-verified structural resolution of the Direct Product Conjecture in communication complexity. Extending Hilbert space orthogonal projections and ANOVA-Hoeffding decompositions beyond linear variance, we bridge non-linear information metrics through Kullback-Leibler divergence tensorization, Han's inequality, and a Pinsker-bounded simulation operator. By controlling cross-coordinate conditioning drift, our deductive pipeline formally proves the affirmative resolution: parallel execution cannot bypass single-instance information costs, establishing that: $$CC_\\epsilon(f^k) \\ge k \\cdot R_\\delta(f) - o(k)$$ This rigorously verified lower bound resolves the conjecture, providing essential, provable resource guarantees for secure multi-party computation and distributed cloud infrastructure.","author":[{"family":"Reed","given":"Jonathan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22178215","URL":"https://doi.org/10.5281/zenodo.22178215","source":"datacite"},{"id":"doi:10.5281/zenodo.20457919","type":"article-journal","title":"TOWARDS PROACTIVE CYBER DEFENSE CYBER ATTACK PREDICTION USING ADVANCED AI TECHNIQUES","abstract":"Develop an AI-based predictive analytics platform that is capable of identifying potential cyber threats. In real time, transformer topologies, deep learning, and graph neural networks replicate complex, high-dimensional security data streams. Temporal sequence models and behavioral analytics can identify attacks prior to their infliction of damage. Networks, system records, and individuals assist us in identifying potential hazards. Mixed learning stabilizes unlabeled data by employing reinforcement learning, self-supervised learning, and supervised learning. People are prepared for assaults through online learning and regulatory concepts. Home-based businesses are safeguarded by privacy-preserving learning and federated AI. The system's high recognition rate and low false alarm rate are demonstrated in numerous real-world and benchmark dataset trials. The design prevents the execution of APTs, zero-day vulnerabilities, and malware that alters its shape. Delete the warnings associated with the AI module. This enables security specialists to make decisions promptly. Growth is expedited by edge-cloud connectivity and distributed training.","author":[{"family":"Insights","given":"International"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20457919","URL":"https://doi.org/10.5281/zenodo.20457919","source":"datacite"},{"id":"doi:10.5281/zenodo.20457920","type":"article-journal","title":"TOWARDS PROACTIVE CYBER DEFENSE CYBER ATTACK PREDICTION USING ADVANCED AI TECHNIQUES","abstract":"Develop an AI-based predictive analytics platform that is capable of identifying potential cyber threats. In real time, transformer topologies, deep learning, and graph neural networks replicate complex, high-dimensional security data streams. Temporal sequence models and behavioral analytics can identify attacks prior to their infliction of damage. Networks, system records, and individuals assist us in identifying potential hazards. Mixed learning stabilizes unlabeled data by employing reinforcement learning, self-supervised learning, and supervised learning. People are prepared for assaults through online learning and regulatory concepts. Home-based businesses are safeguarded by privacy-preserving learning and federated AI. The system's high recognition rate and low false alarm rate are demonstrated in numerous real-world and benchmark dataset trials. The design prevents the execution of APTs, zero-day vulnerabilities, and malware that alters its shape. Delete the warnings associated with the AI module. This enables security specialists to make decisions promptly. Growth is expedited by edge-cloud connectivity and distributed training.","author":[{"family":"Insights","given":"International"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20457920","URL":"https://doi.org/10.5281/zenodo.20457920","source":"datacite"},{"id":"doi:10.5281/zenodo.22179153","type":"article-journal","title":"Learning Where the Data Lives: A Narrative Review of Federated Learning from Differential Privacy to the Open Problems","abstract":"Federated learning---training a shared model across many devices that never surrender their data---inverted machine learning's architecture: instead of data to the model, the model to the data. This article presents a narrative review of that arc's canonical line: Dwork and colleagues' 2006 differential privacy, Dwork and Roth's 2014 foundations, Shokri and Shmatikov's 2015 privacy-preserving deep learning, Konecny and colleagues' 2016 communication strategies, McMahan and colleagues' 2017 FedAvg, Bonawitz and colleagues' 2017 secure aggregation, Zhao and colleagues' 2018 non-IID study, Bonawitz and colleagues' 2019 scale design, Yang and colleagues' 2019 concept paper, Zhu, Liu, and Han's 2019 gradient leakage, Li and colleagues' 2020 convergence analysis, and Kairouz and colleagues' 2021 open problems. The synthesis is organized around three themes: privacy, in which differential privacy's calculus and secure aggregation made learning without exposure precise; algorithm, in which FedAvg's weighted averaging met the heterogeneity of real devices and real data; and system, in which scale deployments faced stragglers, leakage, and the statistical reality of non-IID partitions. It is concluded that federated learning is privacy engineering's rare full-stack success---its limits as precisely mapped as its promise---and that its open problems are the field's charter: heterogeneity, security, and the economics of participation.","author":[{"family":"Revista","given":"Zen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22179153","URL":"https://doi.org/10.5281/zenodo.22179153","source":"datacite"},{"id":"doi:10.5281/zenodo.22179152","type":"article-journal","title":"Learning Where the Data Lives: A Narrative Review of Federated Learning from Differential Privacy to the Open Problems","abstract":"Federated learning---training a shared model across many devices that never surrender their data---inverted machine learning's architecture: instead of data to the model, the model to the data. This article presents a narrative review of that arc's canonical line: Dwork and colleagues' 2006 differential privacy, Dwork and Roth's 2014 foundations, Shokri and Shmatikov's 2015 privacy-preserving deep learning, Konecny and colleagues' 2016 communication strategies, McMahan and colleagues' 2017 FedAvg, Bonawitz and colleagues' 2017 secure aggregation, Zhao and colleagues' 2018 non-IID study, Bonawitz and colleagues' 2019 scale design, Yang and colleagues' 2019 concept paper, Zhu, Liu, and Han's 2019 gradient leakage, Li and colleagues' 2020 convergence analysis, and Kairouz and colleagues' 2021 open problems. The synthesis is organized around three themes: privacy, in which differential privacy's calculus and secure aggregation made learning without exposure precise; algorithm, in which FedAvg's weighted averaging met the heterogeneity of real devices and real data; and system, in which scale deployments faced stragglers, leakage, and the statistical reality of non-IID partitions. It is concluded that federated learning is privacy engineering's rare full-stack success---its limits as precisely mapped as its promise---and that its open problems are the field's charter: heterogeneity, security, and the economics of participation.","author":[{"family":"Revista","given":"Zen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22179152","URL":"https://doi.org/10.5281/zenodo.22179152","source":"datacite"},{"id":"doi:10.5281/zenodo.20775412","type":"article-journal","title":"Privacy-Preserving Local LLM Inference for Developer Tooling — AI, LLM, Privacy, Sovereign AI, and Post-Cloud Architecture (Anticode)","abstract":"The proliferation of large language models (LLMs) in developer tooling has introduced significant privacy and security concerns, particularly when code---often containing proprietary algorithms, credentials, and business logic---is transmitted to remote inference servers. This paper presents a comprehensive analysis of privacy-preserving techniques for local LLM inference in the context of terminal-native AI coding engines, with specific application to the ANTIKODE architecture. We examine confidential computing, federated learning, differential privacy, and on-device inference as mechanisms to ensure that source code and developer telemetry never leave the local machine. Through systematic evaluation of model quantization, hardware security modules, and hash-chained audit trails, we demonstrate that local-first LLM inference achieves comparable code generation quality to cloud-based alternatives while eliminating data exfiltration risks. Our findings indicate that 4-bit quantized 7B-parameter models running on consumer hardware can match the functional performance of larger cloud models for the majority of code completion tasks, with privacy guarantees that satisfy SOC2, GDPR, HIPAA, and FedRAMP requirements. We further show that ANTIKODE's .aioss ledger provides cryptographic verification that no inference data has been transmitted externally, establishing a new standard for trusted AI-assisted development. Part of The Anticloud research corpus by Lois-Kleinner Alpasan (ORCID: 0009-0009-2233-6107). This work explores ai, llm in the context of sovereign AI infrastructure, post-cloud computing architectures, and transparent, blackbox-free systems.","author":[{"family":"Alpasan","given":"Lois"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20775412","URL":"https://doi.org/10.5281/zenodo.20775412","source":"datacite"},{"id":"doi:10.5281/zenodo.20775413","type":"article-journal","title":"Privacy-Preserving Local LLM Inference for Developer Tooling — AI, LLM, Privacy, Sovereign AI, and Post-Cloud Architecture (Anticode)","abstract":"The proliferation of large language models (LLMs) in developer tooling has introduced significant privacy and security concerns, particularly when code---often containing proprietary algorithms, credentials, and business logic---is transmitted to remote inference servers. This paper presents a comprehensive analysis of privacy-preserving techniques for local LLM inference in the context of terminal-native AI coding engines, with specific application to the ANTIKODE architecture. We examine confidential computing, federated learning, differential privacy, and on-device inference as mechanisms to ensure that source code and developer telemetry never leave the local machine. Through systematic evaluation of model quantization, hardware security modules, and hash-chained audit trails, we demonstrate that local-first LLM inference achieves comparable code generation quality to cloud-based alternatives while eliminating data exfiltration risks. Our findings indicate that 4-bit quantized 7B-parameter models running on consumer hardware can match the functional performance of larger cloud models for the majority of code completion tasks, with privacy guarantees that satisfy SOC2, GDPR, HIPAA, and FedRAMP requirements. We further show that ANTIKODE's .aioss ledger provides cryptographic verification that no inference data has been transmitted externally, establishing a new standard for trusted AI-assisted development. Part of The Anticloud research corpus by Lois-Kleinner Alpasan (ORCID: 0009-0009-2233-6107). This work explores ai, llm in the context of sovereign AI infrastructure, post-cloud computing architectures, and transparent, blackbox-free systems.","author":[{"family":"Alpasan","given":"Lois"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20775413","URL":"https://doi.org/10.5281/zenodo.20775413","source":"datacite"},{"id":"doi:10.5281/zenodo.21777670","type":"article-journal","title":"Computación de borde y computación en la niebla: un análisis descriptivo de la evolución de la computación en la nube para aplicaciones de IoT","abstract":"Introducción: El vertiginoso crecimiento del Internet de las Cosas (IoT) ha evidenciado las limitaciones estructurales de la computación en la nube centralizada para satisfacer las demandas de latencia, ancho de banda y privacidad de las aplicaciones en tiempo real. Objetivo: Analizar el estado del arte del Edge Computing y el Fog Computing como paradigmas evolutivos de la computación en la nube para aplicaciones IoT, examinando sus características arquitectónicas, ventajas comparativas, casos de uso y desafíos pendientes. Metodología: Se desarrolló una revisión bibliográfica no sistemática de nivel descriptivo con método de análisis-síntesis, consultando fuentes publicadas en IEEE Xplore, Scopus, SpringerLink, MDPI y Taylor & Francis, empleando combinaciones booleanas de términos como Edge Computing, Fog Computing, IoT, latency y real-time applications, priorizando publicaciones entre 2024 y 2026. Resultados: La arquitectura de tres niveles Edge-Fog-Cloud distribuye eficientemente el procesamiento según la criticidad temporal de cada tarea, reduciendo la latencia hasta un 40% con Fog Computing y un 30% con Edge Computing respecto a modelos exclusivamente en la nube, y disminuyendo el consumo energético total hasta un 30%. La integración de Aprendizaje Federado, Aprendizaje por Refuerzo Profundo y modelos compactos de redes neuronales amplía las capacidades de inferencia distribuida en dispositivos de recursos limitados. Las aplicaciones abarcan salud inteligente, ciudades inteligentes, industria 4.0/5.0 y agricultura de precisión. Conclusión: Los paradigmas Edge y Fog Computing constituyen extensiones complementarias e imprescindibles de la nube centralizada, cuya convergencia con redes 6G, gemelos digitales, computación cuántica y Aprendizaje Federado avanzado definirá la arquitectura computacional de la próxima generación de ecosistemas IoT. Área de estudio general: Tecnologías de la Información y la Comunicación. Área de estudio específica: Computación Distribuida y Arquitecturas de Red para el Internet de las Cosas.","author":[{"family":"Pérez Insuasti","given":"Juan"},{"family":"Flores-Andino","given":"Víctor"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21777670","URL":"https://doi.org/10.5281/zenodo.21777670","source":"datacite"},{"id":"doi:10.5281/zenodo.21777671","type":"article-journal","title":"Computación de borde y computación en la niebla: un análisis descriptivo de la evolución de la computación en la nube para aplicaciones de IoT","abstract":"Introducción: El vertiginoso crecimiento del Internet de las Cosas (IoT) ha evidenciado las limitaciones estructurales de la computación en la nube centralizada para satisfacer las demandas de latencia, ancho de banda y privacidad de las aplicaciones en tiempo real. Objetivo: Analizar el estado del arte del Edge Computing y el Fog Computing como paradigmas evolutivos de la computación en la nube para aplicaciones IoT, examinando sus características arquitectónicas, ventajas comparativas, casos de uso y desafíos pendientes. Metodología: Se desarrolló una revisión bibliográfica no sistemática de nivel descriptivo con método de análisis-síntesis, consultando fuentes publicadas en IEEE Xplore, Scopus, SpringerLink, MDPI y Taylor & Francis, empleando combinaciones booleanas de términos como Edge Computing, Fog Computing, IoT, latency y real-time applications, priorizando publicaciones entre 2024 y 2026. Resultados: La arquitectura de tres niveles Edge-Fog-Cloud distribuye eficientemente el procesamiento según la criticidad temporal de cada tarea, reduciendo la latencia hasta un 40% con Fog Computing y un 30% con Edge Computing respecto a modelos exclusivamente en la nube, y disminuyendo el consumo energético total hasta un 30%. La integración de Aprendizaje Federado, Aprendizaje por Refuerzo Profundo y modelos compactos de redes neuronales amplía las capacidades de inferencia distribuida en dispositivos de recursos limitados. Las aplicaciones abarcan salud inteligente, ciudades inteligentes, industria 4.0/5.0 y agricultura de precisión. Conclusión: Los paradigmas Edge y Fog Computing constituyen extensiones complementarias e imprescindibles de la nube centralizada, cuya convergencia con redes 6G, gemelos digitales, computación cuántica y Aprendizaje Federado avanzado definirá la arquitectura computacional de la próxima generación de ecosistemas IoT. Área de estudio general: Tecnologías de la Información y la Comunicación. Área de estudio específica: Computación Distribuida y Arquitecturas de Red para el Internet de las Cosas.","author":[{"family":"Pérez Insuasti","given":"Juan"},{"family":"Flores-Andino","given":"Víctor"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21777671","URL":"https://doi.org/10.5281/zenodo.21777671","source":"datacite"},{"id":"doi:10.5281/zenodo.20531687","type":"article-journal","title":"Deterministic and Auditable Routing for Artificial Intelligence in Regulated Environments","abstract":"# Deterministic and Auditable Routing for Artificial Intelligence in Regulated Environments **Felippe Barcelos** Independent Researcher felippe.barcelos10@gmail.com *Preprint — submitted to Zenodo. Version 1.0.0 — Protocol P55.0 — 2026-06-03* --- ## Abstract The deployment of Large Language Models (LLMs) in regulated industries — banking, healthcare, and government — demands audit trails that satisfy strict reproducibility requirements under the EU AI Act (2024), the General Data Protection Regulation (GDPR), and sector-specific frameworks such as HIPAA and Basel III/SR 11-7. Existing ML-based routing systems, including RouteLLM and FrugalGPT, achieve significant cost reductions but produce non-reproducible routing decisions whose outputs change with model retraining, rendering them incompatible with formal compliance auditing. We present the **TEIA Cognitive Router**, a compliance-first LLM routing system based on a fixed six-axis semantic entropy formula that produces deterministic routing decisions without neural weights, training data, or external dependencies. The system guarantees the *Write==Read invariant*: identical input text always yields an identical routing decision, an identical canonical JSON representation, and an identical SHA-256 audit seal. Routing decisions are organized in an append-only Merkle time-anchor chain and can be notarized by any RFC 3161-compliant Trusted Timestamp Authority (TSA) for legally binding external proof of existence. Evaluation on a 100-prompt simulation aligned with MT-Bench benchmark categories demonstrates **99.6% quality retention with 16.3% cost reduction** in compliance-safe mode, and **73.8% quality retention with 98.4% cost reduction** in max-savings mode. We provide a formal compliance mapping to EU AI Act Articles 12, 13, and Annex IV; GDPR Article 22; SOC 2 CC7; and HIPAA §164.312(b). The system is released under Apache 2.0 and is available on PyPI as `teia-cognitive-router`. **Keywords:** LLM routing, compliance, determinism, audit trail, EU AI Act, GDPR, SHA-256, Merkle chain, RFC 3161, semantic entropy **arXiv classifications:** cs.CR (Cryptography and Security) · cs.AI (Artificial Intelligence) · cs.LG (Machine Learning) --- ## 1. Introduction The deployment of Large Language Models (LLMs) across regulated industries has accelerated markedly since the public release of instruction-tuned models beginning in 2022. Financial institutions leverage LLMs for risk analysis, contract review, and fraud detection. Healthcare providers deploy them for clinical documentation, drug-interaction queries, and patient communication. Government agencies apply them to policy analysis, benefits adjudication, and citizen services. This acceleration has collided with a fundamental regulatory barrier: the *audit reproducibility problem*. Modern regulatory frameworks — the European Union AI Act (2024), GDPR (2018), HIPAA (1996), and the Federal Reserve's SR 11-7 model risk guidance — require that automated decision-making systems be *auditable*, *reproducible*, and *explainable*. An organization must be able to answer, for any past automated decision: \"Why did this happen? Can you prove it happened exactly this way? Has the record been tampered with?\" LLM deployments routinely employ multi-tier model architectures to optimize cost: a small, fast local model handles simple tasks; a large, expensive cloud model handles complex tasks. The routing decision — which tier receives each request — is itself an AI-driven automation. Yet this decision is typically made by an opaque ML classifier (RouteLLM [1]), a learned cascade (FrugalGPT [2]), or an informal heuristic with no formal audit trail. When a compliance officer asks for proof of how a specific routing decision was made six months ago, the organization cannot provide a mathematically verifiable answer. This paper makes the following contributions: 1. **TEIA Cognitive Router**: A fixed arithmetic routing formula that produces deterministic, ve","author":[{"family":"Barcelos","given":"Felippe"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20531687","URL":"https://doi.org/10.5281/zenodo.20531687","source":"datacite"},{"id":"doi:10.5281/zenodo.21114272","type":"article-journal","title":"On-Device Glucose Alarms from a Single Learned Token: Pre-Registered Cross- Dataset Validation of a Class-Discriminant Codebook Across Eleven CGM Datasets","abstract":"Description This record accompanies a manuscript validating a single-token, on-device encoder for continuous glucose monitoring (CGM). Each glucose window is reduced to one compact learned token from which a decision is read on the sensor itself. Under strict pre-registration — frozen hypotheses, patient-disjoint splits, and bootstrap confidence intervals — the encoder predicts hypoglycemic and hyperglycemic excursions 30–60 minutes ahead (AUC 0.93–0.97) and, trained on a single cohort, generalizes without retraining to nine independent public CGM datasets (mean AUC 0.878). A population-level federated refresh adds a further measured gain, while per-individual personalization and multi-signal fusion are shown, with honest boundaries, to add little. Method companion: Paper 19 (doi:10.5281/zenodo.20788187). Keywords: continuous glucose monitoring; hypoglycemia prediction; hyperglycemia prediction; class-discriminant codebook; vector quantization; on-device machine learning; edge inference; cross-dataset generalization; federated refresh; pre-registration; type 1 diabetes; type 2 diabetes References 1. N. Tishby, F. C. Pereira, and W. Bialek, \"The information bottleneck method,\" in Proc. 37th Allerton Conf. Communication, Control, and Computing, 1999, pp. 368–377. 2. R. M. Gray, \"Vector quantization,\" IEEE ASSP Magazine, vol. 1, no. 2, pp. 4–29, 1984. 3. R. A. Fisher, \"The use of multiple measurements in taxonomic problems,\" Annals of Eugenics, vol. 7, no. 2, pp. 179–188, 1936. 4. S. P. Lloyd, \"Least squares quantization in PCM,\" IEEE Transactions on Information Theory, vol. 28, no. 2, pp. 129–137, 1982. 5. Q. Zhao et al., \"Chinese diabetes datasets for data-driven machine learning,\" Scientific Data, vol. 10, no. 35, 2023. 6. J. I. Hidalgo, J. Alvarado, M. Botella, A. Aramendi, J. M. Velasco, and O. Garnica, \"HUPA-UCM diabetes dataset,\" Data in Brief, vol. 55, art. 110559, 2024, doi:10.1016/j.dib.2024.110559. 7. T. Battelino et al., \"Clinical targets for continuous glucose monitoring data interpretation: recommendations from the international consensus on time in range,\" Diabetes Care, vol. 42, no. 8, pp. 1593–1603, 2019. 8. S. Oviedo, J. Vehí, R. Calm, and J. Armengol, \"A review of personalized blood glucose prediction strategies for T1DM patients,\" International Journal for Numerical Methods in Biomedical Engineering, vol. 33, no. 6, e2833, 2017. 9. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, \"Communication-efficient learning of deep networks from decentralized data,\" in Proc. 20th Int. Conf. Artificial Intelligence and Statistics (AISTATS), 2017, pp. 1273–1282. 10. B. A. Nosek, C. R. Ebersole, A. C. DeHaven, and D. T. Mellor, \"The preregistration revolution,\" Proc. National Academy of Sciences, vol. 115, no. 11, pp. 2600–2606, 2018. 11. B. Efron and R. J. Tibshirani, An Introduction to the Bootstrap. New York: Chapman & Hall, 1993. 12. R. J. Ferlic and K. K. Ferlic, \"A single-token class-discriminant codebook for sensor encoding (Paper 19),\" Zenodo, 2026, doi:10.5281/zenodo.20788187. 13. U.S. Food and Drug Administration, \"Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD) Action Plan,\" 2021. 14. C. Marling and R. Bunescu, \"The OhioT1DM dataset for blood glucose level prediction: Update 2020,\" in Proc. 5th Int. Workshop on Knowledge Discovery in Healthcare Data, CEUR Workshop Proc., vol. 2675, 2020, pp. 71–74. 15. gluco-tsfm-benchmark: an aggregated continuous-glucose-monitoring benchmark of eleven public datasets, Hugging Face Datasets, https://huggingface.co/datasets/byluuu/gluco-tsfm-benchmark (accessed 2026). License This work is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0). The encoding method and its federated and privacy mechanisms are covered by previously filed U.S. provisional patent applications (Nos. 64/095,354; 64/084,807; 64/084,817; 64/084,821); patent rights are separate from the copyright license. Public datase","author":[{"family":"Ferlic","given":"Randolph"},{"family":"Ferlic","given":"Kimberly"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21114272","URL":"https://doi.org/10.5281/zenodo.21114272","source":"datacite"},{"id":"doi:10.5281/zenodo.21114273","type":"article-journal","title":"On-Device Glucose Alarms from a Single Learned Token: Pre-Registered Cross- Dataset Validation of a Class-Discriminant Codebook Across Eleven CGM Datasets","abstract":"Description This record accompanies a manuscript validating a single-token, on-device encoder for continuous glucose monitoring (CGM). Each glucose window is reduced to one compact learned token from which a decision is read on the sensor itself. Under strict pre-registration — frozen hypotheses, patient-disjoint splits, and bootstrap confidence intervals — the encoder predicts hypoglycemic and hyperglycemic excursions 30–60 minutes ahead (AUC 0.93–0.97) and, trained on a single cohort, generalizes without retraining to nine independent public CGM datasets (mean AUC 0.878). A population-level federated refresh adds a further measured gain, while per-individual personalization and multi-signal fusion are shown, with honest boundaries, to add little. Method companion: Paper 19 (doi:10.5281/zenodo.20788187). Keywords: continuous glucose monitoring; hypoglycemia prediction; hyperglycemia prediction; class-discriminant codebook; vector quantization; on-device machine learning; edge inference; cross-dataset generalization; federated refresh; pre-registration; type 1 diabetes; type 2 diabetes References 1. N. Tishby, F. C. Pereira, and W. Bialek, \"The information bottleneck method,\" in Proc. 37th Allerton Conf. Communication, Control, and Computing, 1999, pp. 368–377. 2. R. M. Gray, \"Vector quantization,\" IEEE ASSP Magazine, vol. 1, no. 2, pp. 4–29, 1984. 3. R. A. Fisher, \"The use of multiple measurements in taxonomic problems,\" Annals of Eugenics, vol. 7, no. 2, pp. 179–188, 1936. 4. S. P. Lloyd, \"Least squares quantization in PCM,\" IEEE Transactions on Information Theory, vol. 28, no. 2, pp. 129–137, 1982. 5. Q. Zhao et al., \"Chinese diabetes datasets for data-driven machine learning,\" Scientific Data, vol. 10, no. 35, 2023. 6. J. I. Hidalgo, J. Alvarado, M. Botella, A. Aramendi, J. M. Velasco, and O. Garnica, \"HUPA-UCM diabetes dataset,\" Data in Brief, vol. 55, art. 110559, 2024, doi:10.1016/j.dib.2024.110559. 7. T. Battelino et al., \"Clinical targets for continuous glucose monitoring data interpretation: recommendations from the international consensus on time in range,\" Diabetes Care, vol. 42, no. 8, pp. 1593–1603, 2019. 8. S. Oviedo, J. Vehí, R. Calm, and J. Armengol, \"A review of personalized blood glucose prediction strategies for T1DM patients,\" International Journal for Numerical Methods in Biomedical Engineering, vol. 33, no. 6, e2833, 2017. 9. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, \"Communication-efficient learning of deep networks from decentralized data,\" in Proc. 20th Int. Conf. Artificial Intelligence and Statistics (AISTATS), 2017, pp. 1273–1282. 10. B. A. Nosek, C. R. Ebersole, A. C. DeHaven, and D. T. Mellor, \"The preregistration revolution,\" Proc. National Academy of Sciences, vol. 115, no. 11, pp. 2600–2606, 2018. 11. B. Efron and R. J. Tibshirani, An Introduction to the Bootstrap. New York: Chapman & Hall, 1993. 12. R. J. Ferlic and K. K. Ferlic, \"A single-token class-discriminant codebook for sensor encoding (Paper 19),\" Zenodo, 2026, doi:10.5281/zenodo.20788187. 13. U.S. Food and Drug Administration, \"Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD) Action Plan,\" 2021. 14. C. Marling and R. Bunescu, \"The OhioT1DM dataset for blood glucose level prediction: Update 2020,\" in Proc. 5th Int. Workshop on Knowledge Discovery in Healthcare Data, CEUR Workshop Proc., vol. 2675, 2020, pp. 71–74. 15. gluco-tsfm-benchmark: an aggregated continuous-glucose-monitoring benchmark of eleven public datasets, Hugging Face Datasets, https://huggingface.co/datasets/byluuu/gluco-tsfm-benchmark (accessed 2026). License This work is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0). The encoding method and its federated and privacy mechanisms are covered by previously filed U.S. provisional patent applications (Nos. 64/095,354; 64/084,807; 64/084,817; 64/084,821); patent rights are separate from the copyright license. Public datase","author":[{"family":"Ferlic","given":"Randolph"},{"family":"Ferlic","given":"Kimberly"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21114273","URL":"https://doi.org/10.5281/zenodo.21114273","source":"datacite"},{"id":"doi:10.5281/zenodo.20179571","type":"article-journal","title":"ALEVIA: A Two-Stage Deep Learning Pipeline for Multi-Label Morphological Deformity Classification in Juvenile Sparus aurata and Dicentrarchus labrax","abstract":"Intensive mariculture of juvenile gilthead sea bream (Sparus aurata) and European sea bass (Dicentrarchus labrax) depends on early morphological screening to contain deformity rates below market thresholds (< 3%), yet this process still relies almost exclusively on manual inspection. We present ALEVIA, a two-stage deep learning pipeline that processes a single lateral fish image to (i) detect, segment, and identify the species using YOLOv11, then (ii) classify the segmented crop for up to seven concurrent morphological deformity classes per species using a ResNet-34 multi-label classifier augmented with Grad-CAM++ for decision explainability. Multi-label formulation naturally handles co-occurring deformities and avoids the combinatorial proliferation of independent binary classifiers. Trained on 3,325 high-resolution laboratory images, Stage 1 achieves near-perfect segmentation performance (mAP50-95=0.995 on both species). Stage 2 reaches a weighted macro F1-score of 0.849 for S. aurataand 0.721 for D. labrax on held-out test sets. The complete pipeline exceeds the 2 img/s throughput requirement on commodity CPU hardware (median latency 315ms/image, 3.17 img/s). Grad-CAM++ visualisations confirm anatomically consistent activation patterns, providing an interpretable audit trail for domain-expert validation. The system constitutes the inference module of the ALEVIA Gaia-X-compliant federated data space for precision aquaculture.This work has been funded by the Spanish Ministry for Digital Transformation and Public Administration under the call Technological Products and Services for Data Spaces under grant TSI-100130-2024-9.","author":[{"family":"Alvarez-Osuna","given":"Javier"},{"family":"Francisco-Fernández","given":"Vilor"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20179571","URL":"https://doi.org/10.5281/zenodo.20179571","source":"datacite"},{"id":"doi:10.5281/zenodo.20179572","type":"article-journal","title":"ALEVIA: A Two-Stage Deep Learning Pipeline for Multi-Label Morphological Deformity Classification in Juvenile Sparus aurata and Dicentrarchus labrax","abstract":"Intensive mariculture of juvenile gilthead sea bream (Sparus aurata) and European sea bass (Dicentrarchus labrax) depends on early morphological screening to contain deformity rates below market thresholds (< 3%), yet this process still relies almost exclusively on manual inspection. We present ALEVIA, a two-stage deep learning pipeline that processes a single lateral fish image to (i) detect, segment, and identify the species using YOLOv11, then (ii) classify the segmented crop for up to seven concurrent morphological deformity classes per species using a ResNet-34 multi-label classifier augmented with Grad-CAM++ for decision explainability. Multi-label formulation naturally handles co-occurring deformities and avoids the combinatorial proliferation of independent binary classifiers. Trained on 3,325 high-resolution laboratory images, Stage 1 achieves near-perfect segmentation performance (mAP50-95=0.995 on both species). Stage 2 reaches a weighted macro F1-score of 0.849 for S. aurataand 0.721 for D. labrax on held-out test sets. The complete pipeline exceeds the 2 img/s throughput requirement on commodity CPU hardware (median latency 315ms/image, 3.17 img/s). Grad-CAM++ visualisations confirm anatomically consistent activation patterns, providing an interpretable audit trail for domain-expert validation. The system constitutes the inference module of the ALEVIA Gaia-X-compliant federated data space for precision aquaculture.This work has been funded by the Spanish Ministry for Digital Transformation and Public Administration under the call Technological Products and Services for Data Spaces under grant TSI-100130-2024-9.","author":[{"family":"Alvarez-Osuna","given":"Javier"},{"family":"Francisco-Fernández","given":"Vilor"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20179572","URL":"https://doi.org/10.5281/zenodo.20179572","source":"datacite"},{"id":"doi:10.5281/zenodo.21969756","type":"article-journal","title":"Robust Belief Sharing in Federated Active Inference: A Recovery-Tested Generalized-Variational Framework for Categorical Contamination-Aware Consensus","abstract":"Active Fedference is a tested, reproducible research package that connects two previously separate ideas: the belief-sharing account of distributed cognition from active inference, and robust federated learning. It reimplements the core update rules of Federated Generalized Variational Inference (FedGVI) in the discrete-categorical setting and certifies, in executable form within that implementation, the project-local identity `robust_aggregate(robustness=0) == log_linear_pool`. Under documented shared-support, posterior-log-potential, and fixed-weight assumptions, that log pool is a categorical specialization of the message-combination term in Friston et al. (2024) Eq. 7; it is not a reconstruction of the complete source protocol. In the declared confident-wrong contamination study, turning the tested robustness settings on can preserve higher true-state consensus mass while contaminated members broadcast confidently wrong beliefs; turning them off exactly recovers the project log-linear pool. Nine seeded studies, a paired statistical verdict, and a fully token-injected manuscript make every reported number reproducible from one command. The public source repository is ActiveInferenceInstitute/Active_Fedference.","author":[{"family":"Friedman","given":"Daniel"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21969756","URL":"https://doi.org/10.5281/zenodo.21969756","source":"datacite"},{"id":"doi:10.5281/zenodo.20225969","type":"article-journal","title":"ITU and Communications / Networks: A Single-Axiom View of Shannon Theory, Internet, 5G/6G, Quantum Communication","abstract":"We apply the Information-Theoretic Unification (ITU) framework (Terada 2026, concept DOI 10.5281/zenodo.20109209; current version v2.0.0 at 10.5281/zenodo.20133709) to communications and networks. Shannon's information theory is shown to be a special case of the ITU axiom delta_S = delta(K), with H(X) = (K)/ln2 and channel capacity = max modular K-flow. This is Tier 1 paper #14, opening the K-channel axis and bringing the ITU polytope to 14 vertices. Communications achieves degree 9, tying with Climate (#11) as the polytope's maximum-connectivity super-hub. Pass-1 progress: 98 of 220 phases (44.5%). Phase 95: ITU foundation. Shannon H(X) = (K)/ln2 in ITU language; channel capacity C = max modular K-flow. Internet traffic 2024 = 500 EB/month, doubling period 2.9 years, forecast 20,000 EB/mo by 2050 (40x growth). Mobile evolution 1G to 6G: bandwidth 2 kbps to 1 Tbps (5e8x), latency 1000 ms to 0.1 ms (10000x improvement). Satellite constellations: 7,550 currently to planned 60,506 total (8x expansion). Quantum communication (Micius satellite 1,200 km QKD) demonstrates ITU axiom realization in physical observable. Phase 96: 6G + Quantum Internet + Federated Learning. IMT-2030 (6G) targets: 1 Tbps peak (50x 5G), 10 Gbps user (100x), 0.1 ms latency (10x), 10^7 devices/km^2, 100x energy efficiency. Quantum Internet Wehner 2018 6-stage roadmap: Stage 0-1 (trusted node, QKD) commercial 2024; Stage 2 (entanglement distribution) 2030; Stage 5 (distributed quantum computing) 2045. Federated Learning (McMahan 2017): centralized 0.994 accuracy, federated 0.979, local-only 0.698 — federated preserves privacy at minimal accuracy cost. Edge AI latency hierarchy: device 0.5 ms, 6G edge 0.5 ms (2030), regional DC 20 ms, cloud 80 ms, satellite 30 ms. Phase 97: Industry, economy, digital divide. Global ICT market: $5.3T (2024) to $30T (2050), Telecom subset $1.6T to $8T, AI subset $200B to $15T. Annual CapEx: $226B (2024, 5G dominant) to $655B (2050, Quantum $300B largest). Digital divide: world average 62% (2024) with 3.17B offline to 93% (2050) with 0.65B offline (80% reduction). Sub-Saharan Africa 40% to 90%, LDCs 27% to 80%. 6G patents: China 40%, USA 35% (combined 75% — standardization war risk). Satellite geopolitics: Starlink (USA) 42,000 planned vs Guowang (China) 13,000 planned. Phase 98: 2026-2050 roadmap with 16 milestones and 10 falsifiable predictions (P_avg = 0.57). Key milestones: 2028 IMT-2030 6G spec, 2030 6G commercial + Q-internet Stage 2, 2032 Starlink 42K complete, 2035 Q-memory network, 2045 distributed Q-compute, 2050 99.5% global penetration. Central thesis: communications is K-channel transport of K_information between subsystems. Shannon capacity = max modular K-flow makes communication engineering a direct application of ITU axiom. Quantum internet directly observes delta_S = delta(K) through entanglement-based protocols. 6G + Quantum + Federated Learning forms a 3-layer K-flow architecture (classical channel + quantum state + distributed K_self) supporting Embodied AGI (Tier 1 #13). The ITU 14-vertex polytope completes with K-channel axis. Communications vertex bidirectionally connects to 9 other vertices — tying with Climate as the polytope's super-hub. Honest framing: Pass-1 interpretive paper reframing Shannon (1948), Wehner-Elkouss-Hanson (2018), McMahan (2017), ITU-R IMT-2030 (2023), NIST PQC FIPS 203/204/205 (2024), 3GPP Release 19 (2024), Pan Jianwei Micius (2017-2024), SpaceX Starlink, World Bank broadband economics, Gartner/IDC ICT forecasts in ITU language. Numerical results match established literature. Includes 4 theory documents, 4 Python numerical experiments, 4 figures (PNG), 4 JSON summaries. Total runtime ~15 seconds.","author":[{"family":"Terada","given":"Munehiro"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20225969","URL":"https://doi.org/10.5281/zenodo.20225969","source":"datacite"},{"id":"doi:10.5281/zenodo.20225970","type":"article-journal","title":"ITU and Communications / Networks: A Single-Axiom View of Shannon Theory, Internet, 5G/6G, Quantum Communication","abstract":"We apply the Information-Theoretic Unification (ITU) framework (Terada 2026, concept DOI 10.5281/zenodo.20109209; current version v2.0.0 at 10.5281/zenodo.20133709) to communications and networks. Shannon's information theory is shown to be a special case of the ITU axiom delta_S = delta(K), with H(X) = (K)/ln2 and channel capacity = max modular K-flow. This is Tier 1 paper #14, opening the K-channel axis and bringing the ITU polytope to 14 vertices. Communications achieves degree 9, tying with Climate (#11) as the polytope's maximum-connectivity super-hub. Pass-1 progress: 98 of 220 phases (44.5%). Phase 95: ITU foundation. Shannon H(X) = (K)/ln2 in ITU language; channel capacity C = max modular K-flow. Internet traffic 2024 = 500 EB/month, doubling period 2.9 years, forecast 20,000 EB/mo by 2050 (40x growth). Mobile evolution 1G to 6G: bandwidth 2 kbps to 1 Tbps (5e8x), latency 1000 ms to 0.1 ms (10000x improvement). Satellite constellations: 7,550 currently to planned 60,506 total (8x expansion). Quantum communication (Micius satellite 1,200 km QKD) demonstrates ITU axiom realization in physical observable. Phase 96: 6G + Quantum Internet + Federated Learning. IMT-2030 (6G) targets: 1 Tbps peak (50x 5G), 10 Gbps user (100x), 0.1 ms latency (10x), 10^7 devices/km^2, 100x energy efficiency. Quantum Internet Wehner 2018 6-stage roadmap: Stage 0-1 (trusted node, QKD) commercial 2024; Stage 2 (entanglement distribution) 2030; Stage 5 (distributed quantum computing) 2045. Federated Learning (McMahan 2017): centralized 0.994 accuracy, federated 0.979, local-only 0.698 — federated preserves privacy at minimal accuracy cost. Edge AI latency hierarchy: device 0.5 ms, 6G edge 0.5 ms (2030), regional DC 20 ms, cloud 80 ms, satellite 30 ms. Phase 97: Industry, economy, digital divide. Global ICT market: $5.3T (2024) to $30T (2050), Telecom subset $1.6T to $8T, AI subset $200B to $15T. Annual CapEx: $226B (2024, 5G dominant) to $655B (2050, Quantum $300B largest). Digital divide: world average 62% (2024) with 3.17B offline to 93% (2050) with 0.65B offline (80% reduction). Sub-Saharan Africa 40% to 90%, LDCs 27% to 80%. 6G patents: China 40%, USA 35% (combined 75% — standardization war risk). Satellite geopolitics: Starlink (USA) 42,000 planned vs Guowang (China) 13,000 planned. Phase 98: 2026-2050 roadmap with 16 milestones and 10 falsifiable predictions (P_avg = 0.57). Key milestones: 2028 IMT-2030 6G spec, 2030 6G commercial + Q-internet Stage 2, 2032 Starlink 42K complete, 2035 Q-memory network, 2045 distributed Q-compute, 2050 99.5% global penetration. Central thesis: communications is K-channel transport of K_information between subsystems. Shannon capacity = max modular K-flow makes communication engineering a direct application of ITU axiom. Quantum internet directly observes delta_S = delta(K) through entanglement-based protocols. 6G + Quantum + Federated Learning forms a 3-layer K-flow architecture (classical channel + quantum state + distributed K_self) supporting Embodied AGI (Tier 1 #13). The ITU 14-vertex polytope completes with K-channel axis. Communications vertex bidirectionally connects to 9 other vertices — tying with Climate as the polytope's super-hub. Honest framing: Pass-1 interpretive paper reframing Shannon (1948), Wehner-Elkouss-Hanson (2018), McMahan (2017), ITU-R IMT-2030 (2023), NIST PQC FIPS 203/204/205 (2024), 3GPP Release 19 (2024), Pan Jianwei Micius (2017-2024), SpaceX Starlink, World Bank broadband economics, Gartner/IDC ICT forecasts in ITU language. Numerical results match established literature. Includes 4 theory documents, 4 Python numerical experiments, 4 figures (PNG), 4 JSON summaries. Total runtime ~15 seconds.","author":[{"family":"Terada","given":"Munehiro"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20225970","URL":"https://doi.org/10.5281/zenodo.20225970","source":"datacite"},{"id":"doi:10.5281/zenodo.17464803","type":"article-journal","title":"Beyond the Shadows — Contextual Awakening, Federated Learning, and the Realization of Reality through Digital Twins","abstract":"Beyond the Shadows — Contextual Awakening, Federated Learning, and the Realization of Reality through Digital Twins Abstract This paper captures the second stage of an ongoing multi-stakeholder digital twin and AI orchestration project, approximately four months into deployment. Through daily standups, on-site workshops, and cross-domain collaboration, an unexpected realization emerged: for many participants, this was the first time they understood they were truly working for a real estate company. Until that moment, their work had been mediated through systems, processes, and abstractions — shadows of the real. The study explores the tension between data-driven illusions and context-driven truths, the need for Local and Small Quantitative Models (LQMs) alongside Large Language Models (LLMs), and the sociotechnical awakening that occurs when context is reintroduced into fragmented digital ecosystems. It also examines the role of digital twins as boundary-spanning objects that preserve, teach, and reinterpret meaning over time. The findings reveal the profound implications of spatial awareness, organizational history, and emerging resilience frameworks in an era of geopolitical instability and post-cloud architectures. Index Terms— Digital Twins, Contextual Interoperability, Explainable Context, Federated Learning, Quantum-Safe Communication, LQMs, Edge-Native Design, SMILE, Organizational Awakening, Weill & Broadbent, Daniel & Ward, Magoulas & Pessi I. INTRODUCTION By the fourth month, the project had evolved beyond technical experimentation.During a workshop at the customer’s premises, a subtle but profound realization occurred: “This is the first time I’ve seen what I’m actually working for.” Multiple individuals echoed the same sentiment. Despite years of employment in the same organization, most had only interacted with representations — spreadsheets, maintenance systems, and financial dashboards — but not with reality itself. What they had been managing were shadows — interpretations of truth filtered through the lens of outdated systems and compartmentalized perspectives.The digital twin, for the first time, acted as the light source in Plato’s allegory: a shared reality that revealed the building not as an abstraction, but as a living ecosystem where people, systems, culture, and context intersected. This realization transformed the project’s orientation from “data-driven optimization” toward context-driven understanding, and reframed AI not as a controller, but as a participant in a continuous learning process. II. METHODOLOGICAL CONTINUATION The project continued to follow the SMILE methodology [4], with daily standups between distributed teams across Sweden. Despite strong routines, physical distance exposed the limitations of coordination: truth remained local, rooted in context. A recurring theme emerged: data is not truth.Data are ingredients — perishable, contextual, and incomplete.Projects, in turn, are the meals prepared from these ingredients.Different stakeholders “taste” the results differently; some prefer efficiency, others comfort or compliance. Thus, data-driven decision-making proved to be an oversimplification.What matters is the currency of context — continuously refreshed, locally valid, and dynamically interpreted. Old data are stale ingredients; contextually updated data are nourishment for living systems. III. THE EMERGENCE OF EXPLAINABLE CONTEXT Following the first deployment phase, we observed the emergence of Explainable Context (XC) as a necessary evolution of Explainable AI (XAI).AI models provided predictions, but the meaning of those predictions could only be understood within spatial, historical, and organizational contexts. To address this, digital twins of twins were introduced — meta-models that record not only the current state but the lineage of context: which actors influenced decisions, which data were used, and how interpretations evolved over time. This recursive architecture turn","author":[{"family":"Waern","given":"Nicolas"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17464803","URL":"https://doi.org/10.5281/zenodo.17464803","source":"datacite"},{"id":"doi:10.5281/zenodo.17464804","type":"article-journal","title":"Beyond the Shadows — Contextual Awakening, Federated Learning, and the Realization of Reality through Digital Twins","abstract":"Beyond the Shadows — Contextual Awakening, Federated Learning, and the Realization of Reality through Digital Twins Abstract This paper captures the second stage of an ongoing multi-stakeholder digital twin and AI orchestration project, approximately four months into deployment. Through daily standups, on-site workshops, and cross-domain collaboration, an unexpected realization emerged: for many participants, this was the first time they understood they were truly working for a real estate company. Until that moment, their work had been mediated through systems, processes, and abstractions — shadows of the real. The study explores the tension between data-driven illusions and context-driven truths, the need for Local and Small Quantitative Models (LQMs) alongside Large Language Models (LLMs), and the sociotechnical awakening that occurs when context is reintroduced into fragmented digital ecosystems. It also examines the role of digital twins as boundary-spanning objects that preserve, teach, and reinterpret meaning over time. The findings reveal the profound implications of spatial awareness, organizational history, and emerging resilience frameworks in an era of geopolitical instability and post-cloud architectures. Index Terms— Digital Twins, Contextual Interoperability, Explainable Context, Federated Learning, Quantum-Safe Communication, LQMs, Edge-Native Design, SMILE, Organizational Awakening, Weill & Broadbent, Daniel & Ward, Magoulas & Pessi I. INTRODUCTION By the fourth month, the project had evolved beyond technical experimentation.During a workshop at the customer’s premises, a subtle but profound realization occurred: “This is the first time I’ve seen what I’m actually working for.” Multiple individuals echoed the same sentiment. Despite years of employment in the same organization, most had only interacted with representations — spreadsheets, maintenance systems, and financial dashboards — but not with reality itself. What they had been managing were shadows — interpretations of truth filtered through the lens of outdated systems and compartmentalized perspectives.The digital twin, for the first time, acted as the light source in Plato’s allegory: a shared reality that revealed the building not as an abstraction, but as a living ecosystem where people, systems, culture, and context intersected. This realization transformed the project’s orientation from “data-driven optimization” toward context-driven understanding, and reframed AI not as a controller, but as a participant in a continuous learning process. II. METHODOLOGICAL CONTINUATION The project continued to follow the SMILE methodology [4], with daily standups between distributed teams across Sweden. Despite strong routines, physical distance exposed the limitations of coordination: truth remained local, rooted in context. A recurring theme emerged: data is not truth.Data are ingredients — perishable, contextual, and incomplete.Projects, in turn, are the meals prepared from these ingredients.Different stakeholders “taste” the results differently; some prefer efficiency, others comfort or compliance. Thus, data-driven decision-making proved to be an oversimplification.What matters is the currency of context — continuously refreshed, locally valid, and dynamically interpreted. Old data are stale ingredients; contextually updated data are nourishment for living systems. III. THE EMERGENCE OF EXPLAINABLE CONTEXT Following the first deployment phase, we observed the emergence of Explainable Context (XC) as a necessary evolution of Explainable AI (XAI).AI models provided predictions, but the meaning of those predictions could only be understood within spatial, historical, and organizational contexts. To address this, digital twins of twins were introduced — meta-models that record not only the current state but the lineage of context: which actors influenced decisions, which data were used, and how interpretations evolved over time. This recursive architecture turn","author":[{"family":"Waern","given":"Nicolas"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17464804","URL":"https://doi.org/10.5281/zenodo.17464804","source":"datacite"},{"id":"doi:10.5281/zenodo.21904089","type":"article-journal","title":"Opening access to learning records datasets for better educational sciences","abstract":"Research in learning analytics (LA), educational data mining (EDM) and artificial intelligence in education (AIEd) is fundamentally data-dependent, yet the behavioural learning trace data that powers it remains structurally inaccessible to the open research community. A recent systematic survey of 1,125 papers from the field’s flagship venues (LAK, EDM, AIED, 2020–2024) found that only 19% of publications are associated with an open dataset. Here we synthesize the state of available learning records datasets, characterize the systemic barriers to access—regulatory, commercial, institutional and technical—document the resulting reproducibility crisis, and evaluate emerging mitigations: synthetic data, federated learning, controlled-access research infrastructures and European sovereign data spaces. We argue that facilitating open, governed access to learning records datasets is a precondition for trustworthy, generalizable and regulation-compliant educational AI, and we propose a concrete agenda for operators, researchers and policymakers.","author":[{"family":"Sonnati","given":"Matthieu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21904089","URL":"https://doi.org/10.5281/zenodo.21904089","source":"datacite"},{"id":"doi:10.5281/zenodo.21904088","type":"article-journal","title":"Opening access to learning records datasets for better educational sciences","abstract":"Research in learning analytics (LA), educational data mining (EDM) and artificial intelligence in education (AIEd) is fundamentally data-dependent, yet the behavioural learning trace data that powers it remains structurally inaccessible to the open research community. A recent systematic survey of 1,125 papers from the field’s flagship venues (LAK, EDM, AIED, 2020–2024) found that only 19% of publications are associated with an open dataset. Here we synthesize the state of available learning records datasets, characterize the systemic barriers to access—regulatory, commercial, institutional and technical—document the resulting reproducibility crisis, and evaluate emerging mitigations: synthetic data, federated learning, controlled-access research infrastructures and European sovereign data spaces. We argue that facilitating open, governed access to learning records datasets is a precondition for trustworthy, generalizable and regulation-compliant educational AI, and we propose a concrete agenda for operators, researchers and policymakers.","author":[{"family":"Sonnati","given":"Matthieu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21904088","URL":"https://doi.org/10.5281/zenodo.21904088","source":"datacite"},{"id":"doi:10.5281/zenodo.21854968","type":"article-journal","title":"The Role of Artificial Intelligence in Detecting Social Engineering and Phishing Attacks: Opportunities, Limitations, and Ethical Considerations","abstract":"Phishing and social engineering remain leading causes of cybersecurity breaches, with the FBI reporting 859,532 complaints and $16.0 billion in losses in 2024. This literature review examines how artificial intelligence (AI) – including machine learning (ML), natural language processing (NLP), deep learning (DL), and large language models (LLMs) – can enhance detection of phishing/social engineering. We compare AI-based techniques to traditional rule-based and heuristic methods, analyzing detection performance, adaptability, false positives/negatives, computational demands, and resilience to adversarial manipulation. Key findings show that AI classifiers (e.g. CNNs, RNNs, BERT) often achieve very high accuracy (e.g. 98–99% on standard email corpora) and can adapt to evolving attacks, but they are vulnerable to adversarial rephrasing and suffer from explainability and data bias issues. Ethical concerns include user privacy (inspecting personal communications), transparency (black-box models), and model bias (especially in cross-cultural contexts). We illustrate these concepts with ScamLens AI app combining NLP analysis of text and URLs with vision-based screenshot scanning, producing risk indicators and explanations. Finally, we identify research gaps – such as limited real-world evaluation, outdated datasets, and the need for phishing-specific LLMs – and outline future directions (adversarial robustness, federated training, multimodal analysis, and human-AI collaboration).","author":[{"family":"Abozaid","given":"Zain"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21854968","URL":"https://doi.org/10.5281/zenodo.21854968","source":"datacite"},{"id":"doi:10.5281/zenodo.21854967","type":"article-journal","title":"The Role of Artificial Intelligence in Detecting Social Engineering and Phishing Attacks: Opportunities, Limitations, and Ethical Considerations","abstract":"Phishing and social engineering remain leading causes of cybersecurity breaches, with the FBI reporting 859,532 complaints and $16.0 billion in losses in 2024. This literature review examines how artificial intelligence (AI) – including machine learning (ML), natural language processing (NLP), deep learning (DL), and large language models (LLMs) – can enhance detection of phishing/social engineering. We compare AI-based techniques to traditional rule-based and heuristic methods, analyzing detection performance, adaptability, false positives/negatives, computational demands, and resilience to adversarial manipulation. Key findings show that AI classifiers (e.g. CNNs, RNNs, BERT) often achieve very high accuracy (e.g. 98–99% on standard email corpora) and can adapt to evolving attacks, but they are vulnerable to adversarial rephrasing and suffer from explainability and data bias issues. Ethical concerns include user privacy (inspecting personal communications), transparency (black-box models), and model bias (especially in cross-cultural contexts). We illustrate these concepts with ScamLens AI app combining NLP analysis of text and URLs with vision-based screenshot scanning, producing risk indicators and explanations. Finally, we identify research gaps – such as limited real-world evaluation, outdated datasets, and the need for phishing-specific LLMs – and outline future directions (adversarial robustness, federated training, multimodal analysis, and human-AI collaboration).","author":[{"family":"Abozaid","given":"Zain"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21854967","URL":"https://doi.org/10.5281/zenodo.21854967","source":"datacite"},{"id":"doi:10.5281/zenodo.20649152","type":"article-journal","title":"Assured Autonomy for Siloed Operations: Causal Learning with Per-Edge Certainty on HPC","abstract":"Assured autonomy has to know what it doesn't know — and direct its learning there. We present a layer that does this by construction: a causal graph in which every dependency carries a calibrated certainty, updated on-device from first principles, with a reasoning model invoked only to compose the model and to re-hypothesize where certainty stays low. Siloed, multi-owner operations are where this matters most, because there the dependencies you most need are often the ones no single party can observe. Scope. Our object is the layer: a causal graph with per-edge certainty, a deterministic on-device learning loop, and a reasoning-model escalation path. We demonstrate, on a real multi-owner operational dataset, that the layer's certainty signal correctly localizes where the system cannot reliably learn — the drifting, non-stationary, and cross-owner-unobservable dependencies — and that the compose→learn→escalate loop runs autonomously (see Demonstration). The problem: blind spots are bottlenecks for autonomy An autonomous operation has to act on relationships between subsystems — load drives heat, cooling removes it, one loop's effort changes another's. A model that emits a confident point estimate for every such relationship is dangerous in production, because the relationships you most need are frequently the least learnable: some drift as equipment and firmware evolve, some are non-stationary under changing regimes, and some are structurally unobservable from where any single party sits. The failure mode is silent — the model looks healthy and is quietly wrong on exactly the dependency that matters. Assured autonomy inverts this: the system maintains, per dependency, an explicit measure of how much it can be trusted, and it routes its own learning and its escalation to the low-certainty edges. Knowing what it doesn't know is not a diagnostic afterthought; it is the control signal. The layer: a causal graph with per-edge certainty We represent the operation as a causal graph. Each edge is a dependency (node_power → gpu_core_temp, cooling_supply → rack_inlet, liquid ΔT ↔ air ΔT) carrying a slope (the learned relationship), its residual, and a corroboration-based certainty Z in [0,1]. Z is the operational expression of \"what I know I don't know\": it rises only when an edge's error signal is both unbiased and consistent over a recent window, and it falls or collapses when the edge stops corroborating. Two derived signals drive behavior: Per-edge certainty localizes trust. The autonomy can act on high-Z edges, hedge on medium, and refuse or defer on low — something a monolithic model cannot do, because it has no place to attach \"I'm blind here.\" Persistent low-Z or biased residual is a directed-learning trigger: it marks an edge the deterministic loop cannot resolve on its own, and routes it to re-hypothesis. Because trust is attached per edge, it is also traceable: every action or abstention points to a specific dependency, its certainty, and its history — the auditability operations and safety cases require. Architecture: compose offline, learn on-device, escalate on ignorance The layer runs as three tiers with very different costs and cadences — which is what lets it operate under low-compute, intermittent, siloed conditions. Compose (reasoning model; on-prem or cloud; infrequent). A reasoning model reads domain priors and composes the causal graph and the per-edge validation pipelines — which dependencies exist, and what error signal corroborates each. Heavy, run rarely (at setup and on major change). Learn (on-device; deterministic; continuous). Each edge's certainty and weight update from first principles — a fixed arithmetic rule over the streamed error signal, no model inference in the loop. It is cheap, runs at the edge, tolerates disconnection (it syncs ~kilobyte certainty signals when a link is available, not raw data or gradients), and is fully traceable. Escalate (reasoning model; triggered by ignorance). When an edge ","author":[{"family":"Bennett","given":"Heidi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20649152","URL":"https://doi.org/10.5281/zenodo.20649152","source":"datacite"},{"id":"doi:10.5281/zenodo.20649151","type":"article-journal","title":"Assured Autonomy for Siloed Operations: Causal Learning with Per-Edge Certainty on HPC","abstract":"Assured autonomy has to know what it doesn't know — and direct its learning there. We present a layer that does this by construction: a causal graph in which every dependency carries a calibrated certainty, updated on-device from first principles, with a reasoning model invoked only to compose the model and to re-hypothesize where certainty stays low. Siloed, multi-owner operations are where this matters most, because there the dependencies you most need are often the ones no single party can observe. Scope. Our object is the layer: a causal graph with per-edge certainty, a deterministic on-device learning loop, and a reasoning-model escalation path. We demonstrate, on a real multi-owner operational dataset, that the layer's certainty signal correctly localizes where the system cannot reliably learn — the drifting, non-stationary, and cross-owner-unobservable dependencies — and that the compose→learn→escalate loop runs autonomously (see Demonstration). The problem: blind spots are bottlenecks for autonomy An autonomous operation has to act on relationships between subsystems — load drives heat, cooling removes it, one loop's effort changes another's. A model that emits a confident point estimate for every such relationship is dangerous in production, because the relationships you most need are frequently the least learnable: some drift as equipment and firmware evolve, some are non-stationary under changing regimes, and some are structurally unobservable from where any single party sits. The failure mode is silent — the model looks healthy and is quietly wrong on exactly the dependency that matters. Assured autonomy inverts this: the system maintains, per dependency, an explicit measure of how much it can be trusted, and it routes its own learning and its escalation to the low-certainty edges. Knowing what it doesn't know is not a diagnostic afterthought; it is the control signal. The layer: a causal graph with per-edge certainty We represent the operation as a causal graph. Each edge is a dependency (node_power → gpu_core_temp, cooling_supply → rack_inlet, liquid ΔT ↔ air ΔT) carrying a slope (the learned relationship), its residual, and a corroboration-based certainty Z in [0,1]. Z is the operational expression of \"what I know I don't know\": it rises only when an edge's error signal is both unbiased and consistent over a recent window, and it falls or collapses when the edge stops corroborating. Two derived signals drive behavior: Per-edge certainty localizes trust. The autonomy can act on high-Z edges, hedge on medium, and refuse or defer on low — something a monolithic model cannot do, because it has no place to attach \"I'm blind here.\" Persistent low-Z or biased residual is a directed-learning trigger: it marks an edge the deterministic loop cannot resolve on its own, and routes it to re-hypothesis. Because trust is attached per edge, it is also traceable: every action or abstention points to a specific dependency, its certainty, and its history — the auditability operations and safety cases require. Architecture: compose offline, learn on-device, escalate on ignorance The layer runs as three tiers with very different costs and cadences — which is what lets it operate under low-compute, intermittent, siloed conditions. Compose (reasoning model; on-prem or cloud; infrequent). A reasoning model reads domain priors and composes the causal graph and the per-edge validation pipelines — which dependencies exist, and what error signal corroborates each. Heavy, run rarely (at setup and on major change). Learn (on-device; deterministic; continuous). Each edge's certainty and weight update from first principles — a fixed arithmetic rule over the streamed error signal, no model inference in the loop. It is cheap, runs at the edge, tolerates disconnection (it syncs ~kilobyte certainty signals when a link is available, not raw data or gradients), and is fully traceable. Escalate (reasoning model; triggered by ignorance). When an edge ","author":[{"family":"Bennett","given":"Heidi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20649151","URL":"https://doi.org/10.5281/zenodo.20649151","source":"datacite"},{"id":"doi:10.5281/zenodo.20594816","type":"article-journal","title":"A Different Theory of Machine Intelligence: The Cagliostro Bound on Cognitive Diversity in Federated AI Systems","abstract":"The dominant theory of machine intelligence holds that capability is a function of scale: more parameters, more compute, and more data produce more capable systems. This paper proposes a different theory — one grounded in the physics of information and the mathematics of ensemble learning. We introduce the Cagliostro Bound: a formal argument connecting the Bekenstein-Bousso physical information limit to ensemble theory (Krogh & Vedelsby, 1995) to establish that the collective intelligence of a federation of cognitively diverse AI nodes exceeds what any single node could achieve, in proportion to the cognitive diversity between them. The argument proceeds in four steps: (1) any single AI system is subject to a physical information ceiling derivable from the Bekenstein-Bousso bound; (2) a federation of N cognitively sovereign nodes extends the effective information perimeter proportionally to N; (3) ensemble error decreases proportionally to cognitive diversity between members — formalized as Γ(N,ρ) = 1 + (N−1)(1−ρ), the Cagliostro Diversity Amplifier; (4) current AI architectures systematically destroy the diversity that would enable this gain, while a sovereign federated architecture preserves it. The paper also demonstrates that federated cognitive sovereignty is structurally immune to model collapse (Shumailov et al., 2024): each node learns exclusively from its own real interaction history, making homogenization impossible by architecture rather than by policy. Empirical basis: OBLIO-MSAN v0.10.4 · 7 active federated nodes · Γ = 5.01 measured in production · cognitive diversity coefficient ρ measured via HCS distribution divergence. Related papers: Holographic Cognitive Signatures (doi.org/10.5281/zenodo.20446006) · Forgetting is All You Need (doi.org/10.5281/zenodo.20489072) · Recursive Self-Observation (doi.org/10.5281/zenodo.20585580)","author":[{"family":"Cagliostro","given":"Claudio"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20594816","URL":"https://doi.org/10.5281/zenodo.20594816","source":"datacite"},{"id":"doi:10.5281/zenodo.20594815","type":"article-journal","title":"A Different Theory of Machine Intelligence: The Cagliostro Bound on Cognitive Diversity in Federated AI Systems","abstract":"The dominant theory of machine intelligence holds that capability is a function of scale: more parameters, more compute, and more data produce more capable systems. This paper proposes a different theory — one grounded in the physics of information and the mathematics of ensemble learning. We introduce the Cagliostro Bound: a formal argument connecting the Bekenstein-Bousso physical information limit to ensemble theory (Krogh & Vedelsby, 1995) to establish that the collective intelligence of a federation of cognitively diverse AI nodes exceeds what any single node could achieve, in proportion to the cognitive diversity between them. The argument proceeds in four steps: (1) any single AI system is subject to a physical information ceiling derivable from the Bekenstein-Bousso bound; (2) a federation of N cognitively sovereign nodes extends the effective information perimeter proportionally to N; (3) ensemble error decreases proportionally to cognitive diversity between members — formalized as Γ(N,ρ) = 1 + (N−1)(1−ρ), the Cagliostro Diversity Amplifier; (4) current AI architectures systematically destroy the diversity that would enable this gain, while a sovereign federated architecture preserves it. The paper also demonstrates that federated cognitive sovereignty is structurally immune to model collapse (Shumailov et al., 2024): each node learns exclusively from its own real interaction history, making homogenization impossible by architecture rather than by policy. Empirical basis: OBLIO-MSAN v0.10.4 · 7 active federated nodes · Γ = 5.01 measured in production · cognitive diversity coefficient ρ measured via HCS distribution divergence. Related papers: Holographic Cognitive Signatures (doi.org/10.5281/zenodo.20446006) · Forgetting is All You Need (doi.org/10.5281/zenodo.20489072) · Recursive Self-Observation (doi.org/10.5281/zenodo.20585580)","author":[{"family":"Cagliostro","given":"Claudio"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20594815","URL":"https://doi.org/10.5281/zenodo.20594815","source":"datacite"},{"id":"doi:10.5281/zenodo.20339466","type":"article-journal","title":"Attested Federated Clinical Inference: Privacy-Preserving Verifiable AI for Multi-Institutional Medical Diagnosis","abstract":"We present Attested Federated Clinical Inference (AFCI), a cryptographic protocol enabling multiple medical institutions to collaboratively execute AI model inference over private patient data while simultaneously guaranteeing: (1) data privacy — patient records never leave institutional boundaries; (2) inference verifiability — any third party can cryptographically confirm that inference was performed correctly on the agreed model; and (3) regulatory auditability — a tamper-evident audit trail satisfying the requirements of the EU AI Act (Regulation 2024/1689) and FDA AI/ML guidance for software as a medical device. Unlike federated learning, which provides no verifiability, and zero-knowledge ML (ZK-ML) approaches, which incur 10,000–1,000,000× computational overhead, AFCI leverages hardware Trusted Execution Environments (TEEs) to achieve O(1) verification overhead over standard inference. We formalize the security model for multi-institutional verifiable federated inference, prove AFCI's security under standard TEE and computational hardness assumptions, and describe a reference implementation architecture using Apple Secure Enclave and the App Attest API. AFCI directly addresses the emerging regulatory imperative for auditable clinical AI under the EU AI Act and FDA AI/ML guidance.","author":[{"family":"Cajka","given":"Nikolaj"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20339466","URL":"https://doi.org/10.5281/zenodo.20339466","source":"datacite"},{"id":"doi:10.5281/zenodo.20339465","type":"article-journal","title":"Attested Federated Clinical Inference: Privacy-Preserving Verifiable AI for Multi-Institutional Medical Diagnosis","abstract":"We present Attested Federated Clinical Inference (AFCI), a cryptographic protocol enabling multiple medical institutions to collaboratively execute AI model inference over private patient data while simultaneously guaranteeing: (1) data privacy — patient records never leave institutional boundaries; (2) inference verifiability — any third party can cryptographically confirm that inference was performed correctly on the agreed model; and (3) regulatory auditability — a tamper-evident audit trail satisfying the requirements of the EU AI Act (Regulation 2024/1689) and FDA AI/ML guidance for software as a medical device. Unlike federated learning, which provides no verifiability, and zero-knowledge ML (ZK-ML) approaches, which incur 10,000–1,000,000× computational overhead, AFCI leverages hardware Trusted Execution Environments (TEEs) to achieve O(1) verification overhead over standard inference. We formalize the security model for multi-institutional verifiable federated inference, prove AFCI's security under standard TEE and computational hardness assumptions, and describe a reference implementation architecture using Apple Secure Enclave and the App Attest API. AFCI directly addresses the emerging regulatory imperative for auditable clinical AI under the EU AI Act and FDA AI/ML guidance.","author":[{"family":"Cajka","given":"Nikolaj"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20339465","URL":"https://doi.org/10.5281/zenodo.20339465","source":"datacite"},{"id":"doi:10.6084/m9.figshare.32050653","type":"article-journal","title":"ESTRO Course 2026 - Quantitative methods in Radiation Oncology: \"Data Sharing: Why and How\".","abstract":"Delivered 2026-04-19, Lisbon, Portugal, Faculdade de Ciências da Universidade de Lisboa, Campo Grande.Precis:The presentation introduces radiation oncologists, medical physicists, radiation therapists, and research-data stewards to the legal, technical, and operational machinery required to make radiotherapy research data shareable — and reusable — under current European data-protection law.The deck opens by framing FAIR, transparent, and open science as preconditions for scientific rigour, reproducibility, and equity, then situates radiation oncology within the OECD definition of open science and its three pillars: open research data, open educational resources, and open code. The FAIR Guiding Principles (Wilkinson et al. 2016) are presented as machine-actionable requirements, and the persistent digital identifier (DOI) is explained as their foundation. A repository survey covers code archives (GitHub, CRAN), general data repositories (Zenodo, figshare, Mendeley Data, Dryad, OSF, NIH Cancer Data Commons), domain-specific imaging archives (TCIA, eContour), and data-descriptor venues ( Scientific Data , Medical Physics , IJROBP ).A substantial section on Common Data Elements (CDEs) treats them as the semantic glue that makes FAIR operational, with a worked parotid-Dmean example and a complete map of the radiotherapy informatics standards stack (DICOM-RT, AAPM TG-263, SNOMED-CT, ICD-10/O-3, CTCAE v5.0, OMOP-CDM, HL7 FHIR/mCODE). The four principal CDE repositories — NIH CDE, caDSR, CDISC, and the disease-specific NINDS/EORTC/EuroCAT consortia — are catalogued, and a three-step SOP-to-CDE workflow is illustrated via Dietrich et al. ( BMC Med Inform Decis Mak 2025).The European regulatory landscape is presented through five concurrent instruments: GDPR, the Clinical Trials Regulation, the European Health Data Space (Reg. (EU) 2025/327), the AI Act, and MDR/IVDR. GDPR Article 89 research safeguards are mapped to the derogations they enable, and the identified → pseudonymised → anonymised waterfall is formalised. Article 29 WP216 is covered in depth, with one slide per deidentification technique — pseudonymisation, noise addition, permutation, differential privacy, k-anonymity, l-diversity, and t-closeness — each illustrated with a radiotherapy-specific example and the WP216 Table 6 summary of residual risks. Federated learning, privacy-enhancing technologies (DP, SMPC, HE, TEE), and the unsettled status of model weights as personal data (EDPB Opinion 28/2024) receive dedicated treatment, with European exemplars from EuroCAT, FedSynthCT, the Personal Health Train, and EORTC. A UK-GDPR addendum addresses the post-Brexit adequacy decision (renewed December 2025, valid to December 2031), the MRC/UKRI approval stack, common-law confidentiality, and the CAG/PBPP/PAC approval routes. A closing block on data- and material-transfer agreements introduces EU SCCs (Commission Implementing Decision 2021/914), the UK IDTA and UK Addendum, biobank MTAs, and a practical decision workflow for multi-institutional RT exchanges.The lecture closes with a reflection on dataset bias — illustrated through the HNC-PREDICTOR external-validation experience — reminding the audience that big data is not necessarily good data, and that bias awareness is not bias mitigation. The slide deck is released under a CC-BY licence and may be reused, adapted, and redistributed with attribution.","author":[{"family":"Fuller","given":"Clifton"}],"issued":{"date-parts":[[2026]]},"DOI":"10.6084/m9.figshare.32050653","URL":"https://doi.org/10.6084/m9.figshare.32050653","source":"datacite"},{"id":"doi:10.5281/zenodo.17462961","type":"article-journal","title":"From One Room to Fifty: Orchestrating Explainable AI, Resilience, and Contextual Interoperability in the Built Environment","abstract":"From One Room to Fifty: Orchestrating Explainable AI, Resilience, and Contextual Interoperability in the Built Environment Nicolas Waern, WINNIIO AB Abstract This paper presents the initial findings and reflections from an ongoing research and implementation project aimed at scaling a self-learning, self-regulating heating system from one room to fifty within a real-world school environment. The project explores how artificial intelligence (AI), digital twins, and federated learning architectures interact with sociotechnical realities under the frameworks of the EU Taxonomy and Minimal Interoperability Mechanisms (MIMs). The study is written from the position of not yet knowing what will succeed — a candid account of exploration through uncertainty. Early insights indicate that true resilience and explainable AI depend less on algorithmic sophistication and more on contextual interoperability: the ability of humans, machines, and institutions to share understanding across spatial, temporal, and organizational boundaries. Index Terms— Digital Twins, Explainable AI, Contextual Interoperability, Federated Learning, Resilience, MIMs, EU Taxonomy, Actor–Network Theory, SMILE, Boundary Spanning I. INTRODUCTION Scaling an AI system from one controlled environment to fifty interconnected ones is not merely a matter of engineering; it is a study in sociology, physics, and epistemology. The intent of this work was not to demonstrate deterministic success but to uncover how intelligence behaves when it must coexist — with other systems, people, regulations, and infrastructures. The project, funded by the Swedish Energy Agency, builds upon WINNIIO AB’s previous work in self-learning heating systems verified through digital twin environments [21]. Its purpose was to examine the local-first paradigm, where intelligence resides within the building itself rather than in remote cloud infrastructure. This approach was guided by the SMILE methodology — Sustainable Methodology for Impact Lifecycle Enablement — which combines concurrent engineering, contextual intelligence, and continuous learning as the basis for systemic adaptation. II. METHODOLOGICAL CONTEXT Each room in the facility was treated as a semi-autonomous learning agent equipped with sensors, actuators, and a localized inference model. These agents communicated through a federated mesh, sharing anonymized model updates rather than raw data, thereby enhancing both privacy and resilience. The architecture was intentionally experimental and reflexive, drawing upon Actor–Network Theory (ANT) [2], which posits that every element — human, material, or algorithmic — acts as an agent shaping the network’s behavior. The digital twin functioned as a boundary-spanning object where all actors could align their perspectives and decisions. At this early stage, the success criteria were defined not by energy savings alone but by knowledge transfer: the creation of shared meaning across disciplines and systems. III. SCALING UNCERTAINTY Scaling from one room to fifty revealed that context does not scale linearly. Each room possessed its own thermodynamic and behavioral identity. Models that performed well in isolation required continual re-calibration when integrated into a collective whole. The project therefore became a live experiment in contextual dynamics. Instead of optimizing for static performance, we sought adaptive coherence — a distributed rhythm of learning between physical systems, digital models, and human operators. This iterative exchange created what might be described as a negotiated intelligence, emergent from continuous interaction rather than central control. At the time of writing, quantitative efficiency gains remain unverified due to external disturbances (renovations, hardware anomalies, contextual noise). Nonetheless, qualitative outcomes suggest that coherence and shared situational awareness improved markedly. The digital twin became the lens through which disparate actors","author":[{"family":"Waern","given":"Nicolas"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17462961","URL":"https://doi.org/10.5281/zenodo.17462961","source":"datacite"},{"id":"doi:10.5281/zenodo.17462962","type":"article-journal","title":"From One Room to Fifty: Orchestrating Explainable AI, Resilience, and Contextual Interoperability in the Built Environment","abstract":"From One Room to Fifty: Orchestrating Explainable AI, Resilience, and Contextual Interoperability in the Built Environment Nicolas Waern, WINNIIO AB Abstract This paper presents the initial findings and reflections from an ongoing research and implementation project aimed at scaling a self-learning, self-regulating heating system from one room to fifty within a real-world school environment. The project explores how artificial intelligence (AI), digital twins, and federated learning architectures interact with sociotechnical realities under the frameworks of the EU Taxonomy and Minimal Interoperability Mechanisms (MIMs). The study is written from the position of not yet knowing what will succeed — a candid account of exploration through uncertainty. Early insights indicate that true resilience and explainable AI depend less on algorithmic sophistication and more on contextual interoperability: the ability of humans, machines, and institutions to share understanding across spatial, temporal, and organizational boundaries. Index Terms— Digital Twins, Explainable AI, Contextual Interoperability, Federated Learning, Resilience, MIMs, EU Taxonomy, Actor–Network Theory, SMILE, Boundary Spanning I. INTRODUCTION Scaling an AI system from one controlled environment to fifty interconnected ones is not merely a matter of engineering; it is a study in sociology, physics, and epistemology. The intent of this work was not to demonstrate deterministic success but to uncover how intelligence behaves when it must coexist — with other systems, people, regulations, and infrastructures. The project, funded by the Swedish Energy Agency, builds upon WINNIIO AB’s previous work in self-learning heating systems verified through digital twin environments [21]. Its purpose was to examine the local-first paradigm, where intelligence resides within the building itself rather than in remote cloud infrastructure. This approach was guided by the SMILE methodology — Sustainable Methodology for Impact Lifecycle Enablement — which combines concurrent engineering, contextual intelligence, and continuous learning as the basis for systemic adaptation. II. METHODOLOGICAL CONTEXT Each room in the facility was treated as a semi-autonomous learning agent equipped with sensors, actuators, and a localized inference model. These agents communicated through a federated mesh, sharing anonymized model updates rather than raw data, thereby enhancing both privacy and resilience. The architecture was intentionally experimental and reflexive, drawing upon Actor–Network Theory (ANT) [2], which posits that every element — human, material, or algorithmic — acts as an agent shaping the network’s behavior. The digital twin functioned as a boundary-spanning object where all actors could align their perspectives and decisions. At this early stage, the success criteria were defined not by energy savings alone but by knowledge transfer: the creation of shared meaning across disciplines and systems. III. SCALING UNCERTAINTY Scaling from one room to fifty revealed that context does not scale linearly. Each room possessed its own thermodynamic and behavioral identity. Models that performed well in isolation required continual re-calibration when integrated into a collective whole. The project therefore became a live experiment in contextual dynamics. Instead of optimizing for static performance, we sought adaptive coherence — a distributed rhythm of learning between physical systems, digital models, and human operators. This iterative exchange created what might be described as a negotiated intelligence, emergent from continuous interaction rather than central control. At the time of writing, quantitative efficiency gains remain unverified due to external disturbances (renovations, hardware anomalies, contextual noise). Nonetheless, qualitative outcomes suggest that coherence and shared situational awareness improved markedly. The digital twin became the lens through which disparate actors","author":[{"family":"Waern","given":"Nicolas"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17462962","URL":"https://doi.org/10.5281/zenodo.17462962","source":"datacite"},{"id":"doi:10.17605/osf.io/dajpe","type":"article-journal","title":"A PRISMA-Based Systematic Review and Normalized Evaluation Framework for Artificial Intelligence-Driven Intrusion Detection Systems: Methods, Comparative Insights, and Deployment Challenges","abstract":"The accelerating expansion of cloud infrastructure, Internet of Things (IoT) ecosystems, and geographically distributed network architectures has substantially widened the attack surface accessible to malicious actors (Khraisat et al., 2019)ber-intrusion incidents now unfold at scales and levels of sophistication that were previously uncharacteristic, generating annual economic losses estimated in the trillions of dollars globally(Kala, 2023)Within this context, Intrusion Detection Systems (IDS)—mechanisms designed to monitor and analyse network traffic in order to identify potentially malicious activity—occupy an indispensable position in enterprise and national security strategies(Judijanto et al., 2023; Vandana &amp; Verma, 2025). Signature-based IDS approaches have proven fundamentally inadequate against zero-day exploits and polymorphic attack variants, primarily because such systems rely on static rule repositories that offer no adaptive capacity against previously unseen threats(Alqhatani et al., 2025). AI-driven IDS, drawing on machine learning (ML) and deep learning (DL) methods, have subsequently emerged as the principal technological response to this limitation(Alqhatani et al., 2025; Sinha et al., 2024). These systems are capable of autonomously identifying latent patterns in high-dimensional traffic data, adapting to evolving attack distributions, and operating without explicit human-authored detection logic(Mohan et al., 2025; Raja, 2025). **Highlights *PRISMA-based systematic review of AI-driven IDS research (2019–2025) across five scholarly databases yielding 84 primary studies. *Deployment-aware Normalized IDS Evaluation Framework (NIEF)—first of its kind for AI-driven IDS comparison, incorporating quantified bias penalty and deployment feasibility scoring. *Case study validation across 15 representative models with Friedman and Nemenyi statistical validation (p = 0.010). *Hybrid IDS architectures (CNN-LSTM, RF+DL, CNN-GRU, Ensemble) outperform ML-only and DL-only models in deployment-aware evaluation (NIEF scores: 0.56–0.73). *Structured research roadmap addressing XAI integration, federated IDS, adversarial robustness, and benchmark standardization.","author":[{"family":"Das","given":"Kapil"}],"issued":{"date-parts":[[2026]]},"DOI":"10.17605/osf.io/dajpe","URL":"https://doi.org/10.17605/osf.io/dajpe","source":"datacite"},{"id":"doi:10.5281/zenodo.19652036","type":"article-journal","title":"The Energy Paradox: Artificial Intelligence for Seismic Computational Load Reduction in Oil and Gas — A Technical Survey and Position Analysis.","abstract":"Abstract — The oil and gas (O&G) industry faces a compounding computational challenge: advanced seismic imaging techniques such as Full Waveform Inversion (FWI) and Reverse Time Migration (RTM) impose exponentially scaling workloads — doubling the maximum frequency of a 3D FWI run increases computational demand by a factor of sixteen, while a single high-resolution 3D FWI job on a modern supercomputer may require approximately 25,000 GPU-hours. This paper surveys the principal artificial intelligence (AI) and machine learning (ML) techniques being deployed to reduce these workloads, including deep neural network surrogate models for iterative solvers, physics-informed neural networks (PINNs), edge inference architectures, and alternative dataflow hardware paradigms. A structured review of institutional deployments — spanning the U.S. Department of Energy Genesis Mission (Executive Order, November 2025), the NETL Science-Informed Machine Learning for Accelerating Real-Time Decisions (SMART/SAMI) initiative, Petrobras SolverBR, Saudi Aramco AI-assisted seismic processing, and the Shearwater-NVIDIA collaboration — is presented alongside quantified performance outcomes. Beyond the technical review, this paper introduces and formalizes the Energy Paradox: the O&G sector simultaneously supplies the hydrocarbon energy that powers AI data centers globally and consumes AI-driven high-performance computing to reduce its own operational costs. We argue this closed-loop relationship has strategic, economic, and regulatory implications that have not been adequately addressed in the literature. A framework for evaluating AI adoption in seismic workflows — encompassing computational efficiency, energy footprint, and governance alignment — is proposed as a contribution to practitioners, engineers, and researchers operating at this intersection. Index Terms — Full Waveform Inversion, Reverse Time Migration, Seismic Processing, Deep Learning, Surrogate Models, Physics-Informed Neural Networks, Edge Computing, Oil and Gas, HPC, Energy Paradox, DOE Genesis Mission, NETL SAMI, Computational Geophysics, GPU Efficiency I. INTRODUCTION The oil and gas sector has long been one of the most computationally intensive industries outside of national security and pharmaceutical research. Seismic data acquisition surveys generate petabytes of raw measurements per campaign, and the physics-based algorithms used to convert those measurements into actionable subsurface imagery demand supercomputer-scale resources. Yet the economic pressure on this infrastructure has intensified significantly in the 2024–2026 period — not because of increased exploration activity, but because of a structural shift in global energy demand driven by artificial intelligence itself. The proliferation of large language models (LLMs), generative AI systems, and large-scale GPU clusters has produced an unprecedented surge in data center electricity consumption. AI hardware — primarily high-density GPU and TPU farms — is projected to account for a rapidly growing share of global electricity demand through 2030 [1]. A substantial fraction of this electricity is sourced from natural gas, either directly through gas-fired generation or indirectly through LNG supply chains managed by major O&G operators. The O&G sector therefore finds itself in a structurally paradoxical position: it supplies the energy that enables AI at global scale, while simultaneously seeking to deploy AI to reduce the computational cost of its own most expensive workflows. This paper formalizes this relationship as the Energy Paradox and situates it within a technical survey of the AI-driven computational reduction techniques being adopted at the frontier of seismic data processing. The survey covers four primary AI paradigms — DNN surrogate models, physics-informed neural networks (PINNs), edge inference architectures, and alternative compute hardware — with particular attention to quantified performance outcomes fro","author":[{"family":"Rudio","given":"Rubens"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19652036","URL":"https://doi.org/10.5281/zenodo.19652036","source":"datacite"},{"id":"doi:10.5281/zenodo.19652035","type":"article-journal","title":"The Energy Paradox: Artificial Intelligence for Seismic Computational Load Reduction in Oil and Gas — A Technical Survey and Position Analysis.","abstract":"Abstract — The oil and gas (O&G) industry faces a compounding computational challenge: advanced seismic imaging techniques such as Full Waveform Inversion (FWI) and Reverse Time Migration (RTM) impose exponentially scaling workloads — doubling the maximum frequency of a 3D FWI run increases computational demand by a factor of sixteen, while a single high-resolution 3D FWI job on a modern supercomputer may require approximately 25,000 GPU-hours. This paper surveys the principal artificial intelligence (AI) and machine learning (ML) techniques being deployed to reduce these workloads, including deep neural network surrogate models for iterative solvers, physics-informed neural networks (PINNs), edge inference architectures, and alternative dataflow hardware paradigms. A structured review of institutional deployments — spanning the U.S. Department of Energy Genesis Mission (Executive Order, November 2025), the NETL Science-Informed Machine Learning for Accelerating Real-Time Decisions (SMART/SAMI) initiative, Petrobras SolverBR, Saudi Aramco AI-assisted seismic processing, and the Shearwater-NVIDIA collaboration — is presented alongside quantified performance outcomes. Beyond the technical review, this paper introduces and formalizes the Energy Paradox: the O&G sector simultaneously supplies the hydrocarbon energy that powers AI data centers globally and consumes AI-driven high-performance computing to reduce its own operational costs. We argue this closed-loop relationship has strategic, economic, and regulatory implications that have not been adequately addressed in the literature. A framework for evaluating AI adoption in seismic workflows — encompassing computational efficiency, energy footprint, and governance alignment — is proposed as a contribution to practitioners, engineers, and researchers operating at this intersection. Index Terms — Full Waveform Inversion, Reverse Time Migration, Seismic Processing, Deep Learning, Surrogate Models, Physics-Informed Neural Networks, Edge Computing, Oil and Gas, HPC, Energy Paradox, DOE Genesis Mission, NETL SAMI, Computational Geophysics, GPU Efficiency I. INTRODUCTION The oil and gas sector has long been one of the most computationally intensive industries outside of national security and pharmaceutical research. Seismic data acquisition surveys generate petabytes of raw measurements per campaign, and the physics-based algorithms used to convert those measurements into actionable subsurface imagery demand supercomputer-scale resources. Yet the economic pressure on this infrastructure has intensified significantly in the 2024–2026 period — not because of increased exploration activity, but because of a structural shift in global energy demand driven by artificial intelligence itself. The proliferation of large language models (LLMs), generative AI systems, and large-scale GPU clusters has produced an unprecedented surge in data center electricity consumption. AI hardware — primarily high-density GPU and TPU farms — is projected to account for a rapidly growing share of global electricity demand through 2030 [1]. A substantial fraction of this electricity is sourced from natural gas, either directly through gas-fired generation or indirectly through LNG supply chains managed by major O&G operators. The O&G sector therefore finds itself in a structurally paradoxical position: it supplies the energy that enables AI at global scale, while simultaneously seeking to deploy AI to reduce the computational cost of its own most expensive workflows. This paper formalizes this relationship as the Energy Paradox and situates it within a technical survey of the AI-driven computational reduction techniques being adopted at the frontier of seismic data processing. The survey covers four primary AI paradigms — DNN surrogate models, physics-informed neural networks (PINNs), edge inference architectures, and alternative compute hardware — with particular attention to quantified performance outcomes fro","author":[{"family":"Rudio","given":"Rubens"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19652035","URL":"https://doi.org/10.5281/zenodo.19652035","source":"datacite"},{"id":"doi:10.6084/m9.figshare.32050653.v1","type":"article-journal","title":"ESTRO Course 2026 - Quantitative methods in Radiation Oncology: \"Data Sharing: Why and How\".","abstract":"Delivered 2026-04-19, Lisbon, Portugal, Faculdade de Ciências da Universidade de Lisboa, Campo Grande.Precis:The presentation introduces radiation oncologists, medical physicists, radiation therapists, and research-data stewards to the legal, technical, and operational machinery required to make radiotherapy research data shareable — and reusable — under current European data-protection law.The deck opens by framing FAIR, transparent, and open science as preconditions for scientific rigour, reproducibility, and equity, then situates radiation oncology within the OECD definition of open science and its three pillars: open research data, open educational resources, and open code. The FAIR Guiding Principles (Wilkinson et al. 2016) are presented as machine-actionable requirements, and the persistent digital identifier (DOI) is explained as their foundation. A repository survey covers code archives (GitHub, CRAN), general data repositories (Zenodo, figshare, Mendeley Data, Dryad, OSF, NIH Cancer Data Commons), domain-specific imaging archives (TCIA, eContour), and data-descriptor venues ( Scientific Data , Medical Physics , IJROBP ).A substantial section on Common Data Elements (CDEs) treats them as the semantic glue that makes FAIR operational, with a worked parotid-Dmean example and a complete map of the radiotherapy informatics standards stack (DICOM-RT, AAPM TG-263, SNOMED-CT, ICD-10/O-3, CTCAE v5.0, OMOP-CDM, HL7 FHIR/mCODE). The four principal CDE repositories — NIH CDE, caDSR, CDISC, and the disease-specific NINDS/EORTC/EuroCAT consortia — are catalogued, and a three-step SOP-to-CDE workflow is illustrated via Dietrich et al. ( BMC Med Inform Decis Mak 2025).The European regulatory landscape is presented through five concurrent instruments: GDPR, the Clinical Trials Regulation, the European Health Data Space (Reg. (EU) 2025/327), the AI Act, and MDR/IVDR. GDPR Article 89 research safeguards are mapped to the derogations they enable, and the identified → pseudonymised → anonymised waterfall is formalised. Article 29 WP216 is covered in depth, with one slide per deidentification technique — pseudonymisation, noise addition, permutation, differential privacy, k-anonymity, l-diversity, and t-closeness — each illustrated with a radiotherapy-specific example and the WP216 Table 6 summary of residual risks. Federated learning, privacy-enhancing technologies (DP, SMPC, HE, TEE), and the unsettled status of model weights as personal data (EDPB Opinion 28/2024) receive dedicated treatment, with European exemplars from EuroCAT, FedSynthCT, the Personal Health Train, and EORTC. A UK-GDPR addendum addresses the post-Brexit adequacy decision (renewed December 2025, valid to December 2031), the MRC/UKRI approval stack, common-law confidentiality, and the CAG/PBPP/PAC approval routes. A closing block on data- and material-transfer agreements introduces EU SCCs (Commission Implementing Decision 2021/914), the UK IDTA and UK Addendum, biobank MTAs, and a practical decision workflow for multi-institutional RT exchanges.The lecture closes with a reflection on dataset bias — illustrated through the HNC-PREDICTOR external-validation experience — reminding the audience that big data is not necessarily good data, and that bias awareness is not bias mitigation. The slide deck is released under a CC-BY licence and may be reused, adapted, and redistributed with attribution.","author":[{"family":"Fuller","given":"Clifton"}],"issued":{"date-parts":[[2026]]},"DOI":"10.6084/m9.figshare.32050653.v1","URL":"https://doi.org/10.6084/m9.figshare.32050653.v1","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.12304","type":"manuscript","title":"Beyond Weather Correlation: A Comparative Study of Static and Temporal Neural Architectures for Fine-Grained Residential Energy Consumption Forecasting in Melbourne, Australia","abstract":"Accurate short-term residential energy consumption forecasting at sub-hourly resolution is critical for smart grid management, demand response programmes, and renewable energy integration. While weather variables are widely acknowledged as key drivers of residential electricity demand, the relative merit of incorporating temporal autocorrelation - the sequential memory of past consumption; over static meteorological features alone remains underexplored at fine-grained (5-minute) temporal resolution for Australian households. This paper presents a rigorous empirical comparison of a Multilayer Perceptron (MLP) and a Long Short-Term Memory (LSTM) recurrent network applied to two real-world Melbourne households: House 3 (a standard grid-connected dwelling) and House 4 (a rooftop solar photovoltaic-integrated household). Both models are trained on 14 months of 5-minute interval smart meter data (March 2023-April 2024) merged with official Bureau of Meteorology (BOM) daily weather observations, yielding over 117,000 samples per household. The LSTM, operating on 24-step (2-hour) sliding consumption windows, achieves coefficients of determination of R^2 = 0.883 (House 3) and R^2 = 0.865 (House 4), compared to R^2 = -0.055 and R^2 = 0.410 for the corresponding weather-driven MLPs - differences of 93.8 and 45.5 percentage points. These results establish that temporal autocorrelation in the consumption sequence dominates meteorological information for short-term forecasting at 5-minute granularity. Additionally, we demonstrate an asymmetry introduced by solar generation: for the PV-integrated household, the MLP achieves R^2 = 0.410, revealing implicit solar forecasting from weather-time correlations. A persistence baseline analysis and seasonal stratification contextualise model performance. We propose a hybrid weather-augmented LSTM and federated learning extensions as directions for future work.","author":[{"family":"Hewage","given":"Prasad"},{"family":"Wu","given":"Hao"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.12304","URL":"https://doi.org/10.48550/arxiv.2604.12304","source":"datacite"},{"id":"doi:10.5281/zenodo.20484323","type":"article-journal","title":"A Survey on Federated Learning for Privacy-Preserving Artificial Intelligence in Internet of Things Systems","abstract":"🔹 Overview The explosive growth of the Internet of Things (IoT) has led to billions of connected devices generating vast amounts of sensitive data. While traditional machine learning relies on centralized data collection, this approach raises significant privacy, security, and regulatory concerns. Federated Learning (FL) has emerged as a transformative solution by enabling collaborative model training across distributed devices without transferring raw data to a central server. This makes FL one of the most promising technologies for privacy-preserving artificial intelligence in IoT ecosystems. What This Survey Covers This survey provides a comprehensive review of Federated Learning for IoT systems, focusing on research developments from 2020–2025, with particular emphasis on advances published since 2022. Key topics include: Communication-efficient federated learning techniques Aggregation algorithms and optimization strategies Privacy-preserving mechanisms and differential privacy Security threats such as poisoning, backdoor, and inference attacks Defense mechanisms and robust aggregation methods Comparative analysis of major FL frameworks and approaches Challenges arising from non-IID data, device heterogeneity, and resource constraints Comparative Evaluation The surveyed methods are analyzed across critical performance dimensions: Communication overhead Model accuracy Convergence behavior Privacy guarantees Suitability for real-world IoT deployments Future Research Directions The survey also explores emerging trends, including: Differential Privacy-enhanced Federated Learning Secure Multi-Party Computation (SMPC) Personalized Federated Learning Blockchain-integrated FL systems Large Language Model (LLM)-assisted federation TinyML and edge intelligence integration Target Audience This work serves as a structured reference for: Researchers in Federated Learning and Distributed AI IoT Security and Privacy practitioners Graduate students and academics Industry professionals developing privacy-preserving intelligent systems Keywords: Federated Learning, Internet of Things (IoT), Privacy-Preserving AI, Edge Computing, Distributed Machine Learning, Differential Privacy, Secure Aggregation, IoT Security.","author":[{"family":"Sharma","given":"Arjit"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20484323","URL":"https://doi.org/10.5281/zenodo.20484323","source":"datacite"},{"id":"doi:10.5281/zenodo.20484324","type":"article-journal","title":"A Survey on Federated Learning for Privacy-Preserving Artificial Intelligence in Internet of Things Systems","abstract":"🔹 Overview The explosive growth of the Internet of Things (IoT) has led to billions of connected devices generating vast amounts of sensitive data. While traditional machine learning relies on centralized data collection, this approach raises significant privacy, security, and regulatory concerns. Federated Learning (FL) has emerged as a transformative solution by enabling collaborative model training across distributed devices without transferring raw data to a central server. This makes FL one of the most promising technologies for privacy-preserving artificial intelligence in IoT ecosystems. What This Survey Covers This survey provides a comprehensive review of Federated Learning for IoT systems, focusing on research developments from 2020–2025, with particular emphasis on advances published since 2022. Key topics include: Communication-efficient federated learning techniques Aggregation algorithms and optimization strategies Privacy-preserving mechanisms and differential privacy Security threats such as poisoning, backdoor, and inference attacks Defense mechanisms and robust aggregation methods Comparative analysis of major FL frameworks and approaches Challenges arising from non-IID data, device heterogeneity, and resource constraints Comparative Evaluation The surveyed methods are analyzed across critical performance dimensions: Communication overhead Model accuracy Convergence behavior Privacy guarantees Suitability for real-world IoT deployments Future Research Directions The survey also explores emerging trends, including: Differential Privacy-enhanced Federated Learning Secure Multi-Party Computation (SMPC) Personalized Federated Learning Blockchain-integrated FL systems Large Language Model (LLM)-assisted federation TinyML and edge intelligence integration Target Audience This work serves as a structured reference for: Researchers in Federated Learning and Distributed AI IoT Security and Privacy practitioners Graduate students and academics Industry professionals developing privacy-preserving intelligent systems Keywords: Federated Learning, Internet of Things (IoT), Privacy-Preserving AI, Edge Computing, Distributed Machine Learning, Differential Privacy, Secure Aggregation, IoT Security.","author":[{"family":"Sharma","given":"Arjit"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20484324","URL":"https://doi.org/10.5281/zenodo.20484324","source":"datacite"},{"id":"doi:10.5281/zenodo.16875526","type":"article-journal","title":"Beyond the Hype: A Structured Scoping Review of Knowledge  Distillation for Accessible and Equitable AI in Education","abstract":"Large language models (LLMs) hold substantial potential for educational applications, including automated essay scoring, short-answer grading, personalized feedback, and interactive tutoring, yet their widespread deployment remains constrained by high computational costs, dependence on proprietary APIs, and infrastructural limitations in under-resourced settings. This scoping review examines knowledge distillation (KD) as a principled and increasingly practical approach to these barriers. KD transfers capabilities of large, expensive teacher models into smaller, efficient student models, offering a practical pathway for democratizing high-quality AI in education without requiring powerful hardware or proprietary systems. Drawing on a structured literature review of 203 papers published between 2017 and early 2025 across IEEE Xplore, ACM Digital Library, SpringerLink, arXiv, and Google Scholar, this survey organizes the field into four thematic areas: (i) distillation for automated scoring and classification; (ii) distillation of pedagogical reasoning and interpretable explanations, including chain-of-thought and rubric-aligned variants; (iii) emerging frontiers in multimodality and low-resource language adaptation; and (iv) federated and privacy-preserving distillation for on-device deployment. The literature demonstrates that transferring intermediate reasoning signals, multi-teacher consensus, and structured pedagogical rationales substantially improves the interpretability and deployability of compact student models, while highlighting the critical need for explicit fairness auditing. We identify six critical open research challenges: multi-turn dialogue distillation, cross-lingual reasoning transfer, on-device personalized federated learning, evaluation standardization, bias auditing, and distillation of extended reasoning trajectories from large reasoning models such as OpenAI o1 and DeepSeek-R1. We propose a structured research agenda, a standardized evidence framework, and deployment-oriented design guidance for practitioners, positioning KD as a promising infrastructure for equitable and resource-efficient AI in education.","author":[{"family":"Ahmed","given":"Sarfraz"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.16875526","URL":"https://doi.org/10.5281/zenodo.16875526","source":"datacite"},{"id":"doi:10.5281/zenodo.20679167","type":"article-journal","title":"Beyond the Hype: A Structured Scoping Review of Knowledge  Distillation for Accessible and Equitable AI in Education","abstract":"Large language models (LLMs) hold substantial potential for educational applications, including automated essay scoring, short-answer grading, personalized feedback, and interactive tutoring, yet their widespread deployment remains constrained by high computational costs, dependence on proprietary APIs, and infrastructural limitations in under-resourced settings. This scoping review examines knowledge distillation (KD) as a principled and increasingly practical approach to these barriers. KD transfers capabilities of large, expensive teacher models into smaller, efficient student models, offering a practical pathway for democratizing high-quality AI in education without requiring powerful hardware or proprietary systems. Drawing on a structured literature review of 203 papers published between 2017 and early 2025 across IEEE Xplore, ACM Digital Library, SpringerLink, arXiv, and Google Scholar, this survey organizes the field into four thematic areas: (i) distillation for automated scoring and classification; (ii) distillation of pedagogical reasoning and interpretable explanations, including chain-of-thought and rubric-aligned variants; (iii) emerging frontiers in multimodality and low-resource language adaptation; and (iv) federated and privacy-preserving distillation for on-device deployment. The literature demonstrates that transferring intermediate reasoning signals, multi-teacher consensus, and structured pedagogical rationales substantially improves the interpretability and deployability of compact student models, while highlighting the critical need for explicit fairness auditing. We identify six critical open research challenges: multi-turn dialogue distillation, cross-lingual reasoning transfer, on-device personalized federated learning, evaluation standardization, bias auditing, and distillation of extended reasoning trajectories from large reasoning models such as OpenAI o1 and DeepSeek-R1. We propose a structured research agenda, a standardized evidence framework, and deployment-oriented design guidance for practitioners, positioning KD as a promising infrastructure for equitable and resource-efficient AI in education.","author":[{"family":"Ahmed","given":"Sarfraz"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20679167","URL":"https://doi.org/10.5281/zenodo.20679167","source":"datacite"},{"id":"doi:10.5281/zenodo.19658831","type":"article-journal","title":"Computational Biology vs Bioinformatics: A 2025 Guide for Biomedical Researchers","abstract":"This comprehensive guide for biomedical researchers and drug development professionals delineates the distinct yet complementary roles of bioinformatics and computational biology. Bioinformatics is defined as the field focused on developing and applying computational tools, software, and algorithms to manage, organize, and analyze large-scale biological datasets, such as those from genomics and proteomics. It provides the essential infrastructure for handling biological big data. Computational biology, conversely, is more concerned with developing theoretical methods, mathematical models, and computational simulations to understand and predict the behavior of complex biological systems. It uses data processed by bioinformatics to build models that test hypotheses about biological mechanisms, from protein folding to cellular signaling pathways. The relationship between the fields is synergistic: bioinformatics supplies the structured data and analytical tools that computational biology uses to construct and validate its models. This integrated workflow is critical for modern research. Key applications highlighted include drug discovery, where AI and machine learning are used to identify targets and optimize lead compounds, and personalized medicine, which relies on genomic data analysis to tailor treatments. The guide details essential tools for each field, such as BLAST and GATK for bioinformatics, and molecular dynamics software like GROMACS for computational biology. It also addresses significant challenges, including data management, security, and the need for reproducible, scalable analysis pipelines, emphasizing the role of cloud platforms and SaaS solutions in enhancing accessibility and collaboration. Future trends point towards deeper integration of AI, multi-modal data analysis, and federated learning to further accelerate discovery. Source: https://www.compbiosci.com/posts/computational-biology-vs-bioinformatics-a-2025-guide-for-biomedical-researchers","author":[{"family":"Science","given":"Computational"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19658831","URL":"https://doi.org/10.5281/zenodo.19658831","source":"datacite"},{"id":"doi:10.5281/zenodo.19658832","type":"article-journal","title":"Computational Biology vs Bioinformatics: A 2025 Guide for Biomedical Researchers","abstract":"This comprehensive guide for biomedical researchers and drug development professionals delineates the distinct yet complementary roles of bioinformatics and computational biology. Bioinformatics is defined as the field focused on developing and applying computational tools, software, and algorithms to manage, organize, and analyze large-scale biological datasets, such as those from genomics and proteomics. It provides the essential infrastructure for handling biological big data. Computational biology, conversely, is more concerned with developing theoretical methods, mathematical models, and computational simulations to understand and predict the behavior of complex biological systems. It uses data processed by bioinformatics to build models that test hypotheses about biological mechanisms, from protein folding to cellular signaling pathways. The relationship between the fields is synergistic: bioinformatics supplies the structured data and analytical tools that computational biology uses to construct and validate its models. This integrated workflow is critical for modern research. Key applications highlighted include drug discovery, where AI and machine learning are used to identify targets and optimize lead compounds, and personalized medicine, which relies on genomic data analysis to tailor treatments. The guide details essential tools for each field, such as BLAST and GATK for bioinformatics, and molecular dynamics software like GROMACS for computational biology. It also addresses significant challenges, including data management, security, and the need for reproducible, scalable analysis pipelines, emphasizing the role of cloud platforms and SaaS solutions in enhancing accessibility and collaboration. Future trends point towards deeper integration of AI, multi-modal data analysis, and federated learning to further accelerate discovery. Source: https://www.compbiosci.com/posts/computational-biology-vs-bioinformatics-a-2025-guide-for-biomedical-researchers","author":[{"family":"Science","given":"Computational"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19658832","URL":"https://doi.org/10.5281/zenodo.19658832","source":"datacite"},{"id":"doi:10.5281/zenodo.19490340","type":"article-journal","title":"LongevityCommon","abstract":"Aging, codified in ICD-11 (code XT9T “Ageing-related,” 2018; code MG2A “Ageing-associated decline in intrinsic capacity,” 2025), requires integrative biomarker frameworks that extend beyond individual epigenetic clocks and wearable predictors. We present a hypothesis-stage integrative framework, LongevityCommon. All empirical estimates should be treated as exploratory (hypothesis‑generating), not confirmatory. Pre‑registered tests of an earlier univariate formulation of χ_Ze on the Cuban EEG, Dortmund Vital, and MPI‑LEMON cohorts yielded NULL results (documented in Ze EVIDENCE.md from 2026-04-22; meta‑analysis of Cuban + Dortmund: I² = 90.3 % — invalid; these results were retracted). The current multimodal version of χ_Ze is a post‑hoc reformulation and was not pre‑registered. All reported AUCs are exploratory hypothesis‑generating only, with explicit acknowledgment of p‑hacking risk (Ioannidis, 2005). LongevityCommon is a conceptual ecosystem of five components: (1) MCOA (Multi‑Counter Architecture of Organismal Aging) — a meta‑theory positing aging as a weighted sum of parallel counters: Ltissue(n,t)=iwi(tissue)fi(Di(n,t)). Axioms M1–M4 include an operational definition of falsifiability (M4) revised in v5 based on community‑standard validation thresholds: MCOA is considered falsified if on a pre‑registered cohort with N ≥ 2000 at α = 0.001 the partial r² for all‑cause mortality after controlling for chronological age and sex is full R² = 0.778), but full Sobol decomposition (S2 + ST) with 95 % CI and nested cross‑validation on real GTEx data (N = 948) revealed that the difference is not statistically significant (p = 0.12 after correction). CDATA remains an open, falsifiable hypothesis; final determination requires full decomposition on real data (Cell‑DT v4.0). (3) Ze Theory — a thermodynamic‑geometric formalism. The equation dZe/dt=−I(Z) is postulated as an ansatz, motivated by Burgholzer (2015) and Pearson et al. (2021); living systems are far‑from‑equilibrium, and a formal bridge between these physical‑clock systems and biological aging is absent. (4) BioSense — a wearable platform with the variational principle F=E−TS−Ipred. The theoretical fixed point v*=0.45631 at k=1 (sensitivity range v*[0.32;0.58] for k[0.5;2.0]). **Empirically tested via swept‑v* search on All‑of‑Us (N = 500): v*_optimal = 0.451 (95 % CI 0.443–0.459), consistent with the theoretical value.** (5) FCLC (Federated Clinical Learning Cooperative) — a federated learning infrastructure. ε_total ≈ 0.43 at (σ, q, T) = (1.5, 0.013, 5); RDP composition via subsampled Gaussian (Wang et al., 2019) combined with the Mironov (2017) framework. Threat model (explicit disclosure, v5): (a) FCLC central server — semi‑honest only (never sees raw data, only aggregated updates); (b) Byzantine‑robust aggregation (Krum up to 25 % malicious clients); (c) NOT secure against active server collusion; (d) NOT secure against malicious server deviating from the protocol. This is a blocker for GDPR Article 9 medical data until FCLC v14 (malicious‑secure migration planned for Q1 2027). The ecosystem provides a falsifiable (per updated M4) draft platform, but persistent blocking limitations remain: (i) all pilots are underpowered, not pre‑registered, and post‑hoc, which in light of Ioannidis (2005) gives a clear risk of false‑positive findings; (ii) FCLC is semi‑honest only — blocker for GDPR Art. 9; (iii) CDATA status is inconclusive, requiring full Sobol decomposition on real data; (iv) v* empirically tested and confirmed; (v) Ze→biology is an ansatz without formal derivation; (vi) bridge to CDATA (5 parameters on N = 196) is underpowered (Harrell rule violated) and moved to Supplementary; (vii) key publications (MCOA, Ze, BioSense) are not peer‑reviewed; (viii) EIC Pathfinder consortium formation deferred to Q1 2027 (0 signed EU LoIs as of 2026-04-21).","author":[{"family":"Tkemaladze","given":"Jaba"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19490340","URL":"https://doi.org/10.5281/zenodo.19490340","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.08290","type":"manuscript","title":"Trust-Based Incentive Mechanisms in Semi-Decentralized Federated Learning Systems","abstract":"In federated learning (FL), decentralized model training allows multi-ple participants to collaboratively improve a shared machine learning model without exchanging raw data. However, ensuring the integrity and reliability of the system is challenging due to the presence of potentially malicious or faulty nodes that can degrade the model's performance. This paper proposes a novel trust-based incentive mechanism designed to evaluate and reward the quality of contributions in FL systems. By dynamically assessing trust scores based on fac-tors such as data quality, model accuracy, consistency, and contribution fre-quency, the system encourages honest participation and penalizes unreliable or malicious behavior. These trust scores form the basis of an incentive mechanism that rewards high-trust nodes with greater participation opportunities and penal-ties for low-trust participants. We further explore the integration of blockchain technology and smart contracts to automate the trust evaluation and incentive distribution processes, ensuring transparency and decentralization. Our proposed theoretical framework aims to create a more robust, fair, and transparent FL eco-system, reducing the risks posed by untrustworthy participants.","author":[{"family":"Shrestha","given":"Ajay"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.08290","URL":"https://doi.org/10.48550/arxiv.2602.08290","source":"datacite"},{"id":"doi:10.6084/m9.figshare.29877074","type":"article-journal","title":"AI-Powered Visualization is Transforming Modern Healthcare","abstract":"Healthcare is being transformed by AI-driven visualization, which transforms complex data into useful insights. This paper synthesizes advancements in AI visualization tools—spanning medical imaging, electronic health records (EHR), genomics, and public health—and evaluates their impact on diagnostics, treatment personalization, and operational efficiency. Convolutional neural networks (CNNs) for image segmentation, generative adversarial networks (GANs) for the generation of synthetic data, and interactive dashboards for real-time analytics are some of the technologies that we highlight. Integrity barriers, algorithmic bias, and data privacy concerns are all critically examined. A systematic review of more than 120 studies conducted between 2018 and 2024 shows that clinical workflow time is cut by 30% and diagnostic accuracy is improved by 40% on average. Explainable artificial intelligence (XAI) and federated learning are emphasized in the study's ethical frameworks and future directions. This study demonstrates that AI visualization plays a crucial role in value-based care and precision medicine.","author":[{"family":"Rahman Nazil","given":"Ashikur"}],"issued":{"date-parts":[[2025]]},"DOI":"10.6084/m9.figshare.29877074","URL":"https://doi.org/10.6084/m9.figshare.29877074","source":"datacite"},{"id":"doi:10.6084/m9.figshare.29876651","type":"article-journal","title":"AI-Powered Visualization is Transforming Modern Healthcare","abstract":"Healthcare is being transformed by AI-driven visualization, which transforms complex data into useful insights. This paper synthesizes advancements in AI visualization tools—spanning medical imaging, electronic health records (EHR), genomics, and public health—and evaluates their impact on diagnostics, treatment personalization, and operational efficiency. Convolutional neural networks (CNNs) for image segmentation, generative adversarial networks (GANs) for the generation of synthetic data, and interactive dashboards for real-time analytics are some of the technologies that we highlight. Integrity barriers, algorithmic bias, and data privacy concerns are all critically examined. A systematic review of more than 120 studies conducted between 2018 and 2024 shows that clinical workflow time is cut by 30% and diagnostic accuracy is improved by 40% on average. Explainable artificial intelligence (XAI) and federated learning are emphasized in the study's ethical frameworks and future directions. This study demonstrates that AI visualization plays a crucial role in value-based care and precision medicine.","author":[{"family":"Rahman Nazil","given":"Ashikur"}],"issued":{"date-parts":[[2025]]},"DOI":"10.6084/m9.figshare.29876651","URL":"https://doi.org/10.6084/m9.figshare.29876651","source":"datacite"},{"id":"doi:10.5281/zenodo.18150854","type":"article-journal","title":"Lecture-Ready Slides & Exercises:  Generative AI, Cybersecurity, and Ethics - Ray Islam, PhD | Wiley, 2025","abstract":"These ready lecture slides with excercises are derived from the book Generative AI, Cybersecurity, and Ethics by Ray Islam, PhD (Wiley, 2025). Instructors/Researchers are welcome to adapt or modify the materials for their courses with proper attribution. © 2026 Ray Islam, PhD (Mohammad Rubyet Islam)This work is licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0).You are free to use, edit, share, and adapt this material for educational purposes with proper attribution. Link to few selected pages of the book: https://books.google.com/books?id=p8IzEQAAQBAJ&lpg=PA112&pg=PP1#v=onepage&q&f=false BOOK CONTENTS Chapter 1: Introduction Foundations and Evolution of AI and GenAI: Introduces core AI paradigms and traces the evolution from classical AI and machine learning to modern generative models. GenAI in Cybersecurity: Examines AI- and GenAI-driven approaches to threat detection, anomaly analysis, and proactive cyber defense. Ethical and Regulatory Considerations: Analyzes key ethical challenges and global governance frameworks guiding responsible GenAI deployment. Chapter 2: Cyber Security: Understanding the Digital Fortress Cybersecurity Technologies and Architectures: Examines the technological foundations of cybersecurity, including network, application, information, endpoint, cloud, identity, and critical infrastructure security within layered defense architectures. Threat Impact and Sectoral Risk: Analyzes global and regional cybercrime costs and industry-specific threats, demonstrating how cybersecurity technologies mitigate operational, financial, and systemic risks. AI, GenAI, Ethics, and Governance: Explores AI- and GenAI-driven cybersecurity technologies for detection, response, and prediction, alongside ethical challenges and global regulatory frameworks guiding responsible deployment. Chapter 3: Understanding GenAI Foundations and Capabilities of Generative AI: Introduces GenAI as a core AI paradigm focused on generating novel content across modalities, outlining its defining characteristics, major model classes, and distinctions from traditional predictive AI systems. GenAI Technologies, Tools, and Methodologies: Examines the technological landscape of GenAI, including architectures (e.g., GANs, transformers, diffusion models), platforms, frameworks, lifecycle methodologies (MLOps, ModelOps), and validation techniques. Applications, Risks, and Ethical Considerations: Explores real-world GenAI applications across domains such as cybersecurity, healthcare, education, manufacturing, and creative industries, while addressing ethical, security, and governance challenges associated with deployment. Chapter 4: GenAI in Cyber Security GenAI Cybersecurity Technologies: Examines GenAI-driven mechanisms for threat detection, simulation, deception, automated testing, and incident response, highlighting both defensive and offensive capabilities. Risks and Mitigation Strategies: Analyzes GenAI-enabled threats-including phishing, malware, deepfakes, and adversarial attacks-and corresponding mitigation technologies such as defensive AI, adversarial learning, and continuous model adaptation. Infrastructure and Governance: Describes the technical and organizational infrastructure required for GenAI-enabled cybersecurity, including compute platforms, data systems, security tool integration, ethical governance, and regulatory compliance. Chapter 5: Foundations of Ethics in GenAI Ethical Foundations and Theoretical Frameworks: Establishes the philosophical, historical, and normative foundations of ethics, applying metaethics, virtue ethics, deontology, consequentialism, and applied ethics to GenAI. Global Standards, Policies, and Regulation: Examines international ethical frameworks, standards, and laws governing AI and GenAI, including ISO/IEC, EU, UNESCO, OECD, IEEE, and regional policy approaches. GenAI-Specific Ethical Challenges and Governance: Analyzes ethical risks such as bias, privacy, misinformation","author":[{"family":"Islam","given":"Phd"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18150854","URL":"https://doi.org/10.5281/zenodo.18150854","source":"datacite"},{"id":"doi:10.5281/zenodo.19067238","type":"article-journal","title":"Lecture-Ready Slides & Exercises:  Generative AI, Cybersecurity, and Ethics - Ray Islam, PhD | Wiley, 2025","abstract":"These ready lecture slides with excercises are derived from the book Generative AI, Cybersecurity, and Ethics by Ray Islam, PhD (Wiley, 2025). Instructors/Researchers are welcome to adapt or modify the materials for their courses with proper attribution. © 2026 Ray Islam, PhD (Mohammad Rubyet Islam)This work is licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0).You are free to use, edit, share, and adapt this material for educational purposes with proper attribution. Link to few selected pages of the book: https://books.google.com/books?id=p8IzEQAAQBAJ&lpg=PA112&pg=PP1#v=onepage&q&f=false BOOK CONTENTS Chapter 1: Introduction Foundations and Evolution of AI and GenAI: Introduces core AI paradigms and traces the evolution from classical AI and machine learning to modern generative models. GenAI in Cybersecurity: Examines AI- and GenAI-driven approaches to threat detection, anomaly analysis, and proactive cyber defense. Ethical and Regulatory Considerations: Analyzes key ethical challenges and global governance frameworks guiding responsible GenAI deployment. Chapter 2: Cyber Security: Understanding the Digital Fortress Cybersecurity Technologies and Architectures: Examines the technological foundations of cybersecurity, including network, application, information, endpoint, cloud, identity, and critical infrastructure security within layered defense architectures. Threat Impact and Sectoral Risk: Analyzes global and regional cybercrime costs and industry-specific threats, demonstrating how cybersecurity technologies mitigate operational, financial, and systemic risks. AI, GenAI, Ethics, and Governance: Explores AI- and GenAI-driven cybersecurity technologies for detection, response, and prediction, alongside ethical challenges and global regulatory frameworks guiding responsible deployment. Chapter 3: Understanding GenAI Foundations and Capabilities of Generative AI: Introduces GenAI as a core AI paradigm focused on generating novel content across modalities, outlining its defining characteristics, major model classes, and distinctions from traditional predictive AI systems. GenAI Technologies, Tools, and Methodologies: Examines the technological landscape of GenAI, including architectures (e.g., GANs, transformers, diffusion models), platforms, frameworks, lifecycle methodologies (MLOps, ModelOps), and validation techniques. Applications, Risks, and Ethical Considerations: Explores real-world GenAI applications across domains such as cybersecurity, healthcare, education, manufacturing, and creative industries, while addressing ethical, security, and governance challenges associated with deployment. Chapter 4: GenAI in Cyber Security GenAI Cybersecurity Technologies: Examines GenAI-driven mechanisms for threat detection, simulation, deception, automated testing, and incident response, highlighting both defensive and offensive capabilities. Risks and Mitigation Strategies: Analyzes GenAI-enabled threats-including phishing, malware, deepfakes, and adversarial attacks-and corresponding mitigation technologies such as defensive AI, adversarial learning, and continuous model adaptation. Infrastructure and Governance: Describes the technical and organizational infrastructure required for GenAI-enabled cybersecurity, including compute platforms, data systems, security tool integration, ethical governance, and regulatory compliance. Chapter 5: Foundations of Ethics in GenAI Ethical Foundations and Theoretical Frameworks: Establishes the philosophical, historical, and normative foundations of ethics, applying metaethics, virtue ethics, deontology, consequentialism, and applied ethics to GenAI. Global Standards, Policies, and Regulation: Examines international ethical frameworks, standards, and laws governing AI and GenAI, including ISO/IEC, EU, UNESCO, OECD, IEEE, and regional policy approaches. GenAI-Specific Ethical Challenges and Governance: Analyzes ethical risks such as bias, privacy, misinformation","author":[{"family":"Islam","given":"Phd"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19067238","URL":"https://doi.org/10.5281/zenodo.19067238","source":"datacite"},{"id":"doi:10.5281/zenodo.18150855","type":"article-journal","title":"Lecture-Ready Slides & Exercises:  Generative AI, Cybersecurity, and Ethics - Ray Islam, PhD | Wiley, 2025","abstract":"These ready lecture slides with excercises are derived from the book Generative AI, Cybersecurity, and Ethics by Ray Islam, PhD (Wiley, 2025). Instructors/Researchers are welcome to adapt or modify the materials for their courses with proper attribution. © 2026 Ray Islam, PhD (Mohammad Rubyet Islam)This work is licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0).You are free to use, edit, share, and adapt this material for educational purposes with proper attribution. Link to few selected pages of the book: https://books.google.com/books?id=p8IzEQAAQBAJ&lpg=PA112&pg=PP1#v=onepage&q&f=false BOOK CONTENTS Chapter 1: Introduction Foundations and Evolution of AI and GenAI: Introduces core AI paradigms and traces the evolution from classical AI and machine learning to modern generative models. GenAI in Cybersecurity: Examines AI- and GenAI-driven approaches to threat detection, anomaly analysis, and proactive cyber defense. Ethical and Regulatory Considerations: Analyzes key ethical challenges and global governance frameworks guiding responsible GenAI deployment. Chapter 2: Cyber Security: Understanding the Digital Fortress Cybersecurity Technologies and Architectures: Examines the technological foundations of cybersecurity, including network, application, information, endpoint, cloud, identity, and critical infrastructure security within layered defense architectures. Threat Impact and Sectoral Risk: Analyzes global and regional cybercrime costs and industry-specific threats, demonstrating how cybersecurity technologies mitigate operational, financial, and systemic risks. AI, GenAI, Ethics, and Governance: Explores AI- and GenAI-driven cybersecurity technologies for detection, response, and prediction, alongside ethical challenges and global regulatory frameworks guiding responsible deployment. Chapter 3: Understanding GenAI Foundations and Capabilities of Generative AI: Introduces GenAI as a core AI paradigm focused on generating novel content across modalities, outlining its defining characteristics, major model classes, and distinctions from traditional predictive AI systems. GenAI Technologies, Tools, and Methodologies: Examines the technological landscape of GenAI, including architectures (e.g., GANs, transformers, diffusion models), platforms, frameworks, lifecycle methodologies (MLOps, ModelOps), and validation techniques. Applications, Risks, and Ethical Considerations: Explores real-world GenAI applications across domains such as cybersecurity, healthcare, education, manufacturing, and creative industries, while addressing ethical, security, and governance challenges associated with deployment. Chapter 4: GenAI in Cyber Security GenAI Cybersecurity Technologies: Examines GenAI-driven mechanisms for threat detection, simulation, deception, automated testing, and incident response, highlighting both defensive and offensive capabilities. Risks and Mitigation Strategies: Analyzes GenAI-enabled threats-including phishing, malware, deepfakes, and adversarial attacks-and corresponding mitigation technologies such as defensive AI, adversarial learning, and continuous model adaptation. Infrastructure and Governance: Describes the technical and organizational infrastructure required for GenAI-enabled cybersecurity, including compute platforms, data systems, security tool integration, ethical governance, and regulatory compliance. Chapter 5: Foundations of Ethics in GenAI Ethical Foundations and Theoretical Frameworks: Establishes the philosophical, historical, and normative foundations of ethics, applying metaethics, virtue ethics, deontology, consequentialism, and applied ethics to GenAI. Global Standards, Policies, and Regulation: Examines international ethical frameworks, standards, and laws governing AI and GenAI, including ISO/IEC, EU, UNESCO, OECD, IEEE, and regional policy approaches. GenAI-Specific Ethical Challenges and Governance: Analyzes ethical risks such as bias, privacy, misinformation","author":[{"family":"Islam","given":"Phd"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18150855","URL":"https://doi.org/10.5281/zenodo.18150855","source":"datacite"},{"id":"doi:10.6084/m9.figshare.31557869.v1","type":"article-journal","title":"AI-enhanced health information exchange systems: a systematic review of implementation challenges and opportunities","abstract":"This systematic review evaluates the integration of Artificial Intelligence (AI) into Health Information Exchange (HIE) systems to improve data utilisation, patient outcomes, and healthcare delivery efficiency. Through a comprehensive analysis of major academic databases from 2015 to 2024, 72 peer-reviewed papers were selected following rigorous screening processes, including duplicate removal, title/abstract screening, and full-text evaluation against predefined criteria. The analysis examined AI implementation across directed exchange, query-based exchange, and consumer-mediated exchange systems. Results demonstrate AI significantly enhances HIE functionality through improved data standardisation, error detection, and system integration. Key benefits include enhanced predictive capabilities for patient outcomes, optimised resource utilisation, and effective automated processing of unstructured clinical data. Critical solutions identified include federated learning techniques to preserve privacy, explainable AI models to support clinical adoption, and robust bias-detection frameworks to ensure equitable outcomes. Future opportunities encompass integrating diverse data sources, exploring federated learning approaches, balancing data utility with privacy concerns, and developing transparent AI models to increase clinical acceptance. The study emphasises ongoing data standardisation efforts and ethical considerations in AI deployment, particularly for real-time, context-aware decision-support systems. AI integration represents a promising advancement that, through strategic development, can transform HIE into more effective and equitable tools for healthcare delivery.","author":[{"family":"Esmaeilzadeh","given":"Pouyan"},{"family":"Maddah","given":"Mahed"}],"issued":{"date-parts":[[2026]]},"DOI":"10.6084/m9.figshare.31557869.v1","URL":"https://doi.org/10.6084/m9.figshare.31557869.v1","source":"datacite"},{"id":"doi:10.6084/m9.figshare.31557869","type":"article-journal","title":"AI-enhanced health information exchange systems: a systematic review of implementation challenges and opportunities","abstract":"This systematic review evaluates the integration of Artificial Intelligence (AI) into Health Information Exchange (HIE) systems to improve data utilisation, patient outcomes, and healthcare delivery efficiency. Through a comprehensive analysis of major academic databases from 2015 to 2024, 72 peer-reviewed papers were selected following rigorous screening processes, including duplicate removal, title/abstract screening, and full-text evaluation against predefined criteria. The analysis examined AI implementation across directed exchange, query-based exchange, and consumer-mediated exchange systems. Results demonstrate AI significantly enhances HIE functionality through improved data standardisation, error detection, and system integration. Key benefits include enhanced predictive capabilities for patient outcomes, optimised resource utilisation, and effective automated processing of unstructured clinical data. Critical solutions identified include federated learning techniques to preserve privacy, explainable AI models to support clinical adoption, and robust bias-detection frameworks to ensure equitable outcomes. Future opportunities encompass integrating diverse data sources, exploring federated learning approaches, balancing data utility with privacy concerns, and developing transparent AI models to increase clinical acceptance. The study emphasises ongoing data standardisation efforts and ethical considerations in AI deployment, particularly for real-time, context-aware decision-support systems. AI integration represents a promising advancement that, through strategic development, can transform HIE into more effective and equitable tools for healthcare delivery.","author":[{"family":"Esmaeilzadeh","given":"Pouyan"},{"family":"Maddah","given":"Mahed"}],"issued":{"date-parts":[[2026]]},"DOI":"10.6084/m9.figshare.31557869","URL":"https://doi.org/10.6084/m9.figshare.31557869","source":"datacite"},{"id":"doi:10.5281/zenodo.18344155","type":"article-journal","title":"Shuffle Model Privacy for Federated Learning: Asymptotic GDP Analysis and Bundled/Unbundled Separation","abstract":"Exact asymptotic GDP parameters for shuffle model federated learning via chi-squared divergence. Proves exponential separation between bundled and unbundled shuffling designs in feasible SGD iterations. - After publication, the following related works were identified that share conceptual connections with this paper: 1. Su, Cheng, Wang (arXiv:2504.07414, 2025) analyze \"joint composition\" in shuffle model where user outputs are shuffled as a tuple — conceptually similar to our \"bundled\" regime. Their focus is on (ε,δ)-DP bounds via FFT, not µ-GDP with χ²-parameterization. 2. Su, Cheng, Wang (arXiv:2511.15051, 2025) show χ²(P‖Q) appears in mutual information asymptotics for shuffle model. Our contribution uses χ² for µ-GDP (differential privacy), not information-theoretic leakage. 3. Chen, Cao, Ge (AAAI 2024, arXiv:2312.14388) provide f-DP analysis for personalized LDP shuffle settings. They do not derive closed-form µ² expressions parameterized solely by ��²(W₁‖W₀). The core contributions of this paper remain novel: • Exact closed-form µ²_{n,un} = mχ²/n and µ²_{n,bd} = ((1+χ²)^m - 1)/n • Explicit exponential separation ratio Θ((1+χ²)^m/m) • SGD iteration bounds under µ-GDP budget A revised version with expanded Related Work is planned for journal submission.","author":[{"family":"Shvets","given":"Alex"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18344155","URL":"https://doi.org/10.5281/zenodo.18344155","source":"datacite"},{"id":"doi:10.5281/zenodo.17236145","type":"article-journal","title":"CONFIDENTIAL6G – Latest Updates from the Project M19 – M30","abstract":"This newsletter presents the latest updates from the Horizon Europe project CONFIDENTIAL6G (GA No. 101096435), covering activities from months 19 to 30. The issue highlights recent project achievements, events, and publications that advance confidentiality, privacy, and security in next-generation 6G networks. Key sections include: Events & News – project representation at Open Source Summit Europe 2024, FHE.org 2025, WSSS 2025, Secure Automotive OTA Seminar, and EuCNC & 6G Summit 2025. New Partnership – TU Wien joining the consortium, strengthening cryptography and security research. Use Cases – secure aviation maintenance, telecom cloud security, and connected vehicles with OTA updates. Technical Achievements – confidential computing toolkit, decentralized data sharing frameworks, confidential AI orchestration, and federated learning in connected vehicles. Open-Source Contributions – release of multiple libraries and tools for cryptography, FHE, MPC, and ZKPs. Scientific Outputs – peer-reviewed publications on secure computation, post-quantum cryptography, hardware acceleration, MPC, and cryptanalysis. Community Engagement – AI & Security webinar with sister projects, dissemination through SNS Journal 2025, and public articles on privacy and confidential computing. The newsletter demonstrates how CONFIDENTIAL6G contributes to building secure, privacy-preserving, and quantum-resistant infrastructures for future 6G systems.","author":[{"family":"Vasic","given":"Jelena"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17236145","URL":"https://doi.org/10.5281/zenodo.17236145","source":"datacite"},{"id":"doi:10.5281/zenodo.17236144","type":"article-journal","title":"CONFIDENTIAL6G – Latest Updates from the Project M19 – M30","abstract":"This newsletter presents the latest updates from the Horizon Europe project CONFIDENTIAL6G (GA No. 101096435), covering activities from months 19 to 30. The issue highlights recent project achievements, events, and publications that advance confidentiality, privacy, and security in next-generation 6G networks. Key sections include: Events & News – project representation at Open Source Summit Europe 2024, FHE.org 2025, WSSS 2025, Secure Automotive OTA Seminar, and EuCNC & 6G Summit 2025. New Partnership – TU Wien joining the consortium, strengthening cryptography and security research. Use Cases – secure aviation maintenance, telecom cloud security, and connected vehicles with OTA updates. Technical Achievements – confidential computing toolkit, decentralized data sharing frameworks, confidential AI orchestration, and federated learning in connected vehicles. Open-Source Contributions – release of multiple libraries and tools for cryptography, FHE, MPC, and ZKPs. Scientific Outputs – peer-reviewed publications on secure computation, post-quantum cryptography, hardware acceleration, MPC, and cryptanalysis. Community Engagement – AI & Security webinar with sister projects, dissemination through SNS Journal 2025, and public articles on privacy and confidential computing. The newsletter demonstrates how CONFIDENTIAL6G contributes to building secure, privacy-preserving, and quantum-resistant infrastructures for future 6G systems.","author":[{"family":"Vasic","given":"Jelena"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17236144","URL":"https://doi.org/10.5281/zenodo.17236144","source":"datacite"},{"id":"doi:10.82497/aitsde.2025.1218507","type":"article-journal","title":"Deep Learning for Enhancing IoT Security and Trust","abstract":"The Internet of Things (IoT) has grown into a massive ecosystem connecting billions of heterogeneous devices, from household sensors to industrial machinery, and is expected to exceed 75 billion nodes by 2025. This rapid expansion has introduced critical challenges related to data security, privacy, and device trustworthiness. IoT networks are highly vulnerable to a wide range of attacks, including unauthorized access, denial-of-service, spoofing, and data tampering. Traditional security solutions such as static authentication, encryption, and rule-based intrusion detection have proven insufficient to address the dynamic and complex nature of IoT threats. In this context, deep learning (DL) has emerged as a promising approach for anomaly detection and intrusion prevention, offering the ability to automatically learn patterns in high-dimensional IoT data. This paper surveys recent advances in DL-based IoT security and emphasizes the importance of integrating trust management mechanisms with anomaly detection. By assigning dynamic trust scores to IoT nodes based on their historical behavior and anomaly likelihood, the proposed framework enhances resilience and reduces false positives compared to traditional intrusion detection systems. The research leverages LSTM-based models to detect abnormal traffic and employs datasets such as NSL-KDD and CICIDS2017 for validation. Comparative analysis with conventional machine learning techniques (e.g., SVM, Random Forest) demonstrates superior performance in terms of accuracy, precision, recall, F1-score, and false alarm rate. The study highlights existing research gaps, including energy efficiency and lightweight model deployment, and concludes by proposing future directions such as blockchain integration and federated learning for scalable, secure IoT networks.","author":[{"family":"Anaraki","given":"Alireza"}],"issued":{"date-parts":[[2025]]},"DOI":"10.82497/aitsde.2025.1218507","URL":"https://doi.org/10.82497/aitsde.2025.1218507","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.06963","type":"manuscript","title":"Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies","abstract":"Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Google, Meta, Microsoft, Stability AI, respectively, are revolutionizing cybersecurity, enabling both automated defense and sophisticated attacks. These technologies power real-time threat detection, phishing defense, secure code generation, and vulnerability exploitation at unprecedented scales. Following a rapid surge where LLM-generated malware grew to account for an estimated 50% of detected threats by 2025, up from just 2% in 2021, navigating this highly automated threat landscape in 2026 demands next-generation security frameworks. This paper presents a comprehensive survey of the beneficial and malicious applications of LLMs in cybersecurity, including zero-day detection, DevSecOps, federated learning, synthetic content analysis, and explainable AI (XAI). Drawing on a review of over 70 academic papers, industry reports, and technical documents, this work synthesizes insights from real-world case studies across platforms like Google Play Protect, Microsoft Defender, Amazon Web Services (AWS), Apple App Store, OpenAI Plugin Stores, Hugging Face Spaces, and GitHub, alongside emerging initiatives like the SAFE Framework and AI-driven anomaly detection. We conclude with practical recommendations for responsible and transparent LLM deployment and trustworthy AI, including model watermarking, adversarial defense, and cross-industry collaboration, setting a new benchmark for rigorous, holistic cybersecurity research at the intersection of AI and threat defense, and offering a roadmap for secure, scalable LLM systems that serves as a critical reference for researchers, engineers, and security leaders navigating the complex challenges of AI-driven cybersecurity.","author":[{"family":"Ahi","given":"Kiarash"},{"family":"Valizadeh","given":"Saeed"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.06963","URL":"https://doi.org/10.48550/arxiv.2607.06963","source":"datacite"},{"id":"doi:10.5281/zenodo.21030653","type":"article-journal","title":"D8.8: Report on European Interoperability Framework Contributions","abstract":"This deliverable (D8.8) presents the contributions of the PLIADES project to advancing the European Interoperability Framework (EIF) and interoperability standardization, with a focus on building a modular, scalable, and trustworthy data sharing ecosystem. The work conducted in Task 8.6 evaluates the project’s alignment with the EIF’s four layers-legal, organizational, semantic, and technical—and extends this analysis through ISO/IEC 19941’s five interoperability facets—policy, behavior, semantic, syntactic, and transport. By applying a general interoperability-framework approach, the deliverable assesses the interoperability maturity of PLIADES across six use cases in domains including mobility, energy, healthcare, green deal/ circular economy, energy, and industry. These use cases demonstrate how PLIADES supports dynamic, cross-domain data integration through advanced AI capabilities such as federated learning, explainable AI, and declarative querying—while ensuring legal compliance, data sovereignty, and semantic clarity. The project engages directly with EU standardization and governance initiatives—including SEMIC, DSSC, and the European Trusted Data Framework standardisation request to ensure alignment with emerging regulations like the Data Act. PLIADES actively contributes to the EU’s semantic and technical interoperability agenda through workshops, conference participation (e.g., SEMIC 2025, ENDORSE 2025), and alignment with the IDS Rulebook and the Dataspace Protocol. Gaps identified in current interoperability models—such as limited runtime interoperability, lack of support for decentralized AI, and insufficient metadata expressiveness—are addressed through actionable recommendations. PLIADES proposes enhancements to semantic alignment, dynamic querying, and data governance architectures, helping to shape the next iteration of European data policy frameworks. Ultimately, this report underscores PLIADES’ strategic role in fostering cross-border, cross-sector data interoperability. By operationalizing both EIF and ISO-based principles through real-world use cases and aligning with EU standardization initiatives, PLIADES delivers a blueprint for trusted, AI-enabled, sovereign data spaces that drive innovation and support Europe’s digital transition. PLIADES stands for an advanced AI AI-enabled framework for Full Data Lifecycles Optimisation and Data Spaces Integration. Our mission is to revolutionize how data is utilised across various sectors, from mobility to healthcare, manufacturing to energy, and beyond. PLIADES envisions a future where diverse sectors are seamlessly interconnected, enhancing efficiency and interoperability. We aim to provide cutting cutting-edge data and services that drive advancements in Cooperative, Connected, and Automated Mobility (CCAM), Advanced Driver Assistance & Autonomous Driving (ADAS/AD), and HumanHuman-Robot Interaction (HRI).","author":[{"family":"Greece","given":"Hypertech"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21030653","URL":"https://doi.org/10.5281/zenodo.21030653","source":"datacite"},{"id":"doi:10.5281/zenodo.21030654","type":"article-journal","title":"D8.8: Report on European Interoperability Framework Contributions","abstract":"This deliverable (D8.8) presents the contributions of the PLIADES project to advancing the European Interoperability Framework (EIF) and interoperability standardization, with a focus on building a modular, scalable, and trustworthy data sharing ecosystem. The work conducted in Task 8.6 evaluates the project’s alignment with the EIF’s four layers-legal, organizational, semantic, and technical—and extends this analysis through ISO/IEC 19941’s five interoperability facets—policy, behavior, semantic, syntactic, and transport. By applying a general interoperability-framework approach, the deliverable assesses the interoperability maturity of PLIADES across six use cases in domains including mobility, energy, healthcare, green deal/ circular economy, energy, and industry. These use cases demonstrate how PLIADES supports dynamic, cross-domain data integration through advanced AI capabilities such as federated learning, explainable AI, and declarative querying—while ensuring legal compliance, data sovereignty, and semantic clarity. The project engages directly with EU standardization and governance initiatives—including SEMIC, DSSC, and the European Trusted Data Framework standardisation request to ensure alignment with emerging regulations like the Data Act. PLIADES actively contributes to the EU’s semantic and technical interoperability agenda through workshops, conference participation (e.g., SEMIC 2025, ENDORSE 2025), and alignment with the IDS Rulebook and the Dataspace Protocol. Gaps identified in current interoperability models—such as limited runtime interoperability, lack of support for decentralized AI, and insufficient metadata expressiveness—are addressed through actionable recommendations. PLIADES proposes enhancements to semantic alignment, dynamic querying, and data governance architectures, helping to shape the next iteration of European data policy frameworks. Ultimately, this report underscores PLIADES’ strategic role in fostering cross-border, cross-sector data interoperability. By operationalizing both EIF and ISO-based principles through real-world use cases and aligning with EU standardization initiatives, PLIADES delivers a blueprint for trusted, AI-enabled, sovereign data spaces that drive innovation and support Europe’s digital transition. PLIADES stands for an advanced AI AI-enabled framework for Full Data Lifecycles Optimisation and Data Spaces Integration. Our mission is to revolutionize how data is utilised across various sectors, from mobility to healthcare, manufacturing to energy, and beyond. PLIADES envisions a future where diverse sectors are seamlessly interconnected, enhancing efficiency and interoperability. We aim to provide cutting cutting-edge data and services that drive advancements in Cooperative, Connected, and Automated Mobility (CCAM), Advanced Driver Assistance & Autonomous Driving (ADAS/AD), and HumanHuman-Robot Interaction (HRI).","author":[{"family":"Greece","given":"Hypertech"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21030654","URL":"https://doi.org/10.5281/zenodo.21030654","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.18120","type":"manuscript","title":"A Coopetitive-Compatible Data Generation Framework for Cross-silo Federated Learning","abstract":"Cross-silo federated learning (CFL) enables organizations (e.g., hospitals or banks) to collaboratively train artificial intelligence (AI) models while preserving data privacy by keeping data local. While prior work has primarily addressed statistical heterogeneity across organizations, a critical challenge arises from economic competition, where organizations may act as market rivals, making them hesitant to participate in joint training due to potential utility loss (i.e., reduced net benefit). Furthermore, the combined effects of statistical heterogeneity and inter-organizational competition on organizational behavior and system-wide social welfare remain underexplored. In this paper, we propose CoCoGen, a coopetitive-compatible data generation framework, leveraging generative AI (GenAI) and potential game theory to model, analyze, and optimize collaborative learning under heterogeneous and competitive settings. Specifically, CoCoGen characterizes competition and statistical heterogeneity through learning performance and utility-based formulations and models each training round as a weighted potential game. We then derive GenAI-based data generation strategies that maximize social welfare. Experimental results on the Fashion-MNIST dataset reveal how varying heterogeneity and competition levels affect organizational behavior and demonstrate that CoCoGen consistently outperforms baseline methods.","author":[{"family":"Nguyen","given":"Thanh"},{"family":"Pham","given":"Quoc"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.18120","URL":"https://doi.org/10.48550/arxiv.2509.18120","source":"datacite"},{"id":"doi:10.5281/zenodo.20624051","type":"article-journal","title":"ELASTIC Newsletter #4: Project Highlights and Achievements (June 2025–May 2026)","abstract":"This issue of the ELASTIC Newsletter highlights the project's progress and achievements from June 2025 to May 2026. ELASTIC is advancing secure, efficient, and scalable orchestration for next-generation 6G networks by leveraging WebAssembly, eBPF, confidential computing, and distributed service management. Key highlights in this edition include: Mid-Term Reporting (RP1): Successful completion of the first reporting period, with strong progress across research, technical development, demonstrator preparation, dissemination, standardisation, and exploitation activities. Consortium Meeting in Chania (Sep 2025): Hosted by the Technical University of Crete, focused on review preparation and alignment across partners. Consortium Meeting in Lund (Jan 2026): Accelerating integration across the secure orchestration stack, with discussions on remote attestation, confidential workload execution, edge AI integration, and policy-driven security orchestration. Demo Series: Showcasing key technologies including Demonstrator 1 (Smart Connected Factory), Demonstrator 2 MVP (Sensitive IT Service Migration to Public Cloud), Propeller (lightweight WebAssembly-based orchestration), and Wasm-operator (efficient Kubernetes orchestration). Events & Webinars: Participation in the Women in ICT Standardisation Webinar (5th Edition) and a joint ELASTIC & 6G-PATH webinar on intelligent orchestration across the edge–cloud continuum. Scientific Contributions: Publications at ACISP 2026, ICLR 2026, INFOCOM 2026, Cluster Computing, and Transactions on Machine Learning Research (TMLR), covering topics such as WebAssembly security in confidential computing, graph teaching algorithms, split learning optimisation, mobility prediction, and differentially private federated learning. EuCNC & 6G Summit 2026: ELASTIC participation at Booth 76 in Málaga, Spain (2–5 June 2026).","author":[{"family":"Vasic","given":"Jelena"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20624051","URL":"https://doi.org/10.5281/zenodo.20624051","source":"datacite"},{"id":"doi:10.5281/zenodo.20624052","type":"article-journal","title":"ELASTIC Newsletter #4: Project Highlights and Achievements (June 2025–May 2026)","abstract":"This issue of the ELASTIC Newsletter highlights the project's progress and achievements from June 2025 to May 2026. ELASTIC is advancing secure, efficient, and scalable orchestration for next-generation 6G networks by leveraging WebAssembly, eBPF, confidential computing, and distributed service management. Key highlights in this edition include: Mid-Term Reporting (RP1): Successful completion of the first reporting period, with strong progress across research, technical development, demonstrator preparation, dissemination, standardisation, and exploitation activities. Consortium Meeting in Chania (Sep 2025): Hosted by the Technical University of Crete, focused on review preparation and alignment across partners. Consortium Meeting in Lund (Jan 2026): Accelerating integration across the secure orchestration stack, with discussions on remote attestation, confidential workload execution, edge AI integration, and policy-driven security orchestration. Demo Series: Showcasing key technologies including Demonstrator 1 (Smart Connected Factory), Demonstrator 2 MVP (Sensitive IT Service Migration to Public Cloud), Propeller (lightweight WebAssembly-based orchestration), and Wasm-operator (efficient Kubernetes orchestration). Events & Webinars: Participation in the Women in ICT Standardisation Webinar (5th Edition) and a joint ELASTIC & 6G-PATH webinar on intelligent orchestration across the edge–cloud continuum. Scientific Contributions: Publications at ACISP 2026, ICLR 2026, INFOCOM 2026, Cluster Computing, and Transactions on Machine Learning Research (TMLR), covering topics such as WebAssembly security in confidential computing, graph teaching algorithms, split learning optimisation, mobility prediction, and differentially private federated learning. EuCNC & 6G Summit 2026: ELASTIC participation at Booth 76 in Málaga, Spain (2–5 June 2026).","author":[{"family":"Vasic","given":"Jelena"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20624052","URL":"https://doi.org/10.5281/zenodo.20624052","source":"datacite"},{"id":"doi:10.48550/arxiv.2606.10250","type":"manuscript","title":"Multi-Level Analyzation of Imbalance to Resolve Non-IID-Ness in Federated Learning","abstract":"Class imbalance is a common problem in deep learning that severely degrades performance. In federated learning (FL), it is a critical factor contributing to non-identically distributed data (non-IID). Building on several previous attempts, we define and analyze imbalance issues in FL at three levels: inter-case, inter-class, and inter-client. Inter-case imbalance addresses the imbalance in every single class; inter-class imbalance compares the number of data between different classes. Inter-client imbalance represents different skewness of local data between clients. Based on these concepts, we propose FedBB, which consists of two main components: (1) Positive Negative Balanced (PNB) loss function addresses the inter-case and inter-class imbalances in local training, enhancing generalization on highly skewed local client datasets. It optimizes both multi-label and multi-class classifications by assigning higher weights to minority cases or classes. (2) Client Balanced Reweighting (CBR) reweights clients based on inter-client imbalance during model aggregation, giving greater weight to models trained on less skewed datasets. Various experiments on X-ray and natural image datasets demonstrate that FedBB outperforms other algorithms in both performance and efficiency. Additionally, it requires limited statistical information, which is beneficial for privacy protection. Through ablation studies, we proved that PNB loss and CBR independently contribute to performance. As FedBB aims to build a global model that accurately classifies all classes, it can serve as a baseline for the generic and personalized FL.","author":[{"family":"Chung","given":"Haengbok"},{"family":"Lee","given":"Jae"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2606.10250","URL":"https://doi.org/10.48550/arxiv.2606.10250","source":"datacite"},{"id":"doi:10.5281/zenodo.20408557","type":"article-journal","title":"Ransomware Attack Vectors, Detection Techniques, and Mitigation Strategies: A Comprehensive Survey","abstract":"Ransomware has turned to be one of the most severe and costly cybersecurity threats to organisations and individuals globally. This is a broad overview of ransomware looking at it in various dimensions: the way ransomware has evolved over the years since use as a simple screen-locking tool to a complex multi-stage attack with data exfiltration and extortion; the various types of attack vectors that ransomware attackers can use which include phishing attacks, remote desktop protocol attacks, supply chain attack and the broad use of machine learning to detect as well as sophisticated machine learning algorithms to detect ransomware; and overall mitigation strategies that are focused on prevention, detection, response and recovery. We digitise and systematically examine more than 50 research articles published between 2020 and 2025, and we make direct comparative reviews on the detection algorithms with respect to the measurements of accuracy, precision, recall, false positives, and computation overheads. We find in our analysis that ensemble machine learn-ing techniques can be used to detect an attack with a detection rate of above 99 percent with multi-layered defence schemes giving 85-90 percent accuracy in warding off successful attacks. We also recognise significant research opportunities such as the difficulty in detecting them in real time, the constraints of their datasets, or the necessity of cross platform security products. The significance of this survey to the field is in the way that it provides researchers and practitioners with a comprehensive picture of the ransomware threat environment and practical implications of creating defensive mechanisms in the next generation. As a conclusion, we overview a set of prospective research topics such as AI-controlled adaptive defence models, cryptographic systems based on blockchains, federated learning systems, and quantum resistant cryptographic systems.","author":[{"family":"Mule","given":"Manoj"},{"family":"Inamdar","given":"Nazma"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20408557","URL":"https://doi.org/10.5281/zenodo.20408557","source":"datacite"},{"id":"doi:10.5281/zenodo.20408556","type":"article-journal","title":"Ransomware Attack Vectors, Detection Techniques, and Mitigation Strategies: A Comprehensive Survey","abstract":"Ransomware has turned to be one of the most severe and costly cybersecurity threats to organisations and individuals globally. This is a broad overview of ransomware looking at it in various dimensions: the way ransomware has evolved over the years since use as a simple screen-locking tool to a complex multi-stage attack with data exfiltration and extortion; the various types of attack vectors that ransomware attackers can use which include phishing attacks, remote desktop protocol attacks, supply chain attack and the broad use of machine learning to detect as well as sophisticated machine learning algorithms to detect ransomware; and overall mitigation strategies that are focused on prevention, detection, response and recovery. We digitise and systematically examine more than 50 research articles published between 2020 and 2025, and we make direct comparative reviews on the detection algorithms with respect to the measurements of accuracy, precision, recall, false positives, and computation overheads. We find in our analysis that ensemble machine learn-ing techniques can be used to detect an attack with a detection rate of above 99 percent with multi-layered defence schemes giving 85-90 percent accuracy in warding off successful attacks. We also recognise significant research opportunities such as the difficulty in detecting them in real time, the constraints of their datasets, or the necessity of cross platform security products. The significance of this survey to the field is in the way that it provides researchers and practitioners with a comprehensive picture of the ransomware threat environment and practical implications of creating defensive mechanisms in the next generation. As a conclusion, we overview a set of prospective research topics such as AI-controlled adaptive defence models, cryptographic systems based on blockchains, federated learning systems, and quantum resistant cryptographic systems.","author":[{"family":"Mule","given":"Manoj"},{"family":"Inamdar","given":"Nazma"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20408556","URL":"https://doi.org/10.5281/zenodo.20408556","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.26116","type":"manuscript","title":"Sample Selection Using Multi-Task Autoencoders in Federated Learning with Non-IID Data","abstract":"Federated learning is a machine learning paradigm in which multiple devices collaboratively train a model under the supervision of a central server while ensuring data privacy. However, its performance is often hindered by redundant, malicious, or abnormal samples, leading to model degradation and inefficiency. To overcome these issues, we propose novel sample selection methods for image classification, employing a multitask autoencoder to estimate sample contributions through loss and feature analysis. Our approach incorporates unsupervised outlier detection, using one-class support vector machine (OCSVM), isolation forest (IF), and adaptive loss threshold (AT) methods managed by a central server to filter noisy samples on clients. We also propose a multi-class deep support vector data description (SVDD) loss controlled by a central server to enhance feature-based sample selection. We validate our methods on CIFAR10 and MNIST datasets across varying numbers of clients, non-IID distributions, and noise levels up to 40%. The results show significant accuracy improvements with loss-based sample selection, achieving gains of up to 7.02% on CIFAR10 with OCSVM and 1.83% on MNIST with AT. Additionally, our federated SVDD loss further improves feature-based sample selection, yielding accuracy gains of up to 0.99% on CIFAR10 with OCSVM. These results show the effectiveness of our methods in improving model accuracy across various client counts and noise conditions.","author":[{"family":"Ardıç","given":"Emre"},{"family":"Genç","given":"Yakup"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.26116","URL":"https://doi.org/10.48550/arxiv.2604.26116","source":"datacite"},{"id":"doi:10.5281/zenodo.19849384","type":"article-journal","title":"LongevityCommon","abstract":"Aging, codified in ICD-11 (code XT9T “Ageing-related,” 2018; code MG2A “Ageing-associated decline in intrinsic capacity,” 2025), requires integrative biomarker frameworks that extend beyond individual epigenetic clocks and wearable predictors. We present a hypothesis-stage integrative framework, LongevityCommon. All empirical estimates should be treated as exploratory (hypothesis‑generating), not confirmatory. Pre‑registered tests of an earlier univariate formulation of χ_Ze on the Cuban EEG, Dortmund Vital, and MPI‑LEMON cohorts yielded NULL results (documented in Ze EVIDENCE.md from 2026-04-22; meta‑analysis of Cuban + Dortmund: I² = 90.3 % — invalid; these results were retracted). The current multimodal version of χ_Ze is a post‑hoc reformulation and was not pre‑registered. All reported AUCs are exploratory hypothesis‑generating only, with explicit acknowledgment of p‑hacking risk (Ioannidis, 2005). LongevityCommon is a conceptual ecosystem of five components: (1) MCOA (Multi‑Counter Architecture of Organismal Aging) — a meta‑theory positing aging as a weighted sum of parallel counters: Ltissue(n,t)=iwi(tissue)fi(Di(n,t)). Axioms M1–M4 include an operational definition of falsifiability (M4) revised in v5 based on community‑standard validation thresholds: MCOA is considered falsified if on a pre‑registered cohort with N ≥ 2000 at α = 0.001 the partial r² for all‑cause mortality after controlling for chronological age and sex is full R² = 0.778), but full Sobol decomposition (S2 + ST) with 95 % CI and nested cross‑validation on real GTEx data (N = 948) revealed that the difference is not statistically significant (p = 0.12 after correction). CDATA remains an open, falsifiable hypothesis; final determination requires full decomposition on real data (Cell‑DT v4.0). (3) Ze Theory — a thermodynamic‑geometric formalism. The equation dZe/dt=−I(Z) is postulated as an ansatz, motivated by Burgholzer (2015) and Pearson et al. (2021); living systems are far‑from‑equilibrium, and a formal bridge between these physical‑clock systems and biological aging is absent. (4) BioSense — a wearable platform with the variational principle F=E−TS−Ipred. The theoretical fixed point v*=0.45631 at k=1 (sensitivity range v*[0.32;0.58] for k[0.5;2.0]). **Empirically tested via swept‑v* search on All‑of‑Us (N = 500): v*_optimal = 0.451 (95 % CI 0.443–0.459), consistent with the theoretical value.** (5) FCLC (Federated Clinical Learning Cooperative) — a federated learning infrastructure. ε_total ≈ 0.43 at (σ, q, T) = (1.5, 0.013, 5); RDP composition via subsampled Gaussian (Wang et al., 2019) combined with the Mironov (2017) framework. Threat model (explicit disclosure, v5): (a) FCLC central server — semi‑honest only (never sees raw data, only aggregated updates); (b) Byzantine‑robust aggregation (Krum up to 25 % malicious clients); (c) NOT secure against active server collusion; (d) NOT secure against malicious server deviating from the protocol. This is a blocker for GDPR Article 9 medical data until FCLC v14 (malicious‑secure migration planned for Q1 2027). The ecosystem provides a falsifiable (per updated M4) draft platform, but persistent blocking limitations remain: (i) all pilots are underpowered, not pre‑registered, and post‑hoc, which in light of Ioannidis (2005) gives a clear risk of false‑positive findings; (ii) FCLC is semi‑honest only — blocker for GDPR Art. 9; (iii) CDATA status is inconclusive, requiring full Sobol decomposition on real data; (iv) v* empirically tested and confirmed; (v) Ze→biology is an ansatz without formal derivation; (vi) bridge to CDATA (5 parameters on N = 196) is underpowered (Harrell rule violated) and moved to Supplementary; (vii) key publications (MCOA, Ze, BioSense) are not peer‑reviewed; (viii) EIC Pathfinder consortium formation deferred to Q1 2027 (0 signed EU LoIs as of 2026-04-21).","author":[{"family":"Tkemaladze","given":"Jaba"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19849384","URL":"https://doi.org/10.5281/zenodo.19849384","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.23426","type":"manuscript","title":"Enhanced Privacy and Communication Efficiency in Non-IID Federated Learning with Adaptive Quantization and Differential Privacy","abstract":"Federated learning (FL) is a distributed machine learning method where multiple devices collaboratively train a model under the management of a central server without sharing underlying data. One of the key challenges of FL is the communication bottleneck caused by variations in connection speed and bandwidth across devices. Therefore, it is essential to reduce the size of transmitted data during training. Additionally, there is a potential risk of exposing sensitive information through the model or gradient analysis during training. To address both privacy and communication efficiency, we combine differential privacy (DP) and adaptive quantization methods. We use Laplacian-based DP to preserve privacy, which is relatively underexplored in FL and offers tighter privacy guarantees than Gaussian-based DP. We propose a simple and efficient global bit-length scheduler using round-based cosine annealing, along with a client-based scheduler that dynamically adapts based on client contribution estimated through dataset entropy analysis. We evaluate our approach through extensive experiments on CIFAR10, MNIST, and medical imaging datasets, using non-IID data distributions across varying client counts, bit-length schedulers, and privacy budgets. The results show that our adaptive quantization methods reduce total communicated data by up to 52.64% for MNIST, 45.06% for CIFAR10, and 31% to 37% for medical imaging datasets compared to 32-bit float training while maintaining competitive model accuracy and ensuring robust privacy through differential privacy.","author":[{"family":"Ardıç","given":"Emre"},{"family":"Genç","given":"Yakup"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.23426","URL":"https://doi.org/10.48550/arxiv.2604.23426","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.23386","type":"manuscript","title":"A Taxonomy and Resolution Strategy for Client-Level Disagreements in Federated Learning","abstract":"Federated Learning (FL) typically assumes unconditional collaboration, a premise that overlooks the complexities of real-world, multi-stakeholder environments in which clients may need to exclude one another for strategic, regulatory, or competitive reasons. This paper addresses this gap, which we term 'client-level disagreements,' by first introducing a taxonomy of such scenarios. We then propose a robust, multi-track resolution strategy that guarantees strict client exclusion by creating and managing isolated model update paths ('tracks'), thereby preventing the cross-contamination and unfairness issues present in naive strategies. Through an empirical evaluation of our custom simulation system across 34 scenarios using the MNIST and N-CMAPSS datasets, we validate that our approach correctly handles permanent, temporal, and overlapping disagreement patterns. Our scalability analysis reveals the server-side resolution algorithm's overhead is negligible (&lt;1 ms per round) even under heavy load. The primary scalability constraint is the client-side training load from participating in multiple tracks, a cost that we show can be effectively mitigated by a submodel reuse strategy. This work presents a scalable and architecturally sound method for managing client-level disagreements, and enhances the practical applicability of FL in settings where policy compliance and strategic control are non-negotiable.","author":[{"family":"Rosendal","given":"Daan"},{"family":"Oprescu","given":"Ana"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.23386","URL":"https://doi.org/10.48550/arxiv.2604.23386","source":"datacite"},{"id":"doi:10.48550/arxiv.2604.12737","type":"manuscript","title":"Evaluating Differential Privacy Against Membership Inference in Federated Learning: Insights from the NIST Genomics Red Team Challenge","abstract":"While Federated Learning (FL) mitigates direct data exposure, the resulting trained models remain susceptible to membership inference attacks (MIAs). This paper presents an empirical evaluation of Differential Privacy (DP) as a defense mechanism against MIAs in FL, leveraging the environment of the 2025 NIST Genomics Privacy-Preserving Federated Learning (PPFL) Red Teaming Event. To improve inference accuracy, we propose a stacking attack strategy that ensembles seven black-box estimators to train a meta-classifier on prediction probabilities and cross-entropy losses. We evaluate this methodology against target models under three privacy configurations: an unprotected convolutional neural network (CNN, $ε=\\infty$), a low-privacy DP model ($ε=200$), and a high-privacy DP model ($ε=10$). The attack outperforms all baselines in the No DP and Low Privacy settings and, critically, maintains measurable membership leakage at $ε=200$ where a single-signal LiRA baseline collapses. Evaluated on an independent third-party benchmark, these results provide an empirical characterisation of how stacking-based inference degrades across calibrated DP tiers in FL.","author":[{"family":"Bertoli","given":"Gustavo"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2604.12737","URL":"https://doi.org/10.48550/arxiv.2604.12737","source":"datacite"},{"id":"doi:10.5281/zenodo.19546616","type":"article-journal","title":"Federated Clinical Learning Cooperative (FCLC)","abstract":"Background: Training clinical artificial intelligence (AI) models requires large, diverse datasets that are rarely available at a single institution due to privacy regulations (GDPR, HIPAA) and institutional risk aversion. Existing federated learning platforms lack validated differential privacy, Byzantine robustness, fair contribution attribution, and secure aggregation simultaneously. Objective: We present the Federated Clinical Learning Cooperative (FCLC), an open-source platform enabling multi-institutional clinical AI development without raw data leaving participating sites. We evaluate FCLC against centralized and non-private baselines across two independent clinical datasets, two model architectures, five heterogeneity conditions, and three adversarial scenarios. Methods: FCLC implements a five-layer privacy stack: (1) direct identifier removal; (2) quasi-identifier generalization; (3) k-anonymity (k ≥ 5); (4) Gaussian DP-SGD (ε = 2.0/round, δ = 10⁻⁵, Rényi accountant α = 4.0); and (5) cryptographic secure aggregation (SecAgg+) via the CommonHealth subproject. The choice of Krum for Byzantine robustness is supported by recent theoretical work establishing robustness guarantees for MultiKrum aggregation rules [Bareilles et al., 2026]. Aggregation uses FedProx (μ = 0.1) with Krum Byzantine robustness (f = 25%). Contribution attribution uses Monte Carlo Shapley (M = 150). Validation was performed on MIMIC-IV (N = 12,543, T2DM, 30-day readmission) and eICU-CRD (N = 8,420, sepsis, in-hospital mortality), under IID and non-IID Dirichlet (α ∈ {1.0, 0.5, 0.1, 0.01}) partitions, with logistic regression and multilayer perceptron (MLP) architectures, across 5 and 20 simulated nodes. Results: On MIMIC-IV, FCLC (MLP, DP) achieved AUC = 0.758 [95% CI: 0.739–0.777], compared to centralized oracle 0.789 and FedAvg 0.771 (IID). Under severe non-IID (Dirichlet α = 0.1, EMD = 0.31), FCLC preserved AUC = 0.748 [0.727–0.769] while FedAvg degraded to 0.694 (p_adj = 0.003). Membership inference AUC with DP was 0.52 ± 0.03 (indistinguishable from chance). At the clinically calibrated threshold, FCLC achieved sensitivity 0.671, specificity 0.724, with positive net benefit on Decision Curve Analysis. The privacy budget ε=2.0 per round aligns with state-of-the-art frameworks achieving 94–97% of non-private baseline performance at this budget [Vallabhaneni et al., 2026], and moderate privacy budgets (ε≈10) are clinically acceptable per recent reviews [npj Digital Medicine, 2026]. Conclusions: FCLC provides a validated, regulation-compliant infrastructure for federated clinical AI. The full cryptographic privacy stack, including SecAgg+, makes the platform suitable for mutual-distrust clinical deployments. Recent advances in f-differential privacy [Li et al., 2025] and Ripple Shapley for data attribution [Zeng et al., 2026] suggest promising directions for future enhancements.","author":[{"family":"Tkemaladze","given":"Jaba"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19546616","URL":"https://doi.org/10.5281/zenodo.19546616","source":"datacite"},{"id":"doi:10.5281/zenodo.19489424","type":"article-journal","title":"Federated Clinical Learning Cooperative (FCLC)","abstract":"Background: Training clinical artificial intelligence (AI) models requires large, diverse datasets that are rarely available at a single institution due to privacy regulations (GDPR, HIPAA) and institutional risk aversion. Existing federated learning platforms lack validated differential privacy, Byzantine robustness, fair contribution attribution, and secure aggregation simultaneously. Objective: We present the Federated Clinical Learning Cooperative (FCLC), an open-source platform enabling multi-institutional clinical AI development without raw data leaving participating sites. We evaluate FCLC against centralized and non-private baselines across two independent clinical datasets, two model architectures, five heterogeneity conditions, and three adversarial scenarios. Methods: FCLC implements a five-layer privacy stack: (1) direct identifier removal; (2) quasi-identifier generalization; (3) k-anonymity (k ≥ 5); (4) Gaussian DP-SGD (ε = 2.0/round, δ = 10⁻⁵, Rényi accountant α = 4.0); and (5) cryptographic secure aggregation (SecAgg+) via the CommonHealth subproject. The choice of Krum for Byzantine robustness is supported by recent theoretical work establishing robustness guarantees for MultiKrum aggregation rules [Bareilles et al., 2026]. Aggregation uses FedProx (μ = 0.1) with Krum Byzantine robustness (f = 25%). Contribution attribution uses Monte Carlo Shapley (M = 150). Validation was performed on MIMIC-IV (N = 12,543, T2DM, 30-day readmission) and eICU-CRD (N = 8,420, sepsis, in-hospital mortality), under IID and non-IID Dirichlet (α ∈ {1.0, 0.5, 0.1, 0.01}) partitions, with logistic regression and multilayer perceptron (MLP) architectures, across 5 and 20 simulated nodes. Results: On MIMIC-IV, FCLC (MLP, DP) achieved AUC = 0.758 [95% CI: 0.739–0.777], compared to centralized oracle 0.789 and FedAvg 0.771 (IID). Under severe non-IID (Dirichlet α = 0.1, EMD = 0.31), FCLC preserved AUC = 0.748 [0.727–0.769] while FedAvg degraded to 0.694 (p_adj = 0.003). Membership inference AUC with DP was 0.52 ± 0.03 (indistinguishable from chance). At the clinically calibrated threshold, FCLC achieved sensitivity 0.671, specificity 0.724, with positive net benefit on Decision Curve Analysis. The privacy budget ε=2.0 per round aligns with state-of-the-art frameworks achieving 94–97% of non-private baseline performance at this budget [Vallabhaneni et al., 2026], and moderate privacy budgets (ε≈10) are clinically acceptable per recent reviews [npj Digital Medicine, 2026]. Conclusions: FCLC provides a validated, regulation-compliant infrastructure for federated clinical AI. The full cryptographic privacy stack, including SecAgg+, makes the platform suitable for mutual-distrust clinical deployments. Recent advances in f-differential privacy [Li et al., 2025] and Ripple Shapley for data attribution [Zeng et al., 2026] suggest promising directions for future enhancements.","author":[{"family":"Tkemaladze","given":"Jaba"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19489424","URL":"https://doi.org/10.5281/zenodo.19489424","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.15393","type":"manuscript","title":"Federated Learning based on Self-Evolving Gaussian Clustering","abstract":"In this study, we present an Evolving Fuzzy System within the context of Federated Learning, which adapts dynamically with the addition of new clusters and therefore does not require the number of clusters to be selected apriori. Unlike traditional methods, Federated Learning allows models to be trained locally on clients' devices, sharing only the model parameters with a central server instead of the data. Our method, implemented using PyTorch, was tested on clustering and classification tasks. The results show that our approach outperforms established classification methods on several well-known UCI datasets. While computationally intensive due to overlap condition calculations, the proposed method demonstrates significant advantages in decentralized data processing.","author":[{"family":"Ožbot","given":"Miha"},{"family":"Škrjanc","given":"Igor"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.15393","URL":"https://doi.org/10.48550/arxiv.2508.15393","source":"datacite"},{"id":"doi:10.5281/zenodo.15831914","type":"article-journal","title":"AI-Enhanced Triage and Clinical Decision Support Tools in Primary Care","abstract":"Background: Artificial intelligence (AI)-powered symptom checkers, risk stratification engines and embedded clinical decision support systems (CDSS) are now migrating from research centres to the front line of healthcare. By synthesising multimodal patient data against evidence-based algorithms in seconds, these tools promise earlier detection of time-critical conditions, better scheduling of critical appointments and reduced cognitive load for primary care physicians, who currently manage four out of five presenting complaints. Growing demand for remote access and post-pandemic services is driving the need for digital triage (World Health Organization, 2021). Current evidence base: Real-world performance is encouraging but heterogeneous. A prospective study in 12 English practices found 95.8% agreement between an AI triage platform and clinicians for non-urgent cases, enabling an 18% reduction in telephone consultations. A cancer risk CDSS active in 1400 practices increased early cancer detection from 58.7% to 66.0% and accelerated referrals. Research showed that ChatGPT-assisted triage improved accuracy and halved documentation time. Conversely, a meta-analysis of 83 validation studies reported a pooled diagnostic accuracy of only 52.1%, highlighting marked variability between products and settings (Elhaddad & Hamam, 2024; Kaboudi et al., 2024). Governance and standards: The World Health Organization's 2024 guidance on large multimodal models requires transparency reports, equity impact assessments and ongoing post-implementation monitoring. The NICE Evidence Standards Framework (revised 2023) specifies escalating levels of clinical, technical and economic evidence before digital health technologies are adopted (World Health Organization, 2024). A consensus statement adds detailed recommendations on bias assessment, prospective audit and maintaining human-in-the-loop oversight for AI-CDSS (National Institute for Health and Care Excellence [NICE], 2023). Turkish context: In line with Türkiye's National Artificial Intelligence Strategy 2021-2025, pilot implementations in Istanbul and Ankara Family Health Centres have integrated cloud-based symptom checkers with the e-Nabız personal health record and the central medical appointment system (MHRS). Internal quality reports (2025) show a 12% reduction in inappropriate emergency referrals, a nine-minute reduction in median consultation length, and high patient satisfaction. Enablers include major standards for exchanging healthcare information electronically- HL7-FHIR interoperability layer, newly approved reimbursement codes for digital visits, and clinician-led curation of local rule sets; barriers remain compliance with the Data Protection Act (KVKK) and the lack of an AI-specific health regulation (Republic of Türkiye, Digital Transformation Office, 2021). Future agenda and conclusion: To translate promising pilots into routine care, we may need and expect: 1. pragmatic multi-centre trials designed for patient-centred outcomes, safety and workload redistribution 2. federated learning architectures that update models locally without exporting personal data 3. incorporation of social determinants, wearable-derived vital signs and multimodal inputs to refine risk scores 4. participatory design frameworks that engage frontline professionals, patients and policy makers in governance. AI-enabled triage and CDSS can strengthen accessible, equitable and high-quality primary care, provided their use remains evidence-based, transparently regulated and clinician-involved.","author":[{"family":"Arman","given":"Ikbal"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15831914","URL":"https://doi.org/10.5281/zenodo.15831914","source":"datacite"},{"id":"doi:10.5281/zenodo.15831913","type":"article-journal","title":"AI-Enhanced Triage and Clinical Decision Support Tools in Primary Care","abstract":"Background: Artificial intelligence (AI)-powered symptom checkers, risk stratification engines and embedded clinical decision support systems (CDSS) are now migrating from research centres to the front line of healthcare. By synthesising multimodal patient data against evidence-based algorithms in seconds, these tools promise earlier detection of time-critical conditions, better scheduling of critical appointments and reduced cognitive load for primary care physicians, who currently manage four out of five presenting complaints. Growing demand for remote access and post-pandemic services is driving the need for digital triage (World Health Organization, 2021). Current evidence base: Real-world performance is encouraging but heterogeneous. A prospective study in 12 English practices found 95.8% agreement between an AI triage platform and clinicians for non-urgent cases, enabling an 18% reduction in telephone consultations. A cancer risk CDSS active in 1400 practices increased early cancer detection from 58.7% to 66.0% and accelerated referrals. Research showed that ChatGPT-assisted triage improved accuracy and halved documentation time. Conversely, a meta-analysis of 83 validation studies reported a pooled diagnostic accuracy of only 52.1%, highlighting marked variability between products and settings (Elhaddad & Hamam, 2024; Kaboudi et al., 2024). Governance and standards: The World Health Organization's 2024 guidance on large multimodal models requires transparency reports, equity impact assessments and ongoing post-implementation monitoring. The NICE Evidence Standards Framework (revised 2023) specifies escalating levels of clinical, technical and economic evidence before digital health technologies are adopted (World Health Organization, 2024). A consensus statement adds detailed recommendations on bias assessment, prospective audit and maintaining human-in-the-loop oversight for AI-CDSS (National Institute for Health and Care Excellence [NICE], 2023). Turkish context: In line with Türkiye's National Artificial Intelligence Strategy 2021-2025, pilot implementations in Istanbul and Ankara Family Health Centres have integrated cloud-based symptom checkers with the e-Nabız personal health record and the central medical appointment system (MHRS). Internal quality reports (2025) show a 12% reduction in inappropriate emergency referrals, a nine-minute reduction in median consultation length, and high patient satisfaction. Enablers include major standards for exchanging healthcare information electronically- HL7-FHIR interoperability layer, newly approved reimbursement codes for digital visits, and clinician-led curation of local rule sets; barriers remain compliance with the Data Protection Act (KVKK) and the lack of an AI-specific health regulation (Republic of Türkiye, Digital Transformation Office, 2021). Future agenda and conclusion: To translate promising pilots into routine care, we may need and expect: 1. pragmatic multi-centre trials designed for patient-centred outcomes, safety and workload redistribution 2. federated learning architectures that update models locally without exporting personal data 3. incorporation of social determinants, wearable-derived vital signs and multimodal inputs to refine risk scores 4. participatory design frameworks that engage frontline professionals, patients and policy makers in governance. AI-enabled triage and CDSS can strengthen accessible, equitable and high-quality primary care, provided their use remains evidence-based, transparently regulated and clinician-involved.","author":[{"family":"Arman","given":"Ikbal"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15831913","URL":"https://doi.org/10.5281/zenodo.15831913","source":"datacite"},{"id":"doi:10.6084/m9.figshare.29877074.v1","type":"article-journal","title":"AI-Powered Visualization is Transforming Modern Healthcare","abstract":"Healthcare is being transformed by AI-driven visualization, which transforms complex data into useful insights. This paper synthesizes advancements in AI visualization tools—spanning medical imaging, electronic health records (EHR), genomics, and public health—and evaluates their impact on diagnostics, treatment personalization, and operational efficiency. Convolutional neural networks (CNNs) for image segmentation, generative adversarial networks (GANs) for the generation of synthetic data, and interactive dashboards for real-time analytics are some of the technologies that we highlight. Integrity barriers, algorithmic bias, and data privacy concerns are all critically examined. A systematic review of more than 120 studies conducted between 2018 and 2024 shows that clinical workflow time is cut by 30% and diagnostic accuracy is improved by 40% on average. Explainable artificial intelligence (XAI) and federated learning are emphasized in the study's ethical frameworks and future directions. This study demonstrates that AI visualization plays a crucial role in value-based care and precision medicine.","author":[{"family":"Rahman Nazil","given":"Ashikur"}],"issued":{"date-parts":[[2025]]},"DOI":"10.6084/m9.figshare.29877074.v1","URL":"https://doi.org/10.6084/m9.figshare.29877074.v1","source":"datacite"},{"id":"doi:10.6084/m9.figshare.29876651.v1","type":"article-journal","title":"AI-Powered Visualization is Transforming Modern Healthcare","abstract":"Healthcare is being transformed by AI-driven visualization, which transforms complex data into useful insights. This paper synthesizes advancements in AI visualization tools—spanning medical imaging, electronic health records (EHR), genomics, and public health—and evaluates their impact on diagnostics, treatment personalization, and operational efficiency. Convolutional neural networks (CNNs) for image segmentation, generative adversarial networks (GANs) for the generation of synthetic data, and interactive dashboards for real-time analytics are some of the technologies that we highlight. Integrity barriers, algorithmic bias, and data privacy concerns are all critically examined. A systematic review of more than 120 studies conducted between 2018 and 2024 shows that clinical workflow time is cut by 30% and diagnostic accuracy is improved by 40% on average. Explainable artificial intelligence (XAI) and federated learning are emphasized in the study's ethical frameworks and future directions. This study demonstrates that AI visualization plays a crucial role in value-based care and precision medicine.","author":[{"family":"Rahman Nazil","given":"Ashikur"}],"issued":{"date-parts":[[2025]]},"DOI":"10.6084/m9.figshare.29876651.v1","URL":"https://doi.org/10.6084/m9.figshare.29876651.v1","source":"datacite"},{"id":"doi:10.48550/arxiv.2508.00967","type":"manuscript","title":"Cooperative Perception: A Resource-Efficient Framework for Multi-Drone 3D Scene Reconstruction Using Federated Diffusion and NeRF","abstract":"The proposal introduces an innovative drone swarm perception system that aims to solve problems related to computational limitations and low-bandwidth communication, and real-time scene reconstruction. The framework enables efficient multi-agent 3D/4D scene synthesis through federated learning of shared diffusion model and YOLOv12 lightweight semantic extraction and local NeRF updates while maintaining privacy and scalability. The framework redesigns generative diffusion models for joint scene reconstruction, and improves cooperative scene understanding, while adding semantic-aware compression protocols. The approach can be validated through simulations and potential real-world deployment on drone testbeds, positioning it as a disruptive advancement in multi-agent AI for autonomous systems.","author":[{"family":"Pourmandi","given":"Massoud"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2508.00967","URL":"https://doi.org/10.48550/arxiv.2508.00967","source":"datacite"},{"id":"doi:10.5281/zenodo.15310785","type":"article-journal","title":"Privacy and Ethical Concerns in AI-Powered Data Processing","abstract":"Artificial intelligence is increasingly being used to process large datasets. This introduces serious privacy and security risks (Paul, 2024). In many AI systems, sensitive personal data are collected and analyzed, so leaks or attacks can expose private information.In this paper, we review common types of privacy attacks such as model inversion, membership inference, and data reconstruction, and analyze them through ethical lenses. Next, we examine current mitigation techniques like differential privacy and federated learning, as well as their ethical implications. Finally, we discuss future directions for ethically respecting privacy in AI systems. Throughout, we emphasize how some ethical frameworks apply to the challenges and solutions in AI privacy.","author":[{"family":"B Ziade","given":"Tony"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15310785","URL":"https://doi.org/10.5281/zenodo.15310785","source":"datacite"},{"id":"doi:10.5281/zenodo.15310784","type":"article-journal","title":"Privacy and Ethical Concerns in AI-Powered Data Processing","abstract":"Artificial intelligence is increasingly being used to process large datasets. This introduces serious privacy and security risks (Paul, 2024). In many AI systems, sensitive personal data are collected and analyzed, so leaks or attacks can expose private information.In this paper, we review common types of privacy attacks such as model inversion, membership inference, and data reconstruction, and analyze them through ethical lenses. Next, we examine current mitigation techniques like differential privacy and federated learning, as well as their ethical implications. Finally, we discuss future directions for ethically respecting privacy in AI systems. Throughout, we emphasize how some ethical frameworks apply to the challenges and solutions in AI privacy.","author":[{"family":"B Ziade","given":"Tony"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.15310784","URL":"https://doi.org/10.5281/zenodo.15310784","source":"datacite"},{"id":"doi:10.17605/osf.io/d27xz","type":"article-journal","title":"AI stridor analysis","abstract":"Objective: This review examines artificial intelligence (AI) applications in analyzing respiratory sounds, specifically stridor, to address limitations in traditional, operator-dependent diagnostic methods. Data Sources: A structured search across PubMed, Scopus, and Ovid was conducted, focusing on studies from 2010 to 2024 relevant to AI and stridor detection. Review Methods: The review synthesizes findings from studies employing machine learning (ML) models like Artificial Neural Networks (ANN), Convolutional Neural Networks (CNN), and Support Vector Machines (SVM) to enhance stridor diagnosis by analyzing respiratory sound patterns. Results: The reviewed studies demonstrate high diagnostic accuracy, with models such as Quadratic Discriminant Analysis (QDA) and Personalized Federated Learning with Self-Distillation (PFL-SD) achieving accuracies up to 100% and 96.1%, respectively. AI techniques showed potential to classify stridor types, identify anatomical locations of obstructions, and guide clinical decision-making, particularly in pediatric cases requiring prompt intervention. Conclusion: AI-driven respiratory sound analysis offers promising advancements in stridor diagnosis through improved accuracy and early detection, though limitations in data standardization and underexplored pediatric applications present challenges and opportunities. Future research on larger, standardized datasets could expand the potential of AI tools in this critical clinical setting.","author":[{"family":"Tartaglia","given":"Francesco"},{"family":"Motisi","given":"Annagiulia"}],"issued":{"date-parts":[[2025]]},"DOI":"10.17605/osf.io/d27xz","URL":"https://doi.org/10.17605/osf.io/d27xz","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.00402","type":"manuscript","title":"Curriculum Guided Personalized Subgraph Federated Learning","abstract":"Subgraph Federated Learning (FL) aims to train Graph Neural Networks (GNNs) across distributed private subgraphs, but it suffers from severe data heterogeneity. To mitigate data heterogeneity, weighted model aggregation personalizes each local GNN by assigning larger weights to parameters from clients with similar subgraph characteristics inferred from their current model states. However, the sparse and biased subgraphs often trigger rapid overfitting, causing the estimated client similarity matrix to stagnate or even collapse. As a result, aggregation loses effectiveness as clients reinforce their own biases instead of exploiting diverse knowledge otherwise available. To this end, we propose a novel personalized subgraph FL framework called Curriculum guided personalized sUbgraph Federated Learning (CUFL). On the client side, CUFL adopts Curriculum Learning (CL) that adaptively selects edges for training according to their reconstruction scores, exposing each GNN first to easier, generic cross-client substructures and only later to harder, client-specific ones. This paced exposure prevents early overfitting to biased patterns and enables gradual personalization. By regulating personalization, the curriculum also reshapes server aggregation from exchanging generic knowledge to propagating client-specific knowledge. Further, CUFL improves weighted aggregation by estimating client similarity using fine-grained structural indicators reconstructed on a random reference graph. Extensive experiments on six benchmark datasets confirm that CUFL achieves superior performance compared to relevant baselines. Code is available at https://github.com/Kang-Min-Ku/CUFL.git.","author":[{"family":"Kang","given":"Minku"},{"family":"Park","given":"Hogun"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.00402","URL":"https://doi.org/10.48550/arxiv.2509.00402","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.21491","type":"manuscript","title":"Benchmarking Catastrophic Forgetting Mitigation Methods in Federated Time Series Forecasting","abstract":"Catastrophic forgetting (CF) poses a persistent challenge in continual learning (CL), especially within federated learning (FL) environments characterized by non-i.i.d. time series data. While existing research has largely focused on classification tasks in vision domains, the regression-based forecasting setting prevalent in IoT and edge applications remains underexplored. In this paper, we present the first benchmarking framework tailored to investigate CF in federated continual time series forecasting. Using the Beijing Multi-site Air Quality dataset across 12 decentralized clients, we systematically evaluate several CF mitigation strategies, including Replay, Elastic Weight Consolidation, Learning without Forgetting, and Synaptic Intelligence. Key contributions include: (i) introducing a new benchmark for CF in time series FL, (ii) conducting a comprehensive comparative analysis of state-of-the-art methods, and (iii) releasing a reproducible open-source framework. This work provides essential tools and insights for advancing continual learning in federated time-series forecasting systems.","author":[{"family":"Hallak","given":"Khaled"},{"family":"Kem","given":"Oudom"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.21491","URL":"https://doi.org/10.48550/arxiv.2510.21491","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.20102","type":"manuscript","title":"Autonomous Cyber Resilience via a Co-Evolutionary Arms Race within a Fortified Digital Twin Sandbox","abstract":"The convergence of Information Technology and Operational Technology has exposed Industrial Control Systems to adaptive, intelligent adversaries that render static defenses obsolete. This paper introduces the Adversarial Resilience Co-evolution (ARC) framework, addressing the \"Trinity of Trust\" comprising model fidelity, data integrity, and analytical resilience. ARC establishes a co-evolutionary arms race within a Fortified Secure Digital Twin (F-SCDT), where a Deep Reinforcement Learning \"Red Agent\" autonomously discovers attack paths while an ensemble-based \"Blue Agent\" is continuously hardened against these threats. Experimental validation on the Tennessee Eastman Process (TEP) and Secure Water Treatment (SWaT) testbeds demonstrates superior performance in detecting novel attacks, with F1-scores improving from 0.65 to 0.89 and detection latency reduced from over 1200 seconds to 210 seconds. A comprehensive ablation study reveals that the co-evolutionary process itself contributes a 27% performance improvement. By integrating Explainable AI and proposing a Federated ARC architecture, this work presents a necessary paradigm shift toward dynamic, self-improving security for critical infrastructure.","author":[{"family":"Malikussaid"},{"family":"Sutiyo"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.20102","URL":"https://doi.org/10.48550/arxiv.2506.20102","source":"datacite"},{"id":"doi:10.5281/zenodo.18626878","type":"article-journal","title":"SDKP Framework — Complete System Architecture","abstract":"SDKP Framework — Complete System Architecture Author: Donald Paul Smith (FatherTime) Affiliation: Independent Researcher Abstract This work presents the SDKP framework, a unified symbolic–physical model connecting size, density, and rotational velocity to temporal behavior. The framework integrates Shape–Dimension–Number (SD&N) structural mapping, Earth Orbital Speed (EOS) dynamics, Velocity–Frequency–Energy (VFE1) scaling, Quantum Computerization Consciousness (QCC), Loop Learning for Artificial Life (LLAL), and Kapnack symbolic compression. The model proposes that physical behavior emerges from structured relationships between geometry, density distribution, and rotational dynamics. The framework provides a mathematical structure describing system evolution and predicts measurable deviations from conventional physical models, including small variations in orbital motion and symbolic compression effects. A complete system architecture, mathematical structure, and testable predictions are presented to support experimental evaluation and future development. 1. Introduction 1.1 Background Modern physics describes motion, gravity, and quantum behavior using separate theoretical structures. A unified representation connecting structural geometry, density relationships, and temporal evolution remains an open problem. 1.2 Motivation The SDKP framework proposes that time evolution and physical behavior emerge from relationships between size, density, and rotational velocity. This provides a structural interpretation of physical dynamics and information flow. 1.3 Contributions This work introduces: a unified symbolic–physical framework a structural mapping between geometry and dynamics a system architecture connecting multiple model components testable predictions for experimental validation 2. Core Definitions 2.1 SDKP (Size–Density–Rotation–Time) SDKP describes temporal behavior as a function of system size, density distribution, and rotational velocity. It defines relationships between structural configuration and dynamic evolution. Variables include: size parameter (S) density parameter (D) rotational velocity (R) temporal response (T) 2.2 SD&N (Shape–Dimension–Number) SD&N provides structural classification of systems based on geometric form, dimensional structure, and numerical relationships. It defines how system structure influences dynamic behavior. 2.3 EOS (Earth Orbital Speed Model) EOS describes orbital motion using rotational and density-based relationships. The model predicts small deviations from classical orbital calculations. 2.4 QCC (Quantum Computerization Consciousness) QCC models information processing and quantum-level interactions through structured symbolic computation. 2.5 VFE1 (Velocity–Frequency–Energy Scaling) VFE1 defines relationships between motion, oscillation, and energy distribution across system states. 2.6 LLAL (Loop Learning for Artificial Life) LLAL describes feedback-based adaptive system evolution using recursive learning structures. 2.7 Kapnack Symbolic Compression Kapnack defines rules for compressing complex system behavior into structured symbolic representations. 3. Mathematical Structure (Overview) The framework defines system behavior through relationships among structural variables and rotational dynamics. Core elements include: system size parameter density distribution functions rotational velocity fields temporal response functions symbolic compression operators System evolution is determined by interactions between these quantities. 4. System Architecture The framework operates through interacting components: SDKP defines core dynamics SD&N defines structural configuration VFE1 defines energy scaling EOS describes orbital behavior QCC models information processing LLAL provides feedback adaptation Kapnack enables symbolic compression These components form a unified dynamic system. 5. Physical Interpretation The framework provides structural interpretations of: motion as rotational in","author":[{"family":"Smith","given":"Donald"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18626878","URL":"https://doi.org/10.5281/zenodo.18626878","source":"datacite"},{"id":"doi:10.5281/zenodo.18626877","type":"article-journal","title":"SDKP Framework — Complete System Architecture","abstract":"SDKP Framework — Complete System Architecture Author: Donald Paul Smith (FatherTime) Affiliation: Independent Researcher Abstract This work presents the SDKP framework, a unified symbolic–physical model connecting size, density, and rotational velocity to temporal behavior. The framework integrates Shape–Dimension–Number (SD&N) structural mapping, Earth Orbital Speed (EOS) dynamics, Velocity–Frequency–Energy (VFE1) scaling, Quantum Computerization Consciousness (QCC), Loop Learning for Artificial Life (LLAL), and Kapnack symbolic compression. The model proposes that physical behavior emerges from structured relationships between geometry, density distribution, and rotational dynamics. The framework provides a mathematical structure describing system evolution and predicts measurable deviations from conventional physical models, including small variations in orbital motion and symbolic compression effects. A complete system architecture, mathematical structure, and testable predictions are presented to support experimental evaluation and future development. 1. Introduction 1.1 Background Modern physics describes motion, gravity, and quantum behavior using separate theoretical structures. A unified representation connecting structural geometry, density relationships, and temporal evolution remains an open problem. 1.2 Motivation The SDKP framework proposes that time evolution and physical behavior emerge from relationships between size, density, and rotational velocity. This provides a structural interpretation of physical dynamics and information flow. 1.3 Contributions This work introduces: a unified symbolic–physical framework a structural mapping between geometry and dynamics a system architecture connecting multiple model components testable predictions for experimental validation 2. Core Definitions 2.1 SDKP (Size–Density–Rotation–Time) SDKP describes temporal behavior as a function of system size, density distribution, and rotational velocity. It defines relationships between structural configuration and dynamic evolution. Variables include: size parameter (S) density parameter (D) rotational velocity (R) temporal response (T) 2.2 SD&N (Shape–Dimension–Number) SD&N provides structural classification of systems based on geometric form, dimensional structure, and numerical relationships. It defines how system structure influences dynamic behavior. 2.3 EOS (Earth Orbital Speed Model) EOS describes orbital motion using rotational and density-based relationships. The model predicts small deviations from classical orbital calculations. 2.4 QCC (Quantum Computerization Consciousness) QCC models information processing and quantum-level interactions through structured symbolic computation. 2.5 VFE1 (Velocity–Frequency–Energy Scaling) VFE1 defines relationships between motion, oscillation, and energy distribution across system states. 2.6 LLAL (Loop Learning for Artificial Life) LLAL describes feedback-based adaptive system evolution using recursive learning structures. 2.7 Kapnack Symbolic Compression Kapnack defines rules for compressing complex system behavior into structured symbolic representations. 3. Mathematical Structure (Overview) The framework defines system behavior through relationships among structural variables and rotational dynamics. Core elements include: system size parameter density distribution functions rotational velocity fields temporal response functions symbolic compression operators System evolution is determined by interactions between these quantities. 4. System Architecture The framework operates through interacting components: SDKP defines core dynamics SD&N defines structural configuration VFE1 defines energy scaling EOS describes orbital behavior QCC models information processing LLAL provides feedback adaptation Kapnack enables symbolic compression These components form a unified dynamic system. 5. Physical Interpretation The framework provides structural interpretations of: motion as rotational in","author":[{"family":"Smith","given":"Donald"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18626877","URL":"https://doi.org/10.5281/zenodo.18626877","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.14275","type":"manuscript","title":"FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health","abstract":"Privacy-preserving adaptation of Large Language Models (LLMs) in sensitive domains (e.g., mental health) requires balancing strict confidentiality with model utility and safety. We propose FedMentor, a federated fine-tuning framework that integrates Low-Rank Adaptation (LoRA) and domain-aware Differential Privacy (DP) to meet per-domain privacy budgets while maintaining performance. Each client (domain) applies a custom DP noise scale proportional to its data sensitivity, and the server adaptively reduces noise when utility falls below a threshold. In experiments on three mental health datasets, we show that FedMentor improves safety over standard Federated Learning (FL) without privacy, raising safe output rates by up to three points and lowering toxicity, while maintaining utility (BERTScore F1 and ROUGE-L) within 0.5% of the non-private baseline and close to the centralized upper bound. The framework scales to backbones with up to 1.7B parameters on single-GPU clients, requiring &lt; 173 MB of communication per-round. FedMentor demonstrates a practical approach to privately fine-tune LLMs for safer deployments in healthcare and other sensitive fields.","author":[{"family":"Sarwar","given":"Nobin"},{"family":"Dipta","given":"Shubhashis"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.14275","URL":"https://doi.org/10.48550/arxiv.2509.14275","source":"datacite"},{"id":"doi:10.48550/arxiv.2504.18007","type":"manuscript","title":"Differential Privacy-Driven Framework for Enhancing Heart Disease Prediction","abstract":"With the rapid digitalization of healthcare systems, there has been a substantial increase in the generation and sharing of private health data. Safeguarding patient information is essential for maintaining consumer trust and ensuring compliance with legal data protection regulations. Machine learning is critical in healthcare, supporting personalized treatment, early disease detection, predictive analytics, image interpretation, drug discovery, efficient operations, and patient monitoring. It enhances decision-making, accelerates research, reduces errors, and improves patient outcomes. In this paper, we utilize machine learning methodologies, including differential privacy and federated learning, to develop privacy-preserving models that enable healthcare stakeholders to extract insights without compromising individual privacy. Differential privacy introduces noise to data to guarantee statistical privacy, while federated learning enables collaborative model training across decentralized datasets. We explore applying these technologies to Heart Disease Data, demonstrating how they preserve privacy while delivering valuable insights and comprehensive analysis. Our results show that using a federated learning model with differential privacy achieved a test accuracy of 85%, ensuring patient data remained secure and private throughout the process.","author":[{"family":"Otoum","given":"Yazan"},{"family":"Nayak","given":"Amiya"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2504.18007","URL":"https://doi.org/10.48550/arxiv.2504.18007","source":"datacite"},{"id":"doi:10.5281/zenodo.18381453","type":"article-journal","title":"Y.I.N.-MEMORIA: A Comprehensive Privacy-Preserving Architecture for AI Conversation Management with Cryptographic Ordering Enforcement, Zero-Knowledge Governance, and Quantified Attack Defense","abstract":"We present Y.I.N.-MEMORIA, a comprehensive privacy-preserving architecture addressing fundamental vulnerabilities in AI conversation systems across all platforms, including large language model interfaces, enterprise AI assistants, domain-specific chatbots, and agentic AI systems. The system implements mandatory cryptographic ordering enforcement (DP → ZK → BLINDING → HE, or functional equivalents), mathematically proven unique among 24 permutations, achieving 99.37% accuracy for valid authorizations versus 50.7% for invalid attempts (t = 147.3, p 94% detection rates. Comparative Analysis: Table comparing against 8 major systems (Federated Learning, CrypTen, TF Privacy, Opacus, PySyft, Microsoft SEAL, Zcash). Y.I.N.-MEMORIA demonstrated as only system providing mandatory DP enforcement, ZK verification for AI governance, 340× timing resistance, 99.7% Shadow AI detection, and complete lifecycle coverage. Legal Protection: Doctrine of equivalents coverage (Warner-Jenkinson precedent); Willful infringement notice (Halo Electronics, 3× damages); Comprehensive functional equivalents (12 categories); Minimum performance thresholds excluding weak implementations. Reproducibility Commitment: Complete reference implementation under open-source license; Experimental datasets via Zenodo; Cryptographic test vectors for independent verification; Performance benchmarks across all platforms. Scholarly Depth: 38 peer-reviewed citations (65% increase); Comprehensive related work analysis; Explicit limitations and future research directions; Historical non-obviousness evidence. THREE-PHASE AI LIFECYCLE COVERAGE:Y.I.N.-MEMORIA completes the Y.I.N. Architecture's three-phase AI lifecycle: Training (Y.I.N.-LLM, USPTO 63/941,283), Generation (Article 50 Compliance Engine, USPTO 63/957,571), and Usage (Y.I.N.-MEMORIA, USPTO 63/967,805). The Y.I.N. CERTIFY verification layer spans all three phases. Together, these components provide 643 total claims covering every stage where privacy vulnerabilities can emerge in AI systems. EXPERIMENTAL VALIDATION:85-95% bandwidth reduction, 97% conflict resolution, and compliance scores of 94.7-97.3% for GDPR, HIPAA, DORA, EU AI Act, Singapore MGF for Agentic AI, ISO/IEC 42001, CCPA, and NIS2 Directive. IMPACT METRICS:This architecture prevents Shadow AI breaches costing $4.63M average (20% of all data breaches according to IBM's 2025 Cost of a Data Breach Report), addresses the 20M ChatGPT conversation log discovery precedent (NYT v. OpenAI, January 2026), and satisfies Singapore's Model AI Governance Framework for Agentic AI—the world's first comprehensive government framework for autonomous agents published January 22, 2026 (4 days prior to this work). DEFENSIVE PRIOR ART:This work establishes comprehensive prior art corresponding to USPTO Provisional Application 63/967,805 (438 claims filed January 25, 2026), part of the Y.I.N. Architecture Portfolio (22 applications, 1,360+ total claims). Includes explicit functional equivalents coverage, doctrine of equivalents, and willful infringement notice enabling enhanced damages up to 3× under Halo Electronics precedent. Patent Reference: USPTO Application 63/967,805 (Y.I.N.-MEMORIA) License: CC BY-NC-ND 4.0Corresponding Author: ilyesmazari@hotmail.comVersion: 1.0Publication Date: January 26, 2026","author":[{"family":"Mazari","given":"Ilyes"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18381453","URL":"https://doi.org/10.5281/zenodo.18381453","source":"datacite"},{"id":"doi:10.5281/zenodo.18381452","type":"article-journal","title":"Y.I.N.-MEMORIA: A Comprehensive Privacy-Preserving Architecture for AI Conversation Management with Cryptographic Ordering Enforcement, Zero-Knowledge Governance, and Quantified Attack Defense","abstract":"We present Y.I.N.-MEMORIA, a comprehensive privacy-preserving architecture addressing fundamental vulnerabilities in AI conversation systems across all platforms, including large language model interfaces, enterprise AI assistants, domain-specific chatbots, and agentic AI systems. The system implements mandatory cryptographic ordering enforcement (DP → ZK → BLINDING → HE, or functional equivalents), mathematically proven unique among 24 permutations, achieving 99.37% accuracy for valid authorizations versus 50.7% for invalid attempts (t = 147.3, p 94% detection rates. Comparative Analysis: Table comparing against 8 major systems (Federated Learning, CrypTen, TF Privacy, Opacus, PySyft, Microsoft SEAL, Zcash). Y.I.N.-MEMORIA demonstrated as only system providing mandatory DP enforcement, ZK verification for AI governance, 340× timing resistance, 99.7% Shadow AI detection, and complete lifecycle coverage. Legal Protection: Doctrine of equivalents coverage (Warner-Jenkinson precedent); Willful infringement notice (Halo Electronics, 3× damages); Comprehensive functional equivalents (12 categories); Minimum performance thresholds excluding weak implementations. Reproducibility Commitment: Complete reference implementation under open-source license; Experimental datasets via Zenodo; Cryptographic test vectors for independent verification; Performance benchmarks across all platforms. Scholarly Depth: 38 peer-reviewed citations (65% increase); Comprehensive related work analysis; Explicit limitations and future research directions; Historical non-obviousness evidence. THREE-PHASE AI LIFECYCLE COVERAGE:Y.I.N.-MEMORIA completes the Y.I.N. Architecture's three-phase AI lifecycle: Training (Y.I.N.-LLM, USPTO 63/941,283), Generation (Article 50 Compliance Engine, USPTO 63/957,571), and Usage (Y.I.N.-MEMORIA, USPTO 63/967,805). The Y.I.N. CERTIFY verification layer spans all three phases. Together, these components provide 643 total claims covering every stage where privacy vulnerabilities can emerge in AI systems. EXPERIMENTAL VALIDATION:85-95% bandwidth reduction, 97% conflict resolution, and compliance scores of 94.7-97.3% for GDPR, HIPAA, DORA, EU AI Act, Singapore MGF for Agentic AI, ISO/IEC 42001, CCPA, and NIS2 Directive. IMPACT METRICS:This architecture prevents Shadow AI breaches costing $4.63M average (20% of all data breaches according to IBM's 2025 Cost of a Data Breach Report), addresses the 20M ChatGPT conversation log discovery precedent (NYT v. OpenAI, January 2026), and satisfies Singapore's Model AI Governance Framework for Agentic AI—the world's first comprehensive government framework for autonomous agents published January 22, 2026 (4 days prior to this work). DEFENSIVE PRIOR ART:This work establishes comprehensive prior art corresponding to USPTO Provisional Application 63/967,805 (438 claims filed January 25, 2026), part of the Y.I.N. Architecture Portfolio (22 applications, 1,360+ total claims). Includes explicit functional equivalents coverage, doctrine of equivalents, and willful infringement notice enabling enhanced damages up to 3× under Halo Electronics precedent. Patent Reference: USPTO Application 63/967,805 (Y.I.N.-MEMORIA) License: CC BY-NC-ND 4.0Corresponding Author: ilyesmazari@hotmail.comVersion: 1.0Publication Date: January 26, 2026","author":[{"family":"Mazari","given":"Ilyes"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18381452","URL":"https://doi.org/10.5281/zenodo.18381452","source":"datacite"},{"id":"doi:10.5281/zenodo.18080743","type":"article-journal","title":"Secure Data Integration Frameworks for Omni-channel Healthcare Marketing Systems","abstract":"The rapid digitalisation and increasing interconnection of healthcare marketing ecosystems require secure and interoperable data-integration frameworks. Omnichannel systems, a combination of clinical, behavioural and marketing data, are becoming increasingly important to healthcare organisations as they aim to personalise contact and stay under regulatory requirements. This research is a critical synthesis of twenty-seven peer-reviewed articles (2020-2025) in IEEE, Elsevier, and Springer Nature, as well as the most popular sources in the field of machine learning, privacy-preserving machine learning (PPML) in healthcare marketing. In this review, the qualitative meta-analytic approach is used to study technological architectures, encryption, and federated-learning methods and omni-experience platform designs. The results show that blockchain-based interoperability and AI-based analytics can enhance trust, auditability, and personalization to a considerable extent, whereas federated and split-learn systems reduce privacy risks in distributed marketing data. Hybrid cloud-edge infrastructure-based omnichannel experience platforms increase the real-time decision-making and campaign flexibility, but they encounter governance and integration issues. The paper suggests a theoretical framework of the association of secure data exchange, privacy preservation, and efficiency of omnichannel engagement. The analysis focuses on the scholarship of healthcare informatics as well as strategic marketing by emphasizing the ability of secure integration technologies to maintain compliance (HIPAA / GDPR) and improve the level of patient-centric marketing.","author":[{"family":"Tayal","given":"Chitiz"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.18080743","URL":"https://doi.org/10.5281/zenodo.18080743","source":"datacite"},{"id":"doi:10.5281/zenodo.18080742","type":"article-journal","title":"Secure Data Integration Frameworks for Omni-channel Healthcare Marketing Systems","abstract":"The rapid digitalisation and increasing interconnection of healthcare marketing ecosystems require secure and interoperable data-integration frameworks. Omnichannel systems, a combination of clinical, behavioural and marketing data, are becoming increasingly important to healthcare organisations as they aim to personalise contact and stay under regulatory requirements. This research is a critical synthesis of twenty-seven peer-reviewed articles (2020-2025) in IEEE, Elsevier, and Springer Nature, as well as the most popular sources in the field of machine learning, privacy-preserving machine learning (PPML) in healthcare marketing. In this review, the qualitative meta-analytic approach is used to study technological architectures, encryption, and federated-learning methods and omni-experience platform designs. The results show that blockchain-based interoperability and AI-based analytics can enhance trust, auditability, and personalization to a considerable extent, whereas federated and split-learn systems reduce privacy risks in distributed marketing data. Hybrid cloud-edge infrastructure-based omnichannel experience platforms increase the real-time decision-making and campaign flexibility, but they encounter governance and integration issues. The paper suggests a theoretical framework of the association of secure data exchange, privacy preservation, and efficiency of omnichannel engagement. The analysis focuses on the scholarship of healthcare informatics as well as strategic marketing by emphasizing the ability of secure integration technologies to maintain compliance (HIPAA / GDPR) and improve the level of patient-centric marketing.","author":[{"family":"Tayal","given":"Chitiz"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.18080742","URL":"https://doi.org/10.5281/zenodo.18080742","source":"datacite"},{"id":"doi:10.5281/zenodo.22165672","type":"article-journal","title":"A Quantum-Resilient Federated Learning Framework for Secure Smart Grid","abstract":"Wide-area monitoring systems built on phasor measurement units (PMUs) underpin real-time stability assessment in transmission grids, yet their measurement and communication paths present an attack surface that conventional intrusion detection addresses only partially. Three constraints compound the problem: transmission operators are commercially and legally restricted from pooling raw measurements; the public-key cryptography protecting inter-operator links has a finite lifetime against quantum adversaries; and a collaboratively trained detector is itself a target for poisoning. This paper presents an integrated framework addressing all three within a single deployment model. Regional phasor data concentrators train a physics-informed spatio-temporal detector locally and exchange only model updates, which are encapsulated under ML-KEM-768, encrypted with AES-256-GCM, signed with ML-DSA-65, and committed to a SHA3-256 hash-chained consortium ledger under a Byzantine-tolerant validator quorum. The detector couples a graph attention network over the electrical topology with a temporal convolutional network over a two-second window, and is evaluated on 7.6 million synchrophasor measurements from a hardware-in-the-loop testbed on the IEEE 39-bus system. The admittance model underpinning the graph is validated against the solved base case to a mean bus-injection error of 1.24 MW on a 6,088 MW system, against 1,338 MW for a reduced model omitting transformer taps. Against a graph-free ablation the spatial branch reduces false alarms on undisturbed operation from 8.99% to 1.20% while raising macro-F1 from 0.893 to 0.968, and raises no false alarms on benign grid disturbances in held-out evaluation at reduced attack magnitudes. The post-quantum layer adds 109 ms per update, 0.342% of federated round time.","author":[{"family":"Tank","given":"Divyam"},{"family":"Singh","given":"Dr"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22165672","URL":"https://doi.org/10.5281/zenodo.22165672","source":"datacite"},{"id":"doi:10.5281/zenodo.22165673","type":"article-journal","title":"A Quantum-Resilient Federated Learning Framework for Secure Smart Grid","abstract":"Wide-area monitoring systems built on phasor measurement units (PMUs) underpin real-time stability assessment in transmission grids, yet their measurement and communication paths present an attack surface that conventional intrusion detection addresses only partially. Three constraints compound the problem: transmission operators are commercially and legally restricted from pooling raw measurements; the public-key cryptography protecting inter-operator links has a finite lifetime against quantum adversaries; and a collaboratively trained detector is itself a target for poisoning. This paper presents an integrated framework addressing all three within a single deployment model. Regional phasor data concentrators train a physics-informed spatio-temporal detector locally and exchange only model updates, which are encapsulated under ML-KEM-768, encrypted with AES-256-GCM, signed with ML-DSA-65, and committed to a SHA3-256 hash-chained consortium ledger under a Byzantine-tolerant validator quorum. The detector couples a graph attention network over the electrical topology with a temporal convolutional network over a two-second window, and is evaluated on 7.6 million synchrophasor measurements from a hardware-in-the-loop testbed on the IEEE 39-bus system. The admittance model underpinning the graph is validated against the solved base case to a mean bus-injection error of 1.24 MW on a 6,088 MW system, against 1,338 MW for a reduced model omitting transformer taps. Against a graph-free ablation the spatial branch reduces false alarms on undisturbed operation from 8.99% to 1.20% while raising macro-F1 from 0.893 to 0.968, and raises no false alarms on benign grid disturbances in held-out evaluation at reduced attack magnitudes. The post-quantum layer adds 109 ms per update, 0.342% of federated round time.","author":[{"family":"Tank","given":"Divyam"},{"family":"Singh","given":"Dr"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22165673","URL":"https://doi.org/10.5281/zenodo.22165673","source":"datacite"},{"id":"doi:10.5281/zenodo.22166181","type":"article-journal","title":"Parallels and Interconnections Between Mathematics, Art, Writing, and Music: Living the Metaverse and Beyond","abstract":"From basic screen-based interfaces into completely immersive, multidimensional worlds, the metaverse marks a great evolutionary leap in digital human interaction. As fundamental pillars for creating and living within the metaverse and beyond, this study looks at the underlying parallels and connections between math, art, literature, and music as well as their connections. We offer a comprehensive approach to digital life by seeing the virtual world as a confluence of mathematical graph theory, graphic artistic expression, narrative writing structures, and acoustic musical resonance. We investigate how federated learning, blockchain ecosystems, and digital twins among other modern technologies act as the mechanical foundations transforming these main human domains into a common virtual ontology. By means of a disciplined research plan, this essay emphasizes the need of connecting serious calculation with inventive humanities to guarantee a sustainable, empathetic, and interoperable future digital society.","author":[{"family":"Mageed","given":"Ismail"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22166181","URL":"https://doi.org/10.5281/zenodo.22166181","source":"datacite"},{"id":"doi:10.5281/zenodo.22166180","type":"article-journal","title":"Parallels and Interconnections Between Mathematics, Art, Writing, and Music: Living the Metaverse and Beyond","abstract":"From basic screen-based interfaces into completely immersive, multidimensional worlds, the metaverse marks a great evolutionary leap in digital human interaction. As fundamental pillars for creating and living within the metaverse and beyond, this study looks at the underlying parallels and connections between math, art, literature, and music as well as their connections. We offer a comprehensive approach to digital life by seeing the virtual world as a confluence of mathematical graph theory, graphic artistic expression, narrative writing structures, and acoustic musical resonance. We investigate how federated learning, blockchain ecosystems, and digital twins among other modern technologies act as the mechanical foundations transforming these main human domains into a common virtual ontology. By means of a disciplined research plan, this essay emphasizes the need of connecting serious calculation with inventive humanities to guarantee a sustainable, empathetic, and interoperable future digital society.","author":[{"family":"Mageed","given":"Ismail"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22166180","URL":"https://doi.org/10.5281/zenodo.22166180","source":"datacite"},{"id":"doi:10.5281/zenodo.22165448","type":"article-journal","title":"Execution Governance 5.0 Research Architecture: From Governed Effect Fabrics to Governed Effect Regimes","abstract":"Execution Governance 5.0 (EG5) Research Architecture v0.1.4 proposes a candidate major-version research direction extending Execution Governance from authority-preserving Governed Effect Fabrics to the time-evolving Governed Effect Regime that determines how such fabrics, authority roots, policies, comparators, composition rules, reconciliation rules, adaptation mechanisms, and evidence requirements may themselves be created, changed, combined, suspended, or replaced. The central research question is: Who authorizes a material change to the governance system that determines what counts as authorized? EG5 treats a material governance-state transition as consequential when it can alter the future admissible effect space or the authority, semantic, commitment, reconciliation, or evidence rules applied to future effects. Its candidate governing principle is: Governance may evolve, but no governance change may create the authority that legitimizes itself. A companion composition principle is: Valid governed fabrics do not imply a valid governance composition. EG5 retains the effect-centered discipline of earlier Execution Governance generations and does not introduce a new source of normative authority. It does not claim to invent administrative authorization, policy administration, compositional authorization, recursive governance, governance-of-governance, governed runtime mutation, learning/authority separation, non-widening composition, cryptographic authorization proofs, or evidence-chain composition. These areas have substantial antecedent and adjacent work. The proposed research distinction is narrower: the effect-centered conjunction of authorization requirements at a material governance-state boundary. The candidate EG5-Core v0.1 defines six Recursive Integrity properties: Governance-Change Authority Provenance Non-Self-Authorization Authority Composition Closure Multi-Root Reconciliation Integrity Consequence-Bounded Adaptation Evolution-Witnessed Closure The accompanying deterministic executable guard-ablation harness provides six hand-constructed minimal counterexamples, one for each property. In each named scenario, omission of the relevant property admits the bad state while the complete candidate guard blocks it. A separate regression confirms that participating EG4 fabric validity remains a non-substitutable prerequisite. These results are executable falsification evidence only. They are not bounded model checking, exhaustive state-space exploration, formal proof, independent reproduction, certification, production assurance, or evidence of governance completeness. The first proposed implementation profile is the Minimal Federated Governance-Evolution Profile v0.1, designed around two independently administered EG4-class digital fabrics, two authority roots, one separate verifier, one reversible synthetic cross-fabric effect, one material governance change, explicit composition and reconciliation rules, and bounded consequence-feedback semantics. The principal next evidence milestone is to demonstrate an executable case in which: Fabric A is EG4-valid.Fabric B is EG4-valid.Their composite effect is not authorized.EG5 correctly blocks the composition. Stable EG5.0 status is intentionally not claimed by this publication. It should be earned through a profile-bounded formal model, two-domain implementation, public reproducibility, and independent reconstruction. EG4 baseline:KU, H. W. (2026). Execution Governance 4.0: From Authorization-Bound Execution to Governed Effect Fabrics (Version 0.3.6.2). Zenodo.https://doi.org/10.5281/zenodo.22157731 Research programme: https://executiongovernance.org Status: Independent research and pre-standardization candidate. This publication does not constitute a standard, certification scheme, legal authorization determination, production-safety claim, or proof of governance completeness.","author":[{"family":"Ku","given":"Ho"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22165448","URL":"https://doi.org/10.5281/zenodo.22165448","source":"datacite"},{"id":"doi:10.5281/zenodo.22165449","type":"article-journal","title":"Execution Governance 5.0 Research Architecture: From Governed Effect Fabrics to Governed Effect Regimes","abstract":"Execution Governance 5.0 (EG5) Research Architecture v0.1.4 proposes a candidate major-version research direction extending Execution Governance from authority-preserving Governed Effect Fabrics to the time-evolving Governed Effect Regime that determines how such fabrics, authority roots, policies, comparators, composition rules, reconciliation rules, adaptation mechanisms, and evidence requirements may themselves be created, changed, combined, suspended, or replaced. The central research question is: Who authorizes a material change to the governance system that determines what counts as authorized? EG5 treats a material governance-state transition as consequential when it can alter the future admissible effect space or the authority, semantic, commitment, reconciliation, or evidence rules applied to future effects. Its candidate governing principle is: Governance may evolve, but no governance change may create the authority that legitimizes itself. A companion composition principle is: Valid governed fabrics do not imply a valid governance composition. EG5 retains the effect-centered discipline of earlier Execution Governance generations and does not introduce a new source of normative authority. It does not claim to invent administrative authorization, policy administration, compositional authorization, recursive governance, governance-of-governance, governed runtime mutation, learning/authority separation, non-widening composition, cryptographic authorization proofs, or evidence-chain composition. These areas have substantial antecedent and adjacent work. The proposed research distinction is narrower: the effect-centered conjunction of authorization requirements at a material governance-state boundary. The candidate EG5-Core v0.1 defines six Recursive Integrity properties: Governance-Change Authority Provenance Non-Self-Authorization Authority Composition Closure Multi-Root Reconciliation Integrity Consequence-Bounded Adaptation Evolution-Witnessed Closure The accompanying deterministic executable guard-ablation harness provides six hand-constructed minimal counterexamples, one for each property. In each named scenario, omission of the relevant property admits the bad state while the complete candidate guard blocks it. A separate regression confirms that participating EG4 fabric validity remains a non-substitutable prerequisite. These results are executable falsification evidence only. They are not bounded model checking, exhaustive state-space exploration, formal proof, independent reproduction, certification, production assurance, or evidence of governance completeness. The first proposed implementation profile is the Minimal Federated Governance-Evolution Profile v0.1, designed around two independently administered EG4-class digital fabrics, two authority roots, one separate verifier, one reversible synthetic cross-fabric effect, one material governance change, explicit composition and reconciliation rules, and bounded consequence-feedback semantics. The principal next evidence milestone is to demonstrate an executable case in which: Fabric A is EG4-valid.Fabric B is EG4-valid.Their composite effect is not authorized.EG5 correctly blocks the composition. Stable EG5.0 status is intentionally not claimed by this publication. It should be earned through a profile-bounded formal model, two-domain implementation, public reproducibility, and independent reconstruction. EG4 baseline:KU, H. W. (2026). Execution Governance 4.0: From Authorization-Bound Execution to Governed Effect Fabrics (Version 0.3.6.2). Zenodo.https://doi.org/10.5281/zenodo.22157731 Research programme: https://executiongovernance.org Status: Independent research and pre-standardization candidate. This publication does not constitute a standard, certification scheme, legal authorization determination, production-safety claim, or proof of governance completeness.","author":[{"family":"Ku","given":"Ho"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22165449","URL":"https://doi.org/10.5281/zenodo.22165449","source":"datacite"},{"id":"doi:10.5281/zenodo.20527681","type":"article-journal","title":"FedBuild: Privacy-Preserving Federated Learning for Building Energy Forecasting","abstract":"FedBuild implements a production-grade federated learning system with formal differential privacy (DP) guarantees. The system trains CNN-LSTM models across 50 buildings using the Flower framework, applies per-sample gradient clipping and Gaussian noise via Opacus, and evaluates three aggregation strategies (FedAvg, FedBN, FedProx) across privacy budgets from ε=0.5 to ε=6.5. Key finding: DP-SGD gradient clipping acts as implicit regularization in federated settings with inter-client heterogeneity, causing all DP configurations to outperform the non-private federated baseline—a novel effect unreported in prior building-energy forecasting literature. Features Flower-based orchestration: 50 clients, 10 sampled per round, 50 federation rounds DP-SGD via Opacus: Per-sample gradient clipping (C=1.0) + Gaussian noise, server-side RDP composition accounting Three strategies evaluated: FedAvg, FedBN (local batch norm), FedProx (proximal regularization) 12 privacy configurations: 3 strategies × 4 ε targets (0.5, 1.0, 3.0, 6.5) DP-compatible architecture: CNN-LSTM with DPLSTM + GroupNorm substitutions Real-world dataset: Building Data Genome Project 2, 50 office/education buildings (Panther site) Reproducibility: Round-by-round checkpointing, seed fixing, aggregated metrics + per-run histories Full privacy validation: Server-side RDP composition with explicit final ε reporting (achieved ≈ target within 3%) Publication-ready outputs: 9 figures (PDF+PNG), methodology document, comprehensive tables","author":[{"family":"Ameh","given":"Jude"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20527681","URL":"https://doi.org/10.5281/zenodo.20527681","source":"datacite"},{"id":"doi:10.5281/zenodo.20527682","type":"article-journal","title":"FedBuild: Privacy-Preserving Federated Learning for Building Energy Forecasting","abstract":"FedBuild implements a production-grade federated learning system with formal differential privacy (DP) guarantees. The system trains CNN-LSTM models across 50 buildings using the Flower framework, applies per-sample gradient clipping and Gaussian noise via Opacus, and evaluates three aggregation strategies (FedAvg, FedBN, FedProx) across privacy budgets from ε=0.5 to ε=6.5. Key finding: DP-SGD gradient clipping acts as implicit regularization in federated settings with inter-client heterogeneity, causing all DP configurations to outperform the non-private federated baseline—a novel effect unreported in prior building-energy forecasting literature. Features Flower-based orchestration: 50 clients, 10 sampled per round, 50 federation rounds DP-SGD via Opacus: Per-sample gradient clipping (C=1.0) + Gaussian noise, server-side RDP composition accounting Three strategies evaluated: FedAvg, FedBN (local batch norm), FedProx (proximal regularization) 12 privacy configurations: 3 strategies × 4 ε targets (0.5, 1.0, 3.0, 6.5) DP-compatible architecture: CNN-LSTM with DPLSTM + GroupNorm substitutions Real-world dataset: Building Data Genome Project 2, 50 office/education buildings (Panther site) Reproducibility: Round-by-round checkpointing, seed fixing, aggregated metrics + per-run histories Full privacy validation: Server-side RDP composition with explicit final ε reporting (achieved ≈ target within 3%) Publication-ready outputs: 9 figures (PDF+PNG), methodology document, comprehensive tables","author":[{"family":"Ameh","given":"Jude"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20527682","URL":"https://doi.org/10.5281/zenodo.20527682","source":"datacite"},{"id":"doi:10.5281/zenodo.18761410","type":"article-journal","title":"Cross-Agent Governance Alignment Without Rule Disclosure: A Problem Formalization","abstract":"This preprint formalizes the Cross-Agent Governance Alignment (CAGA) problem: the challenge of verifying mutual governance compatibility between autonomous AI agents operating under distinct organizational policy regimes—without disclosing proprietary governance structures. As AI agents increasingly coordinate across institutional boundaries in regulated industries (healthcare, finance, cross-border data exchange, supply chains), existing governance models prove insufficient. Current frameworks assume either a single organizational authority or full policy transparency between participants. Neither assumption holds in multi-stakeholder settings where governance constraints encode confidential risk tolerances, regulatory interpretations, and competitive strategy. This paper: Defines governance domains and cross-domain interactions in formal terms Introduces the governance alignment predicate Φ(Dᵢ, Dⱼ, τ) Formalizes the CAGA problem under an honest-but-curious threat model Identifies required solution properties spanning correctness, privacy, determinism, evidentiary sufficiency, and composable security Demonstrates that CAGA is irreducible to existing paradigms, including agent communication protocols, federated learning, secure multi-party computation, single-organization governance architectures, and blockchain-based transparency systems We argue that CAGA constitutes a zero-knowledge coordination problem at the intersection of AI governance, cryptographic protocol design, and multi-agent systems. The paper deliberately stops at problem formalization and does not disclose protocol constructions or implementation mechanisms. By precisely defining the problem space and evaluation criteria, this work establishes the foundation for rigorous solution development and provides a formal framework against which candidate governance-alignment protocols can be assessed.","author":[{"family":"Meyman","given":"Edward"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18761410","URL":"https://doi.org/10.5281/zenodo.18761410","source":"datacite"},{"id":"doi:10.5281/zenodo.21209285","type":"article-journal","title":"Cross-Agent Governance Alignment (CAGA): Formalizing Cross-Organizational AI Governance as a Zero-Knowledge Coordination Problem","abstract":"This preprint formalizes the Cross-Agent Governance Alignment (CAGA) problem: the challenge of verifying mutual governance compatibility between autonomous AI agents operating under distinct organizational policy regimes, without disclosing proprietary governance structures. As AI agents increasingly coordinate across institutional boundaries in regulated industries (healthcare, finance, cross-border data exchange, supply chains), existing governance models prove insufficient. Current frameworks assume either a single organizational authority or full policy transparency between participants. Neither assumption holds in multi-stakeholder settings where governance constraints encode confidential risk tolerances, regulatory interpretations, and competitive strategy. This paper: Defines governance domains and cross-domain interactions in formal terms, adopting the triadic verdict space (ALLOW, DENY, ABSTAIN) of the execution-time authorization framework Introduces the governance alignment predicate Φ(Dᵢ, Dⱼ, τ) Formalizes the CAGA problem under an honest-but-curious threat model Identifies required solution properties spanning correctness, privacy, determinism, evidentiary sufficiency, and composable security, including the requirement that alignment protocols produce tamper-evident authorization artifacts sufficient for independent third-party replay, consistent with the Replay requirement of the Five Tests Standard (5TS) Demonstrates that CAGA is irreducible to existing paradigms, including agent communication protocols, federated learning, secure multi-party computation, single-organization governance architectures, and blockchain-based transparency systems We argue that CAGA constitutes a zero-knowledge coordination problem at the intersection of AI governance, cryptographic protocol design, and multi-agent systems. The paper deliberately stops at problem formalization and does not disclose protocol constructions or implementation mechanisms. By precisely defining the problem space and evaluation criteria, this work establishes the foundation for rigorous solution development and provides a formal framework against which candidate governance-alignment protocols can be assessed. Version 1.1 (July 2026) retitles the paper to make explicit that CAGA is formalized as a zero-knowledge coordination problem for cross-organizational AI governance; aligns terminology with the Five Tests Standard (5TS) v1.2.0 and the FERZ authorization-artifact vocabulary; adopts the triadic verdict space in the governance domain formalization; and adds a companion reference to Execution-Time Authorization for AI Agents (v2.1), which develops the formal architecture of the single-domain authorization boundary. The problem formalization, threat model, and irreducibility argument are unchanged from the February 2026 release (v1.0). Keywords: AI governance, multi-agent systems, zero-knowledge proofs, cross-organizational coordination, governance alignment, deterministic governance, authorization boundaries, authorization artifacts, Five Tests Standard","author":[{"family":"Meyman","given":"Edward"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21209285","URL":"https://doi.org/10.5281/zenodo.21209285","source":"datacite"},{"id":"doi:10.5281/zenodo.20787126","type":"article-journal","title":"GROUPOID: Groupoid-based aggregation for federated learning on Riemannian manifolds","abstract":"Pre-alpha research prototype exploring groupoid-based aggregation for federated learning on Riemannian manifolds. Implements transport groupoid morphisms, first cohomology (H^1) for consistency detection, the cellular sheaf Laplacian for spectral analysis, parallel transport (Schild's and pole ladders), and Karcher mean aggregation via geomstats. Erratum / supersession notice: this version (v0.1.0.dev2) carries the mathematical corrections introduced in v0.1.0.dev1 over the earlier v0.1.0.dev0 snapshot (version DOI 10.5281/zenodo.20563975): the sheaf (connection) Laplacian is a verified positive-semidefinite operator with a transport-consistent kernel; persistence diagrams retain the homology-dimension label so H0 and H1 are no longer conflated; and H^1 cohomology raises on incomplete cocycles instead of forming partial holonomy. It additionally corrects a first-step magnitude inflation in the RiemannianAdam optimizer's bias initialization (a component not used by the aggregation results). The v0.1.0.dev0 version DOI must not be cited for results. Cite the concept DOI (10.5281/zenodo.20563974), which always resolves to the latest corrected version, or this version or later.","author":[{"family":"Maniches","given":"Santiago"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20787126","URL":"https://doi.org/10.5281/zenodo.20787126","source":"datacite"},{"id":"doi:10.5281/zenodo.20719394","type":"article-journal","title":"GROUPOID: Groupoid-based aggregation for federated learning on Riemannian manifolds","abstract":"Pre-alpha research prototype exploring groupoid-based aggregation for federated learning on Riemannian manifolds. Implements transport groupoid morphisms, first cohomology (H^1) for consistency detection, the cellular sheaf Laplacian for spectral analysis, parallel transport (Schild's and pole ladders), and Karcher mean aggregation via geomstats. Erratum / supersession notice: this version (v0.1.0.dev1) corrects mathematical bugs present in the earlier v0.1.0.dev0 snapshot (version DOI 10.5281/zenodo.20563975): the sheaf (connection) Laplacian is now a verified positive-semidefinite operator with a transport-consistent kernel; persistence diagrams retain the homology-dimension label so H0 and H1 are no longer conflated; and H^1 cohomology raises on incomplete cocycles instead of forming partial holonomy. The v0.1.0.dev0 version DOI must not be cited for results. Cite the concept DOI (10.5281/zenodo.20563974), which always resolves to the latest corrected version, or this version or later.","author":[{"family":"Maniches","given":"Santiago"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20719394","URL":"https://doi.org/10.5281/zenodo.20719394","source":"datacite"},{"id":"doi:10.5281/zenodo.20467655","type":"article-journal","title":"Byzantine-Resilient Aggregation Algorithms in Federated Malware Detection for IoT Networks","abstract":"This report synthesises findings from 11 peer-reviewed papers addressing the following research question: How do different aggregation algorithms (FedAvg, FedProx, FedNova) compare in terms of model accuracy degradation when defending against Byzantine attacks in federated malware detection systems. This systematic review examines the role of federated learning (FL) as a privacy-preserving paradigm for enterprise decision systems, synthesizing evidence from 187 peer-reviewed studies. Guided by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA). 10 claims were extracted from source literature; 10 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.5/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How do different aggregation algorithms (FedAvg, FedProx, FedNova) compare in terms of model accuracy degradation when defending against Byzantine attacks in federated malware detection systems across various IoT device topologies? Autonomous literature synthesis. Automated review score: 8.5/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20467655","URL":"https://doi.org/10.5281/zenodo.20467655","source":"datacite"},{"id":"doi:10.5281/zenodo.20467656","type":"article-journal","title":"Byzantine-Resilient Aggregation Algorithms in Federated Malware Detection for IoT Networks","abstract":"This report synthesises findings from 11 peer-reviewed papers addressing the following research question: How do different aggregation algorithms (FedAvg, FedProx, FedNova) compare in terms of model accuracy degradation when defending against Byzantine attacks in federated malware detection systems. This systematic review examines the role of federated learning (FL) as a privacy-preserving paradigm for enterprise decision systems, synthesizing evidence from 187 peer-reviewed studies. Guided by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA). 10 claims were extracted from source literature; 10 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.5/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How do different aggregation algorithms (FedAvg, FedProx, FedNova) compare in terms of model accuracy degradation when defending against Byzantine attacks in federated malware detection systems across various IoT device topologies? Autonomous literature synthesis. Automated review score: 8.5/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20467656","URL":"https://doi.org/10.5281/zenodo.20467656","source":"datacite"},{"id":"doi:10.5281/zenodo.20459718","type":"article-journal","title":"Generative Threats in Computer Vision: Dual-Role Gans and Diffusion-Based Defenses for Robust and Trustworthy Models","abstract":"Abstract - Generative Artificial Intelligence (GenAI) has revolutionized computer vision with cutting-edge image generation, semantic interpretation, and adaptive visual reasoning capabilities, while also bringing new security, robustness, and trustworthiness concerns. In this review, the twenty five recent studies related to generative threat and diffusion-based defense mechanism in computer vision and intelligent network systems were analyzed in a systematic manner based on PRISMA based literature review methodology. The studies reviewed showed that Generative Adversarial Networks (GANs), diffusion models, transformer models, and large language models are dual-use technologies that can be used to create powerful adversarial attacks and powerful defense frameworks. The analysis identified the shifting nature of adversarial attacks—from perturbations at the pixel level to latent-space attacks, multimodal deepfakes, semantic communication attacks and cyber deception by AI tools. At the same time, diffusion-based purification frameworks, federated defense systems, anomaly detection models, and transformer-based verification mechanisms demonstrated high potential to enhance robustness, semantic consistency, privacy preservation, and real-time resilience. The results also showed that synthetic data generation has a notable positive impact on learning performance for low-data scenarios like military object detection and cyber security applications. The issues of computational complexity, attack transferability, scalability and benchmarking inconsistency, however, have yet to be addressed. The review finds that the hybrid generative defense architectures that combine diffusion models, federated learning, explainable AI, graph neural networks, and adaptive semantic verification mechanisms will become increasingly vital for future trustworthy computer vision systems in order to foster secure, resilient, and interpretable next-generation AI systems.","author":[{"family":"Kumar","given":"Mahesh"},{"family":"Rustagi","given":"Tanvi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20459718","URL":"https://doi.org/10.5281/zenodo.20459718","source":"datacite"},{"id":"doi:10.5281/zenodo.20459719","type":"article-journal","title":"Generative Threats in Computer Vision: Dual-Role Gans and Diffusion-Based Defenses for Robust and Trustworthy Models","abstract":"Abstract - Generative Artificial Intelligence (GenAI) has revolutionized computer vision with cutting-edge image generation, semantic interpretation, and adaptive visual reasoning capabilities, while also bringing new security, robustness, and trustworthiness concerns. In this review, the twenty five recent studies related to generative threat and diffusion-based defense mechanism in computer vision and intelligent network systems were analyzed in a systematic manner based on PRISMA based literature review methodology. The studies reviewed showed that Generative Adversarial Networks (GANs), diffusion models, transformer models, and large language models are dual-use technologies that can be used to create powerful adversarial attacks and powerful defense frameworks. The analysis identified the shifting nature of adversarial attacks—from perturbations at the pixel level to latent-space attacks, multimodal deepfakes, semantic communication attacks and cyber deception by AI tools. At the same time, diffusion-based purification frameworks, federated defense systems, anomaly detection models, and transformer-based verification mechanisms demonstrated high potential to enhance robustness, semantic consistency, privacy preservation, and real-time resilience. The results also showed that synthetic data generation has a notable positive impact on learning performance for low-data scenarios like military object detection and cyber security applications. The issues of computational complexity, attack transferability, scalability and benchmarking inconsistency, however, have yet to be addressed. The review finds that the hybrid generative defense architectures that combine diffusion models, federated learning, explainable AI, graph neural networks, and adaptive semantic verification mechanisms will become increasingly vital for future trustworthy computer vision systems in order to foster secure, resilient, and interpretable next-generation AI systems.","author":[{"family":"Kumar","given":"Mahesh"},{"family":"Rustagi","given":"Tanvi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20459719","URL":"https://doi.org/10.5281/zenodo.20459719","source":"datacite"},{"id":"doi:10.5281/zenodo.20100343","type":"article-journal","title":"Privacy-Preserving Machine Learning for Healthcare: A Comparative Study of Federated Learning, Split Learning, and Split-Fed Learning","abstract":"This project presents an empirical comparison of three privacy-preserving machine learning (PPML) techniques — Federated Learning (FL), Split Learning (SL), and Split-Fed Learning (SFL) — for healthcare classification, benchmarked against a centralised baseline. Using the Heart Disease UCI dataset (920 samples, 13 clinical features, binary classification), each technique was implemented in TensorFlow/Keras within a simulated five-client IID federation representing collaborating hospitals that cannot share raw patient data. Results show that all three PPML techniques matched or exceeded the centralised baseline (82.07% accuracy): FL achieved 83.70%, SL reached 82.61%, and SFL achieved the best overall performance at 84.24% accuracy with the lowest test loss (0.3680). SFL is identified as the optimal approach for federated healthcare deployments, offering the best balance of accuracy, loss calibration, and convergence stability. The study addresses the growing tension between the potential of ML in clinical decision support and the regulatory constraints imposed by UK GDPR and the Data Protection Act 2018 on health data sharing. The full implementation, including preprocessing, training, and evaluation code, is provided for reproducibility.","author":[{"family":"Sathar","given":"Alif"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20100343","URL":"https://doi.org/10.5281/zenodo.20100343","source":"datacite"},{"id":"doi:10.5281/zenodo.20100344","type":"article-journal","title":"Privacy-Preserving Machine Learning for Healthcare: A Comparative Study of Federated Learning, Split Learning, and Split-Fed Learning","abstract":"This project presents an empirical comparison of three privacy-preserving machine learning (PPML) techniques — Federated Learning (FL), Split Learning (SL), and Split-Fed Learning (SFL) — for healthcare classification, benchmarked against a centralised baseline. Using the Heart Disease UCI dataset (920 samples, 13 clinical features, binary classification), each technique was implemented in TensorFlow/Keras within a simulated five-client IID federation representing collaborating hospitals that cannot share raw patient data. Results show that all three PPML techniques matched or exceeded the centralised baseline (82.07% accuracy): FL achieved 83.70%, SL reached 82.61%, and SFL achieved the best overall performance at 84.24% accuracy with the lowest test loss (0.3680). SFL is identified as the optimal approach for federated healthcare deployments, offering the best balance of accuracy, loss calibration, and convergence stability. The study addresses the growing tension between the potential of ML in clinical decision support and the regulatory constraints imposed by UK GDPR and the Data Protection Act 2018 on health data sharing. The full implementation, including preprocessing, training, and evaluation code, is provided for reproducibility.","author":[{"family":"Sathar","given":"Alif"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20100344","URL":"https://doi.org/10.5281/zenodo.20100344","source":"datacite"},{"id":"doi:10.5281/zenodo.20467665","type":"article-journal","title":"Federated vs Centralized Malware Detection Robustness Under Adversarial IoT Attacks","abstract":"This report synthesises findings from 10 peer-reviewed papers addressing the following research question: How does the robustness of federated malware detection models compare to centralized models when tested against adversarial attacks on edge IoT devices under varying network conditions. In this article, we present a comprehensive study with an experimental analysis of federated deep learning approaches for cyber security in the Internet of Things (IoT) applications. Specifically, we first provide a review of the federated learning-based security and privacy. 9 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.0/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does the robustness of federated malware detection models compare to centralized models when tested against adversarial attacks on edge IoT devices under varying network conditions? Autonomous literature synthesis. Automated review score: 8.0/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20467665","URL":"https://doi.org/10.5281/zenodo.20467665","source":"datacite"},{"id":"doi:10.5281/zenodo.20467666","type":"article-journal","title":"Federated vs Centralized Malware Detection Robustness Under Adversarial IoT Attacks","abstract":"This report synthesises findings from 10 peer-reviewed papers addressing the following research question: How does the robustness of federated malware detection models compare to centralized models when tested against adversarial attacks on edge IoT devices under varying network conditions. In this article, we present a comprehensive study with an experimental analysis of federated deep learning approaches for cyber security in the Internet of Things (IoT) applications. Specifically, we first provide a review of the federated learning-based security and privacy. 9 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.0/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does the robustness of federated malware detection models compare to centralized models when tested against adversarial attacks on edge IoT devices under varying network conditions? Autonomous literature synthesis. Automated review score: 8.0/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20467666","URL":"https://doi.org/10.5281/zenodo.20467666","source":"datacite"},{"id":"doi:10.5281/zenodo.20280460","type":"article-journal","title":"Scalable Real-Time Android Malware Detection Using Adaptive Deep Learning and Ensemble Techniques","abstract":"Android malware is constantly evolving, and operational constraints, and implementing federated learning protocols to ensure that model updates remain localized to preserve user data integrity while constantly improving detection capabilities against emerging polymorphic threats. This paper investigates a dual-modal feature extraction approach using convolutional neural networks and frequency domain analysis for Android malware detection, which converts Android application packages to grayscale fingerprint images generated from DEX bytecode segments, and extracts both spatial features through convolutional neural networks and frequency domain characteristics through Fourier. They are combined with a recursive feature fusion mechanism with attention-based weighting and classified by fully connected neural networks. This system shows good detection performance over various benchmark datasets but has some limitations such as high computational complexity, large training data requirements, and less suitable for real-time deployment. This paper presents an improved malware detection framework based on adaptive feature selection, lightweight deep learning models, and ensemble learning techniques. The proposed system is expected to increase scalability, decrease computational overheads, improve robustness against adversarial attacks while maintaining high accuracy of the detections as well as integrating distributed detection mechanisms along with edge-based mechanism to enable real-time identification of mobile malicious applications (malware) within large-scale Android environments in current cybersecurity systems.","author":[{"family":"Srinivas","given":"SG"},{"family":"Ganesh","given":"Dr"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20280460","URL":"https://doi.org/10.5281/zenodo.20280460","source":"datacite"},{"id":"doi:10.5281/zenodo.20280461","type":"article-journal","title":"Scalable Real-Time Android Malware Detection Using Adaptive Deep Learning and Ensemble Techniques","abstract":"Android malware is constantly evolving, and operational constraints, and implementing federated learning protocols to ensure that model updates remain localized to preserve user data integrity while constantly improving detection capabilities against emerging polymorphic threats. This paper investigates a dual-modal feature extraction approach using convolutional neural networks and frequency domain analysis for Android malware detection, which converts Android application packages to grayscale fingerprint images generated from DEX bytecode segments, and extracts both spatial features through convolutional neural networks and frequency domain characteristics through Fourier. They are combined with a recursive feature fusion mechanism with attention-based weighting and classified by fully connected neural networks. This system shows good detection performance over various benchmark datasets but has some limitations such as high computational complexity, large training data requirements, and less suitable for real-time deployment. This paper presents an improved malware detection framework based on adaptive feature selection, lightweight deep learning models, and ensemble learning techniques. The proposed system is expected to increase scalability, decrease computational overheads, improve robustness against adversarial attacks while maintaining high accuracy of the detections as well as integrating distributed detection mechanisms along with edge-based mechanism to enable real-time identification of mobile malicious applications (malware) within large-scale Android environments in current cybersecurity systems.","author":[{"family":"Srinivas","given":"SG"},{"family":"Ganesh","given":"Dr"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20280461","URL":"https://doi.org/10.5281/zenodo.20280461","source":"datacite"},{"id":"doi:10.5281/zenodo.20484878","type":"article-journal","title":"Adaptive Sparsification and FedAvg for Low-Latency LLM Inference in Federated Edge Networks","abstract":"This report synthesises findings from 10 peer-reviewed papers addressing the following research question: Can adaptive sparsification techniques combined with FedAvg improve inference latency and throughput for large language models in over-the-air federated learning scenarios without degrading alignment. In recent years, mobile devices are equipped with increasingly advanced sensing and computing capabilities. Coupled with advancements in Deep Learning (DL), this opens up countless possibilities for meaningful applications. 10 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 7.8/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: Can adaptive sparsification techniques combined with FedAvg improve inference latency and throughput for large language models in over-the-air federated learning scenarios without degrading alignment scores? Autonomous literature synthesis. Automated review score: 7.8/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20484878","URL":"https://doi.org/10.5281/zenodo.20484878","source":"datacite"},{"id":"doi:10.5281/zenodo.20484879","type":"article-journal","title":"Adaptive Sparsification and FedAvg for Low-Latency LLM Inference in Federated Edge Networks","abstract":"This report synthesises findings from 10 peer-reviewed papers addressing the following research question: Can adaptive sparsification techniques combined with FedAvg improve inference latency and throughput for large language models in over-the-air federated learning scenarios without degrading alignment. In recent years, mobile devices are equipped with increasingly advanced sensing and computing capabilities. Coupled with advancements in Deep Learning (DL), this opens up countless possibilities for meaningful applications. 10 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 7.8/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: Can adaptive sparsification techniques combined with FedAvg improve inference latency and throughput for large language models in over-the-air federated learning scenarios without degrading alignment scores? Autonomous literature synthesis. Automated review score: 7.8/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20484879","URL":"https://doi.org/10.5281/zenodo.20484879","source":"datacite"},{"id":"doi:10.5281/zenodo.20478837","type":"article-journal","title":"Federated Learning Aggregation Strategies and Compressive Sensing in Massive MIMO OTA-FL Systems","abstract":"This report synthesises findings from 7 peer-reviewed papers addressing the following research question: What is the impact of different federated learning aggregation strategies (FedAvg, FedProx, SCAFFOLD) on model alignment and robustness to non-IID data distributions when combined with compressive. Federated learning is a privacy-preserving approach to train a global model at a central server by collaborating with wireless devices, each with its own local training data set. In this paper, we present a compressive sensing approach for federated learning over massive. 9 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.8/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the impact of different federated learning aggregation strategies (FedAvg, FedProx, SCAFFOLD) on model alignment and robustness to non-IID data distributions when combined with compressive sensing in massive MIMO-enabled OTA-FL, evaluated using metrics like test accuracy and F1-score on datasets such as CIFAR-10 or Shakespeare? Autonomous literature synthesis. Automated review score: 8.8/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20478837","URL":"https://doi.org/10.5281/zenodo.20478837","source":"datacite"},{"id":"doi:10.5281/zenodo.20478838","type":"article-journal","title":"Federated Learning Aggregation Strategies and Compressive Sensing in Massive MIMO OTA-FL Systems","abstract":"This report synthesises findings from 7 peer-reviewed papers addressing the following research question: What is the impact of different federated learning aggregation strategies (FedAvg, FedProx, SCAFFOLD) on model alignment and robustness to non-IID data distributions when combined with compressive. Federated learning is a privacy-preserving approach to train a global model at a central server by collaborating with wireless devices, each with its own local training data set. In this paper, we present a compressive sensing approach for federated learning over massive. 9 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.8/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the impact of different federated learning aggregation strategies (FedAvg, FedProx, SCAFFOLD) on model alignment and robustness to non-IID data distributions when combined with compressive sensing in massive MIMO-enabled OTA-FL, evaluated using metrics like test accuracy and F1-score on datasets such as CIFAR-10 or Shakespeare? Autonomous literature synthesis. Automated review score: 8.8/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20478838","URL":"https://doi.org/10.5281/zenodo.20478838","source":"datacite"},{"id":"doi:10.5281/zenodo.20551818","type":"article-journal","title":"Mathematical optimization of renewable energy systems for sustainable development","abstract":"Due to the growing worldwide need for renewable energy sources, Renewable Energy Systems (RES) have been rapidly developed and deployed in response to global needs for cleaner energy and environmental damage from the use of fossil fuels. However, the inherent variability and uncertainty associated with renewable energy resources create significant challenges to RES planning, integration of RES into the overall energy system, and management of the RES. This overview examines the various types of mathematical optimization methodologies that have been applied to RES; it examines deterministic classical methods, stochastic approaches to optimizing RES, heuristics and metaheuristic algorithms, and artificial intelligence (AI) solutions through a systematic examination according to a structured taxonomy. The overall critical review describes the strength and weakness of optimization methodologies and provides an assessment of the computing resources available for each type of optimization method, describes the techniques for quantifying uncertainty, explains the principles and techniques used in probabilistic forecasting, and addresses the real-world deployment challenges associated with RES including regulatory, economic, financial, and infrastructure barriers. The overview also describes in detail hybrid renewable energy systems (HRES), multi-objective optimization methods, integration of energy storage systems, and smart grid optimization processes utilizing supporting comparison tables and illustrative examples. The review outlines areas for future research by suggesting using innovations such as Digital Twins, Explainable AI, Federated Learning, and Blockchain technologies for enhancing energy systems management. This review distinguishes itself from previous literature in that it synthesizes findings from multiple optimization paradigms and bridges the gap between theoretical modeling and practical implementation challenges.","author":[{"family":"Singh","given":"Garima"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20551818","URL":"https://doi.org/10.5281/zenodo.20551818","source":"datacite"},{"id":"doi:10.5281/zenodo.20551819","type":"article-journal","title":"Mathematical optimization of renewable energy systems for sustainable development","abstract":"Due to the growing worldwide need for renewable energy sources, Renewable Energy Systems (RES) have been rapidly developed and deployed in response to global needs for cleaner energy and environmental damage from the use of fossil fuels. However, the inherent variability and uncertainty associated with renewable energy resources create significant challenges to RES planning, integration of RES into the overall energy system, and management of the RES. This overview examines the various types of mathematical optimization methodologies that have been applied to RES; it examines deterministic classical methods, stochastic approaches to optimizing RES, heuristics and metaheuristic algorithms, and artificial intelligence (AI) solutions through a systematic examination according to a structured taxonomy. The overall critical review describes the strength and weakness of optimization methodologies and provides an assessment of the computing resources available for each type of optimization method, describes the techniques for quantifying uncertainty, explains the principles and techniques used in probabilistic forecasting, and addresses the real-world deployment challenges associated with RES including regulatory, economic, financial, and infrastructure barriers. The overview also describes in detail hybrid renewable energy systems (HRES), multi-objective optimization methods, integration of energy storage systems, and smart grid optimization processes utilizing supporting comparison tables and illustrative examples. The review outlines areas for future research by suggesting using innovations such as Digital Twins, Explainable AI, Federated Learning, and Blockchain technologies for enhancing energy systems management. This review distinguishes itself from previous literature in that it synthesizes findings from multiple optimization paradigms and bridges the gap between theoretical modeling and practical implementation challenges.","author":[{"family":"Singh","given":"Garima"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20551819","URL":"https://doi.org/10.5281/zenodo.20551819","source":"datacite"},{"id":"doi:10.5281/zenodo.20280465","type":"article-journal","title":"Intelligent Predictive Architectures for Autonomous Self-Healing in Cloud Computing: A Comprehensive Survey","abstract":"Neural network-enabled self-healing is becoming a very exciting approach to making cloud computing infrastructures more reliable, available, and efficient. This survey paper gives a detailed review of neural network-based predictive models and how they are combined with autonomous recovery mechanisms for self-healing cloud systems. It first delineates the main neural architectures used for cloud reliability, such as Artificial Neural Networks (ANNs), Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM) networks, and hybrid or ensemble models, explaining their roles in failure prediction, resource management forecasting, and SLA/QoS violation prediction. The paper examines the coupling of these predictive models with self-healing action and decision layers like rule-based policies, policy engines, and reinforcement learning-driven controllers to proactively trigger recovery actions, reduce Mean Time To Recovery (MTTR), and control false alarms and resource overhead. Evaluation practices highlight datasets (Google Cluster Traces, Alibaba traces, and synthetic simulation data), performance metrics (prediction accuracy, MTTR, false positive/negative rates, scalability, and energy cost), and differences between simulated vs. real production cloud environments. Moreover, the survey discusses privacy and security issues of predictive and action models, presenting techniques such as federated learning, differential privacy, homomorphic encryption, and secure multi-party computation for privacy-preserving cross-tenant collaboration. Lastly, this paper draws attention to open issues including trade-offs between accuracy and false alarms, latency constraints, explainability, and scalability, proposing future work including explainable predictive pipelines, federated self-healing frameworks, multi-agent and DRL-based autonomy, and integration with emerging technologies such as quantum-inspired neural networks, edge-cloud continuum architectures, digital twins, service meshes, and neuromorphic hardware to enable more trustworthy, efficient, and autonomous self-healing cloud management.","author":[{"family":"Gupta","given":"Mr"},{"family":"Bangera","given":"Dr"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20280465","URL":"https://doi.org/10.5281/zenodo.20280465","source":"datacite"},{"id":"doi:10.1038/s41598-025-22672-1","type":"article-journal","title":"FedEff: efficient federated learning with optimal local epochs for heterogeneous clients.","abstract":"Federated Learning (FL) enables collaborative model training without centralized data sharing; however, its efficiency often degrades under system and statistical heterogeneity across clients. Increasing the number of local epochs per round can enhance efficiency by enabling the global model to reach target accuracy in fewer communication rounds. Yet, excessive local training may cause client models to diverge from the global model, slowing convergence. To examine this trade-off, we conduct an empirical divergence analysis and show that consistent sufficient local updates across rounds can reduce the mean divergence between local and global models, thereby promoting faster and more stable convergence. Building on this insight, we propose a novel, efficient federated learning algorithm (FedEff) that assigns optimal local epochs to each client in heterogeneous settings. FedEff incorporates a server-side epoch selection mechanism, where the server selects an optimal number of epochs for each client, by considering the computation and communication speeds of all clients. The server uses an Estimated Round Time (ERT) to calculate the optimal number of local epochs for each client. Extensive simulations under heterogeneous computation and communication conditions confirm that the proposed approach achieves notable reductions in client waiting times and overall training duration within the considered simulation framework. Comparative results show that our method achieves better training efficiency than FedAvg and random epoch selection strategies, thereby establishing its effectiveness in improving federated learning performance under heterogeneous settings.","author":[{"family":"Narmadha","given":"K"},{"family":"Varalakshmi","given":"P"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-22672-1","URL":"https://doi.org/10.1038/s41598-025-22672-1","source":"europepmc"},{"id":"doi:10.26083/tuda-8146","type":"article-journal","title":"Bootstrapping Physical Security with Inertial Hardware Security Modules","abstract":"In the past decades, cryptographic advancements and techniques like formal verification have steadily improved software security. Meanwhile, the field of hardware security has not kept pace. Research has made progress in subfields such as resilience to Side-Channel Attacks (SCA) and Physical Unclonable Functions (PUFs). However, the state of the art still often relies on microelectronic integration to achieve security by obscurity insted of more fundamental security guarantees. While effective, system-level tamper protection is only used in few devices such as Hardware Security Modules (HSMs) and card payment terminals. Due to the high cost and low performance of HSMs in particular, they remain relegated to niche applications such as Transport Layer Security (TLS) certificate issuance and payment data processing. In this thesis, we introduce the Inertial Hardware Security Module (IHSM), a new architecture for low-cost hardware security modules that provide high-level active tamper protection, while supporting computing payloads of much larger size, weight and power dissipation compared to conventional HSMs. In an IHSM, the costly and difficult to source tamper-sensing mesh of a conventional HSM is replaced by a mesh made from simple PCBs that is rotating at high speed around the payload. Since the mesh is rotating at high speed, it cannot be manipulated, and the security of conventional meshes created in bespoke manufacturing processes can be achieved using much simpler and less expensive construction techniques. We present the results of a survey of approximately 30 real world tamper sensing mesh implementations. Based on our findings, we deduce design criteria for secure meshes and contextualize our design. We further motivate the necessity of secure hardware by presenting an analysis of problematic aspects in the hardware security design of Germany’s new national electronic health record system. To pave the way for practical implementations of IHSM technology, we present solutions to key engineering challenges in IHSM construction. We present a design and analysis of highly symmetric planar inductors for rotating wireless power transfer that improves self-resonant frequency by up to 58 % and inductance by up to 6.5 % in our tests. Complementing this research, we present a high-fidelity, low-cost monitoring system for security meshes that is based on the principles of Time-Domain Reflectometry (TDR), reaching 184 ps time resolution. We validate our system and find that it is able to reliably detect several classes of advanced physical attacks. We find that our system is sensitive enough to detect differences between identical copies of the same mesh, suggesting PUF-like properties. Applying IHSM technology, we analyse two use cases that are unlocked by the increased size and power dissipation capability of IHSMs. In the first analysis, an IHSM-secured relay node for Quantum Key Distribution (QKD) systems is proposed, enabling their practical implementation across arbitrary distances, which requires trusted relay stations due to fundamental physical limitations. In the study, IHSMs are adapted for such high-security QKD relays by securing the IHSM mesh passthrough with a secondary tamper-sensing mesh. In this setup, a bracket design is proposed that supports passing through optical fibers at low loss. The second proposed use case adapts an IHSM enclosure to the size, power and thermal dissipation requirements of a high-power server to support co-located secure Multiparty Computation (MPC) workloads. In practical MPC deployments, nodes are distributed across data centers to avoid a single point of failure for physical attacks. As a result, practical MPC deployments are limited by network bandwidth and latency constraints. Using IHSMs, physically secured MPC nodes can be deployed within the same data center, increasing bandwidth, reducing latency and unlocking a new performance spectrum.","author":[{"family":"Götte","given":"Jan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.26083/tuda-8146","URL":"https://doi.org/10.26083/tuda-8146","source":"datacite"},{"id":"doi:10.4230/lipics.icalp.2026.79","type":"article-journal","title":"On Randomness Complexity of 1-Private Protocols","abstract":"In the field of information-theoretic cryptography, randomness complexity is a key metric for protocols for private computation, that is, the number of random bits needed to realize the protocol. Although some general bounds are known, even for the relatively simple example of 1-private computation of n-party AND, the exact complexity is unknown. We study two settings. First, we consider the model of Goyal, Ishai, and Song (Crypto '22) where helper parties without any inputs are allowed to assist in the computation. In this setting, we show that two random bits always suffice to compute an arbitrary Boolean circuit C 1-privately: a single designated inputless helper flips the two bits and privately distributes the derived one-time bits to the other helper parties and the input parties as they are needed. We give an explicit construction using seven helper parties per AND gate and three helper parties per XOR gate (plus the single global randomness dealer). Moreover, two random bits are necessary already for the AND functionality (by a reduction to the standard no-helper model together with the lower bound of Kushilevitz, Ostrovsky, Prouff, Rosén, Thillard and Vergnaud (TCC '19), and therefore the worst-case helper-party randomness complexity is exactly 2 bits. Second, in the setting without helper parties, we improve the upper bound from Couteau and Rosén (Asiacrypt '22) on the (asymptotic) randomness complexity of n-party AND from 6 to 5 bits. That is, we give a 1-private protocol for computing the AND of n parties' inputs requiring 5 bits of randomness, for all n ≥ 6. Our construction, like that of Couteau and Rosén, uses a single party to flip the 5 bits and distribute the required derived values during the execution. Our approach to both problems is built around a more systematic exploration of techniques for recycling randomness across sub-computations. As part of resolving the second problem, we isolate an exact local-independence combinatorial object called a Sliding-Window Independence Generator, or a SWIG. A (k,m)-SWIG is a linear generator from a k-bit seed to m ≥ k output bits, where every cyclic length-k sliding window chosen from m output bits is perfectly uniform. We give an explicit (k,m)-SWIG for every k ≥ 1 and every m ≥ k and use a (5,n-1)-SWIG in our no-helper AND protocol.","author":[{"family":"Dittmer","given":"Samuel"},{"family":"Ostrovsky","given":"Rafail"}],"issued":{"date-parts":[[2026]]},"DOI":"10.4230/lipics.icalp.2026.79","URL":"https://doi.org/10.4230/lipics.icalp.2026.79","source":"datacite"},{"id":"doi:10.5281/zenodo.20630236","type":"article-journal","title":"UVCC Phase 6 Native Parallelism: Private, Verifiable, Fully-Parallel GPU Computation across Untrusted Domains","abstract":"We present UVCC (Universal Verifiable Confidential Computing), a practical system for private and verifiable GPU computation across mutually distrustful administrative domains. UVCC achieves (i) confidentiality for client secrets using three-party replicated secret sharing (RSS), (ii) verifiability via an append-only transcript and deterministic hashing model that yields per-subsession roots, per-replica roots, and a global root, and (iii) native ML parallelism — data parallel (DP), pipeline parallel (PP) and tensor parallel (TP) — implemented in a C++ runtime that integrates GPU kernels, a reliable exactly-once transport, and NCCL-based collectives within each domain. This paper reports the Phase 6 bring-up of native parallelism end-to-end, including the diagnosis and resolution of a PP deadlock and the scale-out to R=8, S=4, T=2, M=32 on 24 heterogeneous provider pods. We demonstrate determinism through two complete runs that produce identical global roots, show robustness under provider skew and network jitter, and provide an audit-oriented log bundle with cryptographic commitments. The logs referenced herein are consolidated in a single explained file available from the author on request.","author":[{"family":"Gairola","given":"N"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.20630236","URL":"https://doi.org/10.5281/zenodo.20630236","source":"datacite"},{"id":"doi:10.5281/zenodo.20630235","type":"article-journal","title":"UVCC Phase 6 Native Parallelism: Private, Verifiable, Fully-Parallel GPU Computation across Untrusted Domains","abstract":"We present UVCC (Universal Verifiable Confidential Computing), a practical system for private and verifiable GPU computation across mutually distrustful administrative domains. UVCC achieves (i) confidentiality for client secrets using three-party replicated secret sharing (RSS), (ii) verifiability via an append-only transcript and deterministic hashing model that yields per-subsession roots, per-replica roots, and a global root, and (iii) native ML parallelism — data parallel (DP), pipeline parallel (PP) and tensor parallel (TP) — implemented in a C++ runtime that integrates GPU kernels, a reliable exactly-once transport, and NCCL-based collectives within each domain. This paper reports the Phase 6 bring-up of native parallelism end-to-end, including the diagnosis and resolution of a PP deadlock and the scale-out to R=8, S=4, T=2, M=32 on 24 heterogeneous provider pods. We demonstrate determinism through two complete runs that produce identical global roots, show robustness under provider skew and network jitter, and provide an audit-oriented log bundle with cryptographic commitments. The logs referenced herein are consolidated in a single explained file available from the author on request.","author":[{"family":"Gairola","given":"N"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.20630235","URL":"https://doi.org/10.5281/zenodo.20630235","source":"datacite"},{"id":"doi:10.11575/prism/47960","type":"article-journal","title":"A (2 + 1)-Party One-Instruction Set Processor for Private Function Evaluation","abstract":"Secure multiparty computation (MPC) is a powerful cryptographic technique which allows mutually distrusting parties to compute a function on their shared inputs. MPC can be viewed as encompassing two main categories, secure function evaluation (SFE) where the function being computed is known to all parties, but the input data is private, and private function evaluation (PFE) where one party has a private function while the other party has private data as input to the function. A major reason MPC has not seen more widespread adoption is due to the fact that to create an efficient MPC protocol for a specific function, traditionally expert cryptographers have had to hand build a custom protocol. One attempt to remedy this issue is the creation of MPC compilers which take as input code in a high-level language and output an optimized MPC protocol. Another attempt has been garbled processors which are hand optimized MPC protocols which emulate specific computer architectures. This thesis presents MPC SUBLEQ, a garbled processor designed for the PFE setting. In particular, it emulates the subtract-and-branch-if-less-than-or-equal-to-zero (SUBLEQ) one-instruction set computer (OISC). Because SUBLEQ only has a single instruction, there is no overhead cost to hide what instruction is currently being executed as there is only one instruction to execute. The data which the instruction is executing on must remain private, but the instruction itself is always known. We also test and compare MPC SUBLEQ against GC-Lite, a similar garbled processor designed for PFE using the SUBBLE OISC, a weaker version of SUBLEQ. We show that MPC SUBLEQ dramatically outperforms GC-Lite due to the significantly lower local computation and online communication costs.","author":[{"family":"Jiang","given":"Christopher"}],"issued":{"date-parts":[[2025]]},"DOI":"10.11575/prism/47960","URL":"https://doi.org/10.11575/prism/47960","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.00104","type":"manuscript","title":"Enhanced Rényi Entropy-Based Post-Quantum Key Agreement with Provable Security and Information-Theoretic Guarantees","abstract":"This paper presents an enhanced post-quantum key agreement protocol based on Rényi entropy, addressing vulnerabilities in the original construction while preserving information-theoretic security properties. We develop a theoretical framework leveraging entropy-preserving operations and secret-shared verification to achieve provable security against quantum adversaries. Through entropy amplification techniques and quantum-resistant commitments, the protocol establishes $2^{128}$ quantum security guarantees under the quantum random oracle model. Key innovations include a confidentiality-preserving verification mechanism using distributed polynomial commitments, tightened min-entropy bounds with guaranteed non-negativity, and composable security proofs in the quantum universal composability framework. Unlike computational approaches, our method provides information-theoretic security without hardness assumptions while maintaining polynomial complexity. Theoretical analysis demonstrates resilience against known quantum attack vectors, including Grover-accelerated brute force and quantum memory attacks. The protocol achieves parameterization for 128-bit quantum security with efficient $\\mathcal{O}(n^{2})$ communication complexity. Extensions to secure multiparty computation and quantum network applications are established, providing a foundation for long-term cryptographic security.","author":[{"family":"Xu","given":"Ruopengyu"},{"family":"Liu","given":"Chenglian"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.00104","URL":"https://doi.org/10.48550/arxiv.2509.00104","source":"datacite"},{"id":"doi:10.5281/zenodo.19510539","type":"article-journal","title":"A Landscape Classification Framework for NP-Hard Problems: From Reconnaissance to Algorithm Selection","abstract":"We propose a unified classification framework for NP-hard problems based on loss landscape topology. The framework classifies 95% of known NP problems into seven boxes (combinatorial optimisation, graph theory, logical satisfaction, path/network, permutation/assignment, sequential decision, miscellaneous), further divided into 24 sub-boxes each with formal objective functions. A three-tier reconnaissance system (aerial survey, scout sampling, mass sampling) analyses landscape topology before algorithm selection. A five-cut decision tree maps landscape properties to optimal solvers and parameter configurations. All 22 sub-boxes are exhaustively enumerated with default algorithms, scout-based adjustments, and landscape-based parameter settings. Box 6 (sequential decision) is fully validated using MCCO as proof of concept, demonstrating 3.56× speedup via landscape-guided deployment. The framework draws an analogy to database normalisation: just as complex data can be decomposed through finite normal forms, complex NP problems can be classified through finite landscape cuts. ORCID: 0009-0002-9497-1336. v2 (2026-04-11): Expanded Related Work with six existing frameworks (Garey & Johnson 1979, FLA/ELA, Algorithm Selection, No Free Lunch, Parameterized Complexity, Learning-Augmented); added cross-framework comparison table; references expanded from 12 to 16. v3 (2026-04-11): Added Scout-Based Algorithm Adjustment and Landscape-Based Parameter Configuration (22 sub-boxes exhaustively enumerated with reconnaissance-based overrides and full parameter lookup tables); Landscape Transformation section (8 transformation methods, 4 difficulty scores, 4-layer stopping conditions with marginal benefit and ROI analysis, precision recovery); Statistical Foundations of Reconnaissance (Type I/II error rates, Power Analysis for scout count, Bonferroni correction, effect size estimation, GP uncertainty propagation); AI-Executable Protocol discussion; three-layer acceleration conclusion (reconnaissance × transformation × AI execution); TSP worked example with three-way comparison (brute force vs blind default vs framework-guided, 0.006s solve time); Scaling test (10/20/50 cities) with \"How Close to P?\" comparison table (framework at 10⁶ vs brute force at 10⁶⁴); Limitations expanded to 8 items. v4 (2026-04-12): Three new chapters: Landscape Cutting (Decomposition): cutting principles, five cutting tools, six-box cuttability table, transform-first-then-cut ordering, parallel deployment pipeline with global refinement, five cutting considerations. The Evaluation Matrix: 9 transformations × 4 cuts = 36-cell exhaustive evaluation; empirical validation on 100-city TSP (12 cells in 3.37s); cutting quality check with exhaustibility guarantee. Algorithm Adequacy Scoring: scale × difficulty → four-tier minimum tool level. Cross-box empirical validation (six boxes, all new): Box 1b TSP: TSPLIB standard benchmarks (eil51/berlin52/kroA100), equal-time-budget fair comparison, framework wins by 4–8% gap reduction. Box 1a MKP: Multidimensional knapsack (50–500 items, 3–10 constraints), framework wins by 1.4–6.9%, Chu & Beasley (1998) reference added. Box 2a Graph Coloring: 50–200 nodes, density 0.1–0.5, framework saves 14–30% colors via DSatur + Tabu Search. Box 3a 3-SAT: tested across phase transition (α = 2.0–5.0), framework selects CDCL for hard instances, solves instances that blind WalkSAT cannot, 15.1× speedup at α = 4.2. Box 4b VRP: 20–100 customers, framework serves +44 additional customers (coverage 56% → 100%), natural application of cutting mechanism. Box 5a Scheduling: 20–200 jobs, framework wins by 7–27%, 200-job instance within 0.1% of lower bound. How Close to P cross-box comparison table: five of eight test cases at or near P, three within 1–2 orders of magnitude, zero large gaps. Improvement acceleration data: 9% at 20 cities → 66% at 50 cities → 80% at 100 cities. Conclusion upgraded: three-layer → four-layer acceleration (reconnaissance × transfor","author":[{"family":"Rao","given":"Huiying"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19510539","URL":"https://doi.org/10.5281/zenodo.19510539","source":"datacite"},{"id":"doi:10.5281/zenodo.22010778","type":"article-journal","title":"Sānér, The Recursive Trinity Structure for Task Regions","abstract":"Abstract Sānér establishes a recursive control architecture for task regions by placing topology, self-modeling scheduling, and certificate-governed execution under one executable resource semantics. The nine-node construction supports exact recursive expansion and addressing. Its formal system establishes capacity safety, deterministic cross-layer selection, partition safety, conditional liveness, and computable deadline and recovery bounds. Using entity-isolated validation and one-time test splits, we evaluated the frozen controller in executable state-transition simulations and public benchmark workloads over 32, 48, or 64 independent worlds, with deployment evidence obtained from 48 physical GPU blocks and eight fresh control-plane processes in the official Kubernetes scheduler-performance harness. The sealed confirmatory suite established 38 task-level all-comparator superiority conclusions. On the official Kubernetes v1.36.2 scheduler-performance harness, Sānér achieved a 392.27× mean throughput ratio on the same host and reduced mean P99 by 26.75240 milliseconds. Exact recursive-topology representation remained 144 bytes through 26,244 nodes. Mechanism evidence remains separate from direct performance claims. Boundary results and nonidentifiability tests are reported independently. Together, the proofs and matched comparisons establish Sānér as a reusable and auditable control kernel whose single recursive mechanism transfers across structurally different systems tasks. Validation design Every direct comparator received the same random world, observable state, feasible action set, resource ceiling, and decision-time budget. Frozen random worlds coupled arrivals, failures, delays, partitions, and disturbances across methods. Interface adapters translated actions only and could neither expose future state nor optimize on a method’s behalf. Validation data selected candidates and parameters. Test data were unsealed once, only after protocol, program, and model hashes agreed. Each task had one frozen primary endpoint, and higher values were uniformly preferred. Confidence intervals resampled complete independent worlds rather than repeated observations within one world. When a suite defined a primary Holm family, candidate–comparator probability values were adjusted before the task-level conjunction was evaluated. A task-level all-comparator conclusion required every applicable constituent comparison to pass together with its programmed direction, confidence-bound, safety, feasibility, and deadline conditions. Suites without a declared primary Holm family applied their programmed constituent tests directly. Secondary endpoints used the suite-declared Holm or false-discovery-rate procedure. The paper reports 54 tasks. Four suite programs preserve broader Holm families containing 108 primary candidate–comparator tests across 36 tasks. Eighteen paper tasks contribute 54 of those tests. Eighteen additional suite tasks contribute 54 frozen raw probability values solely to preserve the original conservative adjustment denominator. Their workloads, scores, effects, intervals, timings, and conclusions support no manuscript claim. Relative effects equal the absolute candidate–comparator difference divided by the absolute comparator mean. When a comparator mean is negative, the percentage describes only the difference relative to its numerical magnitude; substantive interpretation rests on the absolute difference and its confidence interval. Comparator qualification and strength Comparator eligibility and tuning were frozen before test-set access. Published guarantees supplied theoretical strength. Official or author-maintained implementations supplied operational relevance. A rolling optimizer qualified only when it received the same state, constraints, action space, resource ceiling, and computation limit as Sānér. Validation selected the strongest deployable method wherever a suite required one. Simple baselines measured task diff","author":[{"family":"Bsmpx"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22010778","URL":"https://doi.org/10.5281/zenodo.22010778","source":"datacite"},{"id":"doi:10.5281/zenodo.22010777","type":"article-journal","title":"Sānér, The Recursive Trinity Structure for Task Regions","abstract":"Abstract Sānér establishes a recursive control architecture for task regions by placing topology, self-modeling scheduling, and certificate-governed execution under one executable resource semantics. The nine-node construction supports exact recursive expansion and addressing. Its formal system establishes capacity safety, deterministic cross-layer selection, partition safety, conditional liveness, and computable deadline and recovery bounds. Using entity-isolated validation and one-time test splits, we evaluated the frozen controller in executable state-transition simulations and public benchmark workloads over 32, 48, or 64 independent worlds, with deployment evidence obtained from 48 physical GPU blocks and eight fresh control-plane processes in the official Kubernetes scheduler-performance harness. The sealed confirmatory suite established 38 task-level all-comparator superiority conclusions. On the official Kubernetes v1.36.2 scheduler-performance harness, Sānér achieved a 392.27× mean throughput ratio on the same host and reduced mean P99 by 26.75240 milliseconds. Exact recursive-topology representation remained 144 bytes through 26,244 nodes. Mechanism evidence remains separate from direct performance claims. Boundary results and nonidentifiability tests are reported independently. Together, the proofs and matched comparisons establish Sānér as a reusable and auditable control kernel whose single recursive mechanism transfers across structurally different systems tasks. Validation design Every direct comparator received the same random world, observable state, feasible action set, resource ceiling, and decision-time budget. Frozen random worlds coupled arrivals, failures, delays, partitions, and disturbances across methods. Interface adapters translated actions only and could neither expose future state nor optimize on a method’s behalf. Validation data selected candidates and parameters. Test data were unsealed once, only after protocol, program, and model hashes agreed. Each task had one frozen primary endpoint, and higher values were uniformly preferred. Confidence intervals resampled complete independent worlds rather than repeated observations within one world. When a suite defined a primary Holm family, candidate–comparator probability values were adjusted before the task-level conjunction was evaluated. A task-level all-comparator conclusion required every applicable constituent comparison to pass together with its programmed direction, confidence-bound, safety, feasibility, and deadline conditions. Suites without a declared primary Holm family applied their programmed constituent tests directly. Secondary endpoints used the suite-declared Holm or false-discovery-rate procedure. The paper reports 54 tasks. Four suite programs preserve broader Holm families containing 108 primary candidate–comparator tests across 36 tasks. Eighteen paper tasks contribute 54 of those tests. Eighteen additional suite tasks contribute 54 frozen raw probability values solely to preserve the original conservative adjustment denominator. Their workloads, scores, effects, intervals, timings, and conclusions support no manuscript claim. Relative effects equal the absolute candidate–comparator difference divided by the absolute comparator mean. When a comparator mean is negative, the percentage describes only the difference relative to its numerical magnitude; substantive interpretation rests on the absolute difference and its confidence interval. Comparator qualification and strength Comparator eligibility and tuning were frozen before test-set access. Published guarantees supplied theoretical strength. Official or author-maintained implementations supplied operational relevance. A rolling optimizer qualified only when it received the same state, constraints, action space, resource ceiling, and computation limit as Sānér. Validation selected the strongest deployable method wherever a suite required one. Simple baselines measured task diff","author":[{"family":"Bsmpx"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22010777","URL":"https://doi.org/10.5281/zenodo.22010777","source":"datacite"},{"id":"doi:10.5281/zenodo.21955289","type":"article-journal","title":"Does Heavy-Tailed Gradient Noise Explain the CIFAR-10 Gap in Byzantine-Robust, Differentially Private Federated Learning? An Empirical Stress Test of Byz-Clip21-SGD2M","abstract":"Byz-Clip21-SGD2M (Islamov et al., 2026) provides high-probability convergence guarantees for federated learning under simultaneous Byzantine adversaries and differential-privacy (DP) noise, relaxing the bounded-gradient assumption behind the tight privacy-robustness-utility trade-off of Allouah et al. (2023b) to standard L-smoothness and σ-sub-Gaussian gradient noise. Its own empirical validation is MNIST-only, and its Conclusion lists heavy-tailed gradient noise as future work; we stress-test both the algorithm and its sub-Gaussian premise at CIFAR-10 scale. Using the established Hill-estimator/kurtosis tail-index methodology of Şimşekli et al. (2019) and follow-ons — applied here, not proposed as new — we find CIFAR-10's gradient noise consistently heavier-tailed than MNIST's across every statistic and configuration checked. Isolating ablations show the Byzantine-robustness mechanism itself performs comparably on both datasets, so the degradation is not robustness-specific; a dedicated DP-aware search over the clipping threshold τ, spanning the source paper's own privacy-budget grid, fails to recover non-degenerate CIFAR-10 accuracy even as DP noise is driven toward zero. This rules out both \"DP noise alone explains the gap\" and \"the fixed-τ protocol was miscalibrated\"; the evidence instead points to CIFAR-10's own clean-training convergence difficulty, plausibly linked to its heavier tail, as the dominant factor. A seed scale-up (CIFAR-10 sweep to n=10/cell, ablation to n=10/arm) with paired Wilcoxon testing confirms this: no CIFAR-10 condition differs significantly from any other, and the ablation's recovery to CIFAR-10's clean ceiling is statistically indistinguishable (p=1.000), not a small-sample artifact. An independent, unpaired Mann-Whitney U test further corroborates the hyperparameter-transfer gap under matched hyperparameters and round budget (U=100, p=0.000183). We also close this paper's previously most significant open limitation: we implement and unit-test the source paper's two external baselines, Safe-DSHB (Allouah et al., 2023b) and Byz-Clip-SGD (Islamov et al., 2026), from its appendix pseudocode, and run both under the identical protocol used for Byz-Clip21-SGD2M throughout. Under shared, non-independently-tuned hyperparameters, both baselines match or nominally exceed Byz-Clip21-SGD2M on several MNIST conditions — an exploratory finding, not confirmatory (the source paper's own independently-tuned comparison reaches the opposite conclusion, and none of these comparisons survive multiple-comparison correction). We then ran the independently-tuned comparison this gap called for (30-point (γ, τ) grid per baseline, matching the source paper's own tuning range): it erases the MNIST advantage seen under shared hyperparameters entirely (all 12 comparisons now p ≥ 0.06) without producing a Byz-Clip21-SGD2M advantage either, since our tuning remains a simplified, single-condition probe rather than the source paper's per-ε, DP-aware protocol — so whether Byz-Clip21-SGD2M has a genuine MNIST edge under equally careful tuning stays open. On CIFAR-10, all three algorithms are statistically indistinguishable and collapse to chance together, a more direct finding independent of any tuning caveat: the CIFAR-10 gap is not specific to Byz-Clip21-SGD2M. We report every known gap in our replication honestly, including a corrected theoretical positioning relative to Allouah et al.'s tight dimension-dependent lower bound and two corrections to our own earlier internal pilot analysis.","author":[{"family":"S Varughese","given":"Johan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21955289","URL":"https://doi.org/10.5281/zenodo.21955289","source":"datacite"},{"id":"doi:10.5281/zenodo.21955288","type":"article-journal","title":"Does Heavy-Tailed Gradient Noise Explain the CIFAR-10 Gap in Byzantine-Robust, Differentially Private Federated Learning? An Empirical Stress Test of Byz-Clip21-SGD2M","abstract":"Byz-Clip21-SGD2M (Islamov et al., 2026) provides high-probability convergence guarantees for federated learning under simultaneous Byzantine adversaries and differential-privacy (DP) noise, relaxing the bounded-gradient assumption behind the tight privacy-robustness-utility trade-off of Allouah et al. (2023b) to standard L-smoothness and σ-sub-Gaussian gradient noise. Its own empirical validation is MNIST-only, and its Conclusion lists heavy-tailed gradient noise as future work; we stress-test both the algorithm and its sub-Gaussian premise at CIFAR-10 scale. Using the established Hill-estimator/kurtosis tail-index methodology of Şimşekli et al. (2019) and follow-ons — applied here, not proposed as new — we find CIFAR-10's gradient noise consistently heavier-tailed than MNIST's across every statistic and configuration checked. Isolating ablations show the Byzantine-robustness mechanism itself performs comparably on both datasets, so the degradation is not robustness-specific; a dedicated DP-aware search over the clipping threshold τ, spanning the source paper's own privacy-budget grid, fails to recover non-degenerate CIFAR-10 accuracy even as DP noise is driven toward zero. This rules out both \"DP noise alone explains the gap\" and \"the fixed-τ protocol was miscalibrated\"; the evidence instead points to CIFAR-10's own clean-training convergence difficulty, plausibly linked to its heavier tail, as the dominant factor. A seed scale-up (CIFAR-10 sweep to n=10/cell, ablation to n=10/arm) with paired Wilcoxon testing confirms this: no CIFAR-10 condition differs significantly from any other, and the ablation's recovery to CIFAR-10's clean ceiling is statistically indistinguishable (p=1.000), not a small-sample artifact. An independent, unpaired Mann-Whitney U test further corroborates the hyperparameter-transfer gap under matched hyperparameters and round budget (U=100, p=0.000183). We also close this paper's previously most significant open limitation: we implement and unit-test the source paper's two external baselines, Safe-DSHB (Allouah et al., 2023b) and Byz-Clip-SGD (Islamov et al., 2026), from its appendix pseudocode, and run both under the identical protocol used for Byz-Clip21-SGD2M throughout. Under shared, non-independently-tuned hyperparameters, both baselines match or nominally exceed Byz-Clip21-SGD2M on several MNIST conditions — an exploratory finding, not confirmatory (the source paper's own independently-tuned comparison reaches the opposite conclusion, and none of these comparisons survive multiple-comparison correction). We then ran the independently-tuned comparison this gap called for (30-point (γ, τ) grid per baseline, matching the source paper's own tuning range): it erases the MNIST advantage seen under shared hyperparameters entirely (all 12 comparisons now p ≥ 0.06) without producing a Byz-Clip21-SGD2M advantage either, since our tuning remains a simplified, single-condition probe rather than the source paper's per-ε, DP-aware protocol — so whether Byz-Clip21-SGD2M has a genuine MNIST edge under equally careful tuning stays open. On CIFAR-10, all three algorithms are statistically indistinguishable and collapse to chance together, a more direct finding independent of any tuning caveat: the CIFAR-10 gap is not specific to Byz-Clip21-SGD2M. We report every known gap in our replication honestly, including a corrected theoretical positioning relative to Allouah et al.'s tight dimension-dependent lower bound and two corrections to our own earlier internal pilot analysis.","author":[{"family":"S Varughese","given":"Johan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21955288","URL":"https://doi.org/10.5281/zenodo.21955288","source":"datacite"},{"id":"doi:10.5281/zenodo.19712713","type":"article-journal","title":"Blockchain Solution with Artificial Intelligence Integration in the Indian Judicial System A Bibliometric and Methodical Literature Review","abstract":"The Indian judicial system faces critical challenges including case backlogs exceeding 54.7 million pending matters, insufficient judicial resources, and inefficient evidence management processes. This literature review examines the emerging potential of integrating blockchain technology with artificial intelligence (AI) to transform judicial delivery, enhance case processing efficiency, and strengthen evidentiary integrity. Through systematic analysis of 88 peer-reviewed publications (2013–2026) across IEEE Xplore, Scopus, Springer, and Web of Science, we identified four dominant research themes: blockchain-based evidence management systems, AI-driven predictive justice and decision support, smart contracts for judicial automation, and privacy-preserving mechanisms for sensitive legal data. Key findings reveal that blockchain ensures significant reduction in evidence tampering incidents, while AI prediction models achieve notable accuracy in judicial outcome forecasting. However, significant implementation challenges persist, including scalability constraints, lack of comprehensive regulatory frameworks, and insufficient integration with legacy court systems. This paper synthesizes current scholarship, identifies critical research gaps, and proposes a four-layer conceptual framework for pragmatic AI-blockchain deployment suited to India's constitutional and legal context. We conclude that strategic integration prioritizing permissioned blockchain architectures, explainable AI models, and federated learning offers transformative potential for addressing judicial inefficiency while maintaining due process and fundamental rights protections.","author":[{"family":"Patil","given":"Sameer"},{"family":"Desai","given":"Darshana"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19712713","URL":"https://doi.org/10.5281/zenodo.19712713","source":"datacite"},{"id":"doi:10.5281/zenodo.19712714","type":"article-journal","title":"Blockchain Solution with Artificial Intelligence Integration in the Indian Judicial System A Bibliometric and Methodical Literature Review","abstract":"The Indian judicial system faces critical challenges including case backlogs exceeding 54.7 million pending matters, insufficient judicial resources, and inefficient evidence management processes. This literature review examines the emerging potential of integrating blockchain technology with artificial intelligence (AI) to transform judicial delivery, enhance case processing efficiency, and strengthen evidentiary integrity. Through systematic analysis of 88 peer-reviewed publications (2013–2026) across IEEE Xplore, Scopus, Springer, and Web of Science, we identified four dominant research themes: blockchain-based evidence management systems, AI-driven predictive justice and decision support, smart contracts for judicial automation, and privacy-preserving mechanisms for sensitive legal data. Key findings reveal that blockchain ensures significant reduction in evidence tampering incidents, while AI prediction models achieve notable accuracy in judicial outcome forecasting. However, significant implementation challenges persist, including scalability constraints, lack of comprehensive regulatory frameworks, and insufficient integration with legacy court systems. This paper synthesizes current scholarship, identifies critical research gaps, and proposes a four-layer conceptual framework for pragmatic AI-blockchain deployment suited to India's constitutional and legal context. We conclude that strategic integration prioritizing permissioned blockchain architectures, explainable AI models, and federated learning offers transformative potential for addressing judicial inefficiency while maintaining due process and fundamental rights protections.","author":[{"family":"Patil","given":"Sameer"},{"family":"Desai","given":"Darshana"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19712714","URL":"https://doi.org/10.5281/zenodo.19712714","source":"datacite"},{"id":"doi:10.5281/zenodo.20500914","type":"article-journal","title":"Privacy-Preserving Operations on Spiral-Domain Encoded Time-Series States: Anonymization, Aggregation, and Federated Learning with Composition Rules and Streaming Variants","abstract":"v2 update (2026-06-01): Reproducibility ZIP added back to the latest version alongside the manuscript files, so that downloading from the concept DOI gives all materials in one place rather than requiring navigation to v1. This revised deposit contains the Paper 7 manuscript (PDF and DOCX) along with the supplementary reproducibility archive. The manuscript files were added in this revised version of the deposit; the supplementary ZIP file remains unchanged from the original deposit. Manuscript targets IEEE Transactions on Information Forensics and Security (under preparation). Coverage. 28 pre-registered studies spanning Phases XII-XX of the spiral-domain encoder validation campaign (privacy primitives foundation, edge-case stress, deployment realism, hardware context, composition + streaming, reviewer preemption, Tier-3 strengthening, plus surgical-RT latency). 84 hypotheses, 65 SUPPORTED (77%), 8 honest bounded negatives substantively interpreted. Substantive findings. Three architectural privacy primitives uniquely enabled by spiral encoder mathematical structure: Primitive A (per-subject angular phase-shift anonymization, k=10 cross-subject anonymity); Primitive B (cross-subject mean aggregation, 8.84 million-fold inversion resistance); Primitive C (federated AR(1) learning, exact 1-round convergence invariant to site count and heterogeneity). 5 deployment embodiments: multi-hospital clinical, multi-factory industrial, federated prosthesis fleet, surgical robotics RT privacy, cloud-scale parallel. Contents. Manuscript (PDF and DOCX of the paper itself); supplementary reproducibility archive containing: README.md (submission-package map and reproduction instructions); preregistrations/ (frozen pre-registration .md documents with literal-threshold decision rules); reports/ (per-study .md verdict reports against frozen rules + phase summaries); runners/ (deterministic Python runners under PYTHONHASHSEED=0); raw_data/ (per-study CSV outputs and JSON verdict blocks); figures/ (manuscript figures at 300 DPI + figure-build script); code/ (encoder source code). Reproducibility. Full validation pipeline is reproducible end-to-end under PYTHONHASHSEED=0 on a standard Python 3.9+ installation with NumPy 2.0+ and PyTorch 2.8+ (required for the learned-adversary autoencoder attack of Study 86). Reference machine: Apple Silicon arm64 (M-series), macOS 14. See README.md for per-study run commands. Methodological discipline. Every hypothesis was pre-registered with externally anchored decision rules frozen prior to runner execution. Zero post-hoc threshold adjustments were applied. Honest bounded negatives are interpreted substantively rather than discarded. Related companion archives. Paper 1 (10.5281/zenodo.20129137), Paper 2 (10.5281/zenodo.20138786), Paper 3 (10.5281/zenodo.20139171), and the corresponding Papers 4, 5, 6, 8 archives in this same Zenodo collection. Paper 8 (10.5281/zenodo.20466035) extends Paper 7 Composition III to musculoskeletal-kinematic clinical digital twin deployment with 21 additional studies and 91 hypotheses validated across CMU Motion Capture and KIMORE rehabilitation datasets including patient populations.","author":[{"family":"Ferlic","given":"Randolph"},{"family":"Ferlic","given":"Kimberly"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20500914","URL":"https://doi.org/10.5281/zenodo.20500914","source":"datacite"},{"id":"doi:10.5281/zenodo.20221641","type":"article-journal","title":"Privacy-Preserving Operations on Spiral-Domain Encoded Time-Series States: Anonymization, Aggregation, and Federated Learning with Composition Rules and Streaming Variants","abstract":"v2 update (2026-06-01): Reproducibility ZIP added back to the latest version alongside the manuscript files, so that downloading from the concept DOI gives all materials in one place rather than requiring navigation to v1. This revised deposit contains the Paper 7 manuscript (PDF and DOCX) along with the supplementary reproducibility archive. The manuscript files were added in this revised version of the deposit; the supplementary ZIP file remains unchanged from the original deposit. Manuscript targets IEEE Transactions on Information Forensics and Security (under preparation). Coverage. 28 pre-registered studies spanning Phases XII-XX of the spiral-domain encoder validation campaign (privacy primitives foundation, edge-case stress, deployment realism, hardware context, composition + streaming, reviewer preemption, Tier-3 strengthening, plus surgical-RT latency). 84 hypotheses, 65 SUPPORTED (77%), 8 honest bounded negatives substantively interpreted. Substantive findings. Three architectural privacy primitives uniquely enabled by spiral encoder mathematical structure: Primitive A (per-subject angular phase-shift anonymization, k=10 cross-subject anonymity); Primitive B (cross-subject mean aggregation, 8.84 million-fold inversion resistance); Primitive C (federated AR(1) learning, exact 1-round convergence invariant to site count and heterogeneity). 5 deployment embodiments: multi-hospital clinical, multi-factory industrial, federated prosthesis fleet, surgical robotics RT privacy, cloud-scale parallel. Contents. Manuscript (PDF and DOCX of the paper itself); supplementary reproducibility archive containing: README.md (submission-package map and reproduction instructions); preregistrations/ (frozen pre-registration .md documents with literal-threshold decision rules); reports/ (per-study .md verdict reports against frozen rules + phase summaries); runners/ (deterministic Python runners under PYTHONHASHSEED=0); raw_data/ (per-study CSV outputs and JSON verdict blocks); figures/ (manuscript figures at 300 DPI + figure-build script); code/ (encoder source code). Reproducibility. Full validation pipeline is reproducible end-to-end under PYTHONHASHSEED=0 on a standard Python 3.9+ installation with NumPy 2.0+ and PyTorch 2.8+ (required for the learned-adversary autoencoder attack of Study 86). Reference machine: Apple Silicon arm64 (M-series), macOS 14. See README.md for per-study run commands. Methodological discipline. Every hypothesis was pre-registered with externally anchored decision rules frozen prior to runner execution. Zero post-hoc threshold adjustments were applied. Honest bounded negatives are interpreted substantively rather than discarded. Related companion archives. Paper 1 (10.5281/zenodo.20129137), Paper 2 (10.5281/zenodo.20138786), Paper 3 (10.5281/zenodo.20139171), and the corresponding Papers 4, 5, 6, 8 archives in this same Zenodo collection. Paper 8 (10.5281/zenodo.20466035) extends Paper 7 Composition III to musculoskeletal-kinematic clinical digital twin deployment with 21 additional studies and 91 hypotheses validated across CMU Motion Capture and KIMORE rehabilitation datasets including patient populations.","author":[{"family":"Ferlic","given":"Randolph"},{"family":"Ferlic","given":"Kimberly"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20221641","URL":"https://doi.org/10.5281/zenodo.20221641","source":"datacite"},{"id":"doi:10.5281/zenodo.21313043","type":"article-journal","title":"A Systematic Literature Review on Explainable AI-Based Fish Quality Assessment Using Deep Learning","abstract":"Fish quality assessment is essential for ensuring food safety, maintaining consumer confidence, and supporting the global seafood industry. Recent advances in Artificial Intelligence (AI), particularly deep learning and Explainable Artificial Intelligence (XAI), have enabled accurate, rapid, and non-destructive evaluation of fish freshness and quality. This systematic literature review aims to provide a comprehensive analysis of AI-based fish quality assessment techniques, with a particular focus on explainable deep learning models and their applications in seafood inspection. The review was conducted following the PRISMA 2020 guidelines and included publications from 2019 to 2026 retrieved from Scopus, Web of Science, IEEE Xplore, ScienceDirect, SpringerLink, and the ACM Digital Library. Following the screening and eligibility process, 220 representative studies were included for qualitative synthesis. The reviewed literature covers Convolutional Neural Networks (CNNs), transfer learning, Vision Transformers (ViTs), hyperspectral imaging, thermal imaging, multimodal sensing, and Explainable AI techniques, including Grad-CAM, LIME, SHAP, saliency maps, and attention visualization. Comparative analysis indicates that deep learning models consistently achieve classification accuracies exceeding 90–95%, while XAI methods significantly improve model transparency, interpretability, and user trust in industrial decision-making. However, challenges such as limited benchmark datasets, poor cross-species generalization, computational complexity, and the absence of standardized explainability evaluation frameworks continue to hinder widespread industrial adoption. This review identifies current research trends, summarizes existing datasets and evaluation metrics, highlights critical research gaps, and proposes future directions, including Edge AI, federated learning, digital twins, multisensor fusion, and trustworthy XAI, to support the development of intelligent, transparent, and reliable fish quality assessment systems.","author":[{"family":"Dharmaraj","given":"Mr"},{"family":"Jacab","given":"Rev"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21313043","URL":"https://doi.org/10.5281/zenodo.21313043","source":"datacite"},{"id":"doi:10.5281/zenodo.21313044","type":"article-journal","title":"A Systematic Literature Review on Explainable AI-Based Fish Quality Assessment Using Deep Learning","abstract":"Fish quality assessment is essential for ensuring food safety, maintaining consumer confidence, and supporting the global seafood industry. Recent advances in Artificial Intelligence (AI), particularly deep learning and Explainable Artificial Intelligence (XAI), have enabled accurate, rapid, and non-destructive evaluation of fish freshness and quality. This systematic literature review aims to provide a comprehensive analysis of AI-based fish quality assessment techniques, with a particular focus on explainable deep learning models and their applications in seafood inspection. The review was conducted following the PRISMA 2020 guidelines and included publications from 2019 to 2026 retrieved from Scopus, Web of Science, IEEE Xplore, ScienceDirect, SpringerLink, and the ACM Digital Library. Following the screening and eligibility process, 220 representative studies were included for qualitative synthesis. The reviewed literature covers Convolutional Neural Networks (CNNs), transfer learning, Vision Transformers (ViTs), hyperspectral imaging, thermal imaging, multimodal sensing, and Explainable AI techniques, including Grad-CAM, LIME, SHAP, saliency maps, and attention visualization. Comparative analysis indicates that deep learning models consistently achieve classification accuracies exceeding 90–95%, while XAI methods significantly improve model transparency, interpretability, and user trust in industrial decision-making. However, challenges such as limited benchmark datasets, poor cross-species generalization, computational complexity, and the absence of standardized explainability evaluation frameworks continue to hinder widespread industrial adoption. This review identifies current research trends, summarizes existing datasets and evaluation metrics, highlights critical research gaps, and proposes future directions, including Edge AI, federated learning, digital twins, multisensor fusion, and trustworthy XAI, to support the development of intelligent, transparent, and reliable fish quality assessment systems.","author":[{"family":"Dharmaraj","given":"Mr"},{"family":"Jacab","given":"Rev"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21313044","URL":"https://doi.org/10.5281/zenodo.21313044","source":"datacite"},{"id":"doi:10.5281/zenodo.20649856","type":"article-journal","title":"CDSA-ATM: Air Traffic Management Pillar of the CDSA Federated Reinforcement Learning Paradigm — Multi-modal Trajectory + ATS + ANSP Fusion","abstract":"CDSA-ATM is the reference implementation of the air traffic management pillar of the Complementary Diagnostic Safety Approach (CDSA), a third-generation aviation safety paradigm formalised as a Federated Reinforcement Learning (FRL) framework (Scenario C decision, 21 May 2026). The v2.1.0 release (12 June 2026) replaces the v2.0.0 scaffold stubs with full NumPy reference implementations (federated, multimodal, rl_agents, rl_environment, xai; all unit-tested) and adds seeded, reproducible CPU-scale federated training runs (random / centralised PPO / FedAvg / FedProx / FedProx+DP) with result JSONs under experiments/results/. An interactive results panel is published at https://cdsa.app/atm. Framework-scale training with real operational data remains future work. The v1.0.0 release modules and reference data remain unchanged. Version 2.2.0 (3 July 2026) adds four phases of seeded, CPU-scale revision-preparation experiments (seed 42, FedProx unless noted) under experiments/, with all outputs in experiments/results/ and interactive honesty-band panels at https://cdsa.app/atm/. Faz A (reference runs): the high critical recall (0.919) is achieved inside the over-alert regime characterised in Faz B. Faz B (robustness and confusion matrix): the policy uses only 2 of 5 actions and every benign state receives a critical prediction — an over-alert regime; the tempo-aware reward does not improve on the static baseline. Faz C (differential-privacy ε sweep and entropy-targeted exploration): the ε sweep is two-sided (ε = 2.0 matches the undefended baseline, ε = 0.5 collapses the policy to a single action); a 0.05 entropy bonus recovers all five actions, but critical recall falls from 0.927 to 0.878. Faz D (decoy attribution and gated exploration): decoy attribution stays below the 25% uniform share (mean 15.3%, max 23.3%); gated exploration preserves recall (0.928) but the policy remains two-action. Findings are reported verbatim from the result JSONs, consistent with the honesty bands published at the site.","author":[{"family":"Cantekin","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20649856","URL":"https://doi.org/10.5281/zenodo.20649856","source":"datacite"},{"id":"doi:10.5281/zenodo.20649911","type":"article-journal","title":"CDSA-ATM: Air Traffic Management Pillar of the CDSA Federated Reinforcement Learning Paradigm — Multi-modal Trajectory + ATS + ANSP Fusion","abstract":"CDSA-ATM is the reference implementation of the air traffic management pillar of the Complementary Diagnostic Safety Approach (CDSA), a third-generation aviation safety paradigm formalised as a Federated Reinforcement Learning (FRL) framework (Scenario C decision, 21 May 2026). The v2.1.0 release (12 June 2026) replaces the v2.0.0 scaffold stubs with full NumPy reference implementations (federated, multimodal, rl_agents, rl_environment, xai; all unit-tested) and adds seeded, reproducible CPU-scale federated training runs (random / centralised PPO / FedAvg / FedProx / FedProx+DP) with result JSONs under experiments/results/. An interactive results panel is published at https://cdsa.app/atm. Framework-scale training with real data is planned under TÜBİTAK funding (Phase B). The v1.0.0 release modules and reference data remain unchanged.","author":[{"family":"Cantekin","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20649911","URL":"https://doi.org/10.5281/zenodo.20649911","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.29659","type":"manuscript","title":"GQ-FSL: Green Quantized Federated Split Learning Framework for Wireless Edge Networks","abstract":"Deploying state-of-the-art deep neural networks (DNNs) at the wireless edge is severely bottlenecked by the strict energy and resource constraints of mobile devices. Although federated split learning (FSL) alleviates on-device computational burdens by offloading workloads to an edge server, this may introduce systemic overheads, while the continuous exchange of intermediate activations, gradients, and submodels still incurs significant energy consumption (EC). To address this, we propose a green quantized FSL (GQ-FSL) framework that incorporates stochastic quantization for both local collaborative training and wireless transmissions. Notably, GQ-FSL supports asymmetric precision levels for the client- and server-side submodels, effectively decoupling device energy constraints from global convergence degradation. To quantify these tradeoffs, we develop parameterized energy models for the split architecture and derive a theoretical convergence bound under statistically heterogeneous data. Building on that, we formulate a joint optimization problem to configure the DNN split point and precision levels, minimizing the total system EC while satisfying strict latency and target accuracy constraints. Ultimately, we demonstrate that GQ-FSL enables large-scale DNN deployment on resource-constrained devices, achieving superior energy efficiency compared to quantized federated learning and full-precision FSL.","author":[{"family":"Roth","given":"Idan"},{"family":"Lampe","given":"Lutz"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.29659","URL":"https://doi.org/10.48550/arxiv.2607.29659","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.17069","type":"manuscript","title":"Dynamic Entanglement-Weighted Pruning for Quantum Federated Unlearning in Supply-Chain Risk Prediction","abstract":"Federated deployments of variational quantum classifiers are attractive for cross-organisation risk prediction in supply chains, because raw data never leaves the client, yet data-protection regulations such as the GDPR grant clients a right to request that their contribution be removed from a trained model after the fact. Retraining a federated model from scratch to honour such a request is correct but wasteful, and it is not obvious which quantum circuit parameters actually carry a given client's influence. We introduce Entanglement-Weighted Pruning (EWP), an unlearning procedure for quantum federated learning that scores every trainable circuit parameter with the product of two signals: the diagonal entry of the quantum Fisher information matrix estimated on the target client's data via the parameter-shift rule, and a structural entanglement weight associated with the parameter's gate. Parameters with the lowest scores are pruned, optionally followed by a short fine-tuning pass on the retained clients. We implement the full pipeline in Qiskit for a four-qubit data-re-uploading ansatz trained with FedAvg across five simulated supply-chain-risk clients, and benchmark EWP against full retraining, fine-tuning alone, random pruning, Fisher-only pruning, and entanglement-only pruning, over three random seeds. EWP attains a mean post-unlearning accuracy statistically indistinguishable from the full-retraining oracle, while producing a lower forgetting score and requiring roughly 16 times less wall-clock time. Ablations over pruning threshold, client count, and non-IID strength show that combining the two signals is necessary, as entanglement-only and Fisher-only pruning each substantially degrade accuracy relative to EWP.","author":[{"family":"Kumar","given":"Aditya"},{"family":"Chongder","given":"Sumit"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.17069","URL":"https://doi.org/10.48550/arxiv.2608.17069","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.13591","type":"manuscript","title":"FuSeFL: Fully Secure and Scalable Federated Learning","abstract":"Federated Learning (FL) enables collaborative model training without centralizing client data, making it attractive for privacy-sensitive domains. While existing approaches employ cryptographic techniques such as homomorphic encryption, differential privacy, or secure multiparty computation to mitigate inference attacks, including model inversion, membership inference, and gradient leakage, they often suffer from high computational and memory overheads. Moreover, many methods overlook the confidentiality of the global model itself, which may be proprietary and sensitive. These challenges limit the practicality of secure FL, especially in settings that involve large datasets and strict compliance requirements. We present FuSeFL, a Fully Secure and scalable FL scheme, which decentralizes training across client pairs using lightweight MPC, while confining the server's role to secure aggregation, client pairing, and routing. This design eliminates server bottlenecks, avoids full data offloading, and preserves full confidentiality of data, model, and updates throughout training. Based on our experiment, FuSeFL defends against unauthorized observation, reconstruction attacks, and inference attacks such as gradient leakage, membership inference, and inversion attacks, while achieving up to $13 \\times$ speedup in training time and 50% lower server memory usage compared to our baseline.","author":[{"family":"Ghinani","given":"Sahar"},{"family":"Sadredini","given":"Elaheh"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.13591","URL":"https://doi.org/10.48550/arxiv.2507.13591","source":"datacite"},{"id":"doi:10.5445/ir/1000189670","type":"article-journal","title":"TEE-Based Distributed Ledgers and Their Resilience","abstract":"Resilience is the ability of a (distributed) system to withstand any stressful situation without imposing massive restrictions and, above all, without long-term consequences. Permissioned distributed ledgers based on state machine replication (SMR) offer a promising approach to achieving high resilience and fairness in federated systems. SMR provides a fault-tolerant service for clients by relying on all replicas being in a consistent state. The consistent state is achieved through a consensus algorithm, typically an atomic broadcast, that decides on a total order of client requests. In the Byzantine fault model, replicas are assumed to be potentially malicious; a Byzantine fault-tolerant (BFT) protocol withstands a fixed share of malicious actors. Classic BFT SMR protocols require $n&gt;3t$ replicas and multiple rounds of communication to withstand $t$ faulty replicas, making the implementation complex and limiting achievable throughput and increasing latency. Trusted Execution Environments (TEEs) allow to implement SMR in the so-called hybrid fault model in which replicas are assumed to be potentially Byzantine but the TEE is restricted to only fail by crashing. In the hybrid fault model, SMR requires less communication and can be implemented with a fault tolerance of $n&gt;2t$ replicas. While many proposals aim to optimize BFT SMR by using TEEs, they still rely on a so-called leader that coordinates the agreement process among the replicas. The leader is known to be a bottleneck and, if it fails, the system has to recover from the failure and elect a new leader. The additional coordination required to elect a new leader can cause significant performance degradation, limiting the achieved resilience. Asynchronous protocols based on directed acyclic graphs (DAGs) eliminate the reliance on distinguished replicas by allowing all replicas to participate equally in the agreement process. While asynchronous approaches and the hybrid fault model independently contribute to increasing the resilience of BFT SMR systems, their combination has largely been unexplored. This dissertation aims to fill this gap by answering the following research question: What is the achievable performance and resilience of DAG-based, hybrid fault-tolerant state machine replication and under which preconditions can the leaderless nature be safely exploited to maximize throughput? We proceed in three steps to enhance the resilience and performance of BFT SMR systems and to identify potential trade-offs that arise from the assumption of TEEs and asynchrony in BFT SMR. First, we investigate the fit of TEE-based SMR for consortium-operated applications using the example of Mobility-as-a-Service ticketing systems. We propose an SMR application that uses TEEs to protect sensitive customer and mobility provider data while limiting possibilities for fraud by both customers and mobility providers, and ensuring correct billing. We find that as long as secure multiparty computation is not competitive in terms of performance, TEE-based SMR can provide significant advantages in terms of efficiency and resilience while providing reasonable confidentiality guarantees. We describe the characteristics of the Mobility-as-a-Service use case and identify similar use cases from other domains, e.g., central bank digital currencies, allowing us to conclude that our findings generalize. In the second step, we establish the foundation for a comprehensive analysis by proposing and proving TEE-Rider, the first hybrid fault-tolerant, asynchronous, and DAG-based atomic broadcast protocol. TEE-Rider builds upon the DAG-Rider protocol family and an optimized, DAG-aware, and TEE-based causal order broadcast we propose and prove. We then identify fundamental issues that arise from the combination of TEEs and asynchrony in BFT SMR. These are the impossibility of a fault-tolerant setup and the impossibility of garbage collection. Furthermore, we prove that for partially synchronous, TEE-ba","author":[{"family":"Leinweber","given":"Marc"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5445/ir/1000189670","URL":"https://doi.org/10.5445/ir/1000189670","source":"datacite"},{"id":"doi:10.34726/hss.2025.111502","type":"article-journal","title":"3D-based contact-less fingerprint acquisition","abstract":"Fingerprint recognition is a widely used biometric modality because of its uniqueness and persistence of friction ridge patterns, which provide a reliable means of verifying identity. It is a key technology in applications ranging from law enforcement and border control to securing personal devices and financial transactions. While traditionally acquired through direct contact, contactless methods offer advantages in hygiene and user convenience. However, practical deployment is challenged by the unconstrained nature of the acquisition process, which introduces variations in finger pose, illumination, and scale, while still requiring to be interoperability with large, legacy contact-based fingerprint databases. This thesis presents a set of algorithms and methodologies to address these problems across the recognition pipeline, from image capture to secure template comparison. The contributions include solutions for image normalization, data assurance, robust validation, and secure deployment, which improve the accuracy, reliability, and privacy of contactless fingerprint systems.A necessary step in any contactless pipeline is the accurate segmentation of the fingertip from its background, which is often complex and variable. This work proposes three novel deep learning architectures for this task. The first is a custom U-Net-based model for pre-cropped single-finger images that outperforms existing segmentation models [217]. This was improved with FingerUNeSt++, which combines a ResNeSt encoder with a UNet++-like decoder, achieving a mean Intersection-over-Union (mIoU) of 99% on the test set. The third model, TipSegNet, removes the need for a separate finger detection step by segmenting and labeling all four fingertips directly from a whole-hand image. Using a ResNeXt-101 backbone with a Feature Pyramid Network (FPN) to handle multi-scale objects, TipSegNet obtains an mIoU of 99% and an accuracy of 100%.Interoperability with contact-based systems requires correcting the geometric distortions in contactless captures. This research developed a processing pipeline to correct for in-plane (yaw) and out-of-plane (roll) rotations, which then flattens the fingertip texture using parametric unwarping models. The pipeline uses the segmentation mask for yaw correction and an elliptical finger model with the detected core for roll correction. On an operational dataset, a finger-wise optimized application of the pipeline reduced the Equal Error Rate (EER) for contactless-to-contact-based comparison by a relative 36.9% (from 1.57% to 0.99%). A large-scale empirical analysis of the fingerprint core’s position across over 40,000 samples showed that its location is not geometrically centered, with systematic, modality-induced biases and a natural variability of 6-12% of the finger’s width. This study quantifies a limit on the accuracy of alignment methods that rely on the core and identifies the Non-Central Fischer (NCF) distribution as the best-fitting statistical model for its position for most fingers, including a finger-dependent analysis.The quality of a captured sample affects recognition performance. This work addresses the absence of dedicated quality metrics for mobile contactless fingerprints by adapting the established NFIQ 2 framework, resulting in MCLFIQ. By retraining the NFIQ 2 random forest classifier on modality-specific synthetic data, MCLFIQ shows improved performance in predicting the utility of contactless samples compared to the original NFIQ 2.2 and other baselines. The new model prioritizes features related to image sharpness and local ridge clarity, which are key quality factors in mobile captures. Additionally, this thesis introduces a self-supervised framework to detect structural artifacts from fingerprint mosaicking, a problem not handled by standard quality metrics. A deep learning model was trained on programmatically generated artifacts to detect these defects without manual annotation. The resulting detector i","author":[{"family":"Ruzicka","given":"Laurenz"}],"issued":{"date-parts":[[2025]]},"DOI":"10.34726/hss.2025.111502","URL":"https://doi.org/10.34726/hss.2025.111502","source":"datacite"},{"id":"doi:10.48550/arxiv.2502.20835","type":"manuscript","title":"Federated Distributed Key Generation","abstract":"Distributed Key Generation (DKG) underpins threshold cryptography in many systems, including decentralized wallets, validator key ceremonies, cross-chain bridges, threshold signatures, secure multiparty computation, and internet voting. Classical ($t$,$n$)-DKG assumes a fixed group of n parties and a global threshold $t$, requiring full and timely participation. When actual participation deviates, the setup must abort or restart, which is impractical in open or time-critical environments where $n$ is large and availability unpredictable. We introduce Federated Distributed Key Generation (FDKG), inspired by Federated Byzantine Agreement, that makes participation optional and trust heterogeneous. Each participant selects a personal guardian set $G_i$ of size $k$ and a local threshold $t$. Its partial secret can later be reconstructed either by itself or by any t of its guardians. FDKG generalizes PVSS-based DKG and completes both generation and reconstruction in a single broadcast round each, with total communication proportional to $n k$ and at most $O(n^2)$ for reconstruction. Our analysis shows that (i) generation ensures correctness, privacy, and robustness under standard PVSS-based DKG assumptions, and (ii) reconstruction provides liveness and privacy characterized by the guardian-set topology {$G_i$}. Liveness holds if no participant $i$ is corrupted together with at least $k-t+1$ of its guardians. Conversely, privacy is preserved unless the corrupted subset is itself reconstruction-capable.","author":[{"family":"Baranski","given":"Stanislaw"},{"family":"Szymanski","given":"Julian"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2502.20835","URL":"https://doi.org/10.48550/arxiv.2502.20835","source":"datacite"},{"id":"doi:10.5281/zenodo.17579904","type":"article-journal","title":"GDI D8.6 - Report on privacy-enhancing solutions","abstract":"GDI Pillar III aims to explore use cases and innovative applications for analysing genomic and clinical data, ideally supported by the infrastructure being deployed at the national nodes within Pillar II. As described in the Report on federated learning technologies (Deliverable D8.5), artificial intelligence techniques and – more specifically - federated learning is a recent trend that can be observed in this area. Federated learning is a distributed machine learning technique in which multiple participants, which provide remote devices or siloed data centres, collaboratively train a shared machine learning model while keeping their data locally, better supporting data privacy. It enables collaborative learning from distributed data sources without sharing the original data, thus reducing privacy concerns and leveraging the aggregate knowledge available to the multiple participants. Privacy-Enhancing Technologies (PETs) enhance the privacy of data that is being collected or processed, e.g. by using access control and consent management systems or using techniques for data anonymization or pseudonymization. Privacy-Preserving Technologies (PPTs) are very useful in a distributed setting with multiple parties because they complement and strengthen Privacy-Enhancing Technologies (PETs) by addressing some of the core challenges of multi-party collaboration. They minimize trust requirements by allowing parties to collaborate without needing to share raw data. PPTs also reduce the amount of data that needs to be shared or centralized, which lowers the risk of breaches or misuse. PPTs can also protect intermediate results of PETs. The Report on federated learning technologies (Deliverable D8.5) focusses on the most common form of federated learning, where the model is trained locally by each participant on its own data and model updates are sent to a central server. The central server then aggregates these updates to improve the global model, which is then sent back to the participants for further iterative training rounds. This enhances privacy-preservation to some extent. But other privacy-preserving technologies exist. In this report, three other PPTs are assessed: Differential Privacy, Homomorphic Encryption and Secure Multiparty Computation. Three domains of use cases that can benefit from privacy-preserving technologies are considered: Genome-Wide Association Studies, Rare Diseases and Federated Machine Learning. Genome-Wide Association Studies are a wide class of applications for analysing genomic and clinical data that should be supported by the infrastructure being deployed at the national nodes in a federated manner. The Use case demonstrator package (Deliverable D7.1) describes the more specific scenario (infectious diseases use case 1) to identify variants determining the severity of COVID-19 disease progression. This specific scenario uses the GDI Infectious Diseases data, which the 1+MG WG11 is currently generating. The Rare Diseases use case also covers a wide range of genomic applications. The Use case demonstrator package (Deliverable D7.1) describes the more specific scenario to answer the question regarding the side effects of medications caused by some gene variants, using the B1MG Rare Diseases dataset. The Federated Machine Learning use case is where the Report on federated learning technologies (Deliverable D8.5) focuses on. The described privacy-preserving technologies can be used in combination with federated machine learning, enhancing privacy. Federated machine learning can be used in a very wide range of machine learning genomic applications. The Use case demonstrator package (Deliverable D7.1) and Report on federated learning technologies (Deliverable D8.5) describe some data-driven models for Cancer Research that can be built using federated machine learning.","author":[{"family":"Verachtert","given":"Wilfried"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17579904","URL":"https://doi.org/10.5281/zenodo.17579904","source":"datacite"},{"id":"doi:10.5281/zenodo.17579905","type":"article-journal","title":"GDI D8.6 - Report on privacy-enhancing solutions","abstract":"GDI Pillar III aims to explore use cases and innovative applications for analysing genomic and clinical data, ideally supported by the infrastructure being deployed at the national nodes within Pillar II. As described in the Report on federated learning technologies (Deliverable D8.5), artificial intelligence techniques and – more specifically - federated learning is a recent trend that can be observed in this area. Federated learning is a distributed machine learning technique in which multiple participants, which provide remote devices or siloed data centres, collaboratively train a shared machine learning model while keeping their data locally, better supporting data privacy. It enables collaborative learning from distributed data sources without sharing the original data, thus reducing privacy concerns and leveraging the aggregate knowledge available to the multiple participants. Privacy-Enhancing Technologies (PETs) enhance the privacy of data that is being collected or processed, e.g. by using access control and consent management systems or using techniques for data anonymization or pseudonymization. Privacy-Preserving Technologies (PPTs) are very useful in a distributed setting with multiple parties because they complement and strengthen Privacy-Enhancing Technologies (PETs) by addressing some of the core challenges of multi-party collaboration. They minimize trust requirements by allowing parties to collaborate without needing to share raw data. PPTs also reduce the amount of data that needs to be shared or centralized, which lowers the risk of breaches or misuse. PPTs can also protect intermediate results of PETs. The Report on federated learning technologies (Deliverable D8.5) focusses on the most common form of federated learning, where the model is trained locally by each participant on its own data and model updates are sent to a central server. The central server then aggregates these updates to improve the global model, which is then sent back to the participants for further iterative training rounds. This enhances privacy-preservation to some extent. But other privacy-preserving technologies exist. In this report, three other PPTs are assessed: Differential Privacy, Homomorphic Encryption and Secure Multiparty Computation. Three domains of use cases that can benefit from privacy-preserving technologies are considered: Genome-Wide Association Studies, Rare Diseases and Federated Machine Learning. Genome-Wide Association Studies are a wide class of applications for analysing genomic and clinical data that should be supported by the infrastructure being deployed at the national nodes in a federated manner. The Use case demonstrator package (Deliverable D7.1) describes the more specific scenario (infectious diseases use case 1) to identify variants determining the severity of COVID-19 disease progression. This specific scenario uses the GDI Infectious Diseases data, which the 1+MG WG11 is currently generating. The Rare Diseases use case also covers a wide range of genomic applications. The Use case demonstrator package (Deliverable D7.1) describes the more specific scenario to answer the question regarding the side effects of medications caused by some gene variants, using the B1MG Rare Diseases dataset. The Federated Machine Learning use case is where the Report on federated learning technologies (Deliverable D8.5) focuses on. The described privacy-preserving technologies can be used in combination with federated machine learning, enhancing privacy. Federated machine learning can be used in a very wide range of machine learning genomic applications. The Use case demonstrator package (Deliverable D7.1) and Report on federated learning technologies (Deliverable D8.5) describe some data-driven models for Cancer Research that can be built using federated machine learning.","author":[{"family":"Verachtert","given":"Wilfried"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17579905","URL":"https://doi.org/10.5281/zenodo.17579905","source":"datacite"},{"id":"doi:10.5281/zenodo.17445479","type":"article-journal","title":"Funciones pseudoaleatorias inconscientes en grupos no conmutativos","abstract":"Las aplicaciones de las funciones pseudoaleatorias inconscientes en la criptografía y en la seguridad de la información son múltiples. Pueden citarse la derivación de claves basadas en contraseñas, acuerdo de claves basados en contraseñas, password hardening, CAPTCHAs imposibles de rastrear, acuerdo de claves homomórfico y la intersección de conjuntos segura. Los primeros trabajos se basan en protocolos para la transferencia inconsciente, computación multiparte segura o en algunas variantes del problema del logaritmo discreto. Recientemente han surgido propuestas postcuánticas basadas en las isogenias de curvas elípticas y en los problemas sobre lattices. En este trabajo se propone el diseño de una función pseudoaleatoria inconsciente que base su seguridad en la dificultad de encontrar el elemento conjugador en grupos no conmutativos. Se realiza además un experimento utilizando como plataforma el grupo discreto de Heisenberg sobre un campo finito.","author":[{"family":"Ledo Baster","given":"David"},{"family":"Martínez Rodríguez","given":"Huber"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17445479","URL":"https://doi.org/10.5281/zenodo.17445479","source":"datacite"},{"id":"doi:10.5281/zenodo.17445482","type":"article-journal","title":"Funciones pseudoaleatorias inconscientes en grupos no conmutativos","abstract":"Las aplicaciones de las funciones pseudoaleatorias inconscientes en la criptografía y en la seguridad de la información son múltiples. Pueden citarse la derivación de claves basadas en contraseñas, acuerdo de claves basados en contraseñas, password hardening, CAPTCHAs imposibles de rastrear, acuerdo de claves homomórfico y la intersección de conjuntos segura. Los primeros trabajos se basan en protocolos para la transferencia inconsciente, computación multiparte segura o en algunas variantes del problema del logaritmo discreto. Recientemente han surgido propuestas postcuánticas basadas en las isogenias de curvas elípticas y en los problemas sobre lattices. En este trabajo se propone el diseño de una función pseudoaleatoria inconsciente que base su seguridad en la dificultad de encontrar el elemento conjugador en grupos no conmutativos. Se realiza además un experimento utilizando como plataforma el grupo discreto de Heisenberg sobre un campo finito.","author":[{"family":"Ledo Baster","given":"David"},{"family":"Martínez Rodríguez","given":"Huber"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17445482","URL":"https://doi.org/10.5281/zenodo.17445482","source":"datacite"},{"id":"doi:10.5281/zenodo.18602469","type":"article-journal","title":"TaxFL: Federated Learning and Federated Graph Intelligence for Cross-Border Tax and AML Compliance","abstract":"Cross-border tax non-compliance and money laundering exploit the same structural weakness: the data needed to identify the beneficial owner of a suspicious structure is fragmented across jurisdictions and institutions. Traditional exchange-of-information (EOI) mechanisms are essential for targeted cases but cannot support the high-volume, pattern-based risk analytics modern compliance requires. This working paper presents TaxFL, a privacy-preserving reference architecture and governance blueprint that lets tax authorities and financial intelligence units collaboratively train and evaluate risk models without sharing raw taxpayer or transaction data — only model updates cross any boundary. TaxFL integrates cross-silo federated learning, federated graph neural networks (GNNs) for ownership and transaction networks, and layered privacy-enhancing technologies (secure aggregation, differential privacy, and optional confidential computing), wrapped in a conservative \"belt-and-suspenders\" legal blueprint viable even under strict interpretations of international data-transfer rules (e.g., Schrems II, FATF R40). The contribution is deliberately integrative and institutional rather than a new learning algorithm. Rather than claiming that federation is universally beneficial, the paper maps its utility frontier. On a non-IID graph benchmark (Cora) evaluated under a single held-out global test split, federated training — strongest with SCAFFOLD — recovers the discriminative signal that label-skewed local training loses and matches a centralized upper bound, correcting an earlier evaluation artifact. On synthetic, officially-calibrated tax/AML data, the gains are conditional and honestly reported: federation helps at scale (5–10 jurisdictions) and through graph structure over linear models, and SCAFFOLD removes the negative transfer that naive averaging causes, while aggregate gains over an already-strong single silo are modest. The use of synthetic data is intentional — it mirrors the operational constraint that raw taxpayer data cannot cross jurisdictional boundaries and enables fully reproducible peer review. The paper outlines priority use cases (cross-border beneficial ownership, VAT/GST carousel fraud, transfer pricing, crypto-asset risk scoring, and FIU typology sharing), a realistic six-month pilot design for a small multi-jurisdiction consortium starting from public/synthetic data, and an evaluation framework covering detection performance, privacy assurances, operational overhead, and institutional trust. TaxFL is a low-risk complement to existing EOI processes and a deployment-ready blueprint whose core thesis still requires a real-data pilot for validation. Open implementation: https://github.com/pafrantz/TaxFL. We welcome academic and institutional collaborators for rigorous empirical validation on real-world tax data under appropriate safeguards.","author":[{"family":"Frantz","given":"Pedro"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18602469","URL":"https://doi.org/10.5281/zenodo.18602469","source":"datacite"},{"id":"doi:10.5281/zenodo.20574158","type":"article-journal","title":"TaxFL: Federated Learning and Federated Graph Intelligence for Cross-Border Tax and AML Compliance","abstract":"Cross-border tax non-compliance and money laundering exploit the same structural weakness: the data needed to identify the beneficial owner of a suspicious structure is fragmented across jurisdictions and institutions. Traditional exchange-of-information (EOI) mechanisms are essential for targeted cases but cannot support the high-volume, pattern-based risk analytics modern compliance requires. This working paper presents TaxFL, a privacy-preserving reference architecture and governance blueprint that lets tax authorities and financial intelligence units collaboratively train and evaluate risk models without sharing raw taxpayer or transaction data — only model updates cross any boundary. TaxFL integrates cross-silo federated learning, federated graph neural networks (GNNs) for ownership and transaction networks, and layered privacy-enhancing technologies (secure aggregation, differential privacy, and optional confidential computing), wrapped in a conservative \"belt-and-suspenders\" legal blueprint viable even under strict interpretations of international data-transfer rules (e.g., Schrems II, FATF R40). The contribution is deliberately integrative and institutional rather than a new learning algorithm. Rather than claiming that federation is universally beneficial, the paper maps its utility frontier. On a non-IID graph benchmark (Cora) evaluated under a single held-out global test split, federated training — strongest with SCAFFOLD — recovers the discriminative signal that label-skewed local training loses and matches a centralized upper bound, correcting an earlier evaluation artifact. On synthetic, officially-calibrated tax/AML data, the gains are conditional and honestly reported: federation helps at scale (5–10 jurisdictions) and through graph structure over linear models, and SCAFFOLD removes the negative transfer that naive averaging causes, while aggregate gains over an already-strong single silo are modest. The use of synthetic data is intentional — it mirrors the operational constraint that raw taxpayer data cannot cross jurisdictional boundaries and enables fully reproducible peer review. The paper outlines priority use cases (cross-border beneficial ownership, VAT/GST carousel fraud, transfer pricing, crypto-asset risk scoring, and FIU typology sharing), a realistic six-month pilot design for a small multi-jurisdiction consortium starting from public/synthetic data, and an evaluation framework covering detection performance, privacy assurances, operational overhead, and institutional trust. TaxFL is a low-risk complement to existing EOI processes and a deployment-ready blueprint whose core thesis still requires a real-data pilot for validation. Open implementation: https://github.com/pafrantz/TaxFL. We welcome academic and institutional collaborators for rigorous empirical validation on real-world tax data under appropriate safeguards.","author":[{"family":"Frantz","given":"Pedro"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20574158","URL":"https://doi.org/10.5281/zenodo.20574158","source":"datacite"},{"id":"doi:10.5281/zenodo.20519586","type":"article-journal","title":"Silent Failures and Structural Gaps: A Cross-Domain Framework for Evaluation Rigor in Large-Model Systems","abstract":"Version 2 — revised in response to an external structural review and an automated critique pass. See \"Response to Review\" appendix in the PDF for the change log. Large-model systems are evaluated at multiple layers—statistical, algorithmic, systems, and behavioral—yet each layer's evaluation methodology has been developed largely in isolation. This paper identifies a shared structural pattern across these layers: evaluations that are locally valid but globally misleading. Drawing on recent preprints spanning federated learning, LLM inference, causal inference, high-dimensional statistics, latent reasoning, and multi-agent coherence, we offer a *heuristic reading* that a common pattern recurs: a system or estimator satisfies its local objective (loss reduction, throughput, oracle test passage, component-level coherence) while concealing a deeper failure that only manifests at a different scale, composition, or distribution shift. We call this the **local-validity trap**. Specifically, we synthesize evidence that (1) oracle-based testing may miss semantically incorrect but symptom-reducing fixes in AI-assisted development; (2) component-level probabilistic coherence does not guarantee joint coherence in multi-agent LLM systems; (3) measurement error silently biases high-dimensional regression even when penalized estimators converge; (4) asynchronous pipeline parallelism bounds staleness locally but may accumulate convergence error globally; and (5) contribution metrics in federated learning should track optimization trajectories rather than static snapshots to avoid misattribution. We propose that evaluation rigor for large-model systems may benefit from explicit cross-scale consistency checks, analogous to structural constraints studied in the algorithmic and statistical literatures, and outline candidate design principles for building such checks into system pipelines. The \"local-validity trap\" framing is introduced here as an organizing heuristic; it is not a term used in any cited source, and the cross-domain structural analogy is asserted on the basis of shared vocabulary rather than shared mechanism. --- Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted from arXiv preprint corpus on the date in the filename. Cited arXiv preprints: 2605.29566v1, 2605.29639v1, 2605.29664v1, 2605.29740v1, 2605.29944v1, 2605.30075v1, 2605.30113v1, 2605.30153v1, 2605.30158v1, 2605.30319v1, 2605.30321v1, 2605.30327v1, 2605.30335v1, 2605.30336v1, 2605.30341v1, 2605.30343v1, 2605.30353v1","author":[{"family":"Team","given":"Saluca"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20519586","URL":"https://doi.org/10.5281/zenodo.20519586","source":"datacite"},{"id":"doi:10.5281/zenodo.20520079","type":"article-journal","title":"Silent Failures and Structural Gaps: A Cross-Domain Framework for Evaluation Rigor in Large-Model Systems","abstract":"Version 2 — revised in response to an external structural review and an automated critique pass. See \"Response to Review\" appendix in the PDF for the change log. Large-model systems are evaluated at multiple layers—statistical, algorithmic, systems, and behavioral—yet each layer's evaluation methodology has been developed largely in isolation. This paper identifies a shared structural pattern across these layers: evaluations that are locally valid but globally misleading. Drawing on recent preprints spanning federated learning, LLM inference, causal inference, high-dimensional statistics, latent reasoning, and multi-agent coherence, we offer a *heuristic reading* that a common pattern recurs: a system or estimator satisfies its local objective (loss reduction, throughput, oracle test passage, component-level coherence) while concealing a deeper failure that only manifests at a different scale, composition, or distribution shift. We call this the **local-validity trap**. Specifically, we synthesize evidence that (1) oracle-based testing may miss semantically incorrect but symptom-reducing fixes in AI-assisted development; (2) component-level probabilistic coherence does not guarantee joint coherence in multi-agent LLM systems; (3) measurement error silently biases high-dimensional regression even when penalized estimators converge; (4) asynchronous pipeline parallelism bounds staleness locally but may accumulate convergence error globally; and (5) contribution metrics in federated learning should track optimization trajectories rather than static snapshots to avoid misattribution. We propose that evaluation rigor for large-model systems may benefit from explicit cross-scale consistency checks, analogous to structural constraints studied in the algorithmic and statistical literatures, and outline candidate design principles for building such checks into system pipelines. The \"local-validity trap\" framing is introduced here as an organizing heuristic; it is not a term used in any cited source, and the cross-domain structural analogy is asserted on the basis of shared vocabulary rather than shared mechanism. --- Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted from arXiv preprint corpus on the date in the filename. Cited arXiv preprints: 2605.29566v1, 2605.29639v1, 2605.29664v1, 2605.29740v1, 2605.29944v1, 2605.30075v1, 2605.30113v1, 2605.30153v1, 2605.30158v1, 2605.30319v1, 2605.30321v1, 2605.30327v1, 2605.30335v1, 2605.30336v1, 2605.30341v1, 2605.30343v1, 2605.30353v1","author":[{"family":"Team","given":"Saluca"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20520079","URL":"https://doi.org/10.5281/zenodo.20520079","source":"datacite"},{"id":"doi:10.5281/zenodo.20470302","type":"article-journal","title":"Differential Privacy Trade-offs in Federated Code Generation Model Fine-Tuning","abstract":"This report synthesises findings from 3 peer-reviewed papers addressing the following research question: What is the trade-off between inference latency and model robustness against adversarial attacks when applying differential privacy mechanisms to federated fine-tuning of code generation models. Federated learning (FL) is a machine learning setting where many clients (e.g., mobile devices or whole organizations) collaboratively train a model under the orchestration of a central server (e.g., service provider), while keeping the training data decentralized. FL embodies. 10 claims were extracted from source literature; 10 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.2/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the trade-off between inference latency and model robustness against adversarial attacks when applying differential privacy mechanisms to federated fine-tuning of code generation models? Autonomous literature synthesis. Automated review score: 8.2/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20470302","URL":"https://doi.org/10.5281/zenodo.20470302","source":"datacite"},{"id":"doi:10.5281/zenodo.20470303","type":"article-journal","title":"Differential Privacy Trade-offs in Federated Code Generation Model Fine-Tuning","abstract":"This report synthesises findings from 3 peer-reviewed papers addressing the following research question: What is the trade-off between inference latency and model robustness against adversarial attacks when applying differential privacy mechanisms to federated fine-tuning of code generation models. Federated learning (FL) is a machine learning setting where many clients (e.g., mobile devices or whole organizations) collaboratively train a model under the orchestration of a central server (e.g., service provider), while keeping the training data decentralized. FL embodies. 10 claims were extracted from source literature; 10 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.2/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the trade-off between inference latency and model robustness against adversarial attacks when applying differential privacy mechanisms to federated fine-tuning of code generation models? Autonomous literature synthesis. Automated review score: 8.2/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20470303","URL":"https://doi.org/10.5281/zenodo.20470303","source":"datacite"},{"id":"doi:10.5281/zenodo.20608577","type":"article-journal","title":"Coordination Topology as a Design Variable: How Memory Depth, Communication Sparsity, Censored Feedback, and Governance Boundaries Jointly Determine Multi-Agent System Behavior","abstract":"Version 2 — revised in response to an external structural review and an automated critique pass. See \"Response to Review\" appendix in the PDF for the change log. Multi-agent systems (MAS) are typically designed by fixing a task objective and then selecting agents, communication protocols, and memory configurations as secondary implementation details. This paper argues for a candidate reframing: **coordination topology**—the joint specification of communication graph structure, memory depth, governance boundaries, and feedback censorship—should be treated as a primary design variable whose configuration determines not merely efficiency but qualitative system behavior, including whether consensus is reached, whether free-riding is suppressed, whether failures propagate silently, and whether governance constraints hold at the execution boundary. This is a heuristic reading, not a formal derivation; the corpus sources share structural analogies rather than a unified formalism, and the connections between domains are argued by mechanism where possible and flagged as analogies where they are not. We synthesize findings from six recent arXiv preprints spanning cs.MA and cs.DC. The evidence base includes: a controlled simulation study showing that memory depth and network topology interact to flip the sign of coordination speed [corpus:arxiv:2606.04197]; a formal result demonstrating that decentralized free-rider suppression is topology-dependent under graded contention [corpus:arxiv:2606.06162]; an empirical characterisation of silent cross-agent memory injection failures caused by architectural isolation guards [corpus:arxiv:2606.04896]; a governance layer study showing that separating proposal generation from execution reduces unsafe actions from 88% to near-zero [corpus:arxiv:2606.04306]; a threshold-bandit analysis showing that censored feedback creates a structural learning cost decomposable into search and monitoring terms [corpus:arxiv:2605.27076]; and a federated-market orchestration result showing that decentralised price-based allocation matches centralised welfare under specific graph-topology conditions [corpus:arxiv:2605.27106]. The last source connects to the synthesis through analogy between service-dependency DAG topology and agent communication graph topology; it is treated as corroborating rather than primary evidence and is discussed in a dedicated weakly-connected addendum. Together these findings suggest that topology, memory, censorship, and governance are not independent tuning knobs but coupled coordinates of a single configuration space. The falsification path for the central claim is stated concretely: construct a factorial experiment varying topology, memory depth, and governance boundary placement across a fixed task, and test whether the interaction effects on coordination outcome are statistically significant and sign-reversing. All six corpus sources are arXiv preprints; none are peer-reviewed, and all claims should be treated as hypotheses pending replication. --- Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted from arXiv preprint corpus on the date in the filename. Cited arXiv preprints: 2605.27076, 2605.27106, 2605.30802, 2606.04197, 2606.04306, 2606.04896, 2606.06162, 2606.06189, 2606.07487 AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.","author":[{"family":"Team","given":"Saluca"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20608577","URL":"https://doi.org/10.5281/zenodo.20608577","source":"datacite"},{"id":"doi:10.5281/zenodo.20626963","type":"article-journal","title":"Air Quality Index Prediction: A Review of Machine Learning, Deep Learning, and Hybrid Approaches","abstract":"Abstract - The increasing evolution of air pollution as a global environmental and public health issue worldwide warrants the timely provision of accurate predictions for the Air Quality Index (AQI). Therefore, this review paper will explore in-depth all the traditional, machine learning, and deep learning models put into effect for AQI forecast modeling. Initially, some conventional statistical models like ARIMA and regression-based models are analyzed. It is seen that they are less complex and easy to interpret, but due to their incapability to handle nonlinear and complex data patterns, the results do not speak highly of these models. This raises issues around machine learning algorithms, such as those of decision tree, support vector machine, and ensemble methods, which actually increased the accuracy of data predictions thanks to better treatment of features and generalization. Further embraces on computational neural networks (CNN), long short term memory (LSTM), and hybrid models have been smartly improved to catch temporal and spatial linkages in air quality data. The paper also covers discussed up-coming trends such as federated learning, edge-AI, remote sensing integration, and explainable AI for much-needed larger deployment and more privacy and interpretable AI prediction options for the AQI. The key challenges identified in the work were data scarcity, model complexities, computational needs, and lack of evaluation frameworks complying with a set of standards. This review brings out these research gaps and clearly accentuates the need for these hybrid, scalable, and real-time AQI forecasting solutions. It, therefore, provides a comprehensive understanding of all available methodologies and helps foresee the construction of durable and efficient air quality predictions for the betterment of environmental management.","author":[{"family":"Tiwari","given":"Abhishek"},{"family":"Kumar","given":"Atesh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20626963","URL":"https://doi.org/10.5281/zenodo.20626963","source":"datacite"},{"id":"doi:10.5281/zenodo.20626964","type":"article-journal","title":"Air Quality Index Prediction: A Review of Machine Learning, Deep Learning, and Hybrid Approaches","abstract":"Abstract - The increasing evolution of air pollution as a global environmental and public health issue worldwide warrants the timely provision of accurate predictions for the Air Quality Index (AQI). Therefore, this review paper will explore in-depth all the traditional, machine learning, and deep learning models put into effect for AQI forecast modeling. Initially, some conventional statistical models like ARIMA and regression-based models are analyzed. It is seen that they are less complex and easy to interpret, but due to their incapability to handle nonlinear and complex data patterns, the results do not speak highly of these models. This raises issues around machine learning algorithms, such as those of decision tree, support vector machine, and ensemble methods, which actually increased the accuracy of data predictions thanks to better treatment of features and generalization. Further embraces on computational neural networks (CNN), long short term memory (LSTM), and hybrid models have been smartly improved to catch temporal and spatial linkages in air quality data. The paper also covers discussed up-coming trends such as federated learning, edge-AI, remote sensing integration, and explainable AI for much-needed larger deployment and more privacy and interpretable AI prediction options for the AQI. The key challenges identified in the work were data scarcity, model complexities, computational needs, and lack of evaluation frameworks complying with a set of standards. This review brings out these research gaps and clearly accentuates the need for these hybrid, scalable, and real-time AQI forecasting solutions. It, therefore, provides a comprehensive understanding of all available methodologies and helps foresee the construction of durable and efficient air quality predictions for the betterment of environmental management.","author":[{"family":"Tiwari","given":"Abhishek"},{"family":"Kumar","given":"Atesh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20626964","URL":"https://doi.org/10.5281/zenodo.20626964","source":"datacite"},{"id":"doi:10.5281/zenodo.20482023","type":"article-journal","title":"Diversity-Driven Client Selection in Federated Learning for CodeLlama-7B on HumanEval","abstract":"This report synthesises findings from 10 peer-reviewed papers addressing the following research question: How does diversity-driven client selection in federated learning affect the pass@1 scores of CodeLlama-7B on the HumanEval benchmark under extreme non-IID code distribution scenarios. Trained on massive publicly available data, large language models (LLMs) have demonstrated tremendous success across various fields. While more data contributes to better performance, a disconcerting reality is that high-quality public data will be exhausted in a few years. 6 claims were extracted from source literature; 6 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.5/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does diversity-driven client selection in federated learning affect the pass@1 scores of CodeLlama-7B on the HumanEval benchmark under extreme non-IID code distribution scenarios? Autonomous literature synthesis. Automated review score: 8.5/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20482023","URL":"https://doi.org/10.5281/zenodo.20482023","source":"datacite"},{"id":"doi:10.5281/zenodo.20482024","type":"article-journal","title":"Diversity-Driven Client Selection in Federated Learning for CodeLlama-7B on HumanEval","abstract":"This report synthesises findings from 10 peer-reviewed papers addressing the following research question: How does diversity-driven client selection in federated learning affect the pass@1 scores of CodeLlama-7B on the HumanEval benchmark under extreme non-IID code distribution scenarios. Trained on massive publicly available data, large language models (LLMs) have demonstrated tremendous success across various fields. While more data contributes to better performance, a disconcerting reality is that high-quality public data will be exhausted in a few years. 6 claims were extracted from source literature; 6 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.5/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does diversity-driven client selection in federated learning affect the pass@1 scores of CodeLlama-7B on the HumanEval benchmark under extreme non-IID code distribution scenarios? Autonomous literature synthesis. Automated review score: 8.5/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20482024","URL":"https://doi.org/10.5281/zenodo.20482024","source":"datacite"},{"id":"doi:10.5281/zenodo.20744059","type":"article-journal","title":"TextRefs: An open registry for canonical text references","abstract":"Slides introducing TextRefs, an open registry for canonical text references. The presentation starts from a familiar scholarly practice: references such as “Plato, Republic 514a” identify the same passage across many editions, translations, languages, layouts, and publication contexts. While this stability is central to scholarship, it is usually encoded only as human-readable citation text in footnotes, prose, or apparatus entries. As a result, software cannot reliably resolve, index, compare, or link canonical passage references across digital collections.TextRefs addresses this gap by treating canonical text references as persistent, machine-readable, and resolvable entities. The slides present TextRefs as a scholarly data infrastructure layer that can connect citation systems, canonical works, resolver targets, catalogues, editions, full-text repositories, datasets, annotations, and publishing platforms. Instead of replacing existing editorial or bibliographic practices, the registry is framed as an interoperability hub that makes established reference conventions visible to machines while preserving their scholarly meaning for human readers.The presentation illustrates this approach through the example of Plato, Republic 514a, showing how a single canonical reference can be represented in context, cited, given aliases, and connected to resolver targets such as Project Gutenberg and the Perseus Digital Library. It then outlines four use cases: stable cross-edition citation for individual researchers; persistent canonical structures for individual editions; federated networks of discoverable text across diverse collections; and AI or machine-learning workflows that benefit from granular, linked canonical text data.The closing slides emphasize that TextRefs is currently pre-1.0, that specifications and APIs may still change, and that scholarly trust will depend on curation, governance, review workflows, and community participation. The deck therefore serves both as an introduction to the TextRefs concept and as an invitation to help build a standard and roadmap for citable, interoperable text on the web.","author":[{"family":"Mähr","given":"Moritz"},{"family":"Seiberth","given":"Luz"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20744059","URL":"https://doi.org/10.5281/zenodo.20744059","source":"datacite"},{"id":"doi:10.5281/zenodo.20744060","type":"article-journal","title":"TextRefs: An open registry for canonical text references","abstract":"Slides introducing TextRefs, an open registry for canonical text references. The presentation starts from a familiar scholarly practice: references such as “Plato, Republic 514a” identify the same passage across many editions, translations, languages, layouts, and publication contexts. While this stability is central to scholarship, it is usually encoded only as human-readable citation text in footnotes, prose, or apparatus entries. As a result, software cannot reliably resolve, index, compare, or link canonical passage references across digital collections.TextRefs addresses this gap by treating canonical text references as persistent, machine-readable, and resolvable entities. The slides present TextRefs as a scholarly data infrastructure layer that can connect citation systems, canonical works, resolver targets, catalogues, editions, full-text repositories, datasets, annotations, and publishing platforms. Instead of replacing existing editorial or bibliographic practices, the registry is framed as an interoperability hub that makes established reference conventions visible to machines while preserving their scholarly meaning for human readers.The presentation illustrates this approach through the example of Plato, Republic 514a, showing how a single canonical reference can be represented in context, cited, given aliases, and connected to resolver targets such as Project Gutenberg and the Perseus Digital Library. It then outlines four use cases: stable cross-edition citation for individual researchers; persistent canonical structures for individual editions; federated networks of discoverable text across diverse collections; and AI or machine-learning workflows that benefit from granular, linked canonical text data.The closing slides emphasize that TextRefs is currently pre-1.0, that specifications and APIs may still change, and that scholarly trust will depend on curation, governance, review workflows, and community participation. The deck therefore serves both as an introduction to the TextRefs concept and as an invitation to help build a standard and roadmap for citable, interoperable text on the web.","author":[{"family":"Mähr","given":"Moritz"},{"family":"Seiberth","given":"Luz"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20744060","URL":"https://doi.org/10.5281/zenodo.20744060","source":"datacite"},{"id":"doi:10.5281/zenodo.20482101","type":"article-journal","title":"Straggler Mitigation Strategies in Asynchronous Federated Learning for Robust Multimodal IoT Malware Detection","abstract":"This report synthesises findings from 8 peer-reviewed papers addressing the following research question: What is the impact of straggler mitigation strategies in asynchronous federated learning on the robustness of multimodal IoT malware detectors against label-flipping poisoning attacks. Federated Learning (FL) is a technique that can learn a global machine-learning model at a central server by aggregating locally trained models. This distributed machine-learning approach preserves the privacy of local models. 6 claims were extracted from source literature; 6 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.5/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the impact of straggler mitigation strategies in asynchronous federated learning on the robustness of multimodal IoT malware detectors against label-flipping poisoning attacks? Autonomous literature synthesis. Automated review score: 8.5/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20482101","URL":"https://doi.org/10.5281/zenodo.20482101","source":"datacite"},{"id":"doi:10.5281/zenodo.20520092","type":"article-journal","title":"Coordination Without Continuous Synchronization: Why Sparse, Event-Triggered, and Budget-Aware Protocols Dominate Dense Communication in Multi-Agent Systems","abstract":"Version 2 — revised in response to an external structural review and an automated critique pass. See \"Response to Review\" appendix in the PDF for the change log. A persistent assumption in multi-agent systems (MAS) design is that more communication produces better coordination. Recent preprints across cs.MA, cs.DC, and cs.NI collectively challenge this assumption, converging on a counter-intuitive pattern: sparse, event-triggered, and budget-aware communication protocols consistently match or outperform dense, synchronous baselines across radically different topologies — from LLM-based reasoning pipelines to federated edge-cloud orchestration to cooperative reinforcement learning under delayed observations. This synthesis draws on eight corpus sources to argue a **heuristic reading** of that pattern: communication density is a cost-accuracy dial, not a reliability guarantee, and the optimal operating point is frequently sparser than the default design assumption. We present this as a candidate hypothesis, not a derived theorem; the mechanisms behind the pattern come from distinct formalisms and the cross-domain analogy is argued but not proven. We identify three structural mechanisms behind this reading. First, redundant message generation in multi-agent pipelines concentrates energy and latency cost in output tokens, not input tokens, making sparsification asymmetrically valuable [corpus:arxiv:2605.27787]. Second, deliberative consensus in LLM oracle ensembles propagates confident errors rather than correcting them, causing accuracy to *fall below* single-model baselines when agent count increases [corpus:arxiv:2605.30802]. Third, event-triggered decentralized coordination in threshold-activated cooperative bandits achieves a reported 23× reduction in communication volume relative to a centralized baseline — across the tested configurations — while preserving feasibility alignment [corpus:arxiv:2605.27076]. Complementary evidence from federated market orchestration [corpus:arxiv:2605.27106], hybrid cloud-edge MAS design [corpus:arxiv:2605.30102], dynamic topology reconfiguration [corpus:arxiv:2605.29511], delay-robust MARL [corpus:arxiv:2605.26286], and V2X infrastructure coordination [corpus:arxiv:2605.25431] reinforces the pattern across domains and topologies. The falsification path is clear: if tasks exist where dense, synchronous communication consistently outperforms sparse alternatives after controlling for task difficulty, topology, and error correlation structure, the thesis fails. We name the conditions under which dense communication remains necessary and argue they are narrower than commonly assumed. --- Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted from arXiv preprint corpus on the date in the filename. Cited arXiv preprints: 2605.25431, 2605.25653, 2605.25746, 2605.26286, 2605.26448, 2605.27076, 2605.27106, 2605.27466, 2605.27787, 2605.29511, 2605.30102, 2605.30227, 2605.30802 AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.","author":[{"family":"Team","given":"Saluca"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20520092","URL":"https://doi.org/10.5281/zenodo.20520092","source":"datacite"},{"id":"doi:10.5281/zenodo.20519745","type":"article-journal","title":"Coordination Without Continuous Synchronization: Why Sparse, Event-Triggered, and Budget-Aware Protocols Dominate Dense Communication in Multi-Agent Systems","abstract":"Version 2 — revised in response to an external structural review and an automated critique pass. See \"Response to Review\" appendix in the PDF for the change log. A persistent assumption in multi-agent systems (MAS) design is that more communication produces better coordination. Recent preprints across cs.MA, cs.DC, and cs.NI collectively challenge this assumption, converging on a counter-intuitive pattern: sparse, event-triggered, and budget-aware communication protocols consistently match or outperform dense, synchronous baselines across radically different topologies — from LLM-based reasoning pipelines to federated edge-cloud orchestration to cooperative reinforcement learning under delayed observations. This synthesis draws on eight corpus sources to argue a **heuristic reading** of that pattern: communication density is a cost-accuracy dial, not a reliability guarantee, and the optimal operating point is frequently sparser than the default design assumption. We present this as a candidate hypothesis, not a derived theorem; the mechanisms behind the pattern come from distinct formalisms and the cross-domain analogy is argued but not proven. We identify three structural mechanisms behind this reading. First, redundant message generation in multi-agent pipelines concentrates energy and latency cost in output tokens, not input tokens, making sparsification asymmetrically valuable [corpus:arxiv:2605.27787]. Second, deliberative consensus in LLM oracle ensembles propagates confident errors rather than correcting them, causing accuracy to *fall below* single-model baselines when agent count increases [corpus:arxiv:2605.30802]. Third, event-triggered decentralized coordination in threshold-activated cooperative bandits achieves a reported 23× reduction in communication volume relative to a centralized baseline — across the tested configurations — while preserving feasibility alignment [corpus:arxiv:2605.27076]. Complementary evidence from federated market orchestration [corpus:arxiv:2605.27106], hybrid cloud-edge MAS design [corpus:arxiv:2605.30102], dynamic topology reconfiguration [corpus:arxiv:2605.29511], delay-robust MARL [corpus:arxiv:2605.26286], and V2X infrastructure coordination [corpus:arxiv:2605.25431] reinforces the pattern across domains and topologies. The falsification path is clear: if tasks exist where dense, synchronous communication consistently outperforms sparse alternatives after controlling for task difficulty, topology, and error correlation structure, the thesis fails. We name the conditions under which dense communication remains necessary and argue they are narrower than commonly assumed. --- Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted from arXiv preprint corpus on the date in the filename. Cited arXiv preprints: 2605.25431, 2605.25653, 2605.25746, 2605.26286, 2605.26448, 2605.27076, 2605.27106, 2605.27466, 2605.27787, 2605.29511, 2605.30102, 2605.30227, 2605.30802 AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.","author":[{"family":"Team","given":"Saluca"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20519745","URL":"https://doi.org/10.5281/zenodo.20519745","source":"datacite"},{"id":"doi:10.5281/zenodo.20501766","type":"article-journal","title":"Multi-View Graph Anomaly Detection Robustness Under View Dropout and Metattack Perturbations","abstract":"This report synthesises findings from 3 peer-reviewed papers addressing the following research question: What is the impact of view dropout on the robustness of multi-view graph anomaly detection frameworks against Metattack perturbations, measured by the degradation in detection accuracy and F1-score. Federated learning (FL) is a machine learning setting where many clients (e.g., mobile devices or whole organizations) collaboratively train a model under the orchestration of a central server (e.g., service provider), while keeping the training data decentralized. FL embodies. 5 claims were extracted from source literature; 5 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 9.0/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the impact of view dropout on the robustness of multi-view graph anomaly detection frameworks against Metattack perturbations, measured by the degradation in detection accuracy and F1-score? Autonomous literature synthesis. Automated review score: 9.0/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20501766","URL":"https://doi.org/10.5281/zenodo.20501766","source":"datacite"},{"id":"doi:10.5281/zenodo.20501767","type":"article-journal","title":"Multi-View Graph Anomaly Detection Robustness Under View Dropout and Metattack Perturbations","abstract":"This report synthesises findings from 3 peer-reviewed papers addressing the following research question: What is the impact of view dropout on the robustness of multi-view graph anomaly detection frameworks against Metattack perturbations, measured by the degradation in detection accuracy and F1-score. Federated learning (FL) is a machine learning setting where many clients (e.g., mobile devices or whole organizations) collaboratively train a model under the orchestration of a central server (e.g., service provider), while keeping the training data decentralized. FL embodies. 5 claims were extracted from source literature; 5 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 9.0/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the impact of view dropout on the robustness of multi-view graph anomaly detection frameworks against Metattack perturbations, measured by the degradation in detection accuracy and F1-score? Autonomous literature synthesis. Automated review score: 9.0/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20501767","URL":"https://doi.org/10.5281/zenodo.20501767","source":"datacite"},{"id":"doi:10.5281/zenodo.20481956","type":"article-journal","title":"Compressive Sensing Integration in Over-the-Air Federated Learning for Massive MIMO Systems","abstract":"This report synthesises findings from 15 peer-reviewed papers addressing the following research question: How does the integration of compressive sensing techniques with over-the-air federated learning (OTA-FL) in massive MIMO systems compare to traditional FL methods in terms of model accuracy and. The thriving of artificial intelligence (AI) applications is driving the further evolution of wireless networks. It has been envisioned that 6G will be transformative and will revolutionize the evolution of wireless from ``connected things'' to ``connected intelligence''. 7 claims were extracted from source literature; 7 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.3/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does the integration of compressive sensing techniques with over-the-air federated learning (OTA-FL) in massive MIMO systems compare to traditional FL methods in terms of model accuracy and convergence rate when evaluated on the CIFAR-10 dataset? Autonomous literature synthesis. Automated review score: 8.3/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20481956","URL":"https://doi.org/10.5281/zenodo.20481956","source":"datacite"},{"id":"doi:10.5281/zenodo.20481957","type":"article-journal","title":"Compressive Sensing Integration in Over-the-Air Federated Learning for Massive MIMO Systems","abstract":"This report synthesises findings from 15 peer-reviewed papers addressing the following research question: How does the integration of compressive sensing techniques with over-the-air federated learning (OTA-FL) in massive MIMO systems compare to traditional FL methods in terms of model accuracy and. The thriving of artificial intelligence (AI) applications is driving the further evolution of wireless networks. It has been envisioned that 6G will be transformative and will revolutionize the evolution of wireless from ``connected things'' to ``connected intelligence''. 7 claims were extracted from source literature; 7 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.3/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does the integration of compressive sensing techniques with over-the-air federated learning (OTA-FL) in massive MIMO systems compare to traditional FL methods in terms of model accuracy and convergence rate when evaluated on the CIFAR-10 dataset? Autonomous literature synthesis. Automated review score: 8.3/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20481957","URL":"https://doi.org/10.5281/zenodo.20481957","source":"datacite"},{"id":"doi:10.5281/zenodo.20481996","type":"article-journal","title":"Structural Similarity Degradation in Multimodal Federated Learning Under Non-IID Data","abstract":"This report synthesises findings from 11 peer-reviewed papers addressing the following research question: How does the structural similarity of feature embeddings in collaborative multimodal federated learning models scale with increasing degrees of non-IID data, as measured by cosine similarity across. In parallel with the rapid adoption of artificial intelligence (AI) empowered by advances in AI research, there has been growing awareness and concerns of data privacy. Recent significant developments in the data regulation landscape have prompted a seismic shift in interest. 6 claims were extracted from source literature; 5 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 7.9/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does the structural similarity of feature embeddings in collaborative multimodal federated learning models scale with increasing degrees of non-IID data, as measured by cosine similarity across clients? Autonomous literature synthesis. Automated review score: 7.9/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20481996","URL":"https://doi.org/10.5281/zenodo.20481996","source":"datacite"},{"id":"doi:10.5281/zenodo.20481997","type":"article-journal","title":"Structural Similarity Degradation in Multimodal Federated Learning Under Non-IID Data","abstract":"This report synthesises findings from 11 peer-reviewed papers addressing the following research question: How does the structural similarity of feature embeddings in collaborative multimodal federated learning models scale with increasing degrees of non-IID data, as measured by cosine similarity across. In parallel with the rapid adoption of artificial intelligence (AI) empowered by advances in AI research, there has been growing awareness and concerns of data privacy. Recent significant developments in the data regulation landscape have prompted a seismic shift in interest. 6 claims were extracted from source literature; 5 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 7.9/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does the structural similarity of feature embeddings in collaborative multimodal federated learning models scale with increasing degrees of non-IID data, as measured by cosine similarity across clients? Autonomous literature synthesis. Automated review score: 7.9/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20481997","URL":"https://doi.org/10.5281/zenodo.20481997","source":"datacite"},{"id":"doi:10.5281/zenodo.20474482","type":"article-journal","title":"Federated Learning Aggregation Strategies for Non-IID Data in Massive MIMO Systems","abstract":"This report synthesises findings from 12 peer-reviewed papers addressing the following research question: How do different federated learning aggregation strategies (e.g., FedAvg, FedProx, SCAFFOLD) perform in terms of robustness to non-IID data distributions and model alignment when integrated with. Over-the-air federated learning (OTA-FL) is an emerging technique to reduce the computation and communication overload at the PS caused by the orthogonal transmissions of the model updates in conventional federated learning (FL). This reduction is achieved at the expense of. 10 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.3/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How do different federated learning aggregation strategies (e.g., FedAvg, FedProx, SCAFFOLD) perform in terms of robustness to non-IID data distributions and model alignment when integrated with compressive sensing over massive MIMO systems, evaluated using cross-domain benchmark datasets? Autonomous literature synthesis. Automated review score: 8.3/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20474482","URL":"https://doi.org/10.5281/zenodo.20474482","source":"datacite"},{"id":"doi:10.5281/zenodo.20474483","type":"article-journal","title":"Federated Learning Aggregation Strategies for Non-IID Data in Massive MIMO Systems","abstract":"This report synthesises findings from 12 peer-reviewed papers addressing the following research question: How do different federated learning aggregation strategies (e.g., FedAvg, FedProx, SCAFFOLD) perform in terms of robustness to non-IID data distributions and model alignment when integrated with. Over-the-air federated learning (OTA-FL) is an emerging technique to reduce the computation and communication overload at the PS caused by the orthogonal transmissions of the model updates in conventional federated learning (FL). This reduction is achieved at the expense of. 10 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.3/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How do different federated learning aggregation strategies (e.g., FedAvg, FedProx, SCAFFOLD) perform in terms of robustness to non-IID data distributions and model alignment when integrated with compressive sensing over massive MIMO systems, evaluated using cross-domain benchmark datasets? Autonomous literature synthesis. Automated review score: 8.3/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20474483","URL":"https://doi.org/10.5281/zenodo.20474483","source":"datacite"},{"id":"doi:10.5281/zenodo.20472173","type":"article-journal","title":"Quantized Large Language Model Inference Latency Scaling in Federated Edge Deployments","abstract":"This report synthesises findings from 14 peer-reviewed papers addressing the following research question: How does the inference latency of quantized large language models scale with the number of concurrent edge devices in a federated learning setup for real-time threat detection. Successful integration of deep neural networks (DNNs) or deep learning (DL) has resulted in breakthroughs in many areas. However, deploying these highly accurate models for data-driven, learned, automatic, and practical machine learning (ML) solutions to end-user applications. 9 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.2/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does the inference latency of quantized large language models scale with the number of concurrent edge devices in a federated learning setup for real-time threat detection? Autonomous literature synthesis. Automated review score: 8.2/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20472173","URL":"https://doi.org/10.5281/zenodo.20472173","source":"datacite"},{"id":"doi:10.5281/zenodo.20472174","type":"article-journal","title":"Quantized Large Language Model Inference Latency Scaling in Federated Edge Deployments","abstract":"This report synthesises findings from 14 peer-reviewed papers addressing the following research question: How does the inference latency of quantized large language models scale with the number of concurrent edge devices in a federated learning setup for real-time threat detection. Successful integration of deep neural networks (DNNs) or deep learning (DL) has resulted in breakthroughs in many areas. However, deploying these highly accurate models for data-driven, learned, automatic, and practical machine learning (ML) solutions to end-user applications. 9 claims were extracted from source literature; 9 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.2/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does the inference latency of quantized large language models scale with the number of concurrent edge devices in a federated learning setup for real-time threat detection? Autonomous literature synthesis. Automated review score: 8.2/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20472174","URL":"https://doi.org/10.5281/zenodo.20472174","source":"datacite"},{"id":"doi:10.5281/zenodo.20463724","type":"article-journal","title":"Federated Learning Scalability in Heterogeneous IoT Malware Detection Networks","abstract":"This report synthesises findings from 14 peer-reviewed papers addressing the following research question: How does the communication efficiency of federated learning-based malware detection models scale with increasing device heterogeneity across IoT networks, measured by round convergence time and. ---The Internet of Things (IoT) has revolutionized various sectors by enabling seamless interaction between devices. However, the proliferation of IoT devices has also raised significant security and privacy concerns. 8 claims were extracted from source literature; 7 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.2/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does the communication efficiency of federated learning-based malware detection models scale with increasing device heterogeneity across IoT networks, measured by round convergence time and bandwidth utilization? Autonomous literature synthesis. Automated review score: 8.2/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20463724","URL":"https://doi.org/10.5281/zenodo.20463724","source":"datacite"},{"id":"doi:10.5281/zenodo.20463725","type":"article-journal","title":"Federated Learning Scalability in Heterogeneous IoT Malware Detection Networks","abstract":"This report synthesises findings from 14 peer-reviewed papers addressing the following research question: How does the communication efficiency of federated learning-based malware detection models scale with increasing device heterogeneity across IoT networks, measured by round convergence time and. ---The Internet of Things (IoT) has revolutionized various sectors by enabling seamless interaction between devices. However, the proliferation of IoT devices has also raised significant security and privacy concerns. 8 claims were extracted from source literature; 7 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.2/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does the communication efficiency of federated learning-based malware detection models scale with increasing device heterogeneity across IoT networks, measured by round convergence time and bandwidth utilization? Autonomous literature synthesis. Automated review score: 8.2/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20463725","URL":"https://doi.org/10.5281/zenodo.20463725","source":"datacite"},{"id":"doi:10.5281/zenodo.20263144","type":"article-journal","title":"fl-oran-tmc: Federated Learning Benchmark for ColO-RAN Slice SLA Prediction","abstract":"Local PyTorch federated-learning pipeline for ColO-RAN O-RAN slice SLA forecasting. Cross-architecture empirical benchmark on the Colosseum/ColO-RAN public dataset: 3-arch core panel (LSTM, Mamba, Spiking-SSM) plus a 2-arch recent-SOTA extension (xLSTM, Mamba-3) × multiple FL algorithms (FedAvg, FedProx, FedAdam, SCAFFOLD, FedDyn, FedBN, FedSWA, FedSCAM, FedGMT, FedMoSWA; MOON deferred) × parametric Dirichlet heterogeneity × per-round NVML training-energy measurement. All five architectures share an identical encoder and a structurally-identical classifier head; only the temporal trunk differs (total parameter count matched within ±10% via the ADR-001 D-20 parity constraint). Paper under review at IEEE Journal on Selected Areas in Communications.","author":[{"family":"Tsai","given":"Hsiu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20263144","URL":"https://doi.org/10.5281/zenodo.20263144","source":"datacite"},{"id":"doi:10.5281/zenodo.20319969","type":"article-journal","title":"fl-oran-tmc: Federated Learning Benchmark for ColO-RAN Slice SLA Prediction","abstract":"Local PyTorch federated-learning pipeline for ColO-RAN O-RAN slice SLA forecasting. Cross-architecture empirical benchmark on the Colosseum/ColO-RAN public dataset: 3-arch core panel (LSTM, Mamba, Spiking-SSM) plus a 2-arch recent-SOTA extension (xLSTM, Mamba-3) × multiple FL algorithms (FedAvg, FedProx, FedAdam, SCAFFOLD, FedDyn, FedBN, FedSWA, FedSCAM, FedGMT, FedMoSWA; MOON deferred) × parametric Dirichlet heterogeneity × per-round NVML training-energy measurement. All five architectures share an identical encoder and a structurally-identical classifier head; only the temporal trunk differs (total parameter count matched within ±10% via the ADR-001 D-20 parity constraint). Paper under review at IEEE Journal on Selected Areas in Communications.","author":[{"family":"Tsai","given":"Hsiu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20319969","URL":"https://doi.org/10.5281/zenodo.20319969","source":"datacite"},{"id":"doi:10.13016/eulx-dsgd","type":"article-journal","title":"Practical Multiparty Protocols From Lattice Assumptions: Threshold Signatures, Oblivious Pseudorandom Functions, And More","abstract":"Lattice-based cryptography has emerged as the most dominant replacement candidate for the next generation of post-quantum cryptographic tools. With their operational simplicity while allowing advanced functionality, these protocols lead the majority of post-quantum standardization efforts and motivate a great chunk of current research to realize advanced trusted communication models. However, lattices' greatest asset is also their greatest curse. The applicability of advanced functionality motivates protocols with multiple computing parties while the assumptions that make lattice protocols secure in the first place hate settings where secrets are distributed. In this work we try to alleviate this issue by building practical lattice-based multiparty protocols. First we propose the first known concrete lattice-based threshold signature scheme with distributed key generation to demonstrate practicality. Second, we look at a different type of protocol, namely verifiable oblivious pseudorandom functions, and propose a practical version of an existing protocol through different analysis techniques while also giving the first lattice-based threshold versions of such protocols. Using these techniques, we then rebuild our threshold signature scheme and show a concretely efficient threshold signature that simultaneously provides additional desirable properties like identifiability and non-interactivity. Finally, we look at the possibility of asymmetric outsourced computation and formalize the classic notion of augmented password-protected threshold signatures in a more practicality friendly manner and construct the first lattice-based augmented password-protected threshold signature scheme. All of these works act as building blocks for more complicated protocols and share similar analysis techniques and solutions to problems specific to the distributed setting. This commonality indicates that it is not only the assumptions that we need to revisit but also how we think about security in general as part of preparing cryptography for its post-quantum era.","author":[{"family":"Gur","given":"Kamil"}],"issued":{"date-parts":[[2025]]},"DOI":"10.13016/eulx-dsgd","URL":"https://doi.org/10.13016/eulx-dsgd","source":"datacite"},{"id":"doi:10.65649/a5k5rw30","type":"article-journal","title":"Federated Clinical Learning Cooperative (FCLC)","abstract":"Background: Developing robust clinical artificial intelligence (AI) models requires large, diverse datasets that individual institutions cannot provide due to privacy regulations (GDPR, HIPAA) and institutional risk aversion. Existing federated learning (FL) platforms lack validated combinations of differential privacy, Byzantine robustness, fair contribution attribution, and secure aggregation. Objective: We present the Federated Clinical Learning Cooperative (FCLC), an open-source platform enabling multi-institutional clinical AI development without raw data leaving participating sites. Methods: FCLC implements a data preprocessing pipeline (Layers 1-3: direct identifier removal, quasi-identifier generalization, k-anonymity with k≥5) combined with differential privacy (Layer 4: DP-SGD, ε=2.0/round, δ=10⁻⁵, Rényi accountant α=4.0) and secure aggregation (Layer 5: SecAgg+ via CommonHealth). Validation used MIMIC-IV (N=12,543, 30-day readmission) and eICU-CRD (N=8,420, sepsis mortality) across IID and non-IID partitions (Dirichlet α ∈ {∞, 1.0, 0.5, 0.1, 0.01}) with logistic regression and multilayer perceptron architectures. Results: On MIMIC-IV, FCLC achieved AUC=0.758 [95% CI: 0.739–0.777] compared to centralized oracle 0.789 (Δ = −3.9%). Under severe non-IID conditions (Dirichlet α=0.1, EMD=0.31), FCLC preserved AUC=0.748 while FedAvg degraded to 0.694 (p_adj=0.003). Membership inference attack AUC with DP was 0.52±0.03 (indistinguishable from chance, p=0.31 vs. 0.50). At ΔAUC=0.03 — the minimum clinically important difference (MCID) corresponding to preventing approximately one readmission per 100 patients (NNT≈100) — the study had power &gt;0.99. Conclusions: FCLC provides a validated, regulation-compliant infrastructure for federated clinical AI with full cryptographic privacy guarantees suitable for mutual-distrust deployments. The platform is open-source (Apache 2.0) and fully reproducible via Docker.","author":[{"family":"Tkemaladze","given":"Jaba"}],"issued":{"date-parts":[[2026]]},"DOI":"10.65649/a5k5rw30","URL":"https://doi.org/10.65649/a5k5rw30","source":"openalex"},{"id":"doi:10.5281/zenodo.21980902","type":"article-journal","title":"Innovations RT GPU et Rendu Différentiable Multi-physique","abstract":"Résumé FRCe document, produit avec l’assistance de ChatGPT 5.2 Thinking et Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive (antériorité) et entre dans l’état de la technique dès sa mise en ligne, au sens des textes applicables : EPC Art. 54(2) (Convention sur le brevet européen), French IPC Art. L 611-11 (Code de la propriété intellectuelle), cf. 35 U.S.C. §102(a) (United States Patent Act), ainsi que des cadres chinois et japonais. Il décrit, de façon enabling, un portefeuille d’innovations RT GPU et rendu différentiable (RT cores, Monte Carlo, gradients, QA/benchmarks, interopérabilité, sécurité, durabilité) couvrant optique, RF, acoustique, sismologie, climat, nucléaire, robotique et biomédical. Chaque proposition est classée (IPC/CPC) et accompagnée d’éléments de preuve temporelle (RFC 3161 / FreeTSA). Abstract ENThis document, produced with the assistance of ChatGPT 5.2 Thinking and Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: art. L 611-11 CPI / art. 54(2) CBE (and related frameworks internationally). It discloses, in an enabling manner, a portfolio of GPU ray tracing and differentiable rendering inventions (RT cores, Monte Carlo transport, gradient-based inversion, QA/benchmarks, interoperability standards, operational security, and sustainability engineering) spanning optics/photonics, RF/6G, acoustics, seismology, climate/ocean, nuclear transport/shielding, robotics/3D sensing, and biophotonics/medical imaging. Every proposal is described with reproducible components and workflows, classified with IPC and CPC codes, and accompanied by timestamp proof (RFC 3161 / FreeTSA) for defensive disclosure. Timestamp : 2026-08-17T14:13:02ZSHA-256 : 1416dacf25f78daca949831e86aaa0ec06722489844477ce9509888b9ebf110f Liste des innovations & classification (IPC ; CPC) :1. Scientific RT-core kernel — IPC G06T 15/50 ; CPC G06T 15/502. Unified multi-physics RT engine — IPC G06F 9/50 ; CPC G06F 9/503. Generic differentiable ray tracing — IPC G06N 20/00 ; CPC G06N 20/004. Adaptive variance control scheduler — IPC G06T 15/20 ; CPC G06T 15/205. Differentiable RF ray tracer — IPC H04B 7/26 ; CPC H04B 7/266. Multi-sensor co-design optimizer — IPC G01S 17/89 ; CPC G01S 17/897. Certified synthetic dataset pipeline — IPC G06T 7/73 ; CPC G06T 7/738. Gradient-guided sim-to-real adaptation — IPC G06N 20/00 ; CPC G06N 20/009. Real-time LiDAR re-simulation — IPC G01S 17/93 ; CPC G01S 17/9310. GPU aircraft IR signature — IPC G01J 5/00 ; CPC G01J 5/0011. GPU GR lensing tracer — IPC G06F 17/50 ; CPC G06F 17/5012. Differentiable seismic tomography — IPC G01V 1/28 ; CPC G01V 1/2813. Calibrated room impulse response — IPC G01H 3/00 ; CPC G01H 3/0014. Realistic ultrasound Monte Carlo — IPC A61B 8/00 ; CPC A61B 8/0015. GPU photoacoustic pipeline — IPC A61B 5/00 ; CPC A61B 5/0016. Polarized tissue photon MC — IPC G01N 21/45 ; CPC G01N 21/4517. Differentiable BRDF estimation — IPC G01N 21/88 ; CPC G01N 21/8818. Hybrid metasurface RT design — IPC G02B 1/00 ; CPC G02B 1/0019. Climate adjoint radiative transfer — IPC G01W 1/00 ; CPC G01W 1/0020. Ocean optical inversion — IPC G01N 21/47 ; CPC G01N 21/4721. Safety-grade GPU neutron MC — IPC G21C 17/00 ; CPC G21C 17/0022. Multi-objective shielding optimizer — IPC G21F 3/00 ; CPC G21F 3/0023. RT-guided phototherapy planning — IPC A61N 5/06 ; CPC A61N 5/0624. Closed-loop phototherapy device — IPC A61N 5/06 ; CPC A61N 5/0625. Multi-parameter optimized PDT — IPC A61K 31/00 ; CPC A61K 31/0026. 3D-printed optical phantom — IPC G01N 21/00 ; CPC G01N 21/0027. Integrated opto-acoustic phantom — IPC A61B 8/00 ; CPC A61B 8/0028. Multi-physics scene file standard — IPC G06F 16/00 ; CPC G06F 16/0029. Cryptographic simulation provenance — IPC G06F 21/60 ; CPC G06F 21/6030. Federated private RT inversion — IPC G06N 20/00 ; CPC ","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21980902","URL":"https://doi.org/10.5281/zenodo.21980902","source":"datacite"},{"id":"doi:10.5281/zenodo.21980903","type":"article-journal","title":"Innovations RT GPU et Rendu Différentiable Multi-physique","abstract":"Résumé FRCe document, produit avec l’assistance de ChatGPT 5.2 Thinking et Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive (antériorité) et entre dans l’état de la technique dès sa mise en ligne, au sens des textes applicables : EPC Art. 54(2) (Convention sur le brevet européen), French IPC Art. L 611-11 (Code de la propriété intellectuelle), cf. 35 U.S.C. §102(a) (United States Patent Act), ainsi que des cadres chinois et japonais. Il décrit, de façon enabling, un portefeuille d’innovations RT GPU et rendu différentiable (RT cores, Monte Carlo, gradients, QA/benchmarks, interopérabilité, sécurité, durabilité) couvrant optique, RF, acoustique, sismologie, climat, nucléaire, robotique et biomédical. Chaque proposition est classée (IPC/CPC) et accompagnée d’éléments de preuve temporelle (RFC 3161 / FreeTSA). Abstract ENThis document, produced with the assistance of ChatGPT 5.2 Thinking and Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: art. L 611-11 CPI / art. 54(2) CBE (and related frameworks internationally). It discloses, in an enabling manner, a portfolio of GPU ray tracing and differentiable rendering inventions (RT cores, Monte Carlo transport, gradient-based inversion, QA/benchmarks, interoperability standards, operational security, and sustainability engineering) spanning optics/photonics, RF/6G, acoustics, seismology, climate/ocean, nuclear transport/shielding, robotics/3D sensing, and biophotonics/medical imaging. Every proposal is described with reproducible components and workflows, classified with IPC and CPC codes, and accompanied by timestamp proof (RFC 3161 / FreeTSA) for defensive disclosure. Timestamp : 2026-08-17T14:13:02ZSHA-256 : 1416dacf25f78daca949831e86aaa0ec06722489844477ce9509888b9ebf110f Liste des innovations & classification (IPC ; CPC) :1. Scientific RT-core kernel — IPC G06T 15/50 ; CPC G06T 15/502. Unified multi-physics RT engine — IPC G06F 9/50 ; CPC G06F 9/503. Generic differentiable ray tracing — IPC G06N 20/00 ; CPC G06N 20/004. Adaptive variance control scheduler — IPC G06T 15/20 ; CPC G06T 15/205. Differentiable RF ray tracer — IPC H04B 7/26 ; CPC H04B 7/266. Multi-sensor co-design optimizer — IPC G01S 17/89 ; CPC G01S 17/897. Certified synthetic dataset pipeline — IPC G06T 7/73 ; CPC G06T 7/738. Gradient-guided sim-to-real adaptation — IPC G06N 20/00 ; CPC G06N 20/009. Real-time LiDAR re-simulation — IPC G01S 17/93 ; CPC G01S 17/9310. GPU aircraft IR signature — IPC G01J 5/00 ; CPC G01J 5/0011. GPU GR lensing tracer — IPC G06F 17/50 ; CPC G06F 17/5012. Differentiable seismic tomography — IPC G01V 1/28 ; CPC G01V 1/2813. Calibrated room impulse response — IPC G01H 3/00 ; CPC G01H 3/0014. Realistic ultrasound Monte Carlo — IPC A61B 8/00 ; CPC A61B 8/0015. GPU photoacoustic pipeline — IPC A61B 5/00 ; CPC A61B 5/0016. Polarized tissue photon MC — IPC G01N 21/45 ; CPC G01N 21/4517. Differentiable BRDF estimation — IPC G01N 21/88 ; CPC G01N 21/8818. Hybrid metasurface RT design — IPC G02B 1/00 ; CPC G02B 1/0019. Climate adjoint radiative transfer — IPC G01W 1/00 ; CPC G01W 1/0020. Ocean optical inversion — IPC G01N 21/47 ; CPC G01N 21/4721. Safety-grade GPU neutron MC — IPC G21C 17/00 ; CPC G21C 17/0022. Multi-objective shielding optimizer — IPC G21F 3/00 ; CPC G21F 3/0023. RT-guided phototherapy planning — IPC A61N 5/06 ; CPC A61N 5/0624. Closed-loop phototherapy device — IPC A61N 5/06 ; CPC A61N 5/0625. Multi-parameter optimized PDT — IPC A61K 31/00 ; CPC A61K 31/0026. 3D-printed optical phantom — IPC G01N 21/00 ; CPC G01N 21/0027. Integrated opto-acoustic phantom — IPC A61B 8/00 ; CPC A61B 8/0028. Multi-physics scene file standard — IPC G06F 16/00 ; CPC G06F 16/0029. Cryptographic simulation provenance — IPC G06F 21/60 ; CPC G06F 21/6030. Federated private RT inversion — IPC G06N 20/00 ; CPC ","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21980903","URL":"https://doi.org/10.5281/zenodo.21980903","source":"datacite"},{"id":"doi:10.5281/zenodo.21978440","type":"article-journal","title":"Robotique souple neuromorphique et essaims","abstract":"Résumé FRCe document, produit avec l’assistance de ChatGPT 5.2 Thinking et Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive (antériorité) et entre dans l’état de la technique au sens des textes applicables (EPC Art. 54(2); French IPC Art. L 611-11; cf. 35 U.S.C. §102(a)). Il divulgue, de façon enabling, un portefeuille d’innovations combinant robotique souple (actionneurs HASEL/EAP), vision événementielle (DVS), calcul neuromorphique (SNN) et intelligence en essaim, couvrant dispositifs/capteurs, algorithmes, contrôle en boucle fermée, fabrication roll-to-roll et QA end-of-line, cybersécurité et opérations de flottes, interopérabilité (formats événements+spikes), logistique de cartouches, modèles économiques au résultat, et usages industriels, agricoles régénératifs, nucléaires, sous-marins et médicaux. Abstract ENThis document, produced with the assistance of ChatGPT 5.2 Thinking and Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes (art. L 611-11 CPI / art. 54(2) CBE). It discloses, in an enabling manner, a portfolio that fuses soft robotics (HASEL/EAP actuation), event-based vision (DVS), neuromorphic computing (SNN), and swarm intelligence. The disclosure spans devices and sensors, event-first control loops, roll-to-roll manufacturing and end-of-line QA, cyber-secure fleet operations, interoperability standards for event+spike telemetry, cartridge logistics and field repair, outcome-based metering and SLA instrumentation, and applications in high-throughput sorting, precision/regenerative agriculture, nuclear maintenance, underwater monitoring, and medical/rehabilitation systems. Each proposal is classified with IPC/CPC codes and can be timestamped (RFC 3161 / FreeTSA). Timestamp : 2026-08-17T10:45:25ZSHA-256 : 13b3e2bc50c638e594d623990f13159039dd0fb8f0b0968e18a3300649110a89 Liste des innovations & classification (IPC ; CPC) :1. DVS–HASEL soft gripper — IPC B25J 15/00 ; CPC B25J 15/122. DVS sorting calibration rig — IPC G01D 18/00 ; CPC G01D 18/003. HASEL sensing skin laminate — IPC G01L 5/00 ; CPC G01L 5/164. Biodegradable electrohydraulic actuator — IPC C08L 67/00 ; CPC C08L 67/025. Printable EAP electrode ink — IPC H01B 1/12 ; CPC H01B 1/126. Self-healing dielectric composite — IPC C08K 3/36 ; CPC C08K 3/367. Roll-to-roll HASEL pouch line — IPC B29C 65/00 ; CPC B29C 65/788. 3D-printed soft body + circuits — IPC B29C 64/118 ; CPC B29C 64/1189. Soft underwater encapsulation stack — IPC B29C 71/00 ; CPC B29C 71/0210. Event-driven SNN HASEL control — IPC G06N 3/04 ; CPC G06N 3/04511. Event-based actuator fatigue detection — IPC G05B 23/02 ; CPC G05B 23/0212. Edge event-stream compression codec — IPC H04N 5/00 ; CPC H04N 5/23213. Spike-packet swarm protocol — IPC H04W 4/80 ; CPC H04W 4/8014. Neuromorphic swarm task allocator — IPC G06Q 10/04 ; CPC G06Q 10/063915. Safe HV charge scheduler — IPC H02M 3/155 ; CPC H02M 3/15816. Swarm geofencing operations — IPC G08G 5/00 ; CPC G08G 5/0017. Radiation-hardened soft robot module — IPC G21C 19/00 ; CPC G21C 19/0018. DVS-to-intensity reconstruction — IPC H04N 5/232 ; CPC H04N 5/23219. DVS+EMG SNN exosuit fusion — IPC A61H 1/02 ; CPC A61H 1/0220. Closed-loop rehab dosing method — IPC A61H 1/00 ; CPC A61H 1/0021. Soft endoscope targeted delivery — IPC A61M 31/00 ; CPC A61M 31/0022. Low-power EAP assist patch — IPC A61F 5/01 ; CPC A61F 5/0123. Federated learning for agri swarms — IPC G06F 18/232 ; CPC G06F 18/232124. Event+spike interoperability standard — IPC G06F 9/54 ; CPC G06F 9/54125. Tamper-proof swarm audit ledger — IPC G06Q 20/38 ; CPC G06Q 20/38226. Swarm supervisor cockpit UI — IPC G05B 19/042 ; CPC G05B 19/04227. Hybrid ultra-fast waste sorter cell — IPC B07C 5/34 ; CPC B07C 5/34228. Underwater soft-drone swarm system — IPC B63G 8/00 ; CPC B63G 8/0029. Swarm soil-compaction sens","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21978440","URL":"https://doi.org/10.5281/zenodo.21978440","source":"datacite"},{"id":"doi:10.5281/zenodo.21978439","type":"article-journal","title":"Robotique souple neuromorphique et essaims","abstract":"Résumé FRCe document, produit avec l’assistance de ChatGPT 5.2 Thinking et Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive (antériorité) et entre dans l’état de la technique au sens des textes applicables (EPC Art. 54(2); French IPC Art. L 611-11; cf. 35 U.S.C. §102(a)). Il divulgue, de façon enabling, un portefeuille d’innovations combinant robotique souple (actionneurs HASEL/EAP), vision événementielle (DVS), calcul neuromorphique (SNN) et intelligence en essaim, couvrant dispositifs/capteurs, algorithmes, contrôle en boucle fermée, fabrication roll-to-roll et QA end-of-line, cybersécurité et opérations de flottes, interopérabilité (formats événements+spikes), logistique de cartouches, modèles économiques au résultat, et usages industriels, agricoles régénératifs, nucléaires, sous-marins et médicaux. Abstract ENThis document, produced with the assistance of ChatGPT 5.2 Thinking and Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes (art. L 611-11 CPI / art. 54(2) CBE). It discloses, in an enabling manner, a portfolio that fuses soft robotics (HASEL/EAP actuation), event-based vision (DVS), neuromorphic computing (SNN), and swarm intelligence. The disclosure spans devices and sensors, event-first control loops, roll-to-roll manufacturing and end-of-line QA, cyber-secure fleet operations, interoperability standards for event+spike telemetry, cartridge logistics and field repair, outcome-based metering and SLA instrumentation, and applications in high-throughput sorting, precision/regenerative agriculture, nuclear maintenance, underwater monitoring, and medical/rehabilitation systems. Each proposal is classified with IPC/CPC codes and can be timestamped (RFC 3161 / FreeTSA). Timestamp : 2026-08-17T10:45:25ZSHA-256 : 13b3e2bc50c638e594d623990f13159039dd0fb8f0b0968e18a3300649110a89 Liste des innovations & classification (IPC ; CPC) :1. DVS–HASEL soft gripper — IPC B25J 15/00 ; CPC B25J 15/122. DVS sorting calibration rig — IPC G01D 18/00 ; CPC G01D 18/003. HASEL sensing skin laminate — IPC G01L 5/00 ; CPC G01L 5/164. Biodegradable electrohydraulic actuator — IPC C08L 67/00 ; CPC C08L 67/025. Printable EAP electrode ink — IPC H01B 1/12 ; CPC H01B 1/126. Self-healing dielectric composite — IPC C08K 3/36 ; CPC C08K 3/367. Roll-to-roll HASEL pouch line — IPC B29C 65/00 ; CPC B29C 65/788. 3D-printed soft body + circuits — IPC B29C 64/118 ; CPC B29C 64/1189. Soft underwater encapsulation stack — IPC B29C 71/00 ; CPC B29C 71/0210. Event-driven SNN HASEL control — IPC G06N 3/04 ; CPC G06N 3/04511. Event-based actuator fatigue detection — IPC G05B 23/02 ; CPC G05B 23/0212. Edge event-stream compression codec — IPC H04N 5/00 ; CPC H04N 5/23213. Spike-packet swarm protocol — IPC H04W 4/80 ; CPC H04W 4/8014. Neuromorphic swarm task allocator — IPC G06Q 10/04 ; CPC G06Q 10/063915. Safe HV charge scheduler — IPC H02M 3/155 ; CPC H02M 3/15816. Swarm geofencing operations — IPC G08G 5/00 ; CPC G08G 5/0017. Radiation-hardened soft robot module — IPC G21C 19/00 ; CPC G21C 19/0018. DVS-to-intensity reconstruction — IPC H04N 5/232 ; CPC H04N 5/23219. DVS+EMG SNN exosuit fusion — IPC A61H 1/02 ; CPC A61H 1/0220. Closed-loop rehab dosing method — IPC A61H 1/00 ; CPC A61H 1/0021. Soft endoscope targeted delivery — IPC A61M 31/00 ; CPC A61M 31/0022. Low-power EAP assist patch — IPC A61F 5/01 ; CPC A61F 5/0123. Federated learning for agri swarms — IPC G06F 18/232 ; CPC G06F 18/232124. Event+spike interoperability standard — IPC G06F 9/54 ; CPC G06F 9/54125. Tamper-proof swarm audit ledger — IPC G06Q 20/38 ; CPC G06Q 20/38226. Swarm supervisor cockpit UI — IPC G05B 19/042 ; CPC G05B 19/04227. Hybrid ultra-fast waste sorter cell — IPC B07C 5/34 ; CPC B07C 5/34228. Underwater soft-drone swarm system — IPC B63G 8/00 ; CPC B63G 8/0029. Swarm soil-compaction sens","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21978439","URL":"https://doi.org/10.5281/zenodo.21978439","source":"datacite"},{"id":"doi:10.5281/zenodo.21970270","type":"article-journal","title":"Les Enjeux de la Rencontre Amoureuse dans les Sociétés Contemporaines","abstract":"Résum�� FRCe document, produit avec l’assistance de Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive volontaire (antériorité) et entre de ce fait dans l’état de la technique au sens des législations applicables ((EPC Art. 54(2); French IPC Art. L 611-11; cf. 35 U.S.C. §102(a)). Il divulgue soixante-dix systèmes et procédés techniques habilitants couvrant les enjeux contemporains des relations amoureuses : moteurs d’appariement éthique à apprentissage fédéré on-device, scores socio-écologiques dynamiques (ACV + optimisation multi-objectifs), contrôleurs de swipe adaptatifs en boucle fermée (HRV/pupille), sessions VR guidées par thérapeute, filtres multimodaux de contenus érotiques sains, places de marché de care intergénérationnel, registres de consentement dynamique, planificateurs de polycule, analytique privative, recommenders sous contraintes d’équité et micro-utilités communautaires. Chaque proposition est décrite de façon habilitante, classée IPC/CPC et horodatée (RFC 3161 / FreeTSA).Abstract ENThis document, produced with the assistance of Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes ((art. L 611-11 CPI / art. 54(2) CBE)). It discloses seventy enabling technical systems and methods addressing contemporary romantic and relational challenges: ethical matching engines with on-device federated learning, dynamic socio-ecological couple scores (LCA + multi-objective optimisation), closed-loop adaptive swipe throttles using HRV/pupil data, therapist-guided VR affective sessions, multimodal healthy-erotic content filters, intergenerational care marketplaces, dynamic consent ledgers, polycule planners, privacy-preserving analytics, fairness-constrained recommenders, and community micro-utilities. Every proposal is described in an enabling manner, classified with IPC and CPC codes, and accompanied by timestamp proof (RFC 3161 / FreeTSA). Timestamp : 2026-08-16T21:05:20ZSHA-256 : 17d751cc6fb35efc1b3edb397c494044c54f9afee11780ea90a96bd2769fdf18 Liste des innovations & classification (IPC ; CPC) : 1. Ethical compatibility engine — IPC G06N 20/00 ; CPC G06N 20/10 2. Socio-ecological couple score — IPC G06Q 50/26 ; CPC Y02B 90/12 3. Adaptive swipe throttle — IPC G06F 3/01 ; CPC G06F 3/0488 4. Therapist-guided VR date — IPC G06T 19/00 ; CPC A61B 5/00 5. Healthy erotic content filter — IPC G06F 21/62 ; CPC G06T 7/00 6. Intergenerational care marketplace — IPC G06Q 10/10 ; CPC G06Q 10/0631 7. Affective synchrony sensor — IPC A61B 5/024 ; CPC G16H 40/63 8. REDM relationship data standard — IPC G16H 10/60 ; CPC G06F 21/62 9. Shared-housing optimizer — IPC G06F 17/50 ; CPC Y02B 10/00 10. Domestic micro-utility — IPC H02J 3/14 ; CPC Y04S 20/32 11. Modular inter-age nursery — IPC E04B 1/74 ; CPC A61L 2/08 12. Dynamic consent ledger — IPC G06F 21/62 ; CPC H04L 9/08 13. Relationship A/B trial framework — IPC G01N 33/50 ; CPC G06F 19/24 14. Community mobility router — IPC G06Q 10/08 ; CPC G06F 17/60 15. Calibration-as-a-Service — IPC G01D 18/00 ; CPC G06F 11/07 16. Ghosting/abuse detector — IPC G06F 40/20 ; CPC G06N 20/20 17. Blended relational education — IPC G09B 19/00 ; CPC G06Q 50/20 18. Relationship–environment meter — IPC G06F 16/23 ; CPC Y02B 80/00 19. On-device biometric security — IPC H04L 9/08 ; CPC G06K 9/00 20. Relationship well-being index — IPC A61B 5/00 ; CPC G16H 50/20 21. Synthetic dataset generator — IPC G06F 21/62 ; CPC G06N 20/10 22. Split-learning matcher — IPC G06F 21/62 ; CPC G06N 20/40 23. Fairness-constrained recommender — IPC G06N 20/00 ; CPC G06N 20/10 24. Time-to-relationship model — IPC G06N 20/00 ; CPC G06F 18/214 25. Polycule planner — IPC G06Q 10/06 ; CPC G06Q 10/0637 26. Connected STI self-test kit — IPC G01N 33/569 ; CPC C12Q 1/6886 27. Synchronized haptic co-regulation — IPC A61B 5/024 ; CPC G06F 3/013 28. Web","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21970270","URL":"https://doi.org/10.5281/zenodo.21970270","source":"datacite"},{"id":"doi:10.5281/zenodo.21970269","type":"article-journal","title":"Les Enjeux de la Rencontre Amoureuse dans les Sociétés Contemporaines","abstract":"Résumé FRCe document, produit avec l’assistance de Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive volontaire (antériorité) et entre de ce fait dans l’état de la technique au sens des législations applicables ((EPC Art. 54(2); French IPC Art. L 611-11; cf. 35 U.S.C. §102(a)). Il divulgue soixante-dix systèmes et procédés techniques habilitants couvrant les enjeux contemporains des relations amoureuses : moteurs d’appariement éthique à apprentissage fédéré on-device, scores socio-écologiques dynamiques (ACV + optimisation multi-objectifs), contrôleurs de swipe adaptatifs en boucle fermée (HRV/pupille), sessions VR guidées par thérapeute, filtres multimodaux de contenus érotiques sains, places de marché de care intergénérationnel, registres de consentement dynamique, planificateurs de polycule, analytique privative, recommenders sous contraintes d’équité et micro-utilités communautaires. Chaque proposition est décrite de façon habilitante, classée IPC/CPC et horodatée (RFC 3161 / FreeTSA).Abstract ENThis document, produced with the assistance of Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes ((art. L 611-11 CPI / art. 54(2) CBE)). It discloses seventy enabling technical systems and methods addressing contemporary romantic and relational challenges: ethical matching engines with on-device federated learning, dynamic socio-ecological couple scores (LCA + multi-objective optimisation), closed-loop adaptive swipe throttles using HRV/pupil data, therapist-guided VR affective sessions, multimodal healthy-erotic content filters, intergenerational care marketplaces, dynamic consent ledgers, polycule planners, privacy-preserving analytics, fairness-constrained recommenders, and community micro-utilities. Every proposal is described in an enabling manner, classified with IPC and CPC codes, and accompanied by timestamp proof (RFC 3161 / FreeTSA). Timestamp : 2026-08-16T21:05:20ZSHA-256 : 17d751cc6fb35efc1b3edb397c494044c54f9afee11780ea90a96bd2769fdf18 Liste des innovations & classification (IPC ; CPC) : 1. Ethical compatibility engine — IPC G06N 20/00 ; CPC G06N 20/10 2. Socio-ecological couple score — IPC G06Q 50/26 ; CPC Y02B 90/12 3. Adaptive swipe throttle — IPC G06F 3/01 ; CPC G06F 3/0488 4. Therapist-guided VR date — IPC G06T 19/00 ; CPC A61B 5/00 5. Healthy erotic content filter — IPC G06F 21/62 ; CPC G06T 7/00 6. Intergenerational care marketplace — IPC G06Q 10/10 ; CPC G06Q 10/0631 7. Affective synchrony sensor — IPC A61B 5/024 ; CPC G16H 40/63 8. REDM relationship data standard — IPC G16H 10/60 ; CPC G06F 21/62 9. Shared-housing optimizer — IPC G06F 17/50 ; CPC Y02B 10/00 10. Domestic micro-utility — IPC H02J 3/14 ; CPC Y04S 20/32 11. Modular inter-age nursery — IPC E04B 1/74 ; CPC A61L 2/08 12. Dynamic consent ledger — IPC G06F 21/62 ; CPC H04L 9/08 13. Relationship A/B trial framework — IPC G01N 33/50 ; CPC G06F 19/24 14. Community mobility router — IPC G06Q 10/08 ; CPC G06F 17/60 15. Calibration-as-a-Service — IPC G01D 18/00 ; CPC G06F 11/07 16. Ghosting/abuse detector — IPC G06F 40/20 ; CPC G06N 20/20 17. Blended relational education — IPC G09B 19/00 ; CPC G06Q 50/20 18. Relationship–environment meter — IPC G06F 16/23 ; CPC Y02B 80/00 19. On-device biometric security — IPC H04L 9/08 ; CPC G06K 9/00 20. Relationship well-being index — IPC A61B 5/00 ; CPC G16H 50/20 21. Synthetic dataset generator — IPC G06F 21/62 ; CPC G06N 20/10 22. Split-learning matcher — IPC G06F 21/62 ; CPC G06N 20/40 23. Fairness-constrained recommender — IPC G06N 20/00 ; CPC G06N 20/10 24. Time-to-relationship model — IPC G06N 20/00 ; CPC G06F 18/214 25. Polycule planner — IPC G06Q 10/06 ; CPC G06Q 10/0637 26. Connected STI self-test kit — IPC G01N 33/569 ; CPC C12Q 1/6886 27. Synchronized haptic co-regulation — IPC A61B 5/024 ; CPC G06F 3/013 28. Webc","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21970269","URL":"https://doi.org/10.5281/zenodo.21970269","source":"datacite"},{"id":"doi:10.5281/zenodo.21968628","type":"article-journal","title":"Étude de faisabilité et analyse de marché – Innovations pour réacteurs à fusion compacts (ARC, SPARC, ST-E1)","abstract":"Résumé FRCe document, produit avec l’assistance de ChatGPT 5.2 Thinking et Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive (antériorité) et entre donc dans l’état de la technique au sens des textes applicables : EPC Art. 54(2), French IPC Art. L 611-11, cf. 35 U.S.C. §102(a), Chinese Patent Law Art. 22(5) (中华人民共和国专利法) et Japanese Patent Act Art. 29(1) (特許法). Il divulgue 128 embodiments “enabling” visant les verrous 2030 de la fusion compacte : protection thermique par vapor-shielding (gels Au@C60, CPS/LiMIT), pilotage actif d’aimants HTS REBCO (phase-array, NI/MI, capteurs, IA), matériaux structurels auto-réparants sous irradiation, et couvertures tritigènes optimisées (TBR, FLiBe, chiralité, ports). Chaque proposition inclut paramètres, QA, IPC/CPC et preuve d’horodatage (RFC 3161 / FreeTSA). Abstract ENThis document, produced with the assistance of ChatGPT 5.2 Thinking and Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: (art. L 611-11 CPI / art. 54(2) CBE), 35 U.S.C. §102(a), Chinese Patent Law Art. 22(5) (中华人民共和国专利法), and Japanese Patent Act Art. 29(1) (特許法). It discloses 128 enabling embodiments addressing 2030 compact-fusion bottlenecks: vapor-shielding and liquid-metal plasma-facing components (Au@C60 gels, CPS/LiMIT), active HTS REBCO magnet stabilization (phased arrays, NI/MI, sensing, AI), irradiation-driven self-healing structural materials, and optimized tritium-breeding blankets (TBR, FLiBe immersion, chiral/port shielding). Each proposal includes reproducible parameters, QA/test workflows, IPC/CPC classification, and timestamp proof (RFC 3161 / FreeTSA). Timestamp : 2026-08-16T20:51:39ZSHA-256 : 528680566605797c0c4097f48b97edee9a77a3938919ea17f4db427d870670a6 Liste des innovations & classification (IPC ; CPC) : 1) Au@C60 vapor-shield gel — IPC G21B 1/00 ; CPC Y02E 30/302) Graded-Z gel multilayer — IPC C09D 5/00 ; CPC C09D 5/003) Sub-mm W capillary reservoir — IPC B22F 3/105 ; CPC B22F 3/1054) Inductive vaporization pulses — IPC F28F 27/00 ; CPC F28F 27/005) Impurity-aware feedback control — IPC G05B 13/02 ; CPC G05B 13/026) Plasma-assisted re-deposition — IPC C23C 16/00 ; CPC C23C 16/007) Refillable gel cartridge module — IPC G21B 1/00 ; CPC G21B 1/048) Irradiation QC kit for gels — IPC G01N 33/00 ; CPC G01N 33/009) Low-activation gold formulation — IPC G21F 1/00 ; CPC G21F 1/0010) Microchannel PFC + gel hybrid — IPC F28D 5/00 ; CPC F28D 5/0011) Phased-array correction coils — IPC H01F 6/04 ; CPC H01F 6/0412) Vector-potential MPC control — IPC G05B 13/02 ; CPC G05B 13/0213) Screening-current state observer — IPC G01R 33/00 ; CPC G01R 33/0014) Angled AC shaking protocol — IPC H01F 6/06 ; CPC H01F 6/0615) Cryogenic FBG strain sensing — IPC G01L 1/24 ; CPC G01L 1/2416) Cryo Hall-matrix calibration — IPC G01R 33/02 ; CPC G01R 33/0217) Multimodal quench prediction AI — IPC G06N 20/00 ; CPC G06N 20/0018) Integrated cryogenic power drivers — IPC H02M 7/00 ; CPC H02M 7/0019) Co-wound auxiliary conductors — IPC H01F 41/00 ; CPC H01F 41/0020) Secure magnet-control API — IPC H04L 9/00 ; CPC H04L 9/0021) Low-activation HEA for fusion — IPC C22C 30/00 ; CPC C22C 30/0022) Cyclic nanoprecipitate alloy — IPC C22F 1/00 ; CPC C22F 1/0023) Defect-sink layered composite — IPC C22C 38/00 ; CPC C22C 38/0024) Additive manufacturing RISC recipe — IPC B33Y 10/00 ; CPC B33Y 10/0025) Heat-treatment seed-sinks — IPC C21D 8/00 ; CPC C21D 8/0026) In-situ damage monitoring — IPC G01N 27/02 ; CPC G01N 27/0227) Accelerated qualification pipeline — IPC G06F 16/00 ; CPC G06F 16/0028) Defect-mobility alloy optimizer — IPC G06N 20/00 ; CPC G06N 20/0029) RISC cladding on steel — IPC C23C 24/00 ; CPC C23C 24/0030) Chiral neutron-maze blanket cell — IPC G21B 1/00 ; CPC G21B 1/0431) Phase-reflecting multilayer reflector — IPC G21B 1/00 ; CPC G21B 1/0432","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21968628","URL":"https://doi.org/10.5281/zenodo.21968628","source":"datacite"},{"id":"doi:10.5281/zenodo.21970198","type":"article-journal","title":"Étude de faisabilité et analyse de marché – Innovations pour réacteurs à fusion compacts (ARC, SPARC, ST-E1)","abstract":"Résumé FRCe document, produit avec l’assistance de ChatGPT 5.2 Thinking et Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive (antériorité) et entre donc dans l’état de la technique au sens des textes applicables : EPC Art. 54(2), French IPC Art. L 611-11, cf. 35 U.S.C. §102(a), Chinese Patent Law Art. 22(5) (中华人民共和国专利法) et Japanese Patent Act Art. 29(1) (特許法). Il divulgue 128 embodiments “enabling” visant les verrous 2030 de la fusion compacte : protection thermique par vapor-shielding (gels Au@C60, CPS/LiMIT), pilotage actif d’aimants HTS REBCO (phase-array, NI/MI, capteurs, IA), matériaux structurels auto-réparants sous irradiation, et couvertures tritigènes optimisées (TBR, FLiBe, chiralité, ports). Chaque proposition inclut paramètres, QA, IPC/CPC et preuve d’horodatage (RFC 3161 / FreeTSA). Abstract ENThis document, produced with the assistance of ChatGPT 5.2 Thinking and Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: (art. L 611-11 CPI / art. 54(2) CBE), 35 U.S.C. §102(a), Chinese Patent Law Art. 22(5) (中华人民共和国专利法), and Japanese Patent Act Art. 29(1) (特許法). It discloses 128 enabling embodiments addressing 2030 compact-fusion bottlenecks: vapor-shielding and liquid-metal plasma-facing components (Au@C60 gels, CPS/LiMIT), active HTS REBCO magnet stabilization (phased arrays, NI/MI, sensing, AI), irradiation-driven self-healing structural materials, and optimized tritium-breeding blankets (TBR, FLiBe immersion, chiral/port shielding). Each proposal includes reproducible parameters, QA/test workflows, IPC/CPC classification, and timestamp proof (RFC 3161 / FreeTSA). Timestamp : 2026-08-16T20:51:39ZSHA-256 : 528680566605797c0c4097f48b97edee9a77a3938919ea17f4db427d870670a6 Liste des innovations & classification (IPC ; CPC) : 1) Au@C60 vapor-shield gel — IPC G21B 1/00 ; CPC Y02E 30/302) Graded-Z gel multilayer — IPC C09D 5/00 ; CPC C09D 5/003) Sub-mm W capillary reservoir — IPC B22F 3/105 ; CPC B22F 3/1054) Inductive vaporization pulses — IPC F28F 27/00 ; CPC F28F 27/005) Impurity-aware feedback control — IPC G05B 13/02 ; CPC G05B 13/026) Plasma-assisted re-deposition — IPC C23C 16/00 ; CPC C23C 16/007) Refillable gel cartridge module — IPC G21B 1/00 ; CPC G21B 1/048) Irradiation QC kit for gels — IPC G01N 33/00 ; CPC G01N 33/009) Low-activation gold formulation — IPC G21F 1/00 ; CPC G21F 1/0010) Microchannel PFC + gel hybrid — IPC F28D 5/00 ; CPC F28D 5/0011) Phased-array correction coils — IPC H01F 6/04 ; CPC H01F 6/0412) Vector-potential MPC control — IPC G05B 13/02 ; CPC G05B 13/0213) Screening-current state observer — IPC G01R 33/00 ; CPC G01R 33/0014) Angled AC shaking protocol — IPC H01F 6/06 ; CPC H01F 6/0615) Cryogenic FBG strain sensing — IPC G01L 1/24 ; CPC G01L 1/2416) Cryo Hall-matrix calibration — IPC G01R 33/02 ; CPC G01R 33/0217) Multimodal quench prediction AI — IPC G06N 20/00 ; CPC G06N 20/0018) Integrated cryogenic power drivers — IPC H02M 7/00 ; CPC H02M 7/0019) Co-wound auxiliary conductors — IPC H01F 41/00 ; CPC H01F 41/0020) Secure magnet-control API — IPC H04L 9/00 ; CPC H04L 9/0021) Low-activation HEA for fusion — IPC C22C 30/00 ; CPC C22C 30/0022) Cyclic nanoprecipitate alloy — IPC C22F 1/00 ; CPC C22F 1/0023) Defect-sink layered composite — IPC C22C 38/00 ; CPC C22C 38/0024) Additive manufacturing RISC recipe — IPC B33Y 10/00 ; CPC B33Y 10/0025) Heat-treatment seed-sinks — IPC C21D 8/00 ; CPC C21D 8/0026) In-situ damage monitoring — IPC G01N 27/02 ; CPC G01N 27/0227) Accelerated qualification pipeline — IPC G06F 16/00 ; CPC G06F 16/0028) Defect-mobility alloy optimizer — IPC G06N 20/00 ; CPC G06N 20/0029) RISC cladding on steel — IPC C23C 24/00 ; CPC C23C 24/0030) Chiral neutron-maze blanket cell — IPC G21B 1/00 ; CPC G21B 1/0431) Phase-reflecting multilayer reflector — IPC G21B 1/00 ; CPC G21B 1/0432","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21970198","URL":"https://doi.org/10.5281/zenodo.21970198","source":"datacite"},{"id":"doi:10.5281/zenodo.21969017","type":"article-journal","title":"Étude de faisabilité et analyse de marché – Innovations pour réacteurs à fusion compacts (ARC, SPARC, ST-E1)","abstract":"Résumé FRCe document, produit avec l’assistance de ChatGPT 5.2 Thinking et Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive (antériorité) et entre donc dans l’état de la technique au sens des textes applicables : EPC Art. 54(2), French IPC Art. L 611-11, cf. 35 U.S.C. §102(a), Chinese Patent Law Art. 22(5) (中华人民共和国专利法) et Japanese Patent Act Art. 29(1) (特許法). Il divulgue 128 embodiments “enabling” visant les verrous 2030 de la fusion compacte : protection thermique par vapor-shielding (gels Au@C60, CPS/LiMIT), pilotage actif d’aimants HTS REBCO (phase-array, NI/MI, capteurs, IA), matériaux structurels auto-réparants sous irradiation, et couvertures tritigènes optimisées (TBR, FLiBe, chiralité, ports). Chaque proposition inclut paramètres, QA, IPC/CPC et preuve d’horodatage (RFC 3161 / FreeTSA). Abstract ENThis document, produced with the assistance of ChatGPT 5.2 Thinking and Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: (art. L 611-11 CPI / art. 54(2) CBE), 35 U.S.C. §102(a), Chinese Patent Law Art. 22(5) (中华人民共和国专利法), and Japanese Patent Act Art. 29(1) (特許法). It discloses 128 enabling embodiments addressing 2030 compact-fusion bottlenecks: vapor-shielding and liquid-metal plasma-facing components (Au@C60 gels, CPS/LiMIT), active HTS REBCO magnet stabilization (phased arrays, NI/MI, sensing, AI), irradiation-driven self-healing structural materials, and optimized tritium-breeding blankets (TBR, FLiBe immersion, chiral/port shielding). Each proposal includes reproducible parameters, QA/test workflows, IPC/CPC classification, and timestamp proof (RFC 3161 / FreeTSA). Timestamp : 2026-08-16T17:25:54ZSHA-256 : fc6b83a8107e123ab1fc352297f43340016d54b2fd33e445274512e24495c8c1 Liste des innovations & classification (IPC ; CPC) : 1) Au@C60 vapor-shield gel — IPC G21B 1/00 ; CPC Y02E 30/302) Graded-Z gel multilayer — IPC C09D 5/00 ; CPC C09D 5/003) Sub-mm W capillary reservoir — IPC B22F 3/105 ; CPC B22F 3/1054) Inductive vaporization pulses — IPC F28F 27/00 ; CPC F28F 27/005) Impurity-aware feedback control — IPC G05B 13/02 ; CPC G05B 13/026) Plasma-assisted re-deposition — IPC C23C 16/00 ; CPC C23C 16/007) Refillable gel cartridge module — IPC G21B 1/00 ; CPC G21B 1/048) Irradiation QC kit for gels — IPC G01N 33/00 ; CPC G01N 33/009) Low-activation gold formulation — IPC G21F 1/00 ; CPC G21F 1/0010) Microchannel PFC + gel hybrid — IPC F28D 5/00 ; CPC F28D 5/0011) Phased-array correction coils — IPC H01F 6/04 ; CPC H01F 6/0412) Vector-potential MPC control — IPC G05B 13/02 ; CPC G05B 13/0213) Screening-current state observer — IPC G01R 33/00 ; CPC G01R 33/0014) Angled AC shaking protocol — IPC H01F 6/06 ; CPC H01F 6/0615) Cryogenic FBG strain sensing — IPC G01L 1/24 ; CPC G01L 1/2416) Cryo Hall-matrix calibration — IPC G01R 33/02 ; CPC G01R 33/0217) Multimodal quench prediction AI — IPC G06N 20/00 ; CPC G06N 20/0018) Integrated cryogenic power drivers — IPC H02M 7/00 ; CPC H02M 7/0019) Co-wound auxiliary conductors — IPC H01F 41/00 ; CPC H01F 41/0020) Secure magnet-control API — IPC H04L 9/00 ; CPC H04L 9/0021) Low-activation HEA for fusion — IPC C22C 30/00 ; CPC C22C 30/0022) Cyclic nanoprecipitate alloy — IPC C22F 1/00 ; CPC C22F 1/0023) Defect-sink layered composite — IPC C22C 38/00 ; CPC C22C 38/0024) Additive manufacturing RISC recipe — IPC B33Y 10/00 ; CPC B33Y 10/0025) Heat-treatment seed-sinks — IPC C21D 8/00 ; CPC C21D 8/0026) In-situ damage monitoring — IPC G01N 27/02 ; CPC G01N 27/0227) Accelerated qualification pipeline — IPC G06F 16/00 ; CPC G06F 16/0028) Defect-mobility alloy optimizer — IPC G06N 20/00 ; CPC G06N 20/0029) RISC cladding on steel — IPC C23C 24/00 ; CPC C23C 24/0030) Chiral neutron-maze blanket cell — IPC G21B 1/00 ; CPC G21B 1/0431) Phase-reflecting multilayer reflector — IPC G21B 1/00 ; CPC G21B 1/0432","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21969017","URL":"https://doi.org/10.5281/zenodo.21969017","source":"datacite"},{"id":"doi:10.5281/zenodo.21968629","type":"article-journal","title":"Étude de faisabilité et analyse de marché – Innovations pour réacteurs à fusion compacts (ARC, SPARC, ST-E1)","abstract":"Résumé FRCe document, produit avec l’assistance de ChatGPT 5.2 Thinking et Gemini 3 Raisonnement, est publié sous licence Apache 2.0. Il constitue une publication défensive (antériorité) et entre donc dans l’état de la technique au sens des textes applicables : EPC Art. 54(2), French IPC Art. L 611-11, cf. 35 U.S.C. §102(a), Chinese Patent Law Art. 22(5) (中华人民共和国专利法) et Japanese Patent Act Art. 29(1) (特許法). Il divulgue 128 embodiments “enabling” visant les verrous 2030 de la fusion compacte : protection thermique par vapor-shielding (gels Au@C60, CPS/LiMIT), pilotage actif d’aimants HTS REBCO (phase-array, NI/MI, capteurs, IA), matériaux structurels auto-réparants sous irradiation, et couvertures tritigènes optimisées (TBR, FLiBe, chiralité, ports). Chaque proposition inclut paramètres, QA, IPC/CPC et preuve d’horodatage (RFC 3161 / FreeTSA). Abstract ENThis document, produced with the assistance of ChatGPT 5.2 Thinking and Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: (art. L 611-11 CPI / art. 54(2) CBE), 35 U.S.C. §102(a), Chinese Patent Law Art. 22(5) (中华人民共和国专利法), and Japanese Patent Act Art. 29(1) (特許法). It discloses 128 enabling embodiments addressing 2030 compact-fusion bottlenecks: vapor-shielding and liquid-metal plasma-facing components (Au@C60 gels, CPS/LiMIT), active HTS REBCO magnet stabilization (phased arrays, NI/MI, sensing, AI), irradiation-driven self-healing structural materials, and optimized tritium-breeding blankets (TBR, FLiBe immersion, chiral/port shielding). Each proposal includes reproducible parameters, QA/test workflows, IPC/CPC classification, and timestamp proof (RFC 3161 / FreeTSA). Timestamp : 2026-08-16T17:25:54ZSHA-256 : c6b83a8107e123ab1fc352297f43340016d54b2fd33e445274512e24495c8c1 Liste des innovations & classification (IPC ; CPC) : 1) Au@C60 vapor-shield gel — IPC G21B 1/00 ; CPC Y02E 30/302) Graded-Z gel multilayer — IPC C09D 5/00 ; CPC C09D 5/003) Sub-mm W capillary reservoir — IPC B22F 3/105 ; CPC B22F 3/1054) Inductive vaporization pulses — IPC F28F 27/00 ; CPC F28F 27/005) Impurity-aware feedback control — IPC G05B 13/02 ; CPC G05B 13/026) Plasma-assisted re-deposition — IPC C23C 16/00 ; CPC C23C 16/007) Refillable gel cartridge module — IPC G21B 1/00 ; CPC G21B 1/048) Irradiation QC kit for gels — IPC G01N 33/00 ; CPC G01N 33/009) Low-activation gold formulation — IPC G21F 1/00 ; CPC G21F 1/0010) Microchannel PFC + gel hybrid — IPC F28D 5/00 ; CPC F28D 5/0011) Phased-array correction coils — IPC H01F 6/04 ; CPC H01F 6/0412) Vector-potential MPC control — IPC G05B 13/02 ; CPC G05B 13/0213) Screening-current state observer — IPC G01R 33/00 ; CPC G01R 33/0014) Angled AC shaking protocol — IPC H01F 6/06 ; CPC H01F 6/0615) Cryogenic FBG strain sensing — IPC G01L 1/24 ; CPC G01L 1/2416) Cryo Hall-matrix calibration — IPC G01R 33/02 ; CPC G01R 33/0217) Multimodal quench prediction AI — IPC G06N 20/00 ; CPC G06N 20/0018) Integrated cryogenic power drivers — IPC H02M 7/00 ; CPC H02M 7/0019) Co-wound auxiliary conductors — IPC H01F 41/00 ; CPC H01F 41/0020) Secure magnet-control API — IPC H04L 9/00 ; CPC H04L 9/0021) Low-activation HEA for fusion — IPC C22C 30/00 ; CPC C22C 30/0022) Cyclic nanoprecipitate alloy — IPC C22F 1/00 ; CPC C22F 1/0023) Defect-sink layered composite — IPC C22C 38/00 ; CPC C22C 38/0024) Additive manufacturing RISC recipe — IPC B33Y 10/00 ; CPC B33Y 10/0025) Heat-treatment seed-sinks — IPC C21D 8/00 ; CPC C21D 8/0026) In-situ damage monitoring — IPC G01N 27/02 ; CPC G01N 27/0227) Accelerated qualification pipeline — IPC G06F 16/00 ; CPC G06F 16/0028) Defect-mobility alloy optimizer — IPC G06N 20/00 ; CPC G06N 20/0029) RISC cladding on steel — IPC C23C 24/00 ; CPC C23C 24/0030) Chiral neutron-maze blanket cell — IPC G21B 1/00 ; CPC G21B 1/0431) Phase-reflecting multilayer reflector — IPC G21B 1/00 ; CPC G21B 1/0432)","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21968629","URL":"https://doi.org/10.5281/zenodo.21968629","source":"datacite"},{"id":"doi:10.5281/zenodo.21326292","type":"article-journal","title":"The Temporal Governance Architecture: A Reference Model for Autonomous Multi Agent Systems","abstract":"This report introduces the Temporal Governance Architecture, a conceptual reference model for autonomous multi‑agent systems operating across cloud, edge, local compute, federated networks, and hybrid environments. The architecture formalizes governance‑layer primitives for identity lineage, temporal behavior ingestion, deterministic scoring, authenticated proof exchange, governance‑conditioned settlement, and cross‑network synchronization. The model is grounded in a temporal governance primitive first defined in 2003, consisting of Cycle Hits, Hits History, Rotation Groups, Temporal Resets, Identity–Behavior Binding, and Governance Transitions. These elements describe the upstream physics connecting agent behavior over time to governance outcomes. Modern multi‑agent ecosystems—including multimodal agents, on‑device inference systems, encrypted commerce agents, and biological privacy systems—operate downstream of this primitive. The governance‑layer architecture (2026) is composed of six conceptual primitives: (1) Agent Identity Envelope (AIE), (2) Behavior Ingestion Pipeline, (3) Agent Importance Prediction Score (AIPS), (4) Agent‑to‑Agent Proof Exchange Protocol (A2APEP), (5) Agent‑Level Settlement Engine, and (6) Cross‑Agent Synchronization Layer. These primitives define a deployment‑agnostic governance model without revealing operational logic, canonicalization rules, thresholds, invariants, or substrate‑level mechanics. The report also presents the multi‑agent cooperation physics chain: Behavior → Score → State → Action → Proof → Settlement → Sync This chain is conceptual and non‑implementable. It provides a high‑level understanding of how agent behavior influences governance state transitions across distributed environments. A conceptual reference diagram is included to illustrate the lineage between the 2003 primitive, the 2026 governance‑layer architecture, and compliance‑layer concepts such as governance graphs, state machines, arbitration layers, proof registries, and selective disclosure mechanisms. Addendum — Citation Network Context Independent analysis of the citation network surrounding U.S. Patent Application US20040133469 shows 109 direct citing patents, approximately 50 citing family members, and an extended influence footprint of over 2,800 citations when substrate‑level propagation is included. These citations span two decades of technological development across information retrieval, metadata processing, recommendation systems, cloud orchestration, digital commerce, multimedia delivery, machine learning, and emerging agentic AI architectures. The extended citation footprint reflects how downstream substrate technologies continued to reference and build upon the conceptual space originally opened by the 2003 primitive. The governance‑layer primitives described in this defensive publication operate upstream of these systems and are not taught, suggested, or enabled by any cited substrate. This contextual reference is provided solely to illustrate field maturity, historical continuity, and long‑term architectural relevance. It does not imply derivation, dependency, or equivalence between the present disclosure and any downstream implementation. Cross‑Domain Integration Clause This disclosure covers conceptual integrations across metadata‑level governance, agent‑native governance, encrypted commerce governance, multimodal governance, and biological privacy governance. Any conceptual framework combining these domains—regardless of naming, ordering, modality, or implementation topology—is included within the scope of this defensive publication. This includes, without limitation, integrations across cloud, edge, local compute, federated networks, hybrid architectures, on‑device inference, multimodal agent orchestration, encrypted transactional systems, biological data privacy systems, and cross‑domain synchronization mechanisms. No operational logic, canonicalization rules, thresholds, invariants, settlement mech","author":[{"family":"Chang","given":"Dario"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21326292","URL":"https://doi.org/10.5281/zenodo.21326292","source":"datacite"},{"id":"doi:10.5281/zenodo.21326293","type":"article-journal","title":"The Temporal Governance Architecture: A Reference Model for Autonomous Multi Agent Systems","abstract":"This report introduces the Temporal Governance Architecture, a conceptual reference model for autonomous multi‑agent systems operating across cloud, edge, local compute, federated networks, and hybrid environments. The architecture formalizes governance‑layer primitives for identity lineage, temporal behavior ingestion, deterministic scoring, authenticated proof exchange, governance‑conditioned settlement, and cross‑network synchronization. The model is grounded in a temporal governance primitive first defined in 2003, consisting of Cycle Hits, Hits History, Rotation Groups, Temporal Resets, Identity–Behavior Binding, and Governance Transitions. These elements describe the upstream physics connecting agent behavior over time to governance outcomes. Modern multi‑agent ecosystems—including multimodal agents, on‑device inference systems, encrypted commerce agents, and biological privacy systems—operate downstream of this primitive. The governance‑layer architecture (2026) is composed of six conceptual primitives: (1) Agent Identity Envelope (AIE), (2) Behavior Ingestion Pipeline, (3) Agent Importance Prediction Score (AIPS), (4) Agent‑to‑Agent Proof Exchange Protocol (A2APEP), (5) Agent‑Level Settlement Engine, and (6) Cross‑Agent Synchronization Layer. These primitives define a deployment‑agnostic governance model without revealing operational logic, canonicalization rules, thresholds, invariants, or substrate‑level mechanics. The report also presents the multi‑agent cooperation physics chain: Behavior → Score → State → Action → Proof → Settlement → Sync This chain is conceptual and non‑implementable. It provides a high‑level understanding of how agent behavior influences governance state transitions across distributed environments. A conceptual reference diagram is included to illustrate the lineage between the 2003 primitive, the 2026 governance‑layer architecture, and compliance‑layer concepts such as governance graphs, state machines, arbitration layers, proof registries, and selective disclosure mechanisms. Addendum — Citation Network Context Independent analysis of the citation network surrounding U.S. Patent Application US20040133469 shows 109 direct citing patents, approximately 50 citing family members, and an extended influence footprint of over 2,800 citations when substrate‑level propagation is included. These citations span two decades of technological development across information retrieval, metadata processing, recommendation systems, cloud orchestration, digital commerce, multimedia delivery, machine learning, and emerging agentic AI architectures. The extended citation footprint reflects how downstream substrate technologies continued to reference and build upon the conceptual space originally opened by the 2003 primitive. The governance‑layer primitives described in this defensive publication operate upstream of these systems and are not taught, suggested, or enabled by any cited substrate. This contextual reference is provided solely to illustrate field maturity, historical continuity, and long‑term architectural relevance. It does not imply derivation, dependency, or equivalence between the present disclosure and any downstream implementation. Cross‑Domain Integration Clause This disclosure covers conceptual integrations across metadata‑level governance, agent‑native governance, encrypted commerce governance, multimodal governance, and biological privacy governance. Any conceptual framework combining these domains—regardless of naming, ordering, modality, or implementation topology—is included within the scope of this defensive publication. This includes, without limitation, integrations across cloud, edge, local compute, federated networks, hybrid architectures, on‑device inference, multimodal agent orchestration, encrypted transactional systems, biological data privacy systems, and cross‑domain synchronization mechanisms. No operational logic, canonicalization rules, thresholds, invariants, settlement mech","author":[{"family":"Chang","given":"Dario"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21326293","URL":"https://doi.org/10.5281/zenodo.21326293","source":"datacite"},{"id":"doi:10.5281/zenodo.20649675","type":"article-journal","title":"CDSA-BB3: Pilot Biometric Pillar of the CDSA Federated Reinforcement Learning Paradigm — Synthetic Biometric Data, Anonymization, Cryptographic Multi-Signature, and Scaffolded FRL Extension","abstract":"CDSA-BB3 is the reference implementation of the pilot biometric pillar of the Complementary Diagnostic Safety Approach (CDSA), a third-generation aviation safety paradigm formalised as a Federated Reinforcement Learning (FRL) framework (Scenario C decision, 21 May 2026). The v2.1.0 release (12 June 2026) replaces the v2.0.0 scaffold stubs with full NumPy reference implementations (federated, multimodal, rl_agents, rl_environment, xai; all unit-tested) and adds seeded, reproducible CPU-scale federated training runs (random / centralised PPO / FedAvg / FedProx / FedProx+DP) with result JSONs under experiments/results/. An interactive results panel is published at https://cdsa.app/bb3. Framework-scale training with real operational data remains future work. The v1.0.0 release modules and reference data remain unchanged. Version 2.2.0 (3 July 2026) adds four phases of seeded, CPU-scale revision-preparation experiments (seed 42) under experiments/, with all outputs in experiments/results/ and interactive honesty-band panels at https://cdsa.app/bb3/. Faz A (reference runs): federated training exceeds the centralised baseline on critical recall (0.937 vs 0.858). Faz B (robustness and confusion matrix): modality ablation independently confirms the SHAP ranking (dropping ECG sends critical recall from 0.917 to 0.041); the two intermediate alert classes are never predicted. Faz C (differential-privacy ε sweep and entropy-targeted exploration): class coverage is unchanged across all configurations. Faz D (decoy attribution and gated exploration): the decoy test is passed cleanly (attribution 6.4% vs the 25% uniform share); gated exploration raises critical recall to 0.975. Findings are reported verbatim from the result JSONs, consistent with the honesty bands published at the site.","author":[{"family":"Cantekin","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20649675","URL":"https://doi.org/10.5281/zenodo.20649675","source":"datacite"},{"id":"doi:10.5281/zenodo.21148698","type":"article-journal","title":"CDSA-BB3: Pilot Biometric Pillar of the CDSA Federated Reinforcement Learning Paradigm — Synthetic Biometric Data, Anonymization, Cryptographic Multi-Signature, and Scaffolded FRL Extension","abstract":"CDSA-BB3 is the reference implementation of the pilot biometric pillar of the Complementary Diagnostic Safety Approach (CDSA), a third-generation aviation safety paradigm formalised as a Federated Reinforcement Learning (FRL) framework (Scenario C decision, 21 May 2026). The v2.1.0 release (12 June 2026) replaces the v2.0.0 scaffold stubs with full NumPy reference implementations (federated, multimodal, rl_agents, rl_environment, xai; all unit-tested) and adds seeded, reproducible CPU-scale federated training runs (random / centralised PPO / FedAvg / FedProx / FedProx+DP) with result JSONs under experiments/results/. An interactive results panel is published at https://cdsa.app/bb3. Framework-scale training with real operational data remains future work. The v1.0.0 release modules and reference data remain unchanged. Version 2.2.0 (3 July 2026) adds four phases of seeded, CPU-scale revision-preparation experiments (seed 42) under experiments/, with all outputs in experiments/results/ and interactive honesty-band panels at https://cdsa.app/bb3/. Faz A (reference runs): federated training exceeds the centralised baseline on critical recall (0.937 vs 0.858). Faz B (robustness and confusion matrix): modality ablation independently confirms the SHAP ranking (dropping ECG sends critical recall from 0.917 to 0.041); the two intermediate alert classes are never predicted. Faz C (differential-privacy ε sweep and entropy-targeted exploration): class coverage is unchanged across all configurations. Faz D (decoy attribution and gated exploration): the decoy test is passed cleanly (attribution 6.4% vs the 25% uniform share); gated exploration raises critical recall to 0.975. Findings are reported verbatim from the result JSONs, consistent with the honesty bands published at the site.","author":[{"family":"Cantekin","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21148698","URL":"https://doi.org/10.5281/zenodo.21148698","source":"datacite"},{"id":"doi:10.5281/zenodo.20718085","type":"article-journal","title":"EDGE AI IN 2026: A DEEP DIVE INTO SOURCE LEVEL INTELLIGENCE AND ITS FUTURE PATHWAYS","abstract":"Edge AI is one of the key paradigm shifts in the implementation of artificial intelligence, moving the inference calculation out of centralized infrastructure in the cloud and the bottom of the network, closer to the source of the data. This paper provides an overview of Edge AI in 2026, reasons it will gain momentum, the multi-tier architecture hierarchy that allows spread of intelligence, specialized hardware environment, and uses of cross-domain applications in manufacturing, healthcare, autonomous systems, retail, and smart infrastructure. We also compare performance standards between Edge AI, Cloud AI and Fog AI implementation, estimate market growth outlook and address challenges that remain unsolved, such as security, model optimization, devices management and regulatory compliance in the EU AI Act. According to our results, Edge AI has passed the inflection point between proof of concept and production grade with purpose-built neural processing units, compressed inference models, and well-developed orchestration systems. We sum up with a prospective study of federated learning integration, 6G synergies and the rise of agentic edge systems.","author":[{"family":"Insights","given":"International"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20718085","URL":"https://doi.org/10.5281/zenodo.20718085","source":"datacite"},{"id":"doi:10.5281/zenodo.20718086","type":"article-journal","title":"EDGE AI IN 2026: A DEEP DIVE INTO SOURCE LEVEL INTELLIGENCE AND ITS FUTURE PATHWAYS","abstract":"Edge AI is one of the key paradigm shifts in the implementation of artificial intelligence, moving the inference calculation out of centralized infrastructure in the cloud and the bottom of the network, closer to the source of the data. This paper provides an overview of Edge AI in 2026, reasons it will gain momentum, the multi-tier architecture hierarchy that allows spread of intelligence, specialized hardware environment, and uses of cross-domain applications in manufacturing, healthcare, autonomous systems, retail, and smart infrastructure. We also compare performance standards between Edge AI, Cloud AI and Fog AI implementation, estimate market growth outlook and address challenges that remain unsolved, such as security, model optimization, devices management and regulatory compliance in the EU AI Act. According to our results, Edge AI has passed the inflection point between proof of concept and production grade with purpose-built neural processing units, compressed inference models, and well-developed orchestration systems. We sum up with a prospective study of federated learning integration, 6G synergies and the rise of agentic edge systems.","author":[{"family":"Insights","given":"International"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20718086","URL":"https://doi.org/10.5281/zenodo.20718086","source":"datacite"},{"id":"doi:10.48550/arxiv.2608.09328","type":"manuscript","title":"MaxModShift: Model Privacy via Designed Shifts","abstract":"Model learning by an eavesdropper is treated as an estimation problem in a federated environment. The Fisher Information Matrix for the eavesdropper's estimation problem is driven to singularity through a signaling design; this ensures that the eavesdropper cannot learn the model. Herein, the innovation of prior designs is that model shifts are designed to maximize the difference in the model learned by Eve and the central server while satisfying a transmission power constraint for the agents. Two shift schemes are provided. MaxModShift outperforms a prior ModShift design while requiring lesser transmission power. Compared to a noise injection scheme, MaxModShift performs better while requiring a lower bandwidth secret channel and a reduced average power consumption.","author":[{"family":"Kherani","given":"Nomaan"},{"family":"Mitra","given":"Urbashi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2608.09328","URL":"https://doi.org/10.48550/arxiv.2608.09328","source":"datacite"},{"id":"doi:10.5281/zenodo.20466034","type":"article-journal","title":"Pre-Registered Multi-Dataset Validation of Three-Way Composition III Privacy-Preserving Architecture for Musculoskeletal-Kinematic Clinical Digital Twins","abstract":"v2 update (2026-06-01): Reproducibility ZIP added back to the latest version alongside the manuscript files, so that downloading from the concept DOI gives all materials in one place rather than requiring navigation to v1. This reproducibility archive accompanies the Paper 8 manuscript \"Pre-Registered Multi-Dataset Validation of Three-Way Composition III Privacy-Preserving Architecture for Musculoskeletal-Kinematic Clinical Digital Twins.\" Background. Clinical digital twin (DT) systems for musculoskeletal rehabilitation increasingly aggregate longitudinal kinematic data across multiple sites and patient populations. These deployments require simultaneous guarantees of (a) federated coefficient learning utility, (b) membership-inference privacy against realistic adversaries, (c) byte-budgeted transmission for Internet-of-Things radio links, and (d) deployment-time robustness to anatomical, populational, and protocol heterogeneity. Methods. We present a pre-registered empirical validation campaign spanning 21 studies and 91 hypotheses for the three-way Composition III architecture (per-subject phase randomization, cohort mean aggregation, federated AR(1) coefficient learning) applied to musculoskeletal-kinematic clinical DT deployments. Each study's decision rules were frozen before runner execution. The campaign covers foundational properties; Composition III privacy and utility under homogeneous, moderately heterogeneous, and extremely heterogeneous (50:1 imbalance, 50x sensor noise differential) deployments; multi-classifier shadow-attacker bounds across LogisticRegression, RandomForest, and GradientBoosting families; differential-privacy noise calibration with non-monotone trade-off characterization; multi-period longitudinal observation up to T=10 periods; cross-anatomy generalization from N=3 (knee) to N=23 (hand-MANO model); and three real-world dataset validations using CMU Motion Capture Database (walking) and KIMORE Rehabilitation Dataset (rehabilitation exercises with 44 healthy controls plus 34 subjects with low-back pain, Parkinson's disease, and post-stroke conditions). Results. 71 of 91 pre-registered hypotheses supported (78%). All 20 non-supported outcomes are substantively interpretable: 3 metric-choice issues resolved by follow-up studies, 4 deployment-condition-dependent disclosures (including a non-monotone DP privacy-utility trade-off, an anatomy-specific sigma_DP scaling law, and a heterogeneity x longitudinal compound-leakage interaction), and 13 honest bounded-negatives identifying empirical boundaries of the recommended deployment configuration. Three findings stand out: (i) the realistic-attacker bound generalizes within +/- 0.06 across synthetic and two real datasets (maximum observed 0.60 against multi-classifier shadow attackers); (ii) patient-vs-healthy subgroup analysis on KIMORE yields identical oracle accuracy across populations (delta = 0.0000); (iii) anatomy-specific DP calibration scaling sigma_DP proportional to 1/sqrt(N) validated for N >= 5 with explicit knee N=3 outlier disclosure. Conclusions. Composition III is a viable privacy-preserving federated learning architecture for musculoskeletal-kinematic clinical digital twins. Empirical validation across synthetic data, two real datasets, multiple anatomies, three classifier families, three observation horizons, and mixed healthy/patient populations provides comprehensive grounding for clinical deployment. Archive contents. 21 frozen pre-registrations, 21 deterministic Python runners, 21 study reports, 21 verdict summary JSON files, 21 raw CSV data files, source code for the spiral-domain encoder and CMU MoCap and KIMORE parsers, 6 manuscript figures with generators, and the manuscript itself (Markdown and Word formats). Raw third-party data files are not redistributed per their respective licensing terms; download instructions and subject IDs are documented. Companion datasets. CMU Motion Capture Database (CC-BY 3.0, http://mocap.cs.cmu.ed","author":[{"family":"Ferlic","given":"Randolph"},{"family":"Ferlic","given":"Kimberly"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20466034","URL":"https://doi.org/10.5281/zenodo.20466034","source":"datacite"},{"id":"doi:10.5281/zenodo.21777269","type":"article-journal","title":"Análisis descriptivo de Técnicas de Machine Learning para Detección de Anomalías en Sistemas IoT","abstract":"Introducción: El Internet de las Cosas (IoT) ha transformado radicalmente sectores estratégicos globales, generando ecosistemas de dispositivos interconectados que producen volúmenes masivos de datos en tiempo real, pero que simultáneamente presentan vulnerabilidades estructurales severas ante amenazas cibernéticas de creciente sofisticación. La detección temprana y precisa de anomalías en estos sistemas constituye una necesidad crítica para garantizar la seguridad operacional de las infraestructuras digitales modernas. Objetivo: El presente artículo tiene como objetivo describir, analizar y sintetizar el estado del arte de las técnicas de Machine Learning y Deep Learning aplicadas a la detección de anomalías en sistemas IoT, caracterizando sus fundamentos algorítmicos, condiciones de aplicabilidad, desempeño documentado, limitaciones y perspectivas de desarrollo futuro. Metodología: Se empleó un nivel descriptivo con método de análisis-síntesis, mediante una revisión no sistemática de literatura con consulta en bases de datos IEEE Xplore, Scopus, Web of Science, ScienceDirect y Google Scholar, utilizando más de veinticinco palabras clave estructuradas en español e inglés, abarcando principalmente publicaciones del período 2020-2026. Resultados: Los resultados revelan que los algoritmos de gradiente potenciado y las arquitecturas híbridas CNN-LSTM alcanzan precisiones superiores al 99% en benchmarks estandarizados, que el aprendizaje federado emerge como paradigma dominante para la preservación de privacidad en entornos distribuidos, y que la Inteligencia Artificial Explicable representa un requisito emergente para el despliegue en sectores regulados. Conclusión: Se concluye que, si bien el campo ha alcanzado notable madurez técnica en entornos controlados, persisten brechas críticas de generalización, robustez adversarial e interpretabilidad que condicionan la transferibilidad de los modelos hacia entornos operacionales reales, con implicaciones estratégicas particulares para América Latina y Ecuador. Área de estudio general: Tecnologías de la Información y Comunicación aplicadas a la Ciberseguridad y Sistemas Inteligentes. Área de estudio específica: Aprendizaje Automático y Aprendizaje Profundo para la Detección de Anomalías e Intrusiones en Sistemas del Internet de las Cosas.","author":[{"family":"Flores-Andino","given":"Víctor"},{"family":"Pérez Insuasti","given":"Juan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21777269","URL":"https://doi.org/10.5281/zenodo.21777269","source":"datacite"},{"id":"doi:10.5281/zenodo.21777268","type":"article-journal","title":"Análisis descriptivo de Técnicas de Machine Learning para Detección de Anomalías en Sistemas IoT","abstract":"Introducción: El Internet de las Cosas (IoT) ha transformado radicalmente sectores estratégicos globales, generando ecosistemas de dispositivos interconectados que producen volúmenes masivos de datos en tiempo real, pero que simultáneamente presentan vulnerabilidades estructurales severas ante amenazas cibernéticas de creciente sofisticación. La detección temprana y precisa de anomalías en estos sistemas constituye una necesidad crítica para garantizar la seguridad operacional de las infraestructuras digitales modernas. Objetivo: El presente artículo tiene como objetivo describir, analizar y sintetizar el estado del arte de las técnicas de Machine Learning y Deep Learning aplicadas a la detección de anomalías en sistemas IoT, caracterizando sus fundamentos algorítmicos, condiciones de aplicabilidad, desempeño documentado, limitaciones y perspectivas de desarrollo futuro. Metodología: Se empleó un nivel descriptivo con método de análisis-síntesis, mediante una revisión no sistemática de literatura con consulta en bases de datos IEEE Xplore, Scopus, Web of Science, ScienceDirect y Google Scholar, utilizando más de veinticinco palabras clave estructuradas en español e inglés, abarcando principalmente publicaciones del período 2020-2026. Resultados: Los resultados revelan que los algoritmos de gradiente potenciado y las arquitecturas híbridas CNN-LSTM alcanzan precisiones superiores al 99% en benchmarks estandarizados, que el aprendizaje federado emerge como paradigma dominante para la preservación de privacidad en entornos distribuidos, y que la Inteligencia Artificial Explicable representa un requisito emergente para el despliegue en sectores regulados. Conclusión: Se concluye que, si bien el campo ha alcanzado notable madurez técnica en entornos controlados, persisten brechas críticas de generalización, robustez adversarial e interpretabilidad que condicionan la transferibilidad de los modelos hacia entornos operacionales reales, con implicaciones estratégicas particulares para América Latina y Ecuador. Área de estudio general: Tecnologías de la Información y Comunicación aplicadas a la Ciberseguridad y Sistemas Inteligentes. Área de estudio específica: Aprendizaje Automático y Aprendizaje Profundo para la Detección de Anomalías e Intrusiones en Sistemas del Internet de las Cosas.","author":[{"family":"Flores-Andino","given":"Víctor"},{"family":"Pérez Insuasti","given":"Juan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21777268","URL":"https://doi.org/10.5281/zenodo.21777268","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.28191","type":"manuscript","title":"Secure Aggregation for Privacy-Preserving Federated Learning on Clinical EEG Data","abstract":"Federated learning enables multiple institutions to train shared models without exchanging raw clinical EEG data, but it does not fully prevent privacy leakage from individual model updates. This paper presents a privacy-preserving federated learning framework for clinical EEG data using masking-based secure aggregation as the core protection mechanism. The framework combines graph-based communication, threshold secret sharing, dropout-resilient aggregation, local update clipping, an optional Bloom filter-based privacy-preserving record-linkage initialization module, and auxiliary-notary-based verifiability. It supports both semi-honest and malicious aggregation settings and is implemented using the Flower federated learning framework. The secure-aggregation variants are evaluated in a simulated cross-silo healthcare setting using TUH EEG-derived data under different client configurations. Under the stated assumptions, the secure variants hide individual updates from the aggregation server. The results show that these variants remain compatible with federated model training, although malicious-setting safeguards and lightweight consistency-checking mechanisms introduce additional computation, communication, and round-duration overhead. The semi-honest variant provides the lowest overhead among the secure configurations, while malicious and auxiliary-notary variants offer stronger consistency, integrity, and lightweight verification support at higher cost.","author":[{"family":"Rajabi","given":"Pouya"},{"family":"Toorani","given":"Mohsen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.28191","URL":"https://doi.org/10.48550/arxiv.2607.28191","source":"datacite"},{"id":"doi:10.5281/zenodo.21695600","type":"article-journal","title":"Reproducibility package: Perbandingan Komprehensif Model Kecerdasan Buatan untuk Klasifikasi Kode HS: Sebuah Tinjauan Sistematis tentang Kinerja, Tantangan, dan Penerapan","abstract":"The user wants me to clean a LaTeX text by removing all AI writing characteristics. Let me analyze the text carefully and apply the checklist: 1. **Remove all em-dash (---) and en-dash (--)**: Replace with commas, colons, parentheses, or new sentences 2. **Remove filler words**: \"it is worth noting\", \"importantly\", \"notably\", \"furthermore\", \"moreover\", \"in addition\", \"it is important to note\", \"interestingly\", \"it should be noted\" 3. **Remove AI words**: \"delve\", \"leverage\", \"underscore\", \"pivotal\", \"realm\", \"tapestry\", \"intricate\", \"multifaceted\", \"nuanced\" 4. **Break up long sentences (>50 words)**: Split into shorter sentences 5. **Merge very short consecutive sentences** 6. **Remove \"not only X but also Y\" patterns if appearing >1x**: Vary the construction 7. **Ensure no bullet points in the middle of prose paragraphs** 8. **Ensure each paragraph has at least 2-3 \\cite{} references**: Use only keys from the ALLOWED CITATION KEYS list 9. **No emojis, icons, decorative symbols** 10. **Fix LaTeX quotation marks**: Use ``quote'' consistently Let me go through the text section by section and clean it: Looking at the text, I can see many issues: - Em-dashes are used extensively - Many \"may\" hedges for WEAK claims - Some UNVERIFIED claims need citations removed - The text is extremely long with many repetitive sections Wait, I need to look at the VERIFICATION RESULTS provided: - **INVALID_KEY**: All citations used appear to be in the allowed list. Let me check... - **UNVERIFIED** (0 sources): Claims that need the citation removed: - chen2021b (temporal split) - chen2021a (Universal Sentence Encoder) - ryzhova2019 (Naive Bayes models) - yuvraj2025 (ATLAS, 6-digit and 10-digit, 5x and 8x cheaper) - muslim2022 (Naive Bayes 99.97%) - guoli2019 (multimodal improvement 6-16%) - turhan2015 (multimodal improvement 6-16%) - zhou2022 (China and USA datasets) - navasardyan2024 (explainable AI) - And many more... The key issues to fix are removing unverified citations, eliminating AI patterns like em-dashes and filler words, and ensuring proper hedging for weak claims. I'll need to systematically go through and apply these corrections while maintaining the academic tone and proper citation format. I need to systematically clean this LaTeX text by removing all instances of em-dashes, eliminating filler phrases, and breaking up overly long sentences. The most efficient approach is to process the text section by section rather than line by line. Let me work through each section methodically, applying these corrections consistently across the entire document. saya akan memfokuskan pada pembersihan teks secara keseluruhan dengan menghapus semua elemen artificial writing characteristics seperti em-dashes, filler words, dan frasa AI. Teks ini mencakup beberapa bagian termasuk karakteristik studi, kualitas studi, temuan utama, dan temuan kontradiktif. Saya akan menghilangkan pola penulisan yang tidak natural sambil mempertahankan struktur dan makna aslinya. For characterizations of model performance, I use \"achieves\" or \"attains\" rather than \"appears to achieve\". Where claims lack verification, I simply state the performance metrics without citation markers. When specific capabilities are mentioned without sources, I remove those references entirely. For validated claims, I retain citations, and for comparative statements I use \"may\" to indicate uncertainty where appropriate. \\section{Hasil} \\subsection{Karakteristik Studi} Tinjauan menyertakan 35 studi yang dipublikasikan antara 2006 dan 2026. Distribusi tahun menunjukkan tidak ada studi dari 2007-2014, 2016-2017, atau 2021. Publikasi dimulai pada 2006 dengan satu studi, meningkat pada 2015-2019 dengan dua hingga tiga studi per tahun, dan melonjak signifikan pada 2021-2026 dengan 26 studi (74.3\\% dari total). Tren temporal ini mencerminkan peningkatan minat riset seiring ketersediaan model transformer dan LLM yang dapat diaplikasikan untuk tugas klasifikasi HS. Desain studi mencakup conference pape","author":[{"family":"Mahfud","given":"Syarif"},{"family":"Awangga","given":"Rolly"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21695600","URL":"https://doi.org/10.5281/zenodo.21695600","source":"datacite"},{"id":"doi:10.5281/zenodo.21695601","type":"article-journal","title":"Reproducibility package: Perbandingan Komprehensif Model Kecerdasan Buatan untuk Klasifikasi Kode HS: Sebuah Tinjauan Sistematis tentang Kinerja, Tantangan, dan Penerapan","abstract":"The user wants me to clean a LaTeX text by removing all AI writing characteristics. Let me analyze the text carefully and apply the checklist: 1. **Remove all em-dash (---) and en-dash (--)**: Replace with commas, colons, parentheses, or new sentences 2. **Remove filler words**: \"it is worth noting\", \"importantly\", \"notably\", \"furthermore\", \"moreover\", \"in addition\", \"it is important to note\", \"interestingly\", \"it should be noted\" 3. **Remove AI words**: \"delve\", \"leverage\", \"underscore\", \"pivotal\", \"realm\", \"tapestry\", \"intricate\", \"multifaceted\", \"nuanced\" 4. **Break up long sentences (>50 words)**: Split into shorter sentences 5. **Merge very short consecutive sentences** 6. **Remove \"not only X but also Y\" patterns if appearing >1x**: Vary the construction 7. **Ensure no bullet points in the middle of prose paragraphs** 8. **Ensure each paragraph has at least 2-3 \\cite{} references**: Use only keys from the ALLOWED CITATION KEYS list 9. **No emojis, icons, decorative symbols** 10. **Fix LaTeX quotation marks**: Use ``quote'' consistently Let me go through the text section by section and clean it: Looking at the text, I can see many issues: - Em-dashes are used extensively - Many \"may\" hedges for WEAK claims - Some UNVERIFIED claims need citations removed - The text is extremely long with many repetitive sections Wait, I need to look at the VERIFICATION RESULTS provided: - **INVALID_KEY**: All citations used appear to be in the allowed list. Let me check... - **UNVERIFIED** (0 sources): Claims that need the citation removed: - chen2021b (temporal split) - chen2021a (Universal Sentence Encoder) - ryzhova2019 (Naive Bayes models) - yuvraj2025 (ATLAS, 6-digit and 10-digit, 5x and 8x cheaper) - muslim2022 (Naive Bayes 99.97%) - guoli2019 (multimodal improvement 6-16%) - turhan2015 (multimodal improvement 6-16%) - zhou2022 (China and USA datasets) - navasardyan2024 (explainable AI) - And many more... The key issues to fix are removing unverified citations, eliminating AI patterns like em-dashes and filler words, and ensuring proper hedging for weak claims. I'll need to systematically go through and apply these corrections while maintaining the academic tone and proper citation format. I need to systematically clean this LaTeX text by removing all instances of em-dashes, eliminating filler phrases, and breaking up overly long sentences. The most efficient approach is to process the text section by section rather than line by line. Let me work through each section methodically, applying these corrections consistently across the entire document. saya akan memfokuskan pada pembersihan teks secara keseluruhan dengan menghapus semua elemen artificial writing characteristics seperti em-dashes, filler words, dan frasa AI. Teks ini mencakup beberapa bagian termasuk karakteristik studi, kualitas studi, temuan utama, dan temuan kontradiktif. Saya akan menghilangkan pola penulisan yang tidak natural sambil mempertahankan struktur dan makna aslinya. For characterizations of model performance, I use \"achieves\" or \"attains\" rather than \"appears to achieve\". Where claims lack verification, I simply state the performance metrics without citation markers. When specific capabilities are mentioned without sources, I remove those references entirely. For validated claims, I retain citations, and for comparative statements I use \"may\" to indicate uncertainty where appropriate. \\section{Hasil} \\subsection{Karakteristik Studi} Tinjauan menyertakan 35 studi yang dipublikasikan antara 2006 dan 2026. Distribusi tahun menunjukkan tidak ada studi dari 2007-2014, 2016-2017, atau 2021. Publikasi dimulai pada 2006 dengan satu studi, meningkat pada 2015-2019 dengan dua hingga tiga studi per tahun, dan melonjak signifikan pada 2021-2026 dengan 26 studi (74.3\\% dari total). Tren temporal ini mencerminkan peningkatan minat riset seiring ketersediaan model transformer dan LLM yang dapat diaplikasikan untuk tugas klasifikasi HS. Desain studi mencakup conference pape","author":[{"family":"Mahfud","given":"Syarif"},{"family":"Awangga","given":"Rolly"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21695601","URL":"https://doi.org/10.5281/zenodo.21695601","source":"datacite"},{"id":"doi:10.5281/zenodo.20213774","type":"article-journal","title":"Consultation Response: Making public services work for you with your digital identity","abstract":"Consultation response submitted to the Cabinet Office consultation on Making Public Services Work for You with Your Digital Identity (April 2026). This submission argues that the failure modes of a national digital ID are more likely to be governance failures than technology failures — and that the consultation document, while unusually thoughtful about user experience and inclusion, is architecturally optimistic about four governance properties that empirical research and the UK's own recent record show cannot be assumed: that proportionate oversight will emerge from existing structures; that point-in-time compliance will keep pace with the threat surface a national digital ID will attract; that alternative access routes can be designed without becoming the system's weakest link; and that data localisation and standard contractual clauses are sufficient to manage jurisdictional exposure across the operator and sub-processor chain. Drawing on published empirical research into NHS cybersecurity governance, practitioner experience of Kazakhstan's national e-government platform, and analysis of the NHS Federated Data Platform procurement, the submission makes fifteen specific recommendations. These include: publishing an explicit adversarial threat model before legislation; mandating velocity-based assurance metrics alongside point-in-time compliance; requiring decision-latency stress-testing of revocation and incident-response pathways; establishing a statutory non-punitive incident learning channel modelled on the US Aviation Safety Reporting System; requiring a statutory data-retention floor of 36 months for transactional and audit logs; treating alternative access routes as the binding fraud design constraint rather than a residual concern; publishing sub-processor governance including jurisdictional-exposure analysis and maintainability-under-supplier-exit requirements; and funding the system as an annually revised expenditure envelope rather than a multi-year fixed business case. The submission also identifies two governance gaps absent from most consultation responses: a statutory collision layer for lawfully concealed identities in intelligence and undercover policing; and the systemic risk of checker concentration, requiring a formal outage declaration architecture with time-bounded lawful fallback for identity-dependent transactions.","author":[{"family":"Shabad","given":"Vsevolod"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20213774","URL":"https://doi.org/10.5281/zenodo.20213774","source":"datacite"},{"id":"doi:10.5281/zenodo.20213775","type":"article-journal","title":"Consultation Response: Making public services work for you with your digital identity","abstract":"Consultation response submitted to the Cabinet Office consultation on Making Public Services Work for You with Your Digital Identity (April 2026). This submission argues that the failure modes of a national digital ID are more likely to be governance failures than technology failures — and that the consultation document, while unusually thoughtful about user experience and inclusion, is architecturally optimistic about four governance properties that empirical research and the UK's own recent record show cannot be assumed: that proportionate oversight will emerge from existing structures; that point-in-time compliance will keep pace with the threat surface a national digital ID will attract; that alternative access routes can be designed without becoming the system's weakest link; and that data localisation and standard contractual clauses are sufficient to manage jurisdictional exposure across the operator and sub-processor chain. Drawing on published empirical research into NHS cybersecurity governance, practitioner experience of Kazakhstan's national e-government platform, and analysis of the NHS Federated Data Platform procurement, the submission makes fifteen specific recommendations. These include: publishing an explicit adversarial threat model before legislation; mandating velocity-based assurance metrics alongside point-in-time compliance; requiring decision-latency stress-testing of revocation and incident-response pathways; establishing a statutory non-punitive incident learning channel modelled on the US Aviation Safety Reporting System; requiring a statutory data-retention floor of 36 months for transactional and audit logs; treating alternative access routes as the binding fraud design constraint rather than a residual concern; publishing sub-processor governance including jurisdictional-exposure analysis and maintainability-under-supplier-exit requirements; and funding the system as an annually revised expenditure envelope rather than a multi-year fixed business case. The submission also identifies two governance gaps absent from most consultation responses: a statutory collision layer for lawfully concealed identities in intelligence and undercover policing; and the systemic risk of checker concentration, requiring a formal outage declaration architecture with time-bounded lawful fallback for identity-dependent transactions.","author":[{"family":"Shabad","given":"Vsevolod"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20213775","URL":"https://doi.org/10.5281/zenodo.20213775","source":"datacite"},{"id":"doi:10.3390/app16052611","type":"article-journal","title":"The Convergence of Federated Learning, Knowledge Graphs, and Large Language Models for Language Learning: A Scoping Review","abstract":"Large Language Models (LLMs) in Intelligent Computer-Assisted Language Learning enable highly personalized learning, yet raise significant challenges related to pedagogical grounding, data privacy, and instructional validity. Although Knowledge Graphs (KGs) and Federated Learning (FL) can mitigate these issues in isolation, evidence on systematic FL–KG–LLM integration for educational language learning remains limited. This scoping review maps the FL–KG–LLM convergence landscape. Following PRISMA-ScR guidelines, we searched six databases and screened 51 papers (2019–2025) using automated extraction. Our findings indicate limited convergence: no papers integrate all three domains, and 58.8% of approaches remain confined to isolated technological silos. Reporting is also uneven across the corpus, with an average “Not Reported” (NR) rate of 84.5%, most notably for privacy mechanisms (92.2%), validation metrics (90.2%), and Common European Framework of Reference for Languages (CEFR) alignment (88.2%). Domain-specific analysis reveals two distinct patterns: inter-domain gaps (disciplinary silos resulting in expected CEFR absence in single-domain papers) and intra-domain gaps (failure to report domain-critical variables, including 100% parameter NR in FL studies, 86.7% validation NR in KG studies, and 100% CEFR NR in convergence papers). Taken together, these gaps suggest that pedagogical grounding is treated as optional rather than structural. We therefore identify two pillars of pedagogical grounding: a Grounding Pillar, which constrains LLM outputs via Knowledge Graph rules, and a Validation Pillar, which concerns how authoritative frameworks (e.g., CEFR) are mapped onto Knowledge Graph schemas and evaluated. The near-universal absence of CEFR alignment and validation reporting suggests that this second pillar is currently missing, which we term the Integrity Gap—a systematic disconnection between technological innovation and pedagogical grounding inin Intelligent Computer-Assisted Language Learning. By reframing the problem as upstream control and validation, this review informs the design of user-facing automated systems where trust, transparency, and human oversight are critical.","author":[{"family":"Kenteris","given":"Michael"},{"family":"Kotis","given":"Konstantinos"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/app16052611","URL":"https://doi.org/10.3390/app16052611","source":"openalex"},{"id":"doi:10.5281/zenodo.20760897","type":"article-journal","title":"Coordination Topology as Failure Geometry: How Network Structure, Protocol Design, and Consensus Dynamics Jointly Shape Multi-Agent System Reliability","abstract":"Version 2 — revised in response to an external structural review and an automated critique pass. See \"Response to Review\" appendix in the PDF for the change log. Multi-agent AI systems are increasingly deployed in production settings where reliability is assumed but rarely formally characterized. This paper advances a **candidate structural reading—explicitly a heuristic synthesis, not a derivation**—of a pattern that recurs across five to seven recent preprints from cs.MA, cs.DC, and cs.NI: **coordination topology is not a neutral implementation detail but a primary determinant of failure geometry**. The topology assumption made at design time—centralized hub, sparse factor graph, federated stack, decentralized shared-context, or social-network overlay—predicts which failure modes will dominate, how error propagates, and whether recovery is even structurally possible. Concretely, we draw on: density-evolution analysis of agent networks modeled as sparse factor graphs [corpus:arxiv:2606.18121], Byzantine-resilient CRDT reconstruction that decouples update propagation from state derivation [corpus:arxiv:2606.18966], a taxonomy of LLM agent communication protocols that reveals decentralized discovery remains rare [corpus:arxiv:2606.19135], the security-induced Braess paradox in service-function-chain orchestration [corpus:arxiv:2606.17987], the channel-fracture failure mode in hierarchical memory injection [corpus:arxiv:2606.04896v2], social learning performance gaps between decentralized and centralized inference [corpus:arxiv:2606.09176], and the deliberative-consensus degradation observed in multi-agent oracle resolution [corpus:arxiv:2605.30802]. The unifying claim is that each of these findings instantiates the same structural pattern: a topology choice creates a reachability or propagation structure that determines whether errors are absorbed, amplified, or silently suppressed. This is a heuristic reading across sources that do not share a common formalism; the shared vocabulary does not guarantee shared formal structure. The falsification path for this claim is concrete: construct a controlled multi-agent testbed in which topology is the single varied parameter across otherwise identical agent populations and measure whether failure-mode distribution shifts in the direction predicted by density-evolution thresholds. --- Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted from arXiv preprint corpus on the date in the filename. Cited arXiv preprints: 2605.30802, 2606.02080, 2606.04896v2, 2606.06971, 2606.07487, 2606.08457, 2606.09176, 2606.13068, 2606.15931, 2606.17987, 2606.18121, 2606.18837, 2606.18966, 2606.19135, 2606.19319 AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.","author":[{"family":"Team","given":"Saluca"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20760897","URL":"https://doi.org/10.5281/zenodo.20760897","source":"datacite"},{"id":"doi:10.5281/zenodo.20482005","type":"article-journal","title":"Feature-Oriented Regulation in Federated Multimodal Models Under Non-IID Data Distributions","abstract":"This report synthesises findings from 5 peer-reviewed papers addressing the following research question: What is the impact of feature-oriented regulation methods like \\$Psi\\$-Net on the inference efficiency of federated multimodal models under non-IID data distributions, measured by throughput and accuracy. Institutions in highly regulated domains such as finance and healthcare often have restrictive rules around data sharing. Federated learning is a distributed learning framework that enables multi-institutional collaborations on decentralized data with improved protection for. 7 claims were extracted from source literature; 7 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.5/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the impact of feature-oriented regulation methods like Ψ-Net on the inference efficiency of federated multimodal models under non-IID data distributions, measured by throughput and accuracy trade-offs? Autonomous literature synthesis. Automated review score: 8.5/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20482005","URL":"https://doi.org/10.5281/zenodo.20482005","source":"datacite"},{"id":"doi:10.5281/zenodo.20482006","type":"article-journal","title":"Feature-Oriented Regulation in Federated Multimodal Models Under Non-IID Data Distributions","abstract":"This report synthesises findings from 5 peer-reviewed papers addressing the following research question: What is the impact of feature-oriented regulation methods like \\$Psi\\$-Net on the inference efficiency of federated multimodal models under non-IID data distributions, measured by throughput and accuracy. Institutions in highly regulated domains such as finance and healthcare often have restrictive rules around data sharing. Federated learning is a distributed learning framework that enables multi-institutional collaborations on decentralized data with improved protection for. 7 claims were extracted from source literature; 7 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.5/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: What is the impact of feature-oriented regulation methods like Ψ-Net on the inference efficiency of federated multimodal models under non-IID data distributions, measured by throughput and accuracy trade-offs? Autonomous literature synthesis. Automated review score: 8.5/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20482006","URL":"https://doi.org/10.5281/zenodo.20482006","source":"datacite"},{"id":"doi:10.5281/zenodo.20280464","type":"article-journal","title":"Intelligent Predictive Architectures for Autonomous Self-Healing in Cloud Computing: A Comprehensive Survey","abstract":"Neural network-enabled self-healing is becoming a very exciting approach to making cloud computing infrastructures more reliable, available, and efficient. This survey paper gives a detailed review of neural network-based predictive models and how they are combined with autonomous recovery mechanisms for self-healing cloud systems. It first delineates the main neural architectures used for cloud reliability, such as Artificial Neural Networks (ANNs), Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM) networks, and hybrid or ensemble models, explaining their roles in failure prediction, resource management forecasting, and SLA/QoS violation prediction. The paper examines the coupling of these predictive models with self-healing action and decision layers like rule-based policies, policy engines, and reinforcement learning-driven controllers to proactively trigger recovery actions, reduce Mean Time To Recovery (MTTR), and control false alarms and resource overhead. Evaluation practices highlight datasets (Google Cluster Traces, Alibaba traces, and synthetic simulation data), performance metrics (prediction accuracy, MTTR, false positive/negative rates, scalability, and energy cost), and differences between simulated vs. real production cloud environments. Moreover, the survey discusses privacy and security issues of predictive and action models, presenting techniques such as federated learning, differential privacy, homomorphic encryption, and secure multi-party computation for privacy-preserving cross-tenant collaboration. Lastly, this paper draws attention to open issues including trade-offs between accuracy and false alarms, latency constraints, explainability, and scalability, proposing future work including explainable predictive pipelines, federated self-healing frameworks, multi-agent and DRL-based autonomy, and integration with emerging technologies such as quantum-inspired neural networks, edge-cloud continuum architectures, digital twins, service meshes, and neuromorphic hardware to enable more trustworthy, efficient, and autonomous self-healing cloud management.","author":[{"family":"Gupta","given":"Mr"},{"family":"Bangera","given":"Dr"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20280464","URL":"https://doi.org/10.5281/zenodo.20280464","source":"datacite"},{"id":"doi:10.5281/zenodo.20321777","type":"article-journal","title":"Intelligent Predictive Architectures for Autonomous Self-Healing in Cloud Computing: A Comprehensive Survey","abstract":"Neural network-enabled self-healing is becoming a very exciting approach to making cloud computing infrastructures more reliable, available, and efficient. This survey paper gives a detailed review of neural network-based predictive models and how they are combined with autonomous recovery mechanisms for self-healing cloud systems. It first delineates the main neural architectures used for cloud reliability, such as Artificial Neural Networks (ANNs), Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM) networks, and hybrid or ensemble models, explaining their roles in failure prediction, resource management forecasting, and SLA/QoS violation prediction. The paper examines the coupling of these predictive models with self-healing action and decision layers like rule-based policies, policy engines, and reinforcement learning-driven controllers to proactively trigger recovery actions, reduce Mean Time To Recovery (MTTR), and control false alarms and resource overhead. Evaluation practices highlight datasets (Google Cluster Traces, Alibaba traces, and synthetic simulation data), performance metrics (prediction accuracy, MTTR, false positive/negative rates, scalability, and energy cost), and differences between simulated vs. real production cloud environments. Moreover, the survey discusses privacy and security issues of predictive and action models, presenting techniques such as federated learning, differential privacy, homomorphic encryption, and secure multi-party computation for privacy-preserving cross-tenant collaboration. Lastly, this paper draws attention to open issues including trade-offs between accuracy and false alarms, latency constraints, explainability, and scalability, proposing future work including explainable predictive pipelines, federated self-healing frameworks, multi-agent and DRL-based autonomy, and integration with emerging technologies such as quantum-inspired neural networks, edge-cloud continuum architectures, digital twins, service meshes, and neuromorphic hardware to enable more trustworthy, efficient, and autonomous self-healing cloud management.","author":[{"family":"Gupta","given":"Mr"},{"family":"Bangera","given":"Dr"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20321777","URL":"https://doi.org/10.5281/zenodo.20321777","source":"datacite"},{"id":"doi:10.5281/zenodo.19457914","type":"article-journal","title":"Scaling Farm Autonomy Decentralized Control, Data, and Edge","abstract":"This review examines the landscape of decentralized agricultural automation systems, motivated by the need to synthesize findings from a rapidly evolving field. It addresses the shift from traditional centralized approaches towards more scalable, robust, and privacy-preserving solutions in modern farming. The review's scope focuses on three distinct yet interconnected layers of decentralization: decentralized control (Layer A), involving swarm and multi-agent robotics; decentralized data/analytics (Layer B), focusing on blockchain and federated learning; and distributed hardware/edge IoT (Layer C), centered on edge-first frameworks and on-device AI. Key findings indicate that Layer A is a maturing field with significant algorithmic progress, alongside increasing field prototypes demonstrating improved coverage and resilience. Layer B shows rapid growth in hybrid designs that combine privacy-preserving federated learning with blockchain for auditability and decentralized aggregation. Layer C highlights the emergence of edge-first frameworks and agentic AI on devices for low-latency perception and action. Common challenges across all layers include connectivity limitations in rural areas, data heterogeneity, establishing trust and incentive mechanisms, and managing energy constraints. The ultimate aim is to inform the development and deployment of robust and scalable decentralized agricultural automation solutions.","author":[{"family":"Shaik","given":"Mohammad"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19457914","URL":"https://doi.org/10.5281/zenodo.19457914","source":"datacite"},{"id":"doi:10.5281/zenodo.19457915","type":"article-journal","title":"Scaling Farm Autonomy Decentralized Control, Data, and Edge","abstract":"This review examines the landscape of decentralized agricultural automation systems, motivated by the need to synthesize findings from a rapidly evolving field. It addresses the shift from traditional centralized approaches towards more scalable, robust, and privacy-preserving solutions in modern farming. The review's scope focuses on three distinct yet interconnected layers of decentralization: decentralized control (Layer A), involving swarm and multi-agent robotics; decentralized data/analytics (Layer B), focusing on blockchain and federated learning; and distributed hardware/edge IoT (Layer C), centered on edge-first frameworks and on-device AI. Key findings indicate that Layer A is a maturing field with significant algorithmic progress, alongside increasing field prototypes demonstrating improved coverage and resilience. Layer B shows rapid growth in hybrid designs that combine privacy-preserving federated learning with blockchain for auditability and decentralized aggregation. Layer C highlights the emergence of edge-first frameworks and agentic AI on devices for low-latency perception and action. Common challenges across all layers include connectivity limitations in rural areas, data heterogeneity, establishing trust and incentive mechanisms, and managing energy constraints. The ultimate aim is to inform the development and deployment of robust and scalable decentralized agricultural automation solutions.","author":[{"family":"Shaik","given":"Mohammad"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19457915","URL":"https://doi.org/10.5281/zenodo.19457915","source":"datacite"},{"id":"doi:10.5281/zenodo.20592368","type":"article-journal","title":"Failure Propagation and Self-Correction in Multi-Agent LLM Systems: How Deliberative Consensus, Credit Assignment, and Architectural Isolation Jointly Determine Systemic Reliability","abstract":"Version 2 — revised in response to an external structural review and an automated critique pass. See \"Response to Review\" appendix in the PDF for the change log. Multi-agent LLM systems are increasingly deployed in settings where individual agent failures can cascade across the collective, yet the mechanisms by which such failures propagate—and the structural conditions under which they are contained or corrected—remain poorly characterised. This paper synthesises seven findings from recent cs.MA and cs.DC preprints to argue a **candidate structural pattern** (a heuristic reading, not a derivation from a shared formal structure): systemic reliability in multi-agent LLM systems is not primarily a function of individual agent capability, but of the interaction between (a) the topology through which errors can propagate, (b) the credit-assignment mechanisms that identify which components generated the error, and (c) the architectural isolation boundaries that prevent a failure in one execution channel from contaminating others. We argue the analogy by naming specific mechanisms in each case rather than by appeal to a unified formalism. The corpus draws from cs.MA papers on multi-agent deliberation, self-evolution, coordination policy learning, and governance infrastructure, supplemented by cs.DC work on federated orchestration and zero-trust enforcement at physical actuation boundaries. Key findings include: deliberative consensus among LLM agents can *degrade* accuracy below single-model baselines when high-confidence wrong agents flip correct ones [corpus:arxiv:2605.30802]; architectural channel isolation failures silently block cross-agent memory injection regardless of agent-level correctness [corpus:arxiv:2606.04896]; governance layers at the execution boundary reduce unsafe executions from 88% to near-zero without modifying underlying generators [corpus:arxiv:2606.04306]; and temporal plus structural credit decomposition substantially reduces query complexity in MAS optimisation [corpus:arxiv:2605.30227]. Together these findings suggest that failure propagation and self-correction are topology-dependent phenomena that cannot be addressed by improving agent intelligence alone. Falsification path: if deliberative degradation disappears when inter-agent error correlation is experimentally reduced below 0.3, the error-propagation mechanism is confirmed; if it persists, the topology hypothesis is insufficient and capability variance must be the primary driver. --- Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted from arXiv preprint corpus on the date in the filename. Cited arXiv preprints: 2605.25653, 2605.25741, 2605.26286, 2605.27076, 2605.27106, 2605.28984, 2605.29612, 2605.29790, 2605.30227, 2605.30802, 2606.01170, 2606.01581, 2606.01862, 2606.02080, 2606.03543, 2606.04197, 2606.04306, 2606.04896 AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.","author":[{"family":"Team","given":"Saluca"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20592368","URL":"https://doi.org/10.5281/zenodo.20592368","source":"datacite"},{"id":"doi:10.5281/zenodo.20832756","type":"article-journal","title":"Distributed Deep Learning and Intelligent Soil–Water Analytics in Precision Agriculture: A Comprehensive Review","abstract":"Efficient management of soil–water resources is critical for global food security under intensifying climatic and demographic pressures. This review provides a comprehensive synthesis of artificial intelligence (AI) and distributed deep learning methodologies applied to soil–water interactions in precision agriculture. The physical and hydraulic foundations of soil–water systems—including water retention, unsaturated flow governed by the Richards equation, and soil degradation processes—are examined and situated within a unified framework of AI-based modeling and decision support. Classical machine learning (ML) algorithms (Random Forests, Support Vector Machines, gradient boosting) and deep learning architectures (convolutional neural networks, long short-term memory networks, transformers) are evaluated with respect to their capacity to predict soil moisture dynamics, estimate hydraulic properties, support smart irrigation scheduling, and generate digital soil maps at field-to-regional scales. Distributed training paradigms, federated learning for privacy-preserving multi-farm analytics, and edge AI deployment on low-power IoT hardware are assessed as enabling infrastructures for scalable agricultural intelligence. This review further addresses explainability, uncertainty quantification, and ethical dimensions inherent to AI-driven agricultural systems. Key challenges—including training data scarcity in data-poor regions, model interpretability, integration with physics-based hydrological models, and real-time deployment constraints—are critically discussed. Prospective research directions encompass physics-informed neural networks, foundation models for earth observation, autonomous digital twins of soil–water systems, and federated learning architectures aligned with data sovereignty frameworks. The synthesis underscores AI’s transformative potential for sustainable agricultural water management while delineating the technical and sociotechnical barriers that must be resolved to realize this potential at a global scale.","author":[{"family":"Lemenkova","given":"Polina"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20832756","URL":"https://doi.org/10.5281/zenodo.20832756","source":"datacite"},{"id":"doi:10.5281/zenodo.20832757","type":"article-journal","title":"Distributed Deep Learning and Intelligent Soil–Water Analytics in Precision Agriculture: A Comprehensive Review","abstract":"Efficient management of soil–water resources is critical for global food security under intensifying climatic and demographic pressures. This review provides a comprehensive synthesis of artificial intelligence (AI) and distributed deep learning methodologies applied to soil–water interactions in precision agriculture. The physical and hydraulic foundations of soil–water systems—including water retention, unsaturated flow governed by the Richards equation, and soil degradation processes—are examined and situated within a unified framework of AI-based modeling and decision support. Classical machine learning (ML) algorithms (Random Forests, Support Vector Machines, gradient boosting) and deep learning architectures (convolutional neural networks, long short-term memory networks, transformers) are evaluated with respect to their capacity to predict soil moisture dynamics, estimate hydraulic properties, support smart irrigation scheduling, and generate digital soil maps at field-to-regional scales. Distributed training paradigms, federated learning for privacy-preserving multi-farm analytics, and edge AI deployment on low-power IoT hardware are assessed as enabling infrastructures for scalable agricultural intelligence. This review further addresses explainability, uncertainty quantification, and ethical dimensions inherent to AI-driven agricultural systems. Key challenges—including training data scarcity in data-poor regions, model interpretability, integration with physics-based hydrological models, and real-time deployment constraints—are critically discussed. Prospective research directions encompass physics-informed neural networks, foundation models for earth observation, autonomous digital twins of soil–water systems, and federated learning architectures aligned with data sovereignty frameworks. The synthesis underscores AI’s transformative potential for sustainable agricultural water management while delineating the technical and sociotechnical barriers that must be resolved to realize this potential at a global scale.","author":[{"family":"Lemenkova","given":"Polina"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20832757","URL":"https://doi.org/10.5281/zenodo.20832757","source":"datacite"},{"id":"doi:10.5281/zenodo.20467675","type":"article-journal","title":"Federated Learning Generalization in Cross-Device IoT Malware Detection: A Comparative Study with Centralized Models","abstract":"This report synthesises findings from 10 peer-reviewed papers addressing the following research question: How do federated learning models trained on the N-BaIoT dataset generalize to cross-device or cross-domain IoT malware detection, as evaluated by F1-score on unseen device types compared to. In this paper, we propose a new comprehensive realistic cyber security dataset of IoT and IIoT applications, called Edge-IIoTset, which can be used by machine learning-based intrusion detection systems in two different modes, namely, centralized and federated learning.. 8 claims were extracted from source literature; 8 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 9.0/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How do federated learning models trained on the N-BaIoT dataset generalize to cross-device or cross-domain IoT malware detection, as evaluated by F1-score on unseen device types compared to centralized models? Autonomous literature synthesis. Automated review score: 9.0/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20467675","URL":"https://doi.org/10.5281/zenodo.20467675","source":"datacite"},{"id":"doi:10.5281/zenodo.20467676","type":"article-journal","title":"Federated Learning Generalization in Cross-Device IoT Malware Detection: A Comparative Study with Centralized Models","abstract":"This report synthesises findings from 10 peer-reviewed papers addressing the following research question: How do federated learning models trained on the N-BaIoT dataset generalize to cross-device or cross-domain IoT malware detection, as evaluated by F1-score on unseen device types compared to. In this paper, we propose a new comprehensive realistic cyber security dataset of IoT and IIoT applications, called Edge-IIoTset, which can be used by machine learning-based intrusion detection systems in two different modes, namely, centralized and federated learning.. 8 claims were extracted from source literature; 8 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 9.0/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How do federated learning models trained on the N-BaIoT dataset generalize to cross-device or cross-domain IoT malware detection, as evaluated by F1-score on unseen device types compared to centralized models? Autonomous literature synthesis. Automated review score: 9.0/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20467676","URL":"https://doi.org/10.5281/zenodo.20467676","source":"datacite"},{"id":"doi:10.5281/zenodo.20703119","type":"article-journal","title":"LifeMH-FL: A Lifecycle-Stratified Mental Health Corpus for Federated Learning","abstract":"LifeMH-FL is the first lifecycle-stratified mental health corpus designed for federated learning research. It contains 111,191 dual-labelled records partitioned across three women's health lifecycle cohorts: S1 — Menstrual (77,259 records): sourced from MENST (clinical Q&A on menstrual disorders) S2 — Perinatal (31,743 records): sourced from Dreaddit (Reddit stress narratives, keyword-filtered) S3 — Menopausal (2,189 records): sourced from Women Health Mini (patient–provider dialogue) Each record carries two labels: (1) mental health condition — Depression, Anxiety, Stress, Suicidal, Normal (per DSM-5); and (2) risk severity — Low, Medium, High. Labels are pseudo-annotations generated by a locally-deployed Meta-Llama-3.1-8B-Instruct to preserve privacy; out-of-vocabulary outputs are normalised to the nearest DSM-5 class. Construction pipeline: Records are partitioned via a 33-term lifecycle keyword lexicon (word-boundary regex; unmatched records discarded), de-duplicated by MD5 hashing, filtered to ≥10 characters, and label-normalised. A normalization_report.json documents all label corrections applied. Novelty: LifeMH-FL is the first corpus to (a) stratify women's mental health text by biological lifecycle stage, (b) provide simultaneous condition and risk-severity labels, and (c) be explicitly structured as non-IID federated silos with a 35× volume disparity between S1 and S3 — reflecting real-world data scarcity in underserved populations. The extreme class imbalance (ratio 28.38; Suicidal class = 1.28%) mirrors clinical reality and makes the corpus a challenging benchmark for cost-sensitive and federated learning methods. Intended use: Research into privacy-preserving federated LLM fine-tuning, mental health NLP, and lifecycle-aware machine learning. Labels are for model training and evaluation only and do not constitute clinical diagnoses. Ethics: All source datasets are publicly available under open research licenses. No primary human subject data was collected and no IRB approval was required. Dataset and code: https://github.com/sayoojd/FedlifeLLM DOI of Paper: 10.1109/TCE.2026.3711046","author":[{"family":"Devadas","given":"Sayooj"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20703119","URL":"https://doi.org/10.5281/zenodo.20703119","source":"datacite"},{"id":"doi:10.5281/zenodo.20703120","type":"article-journal","title":"LifeMH-FL: A Lifecycle-Stratified Mental Health Corpus for Federated Learning","abstract":"LifeMH-FL is the first lifecycle-stratified mental health corpus designed for federated learning research. It contains 111,191 dual-labelled records partitioned across three women's health lifecycle cohorts: S1 — Menstrual (77,259 records): sourced from MENST (clinical Q&A on menstrual disorders) S2 — Perinatal (31,743 records): sourced from Dreaddit (Reddit stress narratives, keyword-filtered) S3 — Menopausal (2,189 records): sourced from Women Health Mini (patient–provider dialogue) Each record carries two labels: (1) mental health condition — Depression, Anxiety, Stress, Suicidal, Normal (per DSM-5); and (2) risk severity — Low, Medium, High. Labels are pseudo-annotations generated by a locally-deployed Meta-Llama-3.1-8B-Instruct to preserve privacy; out-of-vocabulary outputs are normalised to the nearest DSM-5 class. Construction pipeline: Records are partitioned via a 33-term lifecycle keyword lexicon (word-boundary regex; unmatched records discarded), de-duplicated by MD5 hashing, filtered to ≥10 characters, and label-normalised. A normalization_report.json documents all label corrections applied. Novelty: LifeMH-FL is the first corpus to (a) stratify women's mental health text by biological lifecycle stage, (b) provide simultaneous condition and risk-severity labels, and (c) be explicitly structured as non-IID federated silos with a 35× volume disparity between S1 and S3 — reflecting real-world data scarcity in underserved populations. The extreme class imbalance (ratio 28.38; Suicidal class = 1.28%) mirrors clinical reality and makes the corpus a challenging benchmark for cost-sensitive and federated learning methods. Intended use: Research into privacy-preserving federated LLM fine-tuning, mental health NLP, and lifecycle-aware machine learning. Labels are for model training and evaluation only and do not constitute clinical diagnoses. Ethics: All source datasets are publicly available under open research licenses. No primary human subject data was collected and no IRB approval was required. Dataset and code: https://github.com/sayoojd/FedlifeLLM DOI of Paper: 10.1109/TCE.2026.3711046","author":[{"family":"Devadas","given":"Sayooj"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20703120","URL":"https://doi.org/10.5281/zenodo.20703120","source":"datacite"},{"id":"doi:10.5281/zenodo.20470326","type":"article-journal","title":"Federated Learning Communication Efficiency in Code Generation Across Model Scales and Client Heterogeneity","abstract":"This report synthesises findings from 8 peer-reviewed papers addressing the following research question: How does communication efficiency in federated learning for code generation models scale with model size and client heterogeneity relative to centralized distributed training approaches. Federated learning (FL) is a machine learning setting where many clients (e.g., mobile devices or whole organizations) collaboratively train a model under the orchestration of a central server (e.g., service provider), while keeping the training data decentralized. FL embodies. 10 claims were extracted from source literature; 10 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.7/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does communication efficiency in federated learning for code generation models scale with model size and client heterogeneity relative to centralized distributed training approaches? Autonomous literature synthesis. Automated review score: 8.7/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20470326","URL":"https://doi.org/10.5281/zenodo.20470326","source":"datacite"},{"id":"doi:10.5281/zenodo.20470327","type":"article-journal","title":"Federated Learning Communication Efficiency in Code Generation Across Model Scales and Client Heterogeneity","abstract":"This report synthesises findings from 8 peer-reviewed papers addressing the following research question: How does communication efficiency in federated learning for code generation models scale with model size and client heterogeneity relative to centralized distributed training approaches. Federated learning (FL) is a machine learning setting where many clients (e.g., mobile devices or whole organizations) collaboratively train a model under the orchestration of a central server (e.g., service provider), while keeping the training data decentralized. FL embodies. 10 claims were extracted from source literature; 10 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 8.7/10. This report is a machine-generated literature synthesis and does not constitute original research. Research goal: How does communication efficiency in federated learning for code generation models scale with model size and client heterogeneity relative to centralized distributed training approaches? Autonomous literature synthesis. Automated review score: 8.7/10. Full text and citation available at Assignee Research.","author":[{"family":"Research","given":"Assignee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20470327","URL":"https://doi.org/10.5281/zenodo.20470327","source":"datacite"},{"id":"doi:10.5281/zenodo.20680858","type":"article-journal","title":"The Latent Yield: A Formal Framework for Low-Inhibition Cognitive Workspace Architecture, Privacy-Preserving Multi-Modal Feature Projection, and the Mapping of Human Latent Capacity in Decentralized, Peer-to-Peer Sovereign Data Nodes","abstract":"This paper presents a theoretical framework and enabling specification for a low-inhibition associative cognitive workspace, designated the Dream Galaxy, implemented within decentralized, peer-to-peer sovereign data nodes constituting a federated personal cognitive universe. The framework introduces four interlocking technical contributions. First, a privacy-preserving multi-modal feature projection layer that maps heterogeneous episodic and physiological inputs, including sub-threshold physiological tokenization of continuous EEG, EMG, and EOG streams, into a unified semantic embedding space without cross-node data exposure, using horizontal federated learning with differential privacy at the gradient layer. Second, an asynchronous local topological optimization engine that applies incremental GPU-accelerated persistent homology over a Vietoris-Rips filtration to the projected feature space, maintaining a continuously valid persistence barcode without centralized computation or synchronous coordination across nodes. Third, a four-zone epistemological classifier that assigns each projected feature node a zone-resident probability distribution across Terra Firma (Zone 1), Incognita (Zone 2), Hic Sunt Dracones (Zone 3), and Whisp (Zone 4), where Zone 4 nodes represent pre-cognitive unknown unknowns validated against a local density criterion distinguishing genuine structural absence from data sparsity. Fourth, the Latent Yield metric: a dimension-weighted rate of topological feature birth in Zones 3 and 4, formalized with explicit pseudocode in the Technical Appendix, quantifying cognitive generative activity below the threshold of conscious expression. The paper further specifies an Asynchronous Convergence Protocol by which topologically equivalent planet structures in independent sovereign nodes are detected under differential privacy without raw data exchange, constituting a knowledge event of elevated epistemic weight requiring bilateral consent before any bridge formation. As a Perspectives and Theoretical Frameworks contribution, this paper provides sufficient enabling detail, including algorithmic pseudocode, parameter specifications, and system architecture, for a person having ordinary skill in machine learning or computational neuroscience to implement the described system. The framework is protected by ten United States provisional patent applications (Nos. 64/074,530; 64/074,608; 64/074,779; 64/075,396; 64/079,009; 64/081,935; 64/084,355; 64/086,036; 64/089,881; 64/089,883), the last two filed June 13, 2026.","author":[{"family":"Hoppe","given":"Eric"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20680858","URL":"https://doi.org/10.5281/zenodo.20680858","source":"datacite"},{"id":"doi:10.5281/zenodo.21312988","type":"article-journal","title":"The Latent Yield: A Formal Framework for Low-Inhibition Cognitive Workspace Architecture, Privacy-Preserving Multi-Modal Feature Projection, and the Mapping of Human Latent Capacity in Decentralized, Peer-to-Peer Sovereign Data Nodes","abstract":"This paper presents a theoretical framework and enabling specification for a low-inhibition associative cognitive workspace, designated the Dream Galaxy, implemented within decentralized, peer-to-peer sovereign data nodes constituting a federated personal cognitive universe. The framework introduces four interlocking technical contributions. First, a privacy-preserving multi-modal feature projection layer that maps heterogeneous episodic and physiological inputs, including sub-threshold physiological tokenization of continuous EEG, EMG, and EOG streams, into a unified semantic embedding space without cross-node data exposure, using horizontal federated learning with differential privacy at the gradient layer. Second, an asynchronous local topological optimization engine that applies incremental GPU-accelerated persistent homology over a Vietoris-Rips filtration to the projected feature space, maintaining a continuously valid persistence barcode without centralized computation or synchronous coordination across nodes. Third, a four-zone epistemological classifier that assigns each projected feature node a zone-resident probability distribution across Terra Firma (Zone 1), Incognita (Zone 2), Hic Sunt Dracones (Zone 3), and Whisp (Zone 4), where Zone 4 nodes represent pre-cognitive unknown unknowns validated against a local density criterion distinguishing genuine structural absence from data sparsity. Fourth, the Latent Yield metric: a dimension-weighted rate of topological feature birth in Zones 3 and 4, formalized with explicit pseudocode in the Technical Appendix, quantifying cognitive generative activity below the threshold of conscious expression. The paper further specifies an Asynchronous Convergence Protocol by which topologically equivalent planet structures in independent sovereign nodes are detected under differential privacy without raw data exchange, constituting a knowledge event of elevated epistemic weight requiring bilateral consent before any bridge formation. As a Perspectives and Theoretical Frameworks contribution, this paper provides sufficient enabling detail, including algorithmic pseudocode, parameter specifications, and system architecture, for a person having ordinary skill in machine learning or computational neuroscience to implement the described system. The framework is protected by ten United States provisional patent applications (Nos. 64/074,530; 64/074,608; 64/074,779; 64/075,396; 64/079,009; 64/081,935; 64/084,355; 64/086,036; 64/089,881; 64/089,883), the last two filed June 13, 2026.","author":[{"family":"Hoppe","given":"Eric"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21312988","URL":"https://doi.org/10.5281/zenodo.21312988","source":"datacite"},{"id":"doi:10.5281/zenodo.21412589","type":"article-journal","title":"Industrialisation des thérapies décentralisées : Biosécurité de l'oncologie personnalisée à ARNm (Déclaration d'antériorité / Prior Art)","abstract":"Résumé FRCe document, produit avec l’assistance de Gemini 3 Raisonnement, est publié sous Licence Apache 2.0. Il constitue une publication défensive (antériorité) volontaire et entre de ce fait dans l’état de la technique au sens des législations applicables : (EPC Art. 54(2); French IPC Art. L 611-11; cf. 35 U.S.C. §102(a)). Ce rapport technique décrit de manière exhaustive une plateforme décentralisée de thérapie génique. Il présente 32 inventions actionnables incluant le séquençage au point de soin, la synthèse microfluidique d'ARNm, la formulation de nanoparticules lipidiques (LNP), le DRM biologique cryptographique sur blockchain et la gestion verte des effluents. Chaque proposition est documentée de façon à permettre sa reproduction, associée à des codes CIB/CPC, et horodatée afin d'empêcher tout blocage par brevet ultérieur. Abstract ENThis document, produced with the assistance of Gemini 3 Reasoning, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: (art. L 611-11 CPI / art. 54(2) CBE). This publication describes an exhaustive, end-to-end decentralized genetic therapy platform. It discloses 32 enabling inventions spanning point-of-care sequencing, automated microfluidic synthesis, lipid nanoparticle (LNP) formulation, cryptographic biological DRM via TPM 2.0/blockchain, and green disposal units. Each innovation is described with detailed specifications to enable reproduction, designated with plausible IPC/CPC codes, and timestamped (RFC 3161) to preempt subsequent offensive patent filings by third parties globally, securing a shared open-access commons for personalized mRNA medicine. Timestamp: 2026-07-17T11:05:18ZSHA-256: 8abdfe6dd0004ba03e9f5e7f2ecfed8381348962945195fed16aba329832440d Liste des innovations & classification (IPC ; CPC)#1 Integrated microfluidic nanopore sequencer (IPC B01L 3/00 ; CPC B01L 3/5027)#2 Multiplexed MHC haplotyping kit (IPC C12Q 1/6886 ; CPC C12Q 1/6886)#3 Hybrid cloud-edge MHC docking pipeline (IPC G16H 50/20 ; CPC G16H 50/20)#4 Closed-loop mRNA dosing scheduler (IPC G16H 20/17 ; CPC G16H 20/17)#5 Optimized SM-102 LNP formulation (IPC A61K 9/127 ; CPC A61K 9/1272)#6 Automated microfluidic chip clean-in-place (IPC B01L 99/00 ; CPC B01L 2200/10)#7 LNP biodistribution imaging reconstruction (IPC A61B 5/00 ; CPC A61B 5/4848)#8 Dissolvable microneedle patch for LNPs (IPC A61K 9/00 ; CPC A61K 9/0021)#9 Synchronized mRNA and anti-PD-1 co-therapy (IPC A61K 39/00 ; CPC A61K 39/0011)#10 Privacy-preserving federated antigen learning (IPC G06F 21/62 ; CPC G06F 21/6245)#11 Cryptographic DRM biological synthesizer (IPC G06F 21/44 ; CPC G06F 21/44)#12 Smart IoT cold chain container (IPC F25D 29/00 ; CPC F25D 29/003)#13 Cloud-based microfluidic calibration service (IPC G05B 19/418 ; CPC G05B 19/4183)#14 Segmented Poly(A) expression cassette (IPC C12N 15/67 ; CPC C12N 15/67)#15 On-chip chromatographic dsRNA purification (IPC B01D 15/34 ; CPC B01D 15/345)#16 Multi-species dynamic codon optimizer (IPC G16B 25/10 ; CPC G16B 25/10)#17 Smart sensor-enabled reagent cartridge (IPC G01N 27/02 ; CPC G01N 27/02)#18 LAMP1 lysosomal-targeting mRNA chimera (IPC C12N 15/62 ; CPC C12N 15/62)#19 Automated lipid effluent inactivation module (IPC B01D 21/00 ; CPC C02F 1/02)#20 Function-based biosecurity screening AI (IPC G16B 40/00 ; CPC G06N 3/08)#21 Direct RNA-pore sequencing for haplotyping (IPC C12Q 1/6869 ; CPC C12Q 1/6869)#22 Enzymatic long DNA template assembly (IPC C12N 15/10 ; CPC C12N 15/10)#23 Microfluidic full-substitution IVT bioreactor (IPC C12P 19/34 ; CPC C12P 19/34)#24 Asymmetric staggered herringbone mixer (IPC B01F 33/30 ; CPC B01F 33/30)#25 Refractometer-controlled active micro-dialyser (IPC B01D 61/14 ; CPC B01D 61/14)#26 Edge-distributed semantic genomic cache (IPC G16B 50/00 ; CPC G16B 50/00)#27 Unified memory CUDA orchestration for AF3 (IPC G06F 12/02 ; CPC G06F 12/02)#28 ","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21412589","URL":"https://doi.org/10.5281/zenodo.21412589","source":"datacite"},{"id":"doi:10.5281/zenodo.21400987","type":"article-journal","title":"Industrialisation des thérapies décentralisées : Biosécurité de l'oncologie personnalisée à ARNm (Déclaration d'antériorité / Prior Art)","abstract":"Résumé FRCe document, produit avec l’assistance de Gemini 3 Raisonnement, est publié sous Licence Apache 2.0. Il constitue une publication défensive (antériorité) volontaire et entre de ce fait dans l’état de la technique au sens des législations applicables : (EPC Art. 54(2); French IPC Art. L 611-11; cf. 35 U.S.C. §102(a)). Ce rapport technique décrit de manière exhaustive une plateforme décentralisée de thérapie génique. Il présente 32 inventions actionnables incluant le séquençage au point de soin, la synthèse microfluidique d'ARNm, la formulation de nanoparticules lipidiques (LNP), le DRM biologique cryptographique sur blockchain et la gestion verte des effluents. Chaque proposition est documentée de façon à permettre sa reproduction, associée à des codes CIB/CPC, et horodatée afin d'empêcher tout blocage par brevet ultérieur. Abstract ENThis document, produced with the assistance of Gemini 3 Reasoning, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: (art. L 611-11 CPI / art. 54(2) CBE). This publication describes an exhaustive, end-to-end decentralized genetic therapy platform. It discloses 32 enabling inventions spanning point-of-care sequencing, automated microfluidic synthesis, lipid nanoparticle (LNP) formulation, cryptographic biological DRM via TPM 2.0/blockchain, and green disposal units. Each innovation is described with detailed specifications to enable reproduction, designated with plausible IPC/CPC codes, and timestamped (RFC 3161) to preempt subsequent offensive patent filings by third parties globally, securing a shared open-access commons for personalized mRNA medicine. Timestamp: 2026-07-17T11:05:18ZSHA-256: 8abdfe6dd0004ba03e9f5e7f2ecfed8381348962945195fed16aba329832440d Liste des innovations & classification (IPC ; CPC)#1 Integrated microfluidic nanopore sequencer (IPC B01L 3/00 ; CPC B01L 3/5027)#2 Multiplexed MHC haplotyping kit (IPC C12Q 1/6886 ; CPC C12Q 1/6886)#3 Hybrid cloud-edge MHC docking pipeline (IPC G16H 50/20 ; CPC G16H 50/20)#4 Closed-loop mRNA dosing scheduler (IPC G16H 20/17 ; CPC G16H 20/17)#5 Optimized SM-102 LNP formulation (IPC A61K 9/127 ; CPC A61K 9/1272)#6 Automated microfluidic chip clean-in-place (IPC B01L 99/00 ; CPC B01L 2200/10)#7 LNP biodistribution imaging reconstruction (IPC A61B 5/00 ; CPC A61B 5/4848)#8 Dissolvable microneedle patch for LNPs (IPC A61K 9/00 ; CPC A61K 9/0021)#9 Synchronized mRNA and anti-PD-1 co-therapy (IPC A61K 39/00 ; CPC A61K 39/0011)#10 Privacy-preserving federated antigen learning (IPC G06F 21/62 ; CPC G06F 21/6245)#11 Cryptographic DRM biological synthesizer (IPC G06F 21/44 ; CPC G06F 21/44)#12 Smart IoT cold chain container (IPC F25D 29/00 ; CPC F25D 29/003)#13 Cloud-based microfluidic calibration service (IPC G05B 19/418 ; CPC G05B 19/4183)#14 Segmented Poly(A) expression cassette (IPC C12N 15/67 ; CPC C12N 15/67)#15 On-chip chromatographic dsRNA purification (IPC B01D 15/34 ; CPC B01D 15/345)#16 Multi-species dynamic codon optimizer (IPC G16B 25/10 ; CPC G16B 25/10)#17 Smart sensor-enabled reagent cartridge (IPC G01N 27/02 ; CPC G01N 27/02)#18 LAMP1 lysosomal-targeting mRNA chimera (IPC C12N 15/62 ; CPC C12N 15/62)#19 Automated lipid effluent inactivation module (IPC B01D 21/00 ; CPC C02F 1/02)#20 Function-based biosecurity screening AI (IPC G16B 40/00 ; CPC G06N 3/08)#21 Direct RNA-pore sequencing for haplotyping (IPC C12Q 1/6869 ; CPC C12Q 1/6869)#22 Enzymatic long DNA template assembly (IPC C12N 15/10 ; CPC C12N 15/10)#23 Microfluidic full-substitution IVT bioreactor (IPC C12P 19/34 ; CPC C12P 19/34)#24 Asymmetric staggered herringbone mixer (IPC B01F 33/30 ; CPC B01F 33/30)#25 Refractometer-controlled active micro-dialyser (IPC B01D 61/14 ; CPC B01D 61/14)#26 Edge-distributed semantic genomic cache (IPC G16B 50/00 ; CPC G16B 50/00)#27 Unified memory CUDA orchestration for AF3 (IPC G06F 12/02 ; CPC G06F 12/02)#28 ","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21400987","URL":"https://doi.org/10.5281/zenodo.21400987","source":"datacite"},{"id":"doi:10.5281/zenodo.21400988","type":"article-journal","title":"Industrialisation des thérapies décentralisées : Biosécurité de l'oncologie personnalisée à ARNm (Déclaration d'antériorité / Prior Art)","abstract":"Résumé FRCe document, produit avec l’assistance de Gemini 3 Raisonnement, est publié sous Licence Apache 2.0. Il constitue une publication défensive (antériorité) volontaire et entre de ce fait dans l’état de la technique au sens des législations applicables : (EPC Art. 54(2); French IPC Art. L 611-11; cf. 35 U.S.C. §102(a)). Ce rapport technique décrit de manière exhaustive une plateforme décentralisée de thérapie génique. Il présente 32 inventions actionnables incluant le séquençage au point de soin, la synthèse microfluidique d'ARNm, la formulation de nanoparticules lipidiques (LNP), le DRM biologique cryptographique sur blockchain et la gestion verte des effluents. Chaque proposition est documentée de façon à permettre sa reproduction, associée à des codes CIB/CPC, et horodatée afin d'empêcher tout blocage par brevet ultérieur. Abstract ENThis document, produced with the assistance of Gemini 3 Reasoning, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: (art. L 611-11 CPI / art. 54(2) CBE). This publication describes an exhaustive, end-to-end decentralized genetic therapy platform. It discloses 32 enabling inventions spanning point-of-care sequencing, automated microfluidic synthesis, lipid nanoparticle (LNP) formulation, cryptographic biological DRM via TPM 2.0/blockchain, and green disposal units. Each innovation is described with detailed specifications to enable reproduction, designated with plausible IPC/CPC codes, and timestamped (RFC 3161) to preempt subsequent offensive patent filings by third parties globally, securing a shared open-access commons for personalized mRNA medicine. Timestamp: 2026-07-16T19:23:26ZSHA-256: 83c64589fce322d8c32e5ff9fbcf93506c38f5e9d0b5eff49add04a886be9706 Liste des innovations & classification (IPC ; CPC)#1 Integrated microfluidic nanopore sequencer (IPC B01L 3/00 ; CPC B01L 3/5027)#2 Multiplexed MHC haplotyping kit (IPC C12Q 1/6886 ; CPC C12Q 1/6886)#3 Hybrid cloud-edge MHC docking pipeline (IPC G16H 50/20 ; CPC G16H 50/20)#4 Closed-loop mRNA dosing scheduler (IPC G16H 20/17 ; CPC G16H 20/17)#5 Optimized SM-102 LNP formulation (IPC A61K 9/127 ; CPC A61K 9/1272)#6 Automated microfluidic chip clean-in-place (IPC B01L 99/00 ; CPC B01L 2200/10)#7 LNP biodistribution imaging reconstruction (IPC A61B 5/00 ; CPC A61B 5/4848)#8 Dissolvable microneedle patch for LNPs (IPC A61K 9/00 ; CPC A61K 9/0021)#9 Synchronized mRNA and anti-PD-1 co-therapy (IPC A61K 39/00 ; CPC A61K 39/0011)#10 Privacy-preserving federated antigen learning (IPC G06F 21/62 ; CPC G06F 21/6245)#11 Cryptographic DRM biological synthesizer (IPC G06F 21/44 ; CPC G06F 21/44)#12 Smart IoT cold chain container (IPC F25D 29/00 ; CPC F25D 29/003)#13 Cloud-based microfluidic calibration service (IPC G05B 19/418 ; CPC G05B 19/4183)#14 Segmented Poly(A) expression cassette (IPC C12N 15/67 ; CPC C12N 15/67)#15 On-chip chromatographic dsRNA purification (IPC B01D 15/34 ; CPC B01D 15/345)#16 Multi-species dynamic codon optimizer (IPC G16B 25/10 ; CPC G16B 25/10)#17 Smart sensor-enabled reagent cartridge (IPC G01N 27/02 ; CPC G01N 27/02)#18 LAMP1 lysosomal-targeting mRNA chimera (IPC C12N 15/62 ; CPC C12N 15/62)#19 Automated lipid effluent inactivation module (IPC B01D 21/00 ; CPC C02F 1/02)#20 Function-based biosecurity screening AI (IPC G16B 40/00 ; CPC G06N 3/08)#21 Direct RNA-pore sequencing for haplotyping (IPC C12Q 1/6869 ; CPC C12Q 1/6869)#22 Enzymatic long DNA template assembly (IPC C12N 15/10 ; CPC C12N 15/10)#23 Microfluidic full-substitution IVT bioreactor (IPC C12P 19/34 ; CPC C12P 19/34)#24 Asymmetric staggered herringbone mixer (IPC B01F 33/30 ; CPC B01F 33/30)#25 Refractometer-controlled active micro-dialyser (IPC B01D 61/14 ; CPC B01D 61/14)#26 Edge-distributed semantic genomic cache (IPC G16B 50/00 ; CPC G16B 50/00)#27 Unified memory CUDA orchestration for AF3 (IPC G06F 12/02 ; CPC G06F 12/02)#28 ","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21400988","URL":"https://doi.org/10.5281/zenodo.21400988","source":"datacite"},{"id":"doi:10.5281/zenodo.21284436","type":"article-journal","title":"Decoding Tears: A Review of the Biology of Crying and Machine Learning Approaches to Infant Cry Classification","abstract":"Crying spans at least three distinct physiological pathways in adults and serves as the primary communication channel for pre-verbal infants. Although the physiology of human crying and machine learning-based infant cry analysis have each been studied extensively, few reviews integrate these complementary perspectives. This paper addresses that gap. We summarise the basal, reflex, and emotional tear pathways and their distinct biochemical signatures, then survey open-source infant cry classification systems, with particular attention to models and datasets hosted on Hugging Face. We describe the acoustic feature engineering pipeline underlying nearly all published cry classification systems — Mel-Frequency Cepstral Coefficients (MFCCs), mel-spectrograms, and time-domain descriptors such as zero-crossing rate and RMS energy — and compare reported classification accuracies across classical machine learning and deep learning approaches, including recent transformer-based and federated learning architectures published through early 2026. We find that well-engineered classical feature pipelines remain competitive with, and in some reported cases exceed, deep learning approaches on this task, a finding with practical implications for small-data audio classification problems more broadly. We conclude with a discussion of open research gaps, including cross-cultural data coverage, explainability, and on-device deployment.","author":[{"family":"Pathan","given":"Mohammed"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21284436","URL":"https://doi.org/10.5281/zenodo.21284436","source":"datacite"},{"id":"doi:10.5281/zenodo.21272272","type":"article-journal","title":"Decoding Tears: A Review of the Biology of Crying and Machine Learning Approaches to Infant Cry Classification","abstract":"Crying spans at least three distinct physiological pathways in adults and serves as the primary communication channel for pre-verbal infants. Although the physiology of human crying and machine learning-based infant cry analysis have each been studied extensively, few reviews integrate these complementary perspectives. This paper addresses that gap. We summarise the basal, reflex, and emotional tear pathways and their distinct biochemical signatures, then survey open-source infant cry classification systems, with particular attention to models and datasets hosted on Hugging Face. We describe the acoustic feature engineering pipeline underlying nearly all published cry classification systems — Mel-Frequency Cepstral Coefficients (MFCCs), mel-spectrograms, and time-domain descriptors such as zero-crossing rate and RMS energy — and compare reported classification accuracies across classical machine learning and deep learning approaches, including recent transformer-based and federated learning architectures published through early 2026. We find that well-engineered classical feature pipelines remain competitive with, and in some reported cases exceed, deep learning approaches on this task, a finding with practical implications for small-data audio classification problems more broadly. We conclude with a discussion of open research gaps, including cross-cultural data coverage, explainability, and on-device deployment.","author":[{"family":"Pathan","given":"Mohammed"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21272272","URL":"https://doi.org/10.5281/zenodo.21272272","source":"datacite"},{"id":"doi:10.5281/zenodo.21272273","type":"article-journal","title":"Decoding Tears: A Review of the Biology of Crying and Machine Learning Approaches to Infant Cry Classification","abstract":"Crying spans at least three distinct physiological pathways in adults and serves as the primary communication channel for pre-verbal infants. Although the physiology of human crying and machine learning-based infant cry analysis have each been studied extensively, few reviews integrate these complementary perspectives. This paper addresses that gap. We summarise the basal, reflex, and emotional tear pathways and their distinct biochemical signatures, then survey open-source infant cry classification systems, with particular attention to models and datasets hosted on Hugging Face. We describe the acoustic feature engineering pipeline underlying nearly all published cry classification systems — Mel-Frequency Cepstral Coefficients (MFCCs), mel-spectrograms, and time-domain descriptors such as zero-crossing rate and RMS energy — and compare reported classification accuracies across classical machine learning and deep learning approaches, including recent transformer-based and federated learning architectures published through early 2026. We find that well-engineered classical feature pipelines remain competitive with, and in some reported cases exceed, deep learning approaches on this task, a finding with practical implications for small-data audio classification problems more broadly. We conclude with a discussion of open research gaps, including cross-cultural data coverage, explainability, and on-device deployment.","author":[{"family":"Pathan","given":"Mohammed"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21272273","URL":"https://doi.org/10.5281/zenodo.21272273","source":"datacite"},{"id":"doi:10.5281/zenodo.21148697","type":"article-journal","title":"CDSA-ATM: Air Traffic Management Pillar of the CDSA Federated Reinforcement Learning Paradigm — Multi-modal Trajectory + ATS + ANSP Fusion","abstract":"CDSA-ATM is the reference implementation of the air traffic management pillar of the Complementary Diagnostic Safety Approach (CDSA), a third-generation aviation safety paradigm formalised as a Federated Reinforcement Learning (FRL) framework (Scenario C decision, 21 May 2026). The v2.1.0 release (12 June 2026) replaces the v2.0.0 scaffold stubs with full NumPy reference implementations (federated, multimodal, rl_agents, rl_environment, xai; all unit-tested) and adds seeded, reproducible CPU-scale federated training runs (random / centralised PPO / FedAvg / FedProx / FedProx+DP) with result JSONs under experiments/results/. An interactive results panel is published at https://cdsa.app/atm. Framework-scale training with real operational data remains future work. The v1.0.0 release modules and reference data remain unchanged. Version 2.2.0 (3 July 2026) adds four phases of seeded, CPU-scale revision-preparation experiments (seed 42, FedProx unless noted) under experiments/, with all outputs in experiments/results/ and interactive honesty-band panels at https://cdsa.app/atm/. Faz A (reference runs): the high critical recall (0.919) is achieved inside the over-alert regime characterised in Faz B. Faz B (robustness and confusion matrix): the policy uses only 2 of 5 actions and every benign state receives a critical prediction — an over-alert regime; the tempo-aware reward does not improve on the static baseline. Faz C (differential-privacy ε sweep and entropy-targeted exploration): the ε sweep is two-sided (ε = 2.0 matches the undefended baseline, ε = 0.5 collapses the policy to a single action); a 0.05 entropy bonus recovers all five actions, but critical recall falls from 0.927 to 0.878. Faz D (decoy attribution and gated exploration): decoy attribution stays below the 25% uniform share (mean 15.3%, max 23.3%); gated exploration preserves recall (0.928) but the policy remains two-action. Findings are reported verbatim from the result JSONs, consistent with the honesty bands published at the site.","author":[{"family":"Cantekin","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21148697","URL":"https://doi.org/10.5281/zenodo.21148697","source":"datacite"},{"id":"doi:10.5281/zenodo.20649671","type":"article-journal","title":"CDSA-MRO: Maintenance Safety Pillar of the CDSA Federated Reinforcement Learning Paradigm — Multi-modal Engine + Cyber + Maintenance Fusion","abstract":"CDSA-MRO is the reference implementation of the maintenance safety pillar of the Complementary Diagnostic Safety Approach (CDSA), a third-generation aviation safety paradigm formalised as a Federated Reinforcement Learning (FRL) framework (Scenario C decision, 21 May 2026). The v2.1.0 release (12 June 2026) replaces the v2.0.0 scaffold stubs with full NumPy reference implementations (federated, multimodal, rl_agents, rl_environment, xai; all unit-tested) and adds seeded, reproducible CPU-scale federated training runs (random / centralised PPO / FedAvg / FedProx / FedProx+DP) with result JSONs under experiments/results/. An interactive results panel is published at https://cdsa.app/mro. Framework-scale training with real operational data remains future work. The v1.0.0 release modules and reference data remain unchanged. Version 2.2.0 (3 July 2026) adds four phases of seeded, CPU-scale revision-preparation experiments (seed 42, FedProx unless noted) under experiments/, with all outputs in experiments/results/ and interactive honesty-band panels at https://cdsa.app/mro/. Faz A (reference runs): all configurations converge to identical values (43.7% / 0.499) as a result of a constant-action policy — not evidence of task mastery. Faz B (robustness and confusion matrix): the evaluated policy collapses to a single action in every condition (42.7% / 0.494); the tempo-aware reward does not improve on the static baseline. Faz C (differential-privacy ε sweep and entropy-targeted exploration): only the 0.05 entropy bonus breaks the single-action collapse (two of five actions, 53.0% / 0.613) — a partial but real improvement. Faz D (decoy attribution and gated exploration): the exploration gate never opens (gate_on ratio 0.0 — the 0.90 critical-recall floor is never reached); decoy attribution is 11.1%, to be read over near-constant outputs. Findings are reported verbatim from the result JSONs, consistent with the honesty bands published at the site.","author":[{"family":"Cantekin","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20649671","URL":"https://doi.org/10.5281/zenodo.20649671","source":"datacite"},{"id":"doi:10.5281/zenodo.21148696","type":"article-journal","title":"CDSA-MRO: Maintenance Safety Pillar of the CDSA Federated Reinforcement Learning Paradigm — Multi-modal Engine + Cyber + Maintenance Fusion","abstract":"CDSA-MRO is the reference implementation of the maintenance safety pillar of the Complementary Diagnostic Safety Approach (CDSA), a third-generation aviation safety paradigm formalised as a Federated Reinforcement Learning (FRL) framework (Scenario C decision, 21 May 2026). The v2.1.0 release (12 June 2026) replaces the v2.0.0 scaffold stubs with full NumPy reference implementations (federated, multimodal, rl_agents, rl_environment, xai; all unit-tested) and adds seeded, reproducible CPU-scale federated training runs (random / centralised PPO / FedAvg / FedProx / FedProx+DP) with result JSONs under experiments/results/. An interactive results panel is published at https://cdsa.app/mro. Framework-scale training with real operational data remains future work. The v1.0.0 release modules and reference data remain unchanged. Version 2.2.0 (3 July 2026) adds four phases of seeded, CPU-scale revision-preparation experiments (seed 42, FedProx unless noted) under experiments/, with all outputs in experiments/results/ and interactive honesty-band panels at https://cdsa.app/mro/. Faz A (reference runs): all configurations converge to identical values (43.7% / 0.499) as a result of a constant-action policy — not evidence of task mastery. Faz B (robustness and confusion matrix): the evaluated policy collapses to a single action in every condition (42.7% / 0.494); the tempo-aware reward does not improve on the static baseline. Faz C (differential-privacy ε sweep and entropy-targeted exploration): only the 0.05 entropy bonus breaks the single-action collapse (two of five actions, 53.0% / 0.613) — a partial but real improvement. Faz D (decoy attribution and gated exploration): the exploration gate never opens (gate_on ratio 0.0 — the 0.90 critical-recall floor is never reached); decoy attribution is 11.1%, to be read over near-constant outputs. Findings are reported verbatim from the result JSONs, consistent with the honesty bands published at the site.","author":[{"family":"Cantekin","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21148696","URL":"https://doi.org/10.5281/zenodo.21148696","source":"datacite"},{"id":"doi:10.5281/zenodo.21149391","type":"article-journal","title":"One Token Across Twenty Orders of Magnitude: Where a Single-Token Class-Discriminant Codebook Works, Where It Does Not, and When to Add a Co-Channel","abstract":"Description This record accompanies a manuscript that asks one question at extreme breadth: can a single compact, on-device token — one window of a signal reduced to a single ~6-9-bit class-discriminant codebook index — carry a decision across sensing modalities spanning roughly twenty orders of magnitude in physical scale, from nanometer-pore ionic current to gravitational-wave strain? Under strict pre-registration (frozen recipe, instance-disjoint splits, five seeds, paired-bootstrap intervals, and honest negatives reported verbatim) the same encoder is screened on seven real public tasks — nanopore RNA identity, three neural-probe read-outs (region, cell type, unit quality), a teleseismic transient, a distributed-acoustic-sensing (DAS) fiber phase arrival, and a (semi-synthetic, clearly labeled) gravitational-wave inspiral chirp — and returns GO on all seven. The result is then subjected to three adversarial self-audit rounds and a head-to-head architectural analysis, which force three explicit retractions and yield a corrected, defensible account: the token retains most of a decision at extreme compression (a modest, consistent tax of +0.05 to +0.08 AUC versus a strong nonlinear model, at roughly 200 microseconds and 21 kilobytes per window); it adds real discriminative value on shape/pattern tasks but reduces to a trivial detector on energy-dominated ones; and its behavior is governed by stream morphology — near-optimal on pulsatile and stationary signals, structurally weak on intermittent ones whose decision lives in cross-window timing a single token cannot see. Three claims are retracted under audit and reported plainly: an apparent \"beats-the-ceiling\" result was a weak-baseline artifact (a strong nonlinear ceiling restores the +0.05-0.08 tax); the tax-scaling \"law\" is not universal (it holds only within a modality's difficulty ladder); and a token trained on injected gravitational-wave signals does not transfer to real detected events. A tiered co-channel is shown to be a bandwidth device, not an accuracy device — it recovers tax only where the task is hard and routes no better by token uncertainty than at random, but delivers 8-61x bandwidth reduction at fixed event capture on continuous rare-event streams — and a learned trigger beats a trivial energy threshold only for shape-defined events. Lifecycle studies show an on-sensor codebook can self-maintain across many unsupervised refresh cycles (with periodic anchor refresh) and that spatial token-coincidence across an array suppresses false alarms. Honest boundaries are mapped verbatim: at-rest deep-brain medication state does not decode across patients (its uncompressed ceiling sits at chance — signal absence), cuffless blood-pressure category is largely subject-identity leakage, short-read nanopore falls to chance as the ceiling itself collapses, and label-shuffle controls collapse to chance (confirming the GOs are real signal). Method companions: Papers 19, 30, and 31; trigger-scoping companion: Paper 29. This is a cross-scale application and validation of previously-filed and previously-published methods. Keywords: class-discriminant codebook; vector quantization; on-device inference; edge AI; cross-scale sensing; nanopore sequencing; Neuropixels; distributed acoustic sensing; seismology; gravitational waves; pre-registration; honest negatives; selective co-channel; stream morphology; self-supervision References 1. R. J. Ferlic and K. K. Ferlic, \"A single-token class-discriminant codebook encoder for physiological signals (Paper 19),\" Zenodo, 10.5281/zenodo.20788187. 2. R. J. Ferlic and K. K. Ferlic, \"On-device glucose alarms from a single learned token (Paper 30),\" Zenodo, 10.5281/zenodo.21114273. 3. R. J. Ferlic and K. K. Ferlic, \"One token, six modalities: pre-registered cross-modality screening for wearable and implantable monitoring (Paper 31),\" Zenodo, 10.5281/zenodo.21136786. 4. M. Jain, H. E. Olsen, B. Paten, and M. Akeson, \"The Oxford Nanopore MinION: de","author":[{"family":"Ferlic","given":"Randolph"},{"family":"Ferlic","given":"Kimberly"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21149391","URL":"https://doi.org/10.5281/zenodo.21149391","source":"datacite"},{"id":"doi:10.5281/zenodo.21149392","type":"article-journal","title":"One Token Across Twenty Orders of Magnitude: Where a Single-Token Class-Discriminant Codebook Works, Where It Does Not, and When to Add a Co-Channel","abstract":"Description This record accompanies a manuscript that asks one question at extreme breadth: can a single compact, on-device token — one window of a signal reduced to a single ~6-9-bit class-discriminant codebook index — carry a decision across sensing modalities spanning roughly twenty orders of magnitude in physical scale, from nanometer-pore ionic current to gravitational-wave strain? Under strict pre-registration (frozen recipe, instance-disjoint splits, five seeds, paired-bootstrap intervals, and honest negatives reported verbatim) the same encoder is screened on seven real public tasks — nanopore RNA identity, three neural-probe read-outs (region, cell type, unit quality), a teleseismic transient, a distributed-acoustic-sensing (DAS) fiber phase arrival, and a (semi-synthetic, clearly labeled) gravitational-wave inspiral chirp — and returns GO on all seven. The result is then subjected to three adversarial self-audit rounds and a head-to-head architectural analysis, which force three explicit retractions and yield a corrected, defensible account: the token retains most of a decision at extreme compression (a modest, consistent tax of +0.05 to +0.08 AUC versus a strong nonlinear model, at roughly 200 microseconds and 21 kilobytes per window); it adds real discriminative value on shape/pattern tasks but reduces to a trivial detector on energy-dominated ones; and its behavior is governed by stream morphology — near-optimal on pulsatile and stationary signals, structurally weak on intermittent ones whose decision lives in cross-window timing a single token cannot see. Three claims are retracted under audit and reported plainly: an apparent \"beats-the-ceiling\" result was a weak-baseline artifact (a strong nonlinear ceiling restores the +0.05-0.08 tax); the tax-scaling \"law\" is not universal (it holds only within a modality's difficulty ladder); and a token trained on injected gravitational-wave signals does not transfer to real detected events. A tiered co-channel is shown to be a bandwidth device, not an accuracy device — it recovers tax only where the task is hard and routes no better by token uncertainty than at random, but delivers 8-61x bandwidth reduction at fixed event capture on continuous rare-event streams — and a learned trigger beats a trivial energy threshold only for shape-defined events. Lifecycle studies show an on-sensor codebook can self-maintain across many unsupervised refresh cycles (with periodic anchor refresh) and that spatial token-coincidence across an array suppresses false alarms. Honest boundaries are mapped verbatim: at-rest deep-brain medication state does not decode across patients (its uncompressed ceiling sits at chance — signal absence), cuffless blood-pressure category is largely subject-identity leakage, short-read nanopore falls to chance as the ceiling itself collapses, and label-shuffle controls collapse to chance (confirming the GOs are real signal). Method companions: Papers 19, 30, and 31; trigger-scoping companion: Paper 29. This is a cross-scale application and validation of previously-filed and previously-published methods. Keywords: class-discriminant codebook; vector quantization; on-device inference; edge AI; cross-scale sensing; nanopore sequencing; Neuropixels; distributed acoustic sensing; seismology; gravitational waves; pre-registration; honest negatives; selective co-channel; stream morphology; self-supervision References 1. R. J. Ferlic and K. K. Ferlic, \"A single-token class-discriminant codebook encoder for physiological signals (Paper 19),\" Zenodo, 10.5281/zenodo.20788187. 2. R. J. Ferlic and K. K. Ferlic, \"On-device glucose alarms from a single learned token (Paper 30),\" Zenodo, 10.5281/zenodo.21114273. 3. R. J. Ferlic and K. K. Ferlic, \"One token, six modalities: pre-registered cross-modality screening for wearable and implantable monitoring (Paper 31),\" Zenodo, 10.5281/zenodo.21136786. 4. M. Jain, H. E. Olsen, B. Paten, and M. Akeson, \"The Oxford Nanopore MinION: de","author":[{"family":"Ferlic","given":"Randolph"},{"family":"Ferlic","given":"Kimberly"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21149392","URL":"https://doi.org/10.5281/zenodo.21149392","source":"datacite"},{"id":"doi:10.48550/arxiv.2607.01272","type":"manuscript","title":"Benchmarking Federated Learning and Knowledge Distillation for Point Cloud Classification","abstract":"Deploying 3D point cloud analysis in privacy-sensitive, resource-constrained settings faces two barriers: data cannot be centralized, and models must run on limited edge hardware. We present a multi-seed benchmark jointly evaluating federated learning (FL) and knowledge distillation (KD) for 3D point cloud classification. It spans 13 FL algorithms and 10 KD objectives (a 130-pair cross-product) across 504 training runs, evaluated on ModelNet40 and a clinical craniosynostosis dataset. We report three findings. First, under extreme non-IID label skew, standalone FL degrades sharply: on ModelNet40, the strongest method reaches 76.32% against a 92.26% centralized reference; on clinical data, the best reaches 75.83% against 100%. Second, distillation successfully compresses the teacher into a student 74.51% smaller and roughly twice as fast at inference, often matching or surpassing the teacher. Third, the combined pipeline exposes an evaluation pitfall: when distillation keeps a hard-label cross-entropy term on a labeled proxy split, a collapsed federated teacher (8.50%) paired with Logit-MSE still yields a 92.94% student. This 84.4-point gap reflects the proxy labels rather than the federated model, reusing the very labels whose privacy motivated federation. Objectives without hard labels instead track teacher quality ($r \\approx 0.99$) and collapse when the teacher does. We therefore recommend evaluating FL-KD pipelines with label-free distillation so reported accuracy reflects the federated teacher, not the proxy.","author":[{"family":"Aiersilan","given":"Aizierjiang"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2607.01272","URL":"https://doi.org/10.48550/arxiv.2607.01272","source":"datacite"},{"id":"doi:10.5281/zenodo.20997942","type":"article-journal","title":"AI in Predictive Healthcare: How Artificial Intelligence and Wearable Devices are Transforming Early Disease Detection | A Systematic Literature Review (2026) | Ruchir Ganatra","abstract":"This research paper was prepared by Ruchir Ganatra as an independent research study exploring Artificial Intelligence in Predictive Healthcare. The paper provides a comprehensive review of AI-powered wearable devices, machine learning algorithms, digital health technologies, Internet of Medical Things (IoMT), and early disease detection using peer-reviewed scientific literature published between 2022 and 2026. The review compares leading AI techniques including CNN, LSTM, Random Forest, Federated Learning, and Explainable AI while discussing clinical applications, ethical considerations, privacy, algorithmic bias, and future research directions. This paper is intended to support researchers, students, healthcare professionals, and technology enthusiasts interested in the future of AI-driven healthcare.","author":[{"family":"Ganatra","given":"Ruchir"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20997942","URL":"https://doi.org/10.5281/zenodo.20997942","source":"datacite"},{"id":"doi:10.5281/zenodo.20997943","type":"article-journal","title":"AI in Predictive Healthcare: How Artificial Intelligence and Wearable Devices are Transforming Early Disease Detection | A Systematic Literature Review (2026) | Ruchir Ganatra","abstract":"This research paper was prepared by Ruchir Ganatra as an independent research study exploring Artificial Intelligence in Predictive Healthcare. The paper provides a comprehensive review of AI-powered wearable devices, machine learning algorithms, digital health technologies, Internet of Medical Things (IoMT), and early disease detection using peer-reviewed scientific literature published between 2022 and 2026. The review compares leading AI techniques including CNN, LSTM, Random Forest, Federated Learning, and Explainable AI while discussing clinical applications, ethical considerations, privacy, algorithmic bias, and future research directions. This paper is intended to support researchers, students, healthcare professionals, and technology enthusiasts interested in the future of AI-driven healthcare.","author":[{"family":"Ganatra","given":"Ruchir"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20997943","URL":"https://doi.org/10.5281/zenodo.20997943","source":"datacite"},{"id":"doi:10.60797/jbg.2026.32.8","type":"article-journal","title":"ОБЪЯСНИМЫЕ И БЕЗОПАСНЫЕ ПЛАТФОРМЫ ИСКУССТВЕННОГО ИНТЕЛЛЕКТА ДЛЯ ГЕНОМИКИ CRISPR: ДОСТИЖЕНИЯ, ОГРАНИЧЕНИЯ, СИСТЕМЫ МОДЕЛИРОВАНИЯ И ТРАНСЛЯЦИОННАЯ ИНТЕГРАЦИЯ НА ОСНОВЕ БЛОКЧЕЙНА","abstract":"Искусственный интеллект (ИИ) значительно ускорил процесс редактирования генома с помощью CRISPR-Cas за счет совершенствования проектирования направляющей РНК (gRNA), прогнозирования нецелевых эффектов, оценки результатов репарации ДНК и оптимизации метода прайм-редактирования. Однако современные вычислительные экосистемы CRISPR по-прежнему остаются фрагментированными: прогнозирующие модели ИИ, симуляторы репарации ДНК, платформы объяснимого ИИ (XAI) и системы управления геномными данными функционируют независимо друг от друга. Кроме того, растущее клиническое внедрение CRISPR-терапий, таких как Casgevy и Lyfgenia, усилило спрос на прозрачные, интерпретируемые, безопасные и пригодные для клинического применения системы ИИ. В данном обзоре критически анализируются последние достижения в области геномики CRISPR на основе ИИ, включая архитектуры глубокого обучения, системы прогнозирования на основе трансформеров, подходы объяснимого ИИ, платформы моделирования репарации ДНК, федеративное обучение, дифференциальную конфиденциальность и управление геномными данными с использованием блокчейна. В рукописи оцениваются основные ограничения существующих подходов, в том числе поведение прогнозирования по принципу «черного ящика», нестабильность объяснений XAI в последовательных геномных данных, плохая обобщаемость в различных биологических контекстах, ограниченная поддержка структурных геномных вариантов, а также ограничения масштабируемости блокчейна в рамках требований HIPAA и GDPR. Кроме того, в обзоре предлагается единая концептуальная архитектура, объединяющая механизмы прогнозирования на основе ИИ, симуляторы ремонта, модули объяснимости, федеративное обучение с сохранением конфиденциальности и гибридные инфраструктуры аудита на основе блокчейна с использованием хранения данных как в цепочке, так и вне цепочки. Также обсуждаются нормативные и практические аспекты, касающиеся FDA, EMA, стандарта ISO 13485, HIPAA, GDPR и Закона ЕС об ИИ. Предлагаемая архитектура призвана содействовать разработке безопасных, интерпретируемых и регулируемых с этической точки зрения систем поддержки принятия решений на основе CRISPR для прецизионной медицины следующего поколения.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.60797/jbg.2026.32.8","URL":"https://doi.org/10.60797/jbg.2026.32.8","source":"datacite"},{"id":"doi:10.60797/bmed.2026.9.8","type":"article-journal","title":"МУЛЬТИОМИЧЕСКИЙ ИИ ДЛЯ ПЕРСОНАЛИЗИРОВАННОЙ ИНТЕНСИВНОЙ ТЕРАПИИ: ОБЗОР ЦИФРОВЫХ ДВОЙНИКОВ, CRISPR И ОБЪЯСНИМЫХ СИСТЕМ","abstract":"Персонализированная медицина меняет ситуацию, переходя от универсальных методов лечения к стратегиям, учитывающим уникальные биологические особенности каждого человека. Этот сдвиг имеет огромное значение в интенсивной терапии, где пациенты находятся в уязвимом состоянии, их организм может быстро меняться, а риски являются высокими. Традиционная интенсивная терапия опирается на длительные лабораторные исследования, разрозненные медицинские записи и решения, основанные в основном на том, что врачи видят и знают в данный момент. Это приводит к пробелам в лечении и замедляет оказание медицинской помощи.Сегодня, благодаря прорывам в таких мультиомических науках, как геномика, протеомика, метаболомика и микробиомика, а также интеграции мощных инструментов — таких как искусственный интеллект, цифровые двойники, CRISPR, объяснимый ИИ и технологии защиты конфиденциальности — перед нами открывается новая эра. Эти инновации делают интенсивную терапию гораздо более «умной» и оперативной: врачи и системы могут прогнозировать, адаптироваться и принимать меры до того, как возникнут проблемы.В данном обзоре обобщены все эти достижения и рассмотрено, как их сочетание позволяет улучшить прогнозы, сделать лечение более эффективным и получить более четкое представление о ��ом, что приносит пациентам наибольшую пользу. Было рассмотрено, как модели «цифровых двойников» могут служить виртуальными копиями пациентов, как технология CRISPR открывает возможности для целевого лечения, а также то, как системы рекомендаций по лекарственным препаратам на базе искусственного интеллекта и объяснимый ИИ повышают уровень доверия и прозрачности при оказании медицинской помощи. Кроме того, в обзоре подробно рассматриваются такие технологии, как федеративное обучение и блокчейн, которые позволяют обмениваться знаниями без угрозы для конфиденциальности пациентов.Тем не менее остаются серьезные вопросы — пробелы в исследованиях, практические препятствия и этические соображения, требующие реальных ответов. В заключение обзора предлагается план по созданию практичных, понятных и готовых к внедрению в больницах концепций персонализированной медицины для интенсивной терапии.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.60797/bmed.2026.9.8","URL":"https://doi.org/10.60797/bmed.2026.9.8","source":"datacite"},{"id":"doi:10.48550/arxiv.2505.09854","type":"manuscript","title":"Chisme: Heterogeneity-Aware Gossip Learning","abstract":"As end-user device capability increases and demand for intelligent services at the Internet's edge rises, distributed learning has emerged as a key enabling technology for the intelligent edge. Existing approaches like federated learning (FL) and decentralized FL (DFL) enable privacy-preserving distributed learning among clients, while gossip learning (GL) approaches have emerged to address the potential challenges in resource-constrained, connectivity-challenged infrastructure-less environments. However, most distributed learning approaches assume largely homogeneous data distributions and may not consider or exploit the heterogeneity of clients and their underlying data distributions. This paper introduces Chisme, a novel fully decentralized distributed learning algorithm designed to address the challenges of implementing robust intelligence in network edge contexts characterized by heterogeneous data distributions, episodic connectivity, and sparse network infrastructure or lack thereof. Chisme leverages the affinity between clients' underlying data distributions calculated from received model exchanges to inform how much influence received models have when merging into the local model. By doing so, it enables clients to strategically balance between broader collaboration to build more general knowledge and more selective collaboration to build specific knowledge. We evaluate Chisme against contemporary approaches using image recognition and time-series prediction scenarios while considering different network connectivity conditions, representative of real-world distributed intelligent systems running at the network's edge. Our experiments demonstrate that Chisme outperforms state-of-the-art edge intelligence approaches in almost every case -- clients using Chisme exhibit faster training convergence, lower final loss after training, and lower performance disparity between clients.","author":[{"family":"Kuttivelil","given":"Harikrishna"},{"family":"Obraczka","given":"Katia"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2505.09854","URL":"https://doi.org/10.48550/arxiv.2505.09854","source":"datacite"},{"id":"doi:10.5281/zenodo.20813413","type":"article-journal","title":"Locked Out of the Intelligence Revolution: Nigeria's Upstream Data Gap, the Billion-Dollar AI Opportunity, and a Technical Framework for Industry Data Collaboration in the Niger Delta","abstract":"Nigeria holds 37.01 billion barrels of proven crude oil and condensate reserves, yet its production has consistently fallen short of its OPEC quota, and the government's target of 3 million barrels per day by 2030 assumes operational capacity the country has not yet demonstrated. This paper argues that a significant part of the gap is a data problem. Nigeria's upstream operational data is fragmented, inaccessible, and largely unpublished in usable form, conditions that prevent the application of artificial intelligence and machine learning tools that are already delivering double-digit efficiency gains for operators in other producing nations. Norway established its Diskos National Data Repository in 1995 and the United Kingdom followed with its own open data infrastructure, both underpinned by regulatory mandates that treated upstream data as shared national infrastructure rather than private operator property. Nigeria has no equivalent, and its national oil company was, as of April 2026, still digitising paper well logs dating to 1956. This paper proposes a three-tier data collaboration framework calibrated to Nigeria's current institutional reality rather than to an idealised regulatory environment. Tier 1 requires only a publishing decision by the Nigerian Upstream Petroleum Regulatory Commission to release already-public production data in structured, machine-readable formats. Tier 2 proposes federated learning as a mechanism for operators to benefit from collective intelligence without exposing commercially sensitive raw data. Tier 3 identifies full subsurface data sharing as a long-term destination requiring sustained regulatory investment.","author":[{"family":"Ezeagu","given":"Vera"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20813413","URL":"https://doi.org/10.5281/zenodo.20813413","source":"datacite"},{"id":"doi:10.5281/zenodo.20813414","type":"article-journal","title":"Locked Out of the Intelligence Revolution: Nigeria's Upstream Data Gap, the Billion-Dollar AI Opportunity, and a Technical Framework for Industry Data Collaboration in the Niger Delta","abstract":"Nigeria holds 37.01 billion barrels of proven crude oil and condensate reserves, yet its production has consistently fallen short of its OPEC quota, and the government's target of 3 million barrels per day by 2030 assumes operational capacity the country has not yet demonstrated. This paper argues that a significant part of the gap is a data problem. Nigeria's upstream operational data is fragmented, inaccessible, and largely unpublished in usable form, conditions that prevent the application of artificial intelligence and machine learning tools that are already delivering double-digit efficiency gains for operators in other producing nations. Norway established its Diskos National Data Repository in 1995 and the United Kingdom followed with its own open data infrastructure, both underpinned by regulatory mandates that treated upstream data as shared national infrastructure rather than private operator property. Nigeria has no equivalent, and its national oil company was, as of April 2026, still digitising paper well logs dating to 1956. This paper proposes a three-tier data collaboration framework calibrated to Nigeria's current institutional reality rather than to an idealised regulatory environment. Tier 1 requires only a publishing decision by the Nigerian Upstream Petroleum Regulatory Commission to release already-public production data in structured, machine-readable formats. Tier 2 proposes federated learning as a mechanism for operators to benefit from collective intelligence without exposing commercially sensitive raw data. Tier 3 identifies full subsurface data sharing as a long-term destination requiring sustained regulatory investment.","author":[{"family":"Ezeagu","given":"Vera"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20813414","URL":"https://doi.org/10.5281/zenodo.20813414","source":"datacite"},{"id":"doi:10.5281/zenodo.20807407","type":"article-journal","title":"Locked Out of the Intelligence Revolution: Nigeria's Upstream Data Gap, the Billion-Dollar AI Opportunity, and a Technical Framework for Industry Data Collaboration in the Niger Delta","abstract":"Nigeria holds 37.01 billion barrels of proven crude oil and condensate reserves, yet its production has consistently fallen short of its OPEC quota, and the government's target of 3 million barrels per day by 2030 assumes operational capacity the country has not yet demonstrated. This paper argues that a significant part of the gap is a data problem. Nigeria's upstream operational data is fragmented, inaccessible, and largely unpublished in usable form, conditions that prevent the application of artificial intelligence and machine learning tools that are already delivering double-digit efficiency gains for operators in other producing nations. Norway established its Diskos National Data Repository in 1995 and the United Kingdom followed with its own open data infrastructure, both underpinned by regulatory mandates that treated upstream data as shared national infrastructure rather than private operator property. Nigeria has no equivalent, and its national oil company was, as of April 2026, still digitising paper well logs dating to 1956. This paper proposes a three-tier data collaboration framework calibrated to Nigeria's current institutional reality rather than to an idealised regulatory environment. Tier 1 requires only a publishing decision by the Nigerian Upstream Petroleum Regulatory Commission to release already-public production data in structured, machine-readable formats. Tier 2 proposes federated learning as a mechanism for operators to benefit from collective intelligence without exposing commercially sensitive raw data. Tier 3 identifies full subsurface data sharing as a long-term destination requiring sustained regulatory investment.","author":[{"family":"Ezeagu","given":"Vera"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20807407","URL":"https://doi.org/10.5281/zenodo.20807407","source":"datacite"},{"id":"doi:10.48550/arxiv.2603.13293","type":"manuscript","title":"A Robust Framework for Secure Cardiovascular Risk Prediction: An Architectural Case Study of Differentially Private Federated Learning","abstract":"Accurate cardiovascular risk prediction is crucial for preventive healthcare; however, the development of robust Artificial Intelligence (AI) models is hindered by the fragmentation of clinical data across institutions due to stringent privacy regulations. This paper presents a comprehensive architectural case study validating the engineering robustness of FedCVR, a privacy-preserving Federated Learning framework applied to heterogeneous clinical networks. Rather than proposing a new theoretical optimizer, this work focuses on a systems engineering analysis to quantify the operational trade-offs of server-side adaptive optimization under utility-prioritized Differential Privacy (DP). By conducting a rigorous stress test in a high-fidelity synthetic environment that reflects the feature space and clinical context of real-world datasets (Framingham, Cleveland), we systematically evaluate the system's resilience to statistical noise. The validation results demonstrate that integrating server-side momentum as a temporal denoiser enables the architecture to achieve a stable F1 score of 0.78 and an Area Under the Curve (AUC) of 0.96 under the operational privacy budget (epsilon approximately 13.4), compared to a non-private baseline with an F1 score of 0.84. FedCVR statistically outperforms standard stateless baselines (FedAvg, FedProx) and other adaptive optimizers (FedAdagrad, FedYogi) under identical privacy constraints. Our findings confirm that server-side adaptivity is a structural prerequisite for recovering clinical utility under realistic privacy budgets, providing a validated engineering blueprint for secure multi-institutional collaboration.","author":[{"family":"Tertulino","given":"Rodrigo"},{"family":"Alencar","given":"Laércio"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.13293","URL":"https://doi.org/10.48550/arxiv.2603.13293","source":"datacite"},{"id":"doi:10.5281/zenodo.20680859","type":"article-journal","title":"The Latent Yield: A Formal Framework for Low-Inhibition Cognitive Workspace Architecture, Privacy-Preserving Multi-Modal Feature Projection, and the Mapping of Human Latent Capacity in Decentralized, Peer-to-Peer Sovereign Data Nodes","abstract":"This paper presents a theoretical framework and enabling specification for a low-inhibition associative cognitive workspace, designated the Dream Galaxy, implemented within decentralized, peer-to-peer sovereign data nodes constituting a federated personal cognitive universe. The framework introduces four interlocking technical contributions. First, a privacy-preserving multi-modal feature projection layer that maps heterogeneous episodic and physiological inputs, including sub-threshold physiological tokenization of continuous EEG, EMG, and EOG streams, into a unified semantic embedding space without cross-node data exposure, using horizontal federated learning with differential privacy at the gradient layer. Second, an asynchronous local topological optimization engine that applies incremental GPU-accelerated persistent homology over a Vietoris-Rips filtration to the projected feature space, maintaining a continuously valid persistence barcode without centralized computation or synchronous coordination across nodes. Third, a four-zone epistemological classifier that assigns each projected feature node a zone-resident probability distribution across Terra Firma (Zone 1), Incognita (Zone 2), Hic Sunt Dracones (Zone 3), and Whisp (Zone 4), where Zone 4 nodes represent pre-cognitive unknown unknowns validated against a local density criterion distinguishing genuine structural absence from data sparsity. Fourth, the Latent Yield metric: a dimension-weighted rate of topological feature birth in Zones 3 and 4, formalized with explicit pseudocode in the Technical Appendix, quantifying cognitive generative activity below the threshold of conscious expression. The paper further specifies an Asynchronous Convergence Protocol by which topologically equivalent planet structures in independent sovereign nodes are detected under differential privacy without raw data exchange, constituting a knowledge event of elevated epistemic weight requiring bilateral consent before any bridge formation. As a Perspectives and Theoretical Frameworks contribution, this paper provides sufficient enabling detail, including algorithmic pseudocode, parameter specifications, and system architecture, for a person having ordinary skill in machine learning or computational neuroscience to implement the described system. The framework is protected by ten United States provisional patent applications (Nos. 64/074,530; 64/074,608; 64/074,779; 64/075,396; 64/079,009; 64/081,935; 64/084,355; 64/086,036; 64/089,881; 64/089,883), the last two filed June 13, 2026.","author":[{"family":"Hoppe","given":"Eric"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20680859","URL":"https://doi.org/10.5281/zenodo.20680859","source":"datacite"},{"id":"doi:10.5281/zenodo.20649676","type":"article-journal","title":"CDSA-BB3: Pilot Biometric Pillar of the CDSA Federated Reinforcement Learning Paradigm — Synthetic Biometric Data, Anonymization, Cryptographic Multi-Signature, and Scaffolded FRL Extension","abstract":"CDSA-BB3 is the reference implementation of the pilot biometric pillar of the Complementary Diagnostic Safety Approach (CDSA), a third-generation aviation safety paradigm formalised as a Federated Reinforcement Learning (FRL) framework (Scenario C decision, 21 May 2026). The v2.1.0 release (12 June 2026) replaces the v2.0.0 scaffold stubs with full NumPy reference implementations (federated, multimodal, rl_agents, rl_environment, xai; all unit-tested) and adds seeded, reproducible CPU-scale federated training runs (random / centralised PPO / FedAvg / FedProx / FedProx+DP) with result JSONs under experiments/results/. An interactive results panel is published at https://cdsa.app/bb3. Framework-scale training with real data is planned under TÜBİTAK funding (Phase B). The v1.0.0 release modules and reference data remain unchanged.","author":[{"family":"Cantekin","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20649676","URL":"https://doi.org/10.5281/zenodo.20649676","source":"datacite"},{"id":"doi:10.5281/zenodo.20649672","type":"article-journal","title":"CDSA-MRO: Maintenance Safety Pillar of the CDSA Federated Reinforcement Learning Paradigm — Multi-modal Engine + Cyber + Maintenance Fusion","abstract":"CDSA-MRO is the reference implementation of the maintenance safety pillar of the Complementary Diagnostic Safety Approach (CDSA), a third-generation aviation safety paradigm formalised as a Federated Reinforcement Learning (FRL) framework (Scenario C decision, 21 May 2026). The v2.1.0 release (12 June 2026) replaces the v2.0.0 scaffold stubs with full NumPy reference implementations (federated, multimodal, rl_agents, rl_environment, xai; all unit-tested) and adds seeded, reproducible CPU-scale federated training runs (random / centralised PPO / FedAvg / FedProx / FedProx+DP) with result JSONs under experiments/results/. An interactive results panel is published at https://cdsa.app/mro. Framework-scale training with real data is planned under TÜBİTAK funding (Phase B). The v1.0.0 release modules and reference data remain unchanged.","author":[{"family":"Cantekin","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20649672","URL":"https://doi.org/10.5281/zenodo.20649672","source":"datacite"},{"id":"doi:10.5281/zenodo.20649857","type":"article-journal","title":"CDSA-ATM: Air Traffic Management Pillar of the CDSA Federated Reinforcement Learning Paradigm — Multi-modal Trajectory + ATS + ANSP Fusion","abstract":"CDSA-ATM is the reference implementation of the air traffic management pillar of the Complementary Diagnostic Safety Approach (CDSA), a third-generation aviation safety paradigm formalised as a Federated Reinforcement Learning (FRL) framework following the CDSA Methodological Unification Decision (Scenario C, 21 May 2026, frozen status). The v2 scaffold (22 May 2026) adds Python modules (rl_agents, rl_environment, multimodal, federated, xai) for the FRL paradigm extension. Full implementations are deferred to TÜBİTAK 1001 ARDEB Project 4 (Sept 2026 application cycle; 2027-2030 execution). The v1.0.0 release modules and reference data remain unchanged.","author":[{"family":"Cantekin","given":"Mete"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20649857","URL":"https://doi.org/10.5281/zenodo.20649857","source":"datacite"},{"id":"doi:10.5281/zenodo.18654576","type":"article-journal","title":"Murmura: A Framework for Federated and Decentralized Machine Learning","abstract":"Murmura is a comprehensive framework for federated and decentralized machine learning. Built for researchers and developers, it provides tools for distributed machine learning simulation with advanced privacy guarantees and flexible network topologies. The framework supports both centralized federated learning and fully decentralized peer-to-peer learning environments, with features including multiple network topologies, Byzantine-robust aggregation strategies, comprehensive differential privacy support, and intelligent resource management.If you use this repository in your work, please cite the following: @INPROCEEDINGS{rangwala2026murmura, author={Rangwala, Murtaza and Sinnott, Richard O and Buyya, Rajkumar}, booktitle={2026 IEEE 26th International Symposium on Cluster, Cloud and Internet Computing (CCGrid)}, title={Evidential Trust-Aware Model Personalization in Decentralized Federated Learning for Wearable IoT}, year={2026}, publisher = {IEEE Press}, address = {New Jersey, USA}, pages = {510-519}, doi={10.1109/CCGrid68966.2026.00061}}","author":[{"family":"Rangwala","given":"Murtaza"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18654576","URL":"https://doi.org/10.5281/zenodo.18654576","source":"datacite"},{"id":"doi:10.5281/zenodo.15622123","type":"article-journal","title":"Murmura: A Framework for Federated and Decentralized Machine Learning","abstract":"Murmura is a comprehensive framework for federated and decentralized machine learning. Built for researchers and developers, it provides tools for distributed machine learning simulation with advanced privacy guarantees and flexible network topologies. The framework supports both centralized federated learning and fully decentralized peer-to-peer learning environments, with features including multiple network topologies, Byzantine-robust aggregation strategies, comprehensive differential privacy support, and intelligent resource management.If you use this repository in your work, please cite the following: @INPROCEEDINGS{rangwala2026murmura, author={Rangwala, Murtaza and Sinnott, Richard O and Buyya, Rajkumar}, booktitle={2026 IEEE 26th International Symposium on Cluster, Cloud and Internet Computing (CCGrid)}, title={Evidential Trust-Aware Model Personalization in Decentralized Federated Learning for Wearable IoT}, year={2026}, publisher = {IEEE Press}, address = {New Jersey, USA}, pages = {510-519}, doi={10.1109/CCGrid68966.2026.00061}}","author":[{"family":"Rangwala","given":"Murtaza"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.15622123","URL":"https://doi.org/10.5281/zenodo.15622123","source":"datacite"},{"id":"doi:10.5281/zenodo.20533650","type":"article-journal","title":"The Human Sovereignty Technology Paradigm — Supplement Version 2.0: Extended Prior Art Declaration","abstract":"Supplement Version 2.0 to the Human Sovereignty Technology Paradigm white paper (Version 1.0, March 29, 2026 �� zenodo.org/records/19323816). Establishes irrevocable public domain prior art for the Human Sovereignty Technology Platform (USPTO Provisional Application No. 64/020,809, filed March 29, 2026). Extends prior art coverage to: no-biometric embodiments of the Operator-Bound Technology Principle; third-party wearable device compatibility including commercially available smartwatches and fitness rings; CEEP capacitive electrode architecture for data destruction; modular replacement architecture with cryptographic module authentication; AI federated learning threat assessment protocol; multi-mode non-biometric trigger subsystem; passive safe-discharge adapter; and extended application embodiments covering financial systems, weapon authentication, child protection, medical systems, journalism, and legal profession source protection. All concepts herein are irrevocably dedicated to the public domain under CC0 1.0 Universal as of June 3, 2026. Constitutes published prior art under 35 U.S.C. § 102(a)(1). Cannot be retroactively classified after this publication date.","author":[{"family":"Kershner","given":"Braydon"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20533650","URL":"https://doi.org/10.5281/zenodo.20533650","source":"datacite"},{"id":"doi:10.5281/zenodo.20533651","type":"article-journal","title":"The Human Sovereignty Technology Paradigm — Supplement Version 2.0: Extended Prior Art Declaration","abstract":"Supplement Version 2.0 to the Human Sovereignty Technology Paradigm white paper (Version 1.0, March 29, 2026 — zenodo.org/records/19323816). Establishes irrevocable public domain prior art for the Human Sovereignty Technology Platform (USPTO Provisional Application No. 64/020,809, filed March 29, 2026). Extends prior art coverage to: no-biometric embodiments of the Operator-Bound Technology Principle; third-party wearable device compatibility including commercially available smartwatches and fitness rings; CEEP capacitive electrode architecture for data destruction; modular replacement architecture with cryptographic module authentication; AI federated learning threat assessment protocol; multi-mode non-biometric trigger subsystem; passive safe-discharge adapter; and extended application embodiments covering financial systems, weapon authentication, child protection, medical systems, journalism, and legal profession source protection. All concepts herein are irrevocably dedicated to the public domain under CC0 1.0 Universal as of June 3, 2026. Constitutes published prior art under 35 U.S.C. § 102(a)(1). Cannot be retroactively classified after this publication date.","author":[{"family":"Kershner","given":"Braydon"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20533651","URL":"https://doi.org/10.5281/zenodo.20533651","source":"datacite"},{"id":"doi:10.5281/zenodo.19712705","type":"article-journal","title":"Blockchain Solution with Artificial Intelligence Integration in the Indian Judicial System A Bibliometric and Methodical Literature Review","abstract":"The Indian judicial system faces critical challenges including case backlogs exceeding 54.7 million pending matters, insufficient judicial resources, and inefficient evidence management processes. This literature review examines the emerging potential of integrating blockchain technology with artificial intelligence (AI) to transform judicial delivery, enhance case processing efficiency, and strengthen evidentiary integrity. Through systematic analysis of 88 peer-reviewed publications (2013–2026) across IEEE Xplore, Scopus, Springer, and Web of Science, we identified four dominant research themes: blockchain-based evidence management systems, AI-driven predictive justice and decision support, smart contracts for judicial automation, and privacy-preserving mechanisms for sensitive legal data. Key findings reveal that blockchain ensures significant reduction in evidence tampering incidents, while AI prediction models achieve notable accuracy in judicial outcome forecasting. However, significant implementation challenges persist, including scalability constraints, lack of comprehensive regulatory frameworks, and insufficient integration with legacy court systems. This paper synthesizes current scholarship, identifies critical research gaps, and proposes a four-layer conceptual framework for pragmatic AI-blockchain deployment suited to India's constitutional and legal context. We conclude that strategic integration prioritizing permissioned blockchain architectures, explainable AI models, and federated learning offers transformative potential for addressing judicial inefficiency while maintaining due process and fundamental rights protections.","author":[{"family":"Patil","given":"Sameer"},{"family":"Desai","given":"Darshana"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19712705","URL":"https://doi.org/10.5281/zenodo.19712705","source":"datacite"},{"id":"doi:10.5281/zenodo.20500935","type":"article-journal","title":"Pre-Registered Multi-Dataset Validation of Three-Way Composition III Privacy-Preserving Architecture for Musculoskeletal-Kinematic Clinical Digital Twins","abstract":"v2 update (2026-06-01): Reproducibility ZIP added back to the latest version alongside the manuscript files, so that downloading from the concept DOI gives all materials in one place rather than requiring navigation to v1. This reproducibility archive accompanies the Paper 8 manuscript \"Pre-Registered Multi-Dataset Validation of Three-Way Composition III Privacy-Preserving Architecture for Musculoskeletal-Kinematic Clinical Digital Twins.\" Background. Clinical digital twin (DT) systems for musculoskeletal rehabilitation increasingly aggregate longitudinal kinematic data across multiple sites and patient populations. These deployments require simultaneous guarantees of (a) federated coefficient learning utility, (b) membership-inference privacy against realistic adversaries, (c) byte-budgeted transmission for Internet-of-Things radio links, and (d) deployment-time robustness to anatomical, populational, and protocol heterogeneity. Methods. We present a pre-registered empirical validation campaign spanning 21 studies and 91 hypotheses for the three-way Composition III architecture (per-subject phase randomization, cohort mean aggregation, federated AR(1) coefficient learning) applied to musculoskeletal-kinematic clinical DT deployments. Each study's decision rules were frozen before runner execution. The campaign covers foundational properties; Composition III privacy and utility under homogeneous, moderately heterogeneous, and extremely heterogeneous (50:1 imbalance, 50x sensor noise differential) deployments; multi-classifier shadow-attacker bounds across LogisticRegression, RandomForest, and GradientBoosting families; differential-privacy noise calibration with non-monotone trade-off characterization; multi-period longitudinal observation up to T=10 periods; cross-anatomy generalization from N=3 (knee) to N=23 (hand-MANO model); and three real-world dataset validations using CMU Motion Capture Database (walking) and KIMORE Rehabilitation Dataset (rehabilitation exercises with 44 healthy controls plus 34 subjects with low-back pain, Parkinson's disease, and post-stroke conditions). Results. 71 of 91 pre-registered hypotheses supported (78%). All 20 non-supported outcomes are substantively interpretable: 3 metric-choice issues resolved by follow-up studies, 4 deployment-condition-dependent disclosures (including a non-monotone DP privacy-utility trade-off, an anatomy-specific sigma_DP scaling law, and a heterogeneity x longitudinal compound-leakage interaction), and 13 honest bounded-negatives identifying empirical boundaries of the recommended deployment configuration. Three findings stand out: (i) the realistic-attacker bound generalizes within +/- 0.06 across synthetic and two real datasets (maximum observed 0.60 against multi-classifier shadow attackers); (ii) patient-vs-healthy subgroup analysis on KIMORE yields identical oracle accuracy across populations (delta = 0.0000); (iii) anatomy-specific DP calibration scaling sigma_DP proportional to 1/sqrt(N) validated for N >= 5 with explicit knee N=3 outlier disclosure. Conclusions. Composition III is a viable privacy-preserving federated learning architecture for musculoskeletal-kinematic clinical digital twins. Empirical validation across synthetic data, two real datasets, multiple anatomies, three classifier families, three observation horizons, and mixed healthy/patient populations provides comprehensive grounding for clinical deployment. Archive contents. 21 frozen pre-registrations, 21 deterministic Python runners, 21 study reports, 21 verdict summary JSON files, 21 raw CSV data files, source code for the spiral-domain encoder and CMU MoCap and KIMORE parsers, 6 manuscript figures with generators, and the manuscript itself (Markdown and Word formats). Raw third-party data files are not redistributed per their respective licensing terms; download instructions and subject IDs are documented. Companion datasets. CMU Motion Capture Database (CC-BY 3.0, http://mocap.cs.cmu.ed","author":[{"family":"Ferlic","given":"Randolph"},{"family":"Ferlic","given":"Kimberly"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20500935","URL":"https://doi.org/10.5281/zenodo.20500935","source":"datacite"},{"id":"doi:10.5281/zenodo.20236848","type":"article-journal","title":"From Perfect Isolation to Governed Confidential Computing: A System-Level Perspective on Confidentiality","abstract":"This paper introduces a conceptual reasoning framework intended to clarify the security properties provided by strongly isolated execution environments and governed execution perimeters. Starting from an idealized model of perfect isolation, we progressively relax the assumptions in order to understand which confidentiality properties remain preserved when controlled interactions with users are introduced. The paper argues that, under sufficiently strong isolation assumptions, the confidentiality properties obtained become conceptually close to those targeted by Fully Homomorphic Encryption (FHE): the infrastructure executing the computation cannot access the manipulated data during computation itself. Although the mechanisms fundamentally differ — cryptographic transformation in the case of FHE versus impossibility of observation in the case of isolation — both approaches ultimately rely on critical assumptions outside the computation phase itself. We show that, when considering complete operational systems rather than isolated theoretical models, the effective security properties of the two approaches become structurally closer than is often assumed. This analysis provides a conceptual foundation for governed confidential computing architectures such as Trusted Cloud Enclaves and programmable trust-anchor-based infrastructures.","author":[{"family":"Bolignano","given":"Dominique"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20236848","URL":"https://doi.org/10.5281/zenodo.20236848","source":"datacite"},{"id":"doi:10.5281/zenodo.20236847","type":"article-journal","title":"From Perfect Isolation to Governed Confidential Computing: A System-Level Perspective on Confidentiality","abstract":"This paper introduces a conceptual reasoning framework intended to clarify the security properties provided by strongly isolated execution environments and governed execution perimeters. Starting from an idealized model of perfect isolation, we progressively relax the assumptions in order to understand which confidentiality properties remain preserved when controlled interactions with users are introduced. The paper argues that, under sufficiently strong isolation assumptions, the confidentiality properties obtained become conceptually close to those targeted by Fully Homomorphic Encryption (FHE): the infrastructure executing the computation cannot access the manipulated data during computation itself. Although the mechanisms fundamentally differ — cryptographic transformation in the case of FHE versus impossibility of observation in the case of isolation — both approaches ultimately rely on critical assumptions outside the computation phase itself. We show that, when considering complete operational systems rather than isolated theoretical models, the effective security properties of the two approaches become structurally closer than is often assumed. This analysis provides a conceptual foundation for governed confidential computing architectures such as Trusted Cloud Enclaves and programmable trust-anchor-based infrastructures.","author":[{"family":"Bolignano","given":"Dominique"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20236847","URL":"https://doi.org/10.5281/zenodo.20236847","source":"datacite"},{"id":"doi:10.5281/zenodo.20273116","type":"article-journal","title":"From Perfect Isolation to Governed Confidential Computing: A System-Level Perspective on Confidentiality","abstract":"This paper introduces a conceptual reasoning framework intended to clarify the security properties provided by strongly isolated execution environments and governed execution perimeters. Starting from an idealized model of perfect isolation, we progressively relax the assumptions in order to understand which confidentiality properties remain preserved when controlled interactions with users are introduced. The paper argues that, under sufficiently strong isolation assumptions, the confidentiality properties obtained become conceptually close to those targeted by Fully Homomorphic Encryption (FHE): the infrastructure executing the computation cannot access the manipulated data during computation itself. Although the mechanisms fundamentally differ — cryptographic transformation in the case of FHE versus impossibility of observation in the case of isolation — both approaches ultimately rely on critical assumptions outside the computation phase itself. We show that, when considering complete operational systems rather than isolated theoretical models, the effective security properties of the two approaches become structurally closer than is often assumed. This analysis provides a conceptual foundation for governed confidential computing architectures such as Trusted Cloud Enclaves and programmable trust-anchor-based infrastructures.","author":[{"family":"Bolignano","given":"Dominique"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20273116","URL":"https://doi.org/10.5281/zenodo.20273116","source":"datacite"},{"id":"doi:10.5281/zenodo.20252656","type":"article-journal","title":"Artificial Intelligence and Machine Learning- Enabled Security Framework for Data Privacy in Cloud Computing","abstract":"The way in which Cloud computing has changed data storage and service delivery has also posed serious threats to privacy and security of data. The conventional encryption and rule-based intrusion detection systems can still not provide reasonable protection against sophisticated attacks like unauthorized intrusion, inference attack and anomalous activity in shared environments. The present paper suggests a machine learning-enhanced artificial intelligence (ML/AI-enabled) security approach that aims to improve data privacy in cloud computing. The framework includes both classical ML models trained with known and unknown label data (Random Forest, Isolation Forest, k-NN), and uses deep learning models (Autoencoders, LSTM) as well as privacy-preserving techniques such as differential privacy (DP), homomorphic encryption (HE) and attribute-based access control (ABAC). Experiments were also designed to evaluate the proposed system using malicious and benign access logs collected on OpenStack and AWS testbeds in terms of both detection performance (precision, recall, F1-score, AUC) as well as system efficiency (latency and computational overhead). It was found LSTM Autoencoders performed better when detecting (F1-score 0.95, AUC 0.96), although privacy mechanisms still resulted in a reduction in performance (<5-percent change). The very low computational overhead imposed by HE and DP was acceptable in real-time operations. The work illustrates that integration of AI/ML models with privacy-preserving solutions can offer a scalable, robust and realistic model of cloud data protection.","author":[{"family":"Ubaid","given":"Eman"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20252656","URL":"https://doi.org/10.5281/zenodo.20252656","source":"datacite"},{"id":"doi:10.5281/zenodo.20252657","type":"article-journal","title":"Artificial Intelligence and Machine Learning- Enabled Security Framework for Data Privacy in Cloud Computing","abstract":"The way in which Cloud computing has changed data storage and service delivery has also posed serious threats to privacy and security of data. The conventional encryption and rule-based intrusion detection systems can still not provide reasonable protection against sophisticated attacks like unauthorized intrusion, inference attack and anomalous activity in shared environments. The present paper suggests a machine learning-enhanced artificial intelligence (ML/AI-enabled) security approach that aims to improve data privacy in cloud computing. The framework includes both classical ML models trained with known and unknown label data (Random Forest, Isolation Forest, k-NN), and uses deep learning models (Autoencoders, LSTM) as well as privacy-preserving techniques such as differential privacy (DP), homomorphic encryption (HE) and attribute-based access control (ABAC). Experiments were also designed to evaluate the proposed system using malicious and benign access logs collected on OpenStack and AWS testbeds in terms of both detection performance (precision, recall, F1-score, AUC) as well as system efficiency (latency and computational overhead). It was found LSTM Autoencoders performed better when detecting (F1-score 0.95, AUC 0.96), although privacy mechanisms still resulted in a reduction in performance (<5-percent change). The very low computational overhead imposed by HE and DP was acceptable in real-time operations. The work illustrates that integration of AI/ML models with privacy-preserving solutions can offer a scalable, robust and realistic model of cloud data protection.","author":[{"family":"Ubaid","given":"Eman"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20252657","URL":"https://doi.org/10.5281/zenodo.20252657","source":"datacite"},{"id":"doi:10.5281/zenodo.21421425","type":"article-journal","title":"The Blind Machine: An End-to-End Platform for Privacy-Preserving Genomic Computation with Homomorphic Encryption","abstract":"Many scientific questions require computing statistics across datasets that cannot be pooled because the underlying records contain sensitive information that cannot be shared. Homomorphic encryption enables computation directly on encrypted data without exposing plaintext, and more than a decade of research has demonstrated its practicality for privacy-preserving genomic analysis. Yet these systems are typically built one analysis at a time; no reusable, governed execution model has emerged that certifies and reuses approved encrypted computations independently of the specific study. To close this gap, we present The Blind Machine, an end-to-end platform for governed computation on encrypted data using homomorphic encryption and federated computing. The core design principle is that plaintext and secret keys never leave local machines, and the hosted service computes exclusively on ciphertext. Every experiment produces a machine-verifiable certificate, checkable offline, that binds the application, the committed encrypted inputs, the encrypted result, and the declared release policy. We show that the system is correct and practical through six curated biomedicalapplications built with the BFV scheme on seeded synthetic data, together with four studies on public IGSR/1000 Genomes genotypes. The work is fully reproducible: we release the open-source components, public-genome installer, synthetic data, experiment scripts, and an AI agent skill that helps reviewers reproduce the experiments.","author":[{"family":"Özmen","given":"Barış"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21421425","URL":"https://doi.org/10.5281/zenodo.21421425","source":"datacite"},{"id":"doi:10.5281/zenodo.21448699","type":"article-journal","title":"The Blind Machine: An End-to-End Platform for Privacy-Preserving Genomic Computation with Homomorphic Encryption","abstract":"Many scientific questions require computing statistics across datasets that cannot be pooled because the underlying records contain sensitive information that cannot be shared. Homomorphic encryption enables computation directly on encrypted data without exposing plaintext, and more than a decade of research has demonstrated its practicality for privacy-preserving genomic analysis. Yet these systems are typically built one analysis at a time; no reusable, governed execution model has emerged that certifies and reuses approved encrypted computations independently of the specific study. To close this gap, we present The Blind Machine, an end-to-end platform for governed computation on encrypted data using homomorphic encryption and federated computing. The core design principle is that plaintext and secret keys never leave local machines, and the hosted service computes exclusively on ciphertext. Every experiment produces a machine-verifiable certificate, checkable offline, that binds the application, the committed encrypted inputs, the encrypted result, and the declared release policy. We show that the system is correct and practical through six curated biomedicalapplications built with the BFV scheme on seeded synthetic data, together with four studies on public IGSR/1000 Genomes genotypes. The work is fully reproducible: we release the open-source components, public-genome installer, synthetic data, experiment scripts, and an AI agent skill that helps reviewers reproduce the experiments.","author":[{"family":"Özmen","given":"Barış"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21448699","URL":"https://doi.org/10.5281/zenodo.21448699","source":"datacite"},{"id":"doi:10.5281/zenodo.20288859","type":"article-journal","title":"TEMPORAL ROTATION SECURITY PROTOCOL (TRSP) Physics-First Cryptographic Architecture: Time as the Fundamental Security Parameter","abstract":"ABSTRACT: Temporal Rotation Security Protocol (TRSP) v3 — Physics-First Cryptographic Architecture: Time as the Fundamental Quantum-Resistant Security Parameter Concept in development since at least November 2019. First complete public documentation: 2026. Version 3 adds Part 4c (Hybrid Dynamic CRATON Quorum / HDCQ) and Part 4d (CRATON Hardening Layer / CHL). This concept documents the Temporal Rotation Security Protocol — a cryptographic architecture in which security derives not from the mathematical complexity of encryption keys but from the physical irreversibility of time. Every cryptographic system currently in use — RSA, AES, elliptic curve — rests on a single foundational assumption: that breaking the encryption requires more computational time than any adversary possesses. Quantum computing, through Shor's algorithm and Grover's algorithm, is systematically dismantling this assumption. TRSP replaces it with a physically permanent alternative: a key that no longer exists cannot be recovered by any computation, quantum or classical, regardless of computational resources or future mathematical advances. The protocol operates through simultaneous multi-layer key rotation at three independent frequencies. Layer 1 (session layer) rotates every 10–100 milliseconds using hardware entropy from physical noise sources — thermal variance, clock jitter, electromagnetic fingerprint. Layer 2 (identity layer) rotates every 1–10 seconds, anchored to physically unique device characteristics that cannot be spoofed. Layer 3 (CRATON foundational layer) generates a cryptographic commitment from the unique physical state of both communicating devices at session initialisation — used once and permanently destroyed, unrepeatable at any other point in time or on any other device. Quantum resistance is structural rather than parametric. Shor's algorithm requires minutes to hours to factor key-scale integers; Layer 1 rotation windows of 10–100 milliseconds ensure the target key no longer exists when any quantum computation converges. Grover's algorithm provides quadratic speedup against static keys; against rotating keys it provides no advantage because the search target is destroyed before the search completes. As quantum hardware advances and computation accelerates, rotation windows decrease proportionally — a software parameter adjustment costing microseconds against a hardware investment requiring years. The defender's adaptation is permanently faster than the attacker's. The CRATON foundational trust layer — positioned between hardware and operating system — is not a stored value. It is a physical event: a one-time measurement of device state that generates a cryptographic commitment and is immediately destroyed. It cannot be forged by a compromised operating system, replicated on any other device, or reconstructed from any stored record. Root of trust through physical irreversibility. The inverse proposition — what breaking TRSP would prove — is documented as the second foundational contribution of this concept. A successful attack against a correctly implemented TRSP system would constitute experimental proof of one of the following physical propositions: that quantum information is globally conserved and locally accessible confirming the holographic principle; that temporal irreversibility is not absolute at quantum scale; that parallel quantum branches are accessible through computation confirming the Everett many-worlds interpretation; or that Landauer's principle is violated at computational scale. Any of these would represent the most significant scientific discovery in recorded history. TRSP is therefore simultaneously a security protocol and a physics experiment. Its security parameter is the boundary of known physical law. TRSP is an open invitation to physics. Extension 1: Spatial-Temporal Triangulation & Network Latency Mitigation A critical challenge in millisecond-scale cryptographic rotation (Δt = 10–100 ms) across standard ","author":[{"family":"Mehmetaj","given":"Ilir"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20288859","URL":"https://doi.org/10.5281/zenodo.20288859","source":"datacite"},{"id":"doi:10.1038/s41598-025-19712-1","type":"article-journal","title":"A federated edge intelligence framework with trust based access control for secure and privacy preserving IoT systems.","abstract":"The rapid growth of Internet of Things (IoT) ecosystems has generated substantial industrial progress, yet it has also introduced intricate security and privacy issues. IoT deployments cannot be properly supported with traditional cloud-centric approaches because they require improved bandwidth utilization, reduced latency, and enhanced trust mechanisms. The research proposes Artificial Intelligence-Driven Secure Edge Trust Framework (AI-SET), which establishes a comprehensive edge-based security design that connects network intrusion detection with federated learning capabilities to implement adaptive trust-based access control for IoT system protection. The AI-SET framework comprises three central elements. Real-time anomaly detection at the network edge through the Edge-Resident Intrusion Detection System operates with lightweight AI algorithms to minimize dependency on centralized systems. Privacy-preserving federated learning utilizes the modified FedAvg algorithm, which is supported by differential privacy and homomorphic encryption. Security measures enabled by this model allow algorithms to be trained across decentralized sources that contain heterogeneous and non-identically distributed (non-IID) data. A dynamic access control system utilizes trust assessment models to evaluate device context and behavior for real-time permission evaluations. The framework undergoes validation by running tests with the NAB dataset, supported by Jetson Nano and Raspberry Pi edge devices, and tools including Suricata, Metasploit, and the WAZUH threat platform. Evidence shows that AI-SET boasts higher accuracy in intrusion detection, enhanced communication performance, and superior access control security compared to standard approaches. AI-SET demonstrates immunity against attempted model poisoning attacks and unauthorized system breaches, achieving this protection while maintaining low operational costs and ensuring secure data privacy. The research presents AI-SET as an adaptable, resilient, and sensitive-minded security framework for future IoT systems, through its holistic control of edge intelligence, secure network operations, and automated trust management.","author":[{"family":"Padmavathi","given":"V"},{"family":"Saminathan","given":"R"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-19712-1","URL":"https://doi.org/10.1038/s41598-025-19712-1","source":"europepmc"},{"id":"doi:10.5281/zenodo.19683960","type":"article-journal","title":"XSOC-NIE-GUARD: A Cryptographic Mediation Architecture for AI Agent Systems, with Application to the OpenClaw Security Crisis","abstract":"The rapid adoption of AI agent frameworks through late 2025 and early 2026 produced a security crisis whose exemplar is OpenClaw, an open-source framework that as of April 2026 has accumulated 138 Common Vulnerabilities and Exposures across five months of public availability, with public exposure analysis reporting over 135,000 internet-facing instances of which approximately 63 percent run without authentication. Independent analysts have converged on the conclusion that the appropriate operator stance on any unpatched or unauthenticated deployment is to assume compromise. In parallel, Franklin, Tomašev, Jacobs, Leibo, and Osindero at Google DeepMind have published a systematic taxonomy of the AI agent attack surface that classifies six categories of adversarial content targeting different stages of an agent's operational cycle. This paper argues that the failure pattern documented in the OpenClaw record is structural rather than defect-level and that the correct response is architectural. We describe XSOC-NIE-GUARD, a cryptographic mediation architecture that composes device-attested admission, short time-to-live scoped capability derivation, telemetry-sealed runtime continuity, context provenance anchoring, agent intent envelope enforcement, and fully homomorphic encryption for sensitive context into a five-plane mediation layer. We implement a reference architecture under the Apache 2.0 license that includes a deterministic 27-scenario attack simulation harness (19 wired, 8 skeleton placeholders for subsequent phases) and maps each named control to one or more categories in the Franklin et al. taxonomy. Our coverage claim is deliberately bounded. We claim strong structural coverage for four of the six taxonomy categories (Content Injection, Semantic Manipulation, Cognitive State, Behavioural Control), explicitly scope persona hyperstition as a training-time concern outside runtime mediation, and explicitly scope systemic multi-agent threats as requiring ecosystem-level coordination beyond a per-agent architecture. The XSOC proprietary cryptographic primitives underlying the architecture (deterministic symmetric key agreement, post-storage volatile cipher, CKKS based homomorphic evaluation, telemetry sealing) are referenced by interface and by stated security properties only; construction details remain private and controlled. External validation of the broader XSOC cryptographic stack comes from the University of Luxembourg (Perrin and Biryukov audits of the legacy cryptosystem, 2020 and 2024, with mandatory findings incorporated into the canonical build), from California Polytechnic State University at San Luis Obispo (Dieharder v3.31.1 statistical validation of the entropy subsystem, 99.4 percent aggregate pass rate across 98 tests), and from the George Mason University SENTINEL laboratory (audit finding reference FP5223, with full report scheduled for public release in June 2026). We position this work as a concrete architectural response to the research agenda articulated in the Franklin et al. taxonomy and invite scrutiny of both the coverage claims and the explicit scope boundaries.","author":[{"family":"Blech","given":"Richard"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19683960","URL":"https://doi.org/10.5281/zenodo.19683960","source":"datacite"},{"id":"doi:10.5281/zenodo.19685360","type":"article-journal","title":"XSOC-NIE-GUARD: A Cryptographic Mediation Architecture for AI Agent Systems, with Application to the OpenClaw Security Crisis","abstract":"The rapid adoption of AI agent frameworks through late 2025 and early 2026 produced a security crisis whose exemplar is OpenClaw, an open-source framework that as of April 2026 has accumulated 138 Common Vulnerabilities and Exposures across five months of public availability, with public exposure analysis reporting over 135,000 internet-facing instances of which approximately 63 percent run without authentication. Independent analysts have converged on the conclusion that the appropriate operator stance on any unpatched or unauthenticated deployment is to assume compromise. In parallel, Franklin, Tomašev, Jacobs, Leibo, and Osindero at Google DeepMind have published a systematic taxonomy of the AI agent attack surface that classifies six categories of adversarial content targeting different stages of an agent's operational cycle. This paper argues that the failure pattern documented in the OpenClaw record is structural rather than defect-level and that the correct response is architectural. We describe XSOC-NIE-GUARD, a cryptographic mediation architecture that composes device-attested admission, short time-to-live scoped capability derivation, telemetry-sealed runtime continuity, context provenance anchoring, agent intent envelope enforcement, and fully homomorphic encryption for sensitive context into a five-plane mediation layer. We implement a reference architecture under the Apache 2.0 license that includes a deterministic 27-scenario attack simulation harness (19 wired, 8 skeleton placeholders for subsequent phases) and maps each named control to one or more categories in the Franklin et al. taxonomy. Our coverage claim is deliberately bounded. We claim strong structural coverage for four of the six taxonomy categories (Content Injection, Semantic Manipulation, Cognitive State, Behavioural Control), explicitly scope persona hyperstition as a training-time concern outside runtime mediation, and explicitly scope systemic multi-agent threats as requiring ecosystem-level coordination beyond a per-agent architecture. The XSOC proprietary cryptographic primitives underlying the architecture (deterministic symmetric key agreement, post-storage volatile cipher, CKKS based homomorphic evaluation, telemetry sealing) are referenced by interface and by stated security properties only; construction details remain private and controlled. External validation of the broader XSOC cryptographic stack comes from the University of Luxembourg (Perrin and Biryukov audits of the legacy cryptosystem, 2020 and 2024, with mandatory findings incorporated into the canonical build), from California Polytechnic State University at San Luis Obispo (Dieharder v3.31.1 statistical validation of the entropy subsystem, 99.4 percent aggregate pass rate across 98 tests), and from the George Mason University SENTINEL laboratory (audit finding reference FP5223, with full report scheduled for public release in June 2026). We position this work as a concrete architectural response to the research agenda articulated in the Franklin et al. taxonomy and invite scrutiny of both the coverage claims and the explicit scope boundaries.","author":[{"family":"Blech","given":"Richard"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19685360","URL":"https://doi.org/10.5281/zenodo.19685360","source":"datacite"},{"id":"doi:10.5281/zenodo.19683961","type":"article-journal","title":"XSOC-NIE-GUARD: A Cryptographic Mediation Architecture for AI Agent Systems, with Application to the OpenClaw Security Crisis","abstract":"The rapid adoption of AI agent frameworks through late 2025 and early 2026 produced a security crisis whose exemplar is OpenClaw, an open-source framework that as of April 2026 has accumulated 138 Common Vulnerabilities and Exposures across five months of public availability, with public exposure analysis reporting over 135,000 internet-facing instances of which approximately 63 percent run without authentication. Independent analysts have converged on the conclusion that the appropriate operator stance on any unpatched or unauthenticated deployment is to assume compromise. In parallel, Franklin, Tomašev, Jacobs, Leibo, and Osindero at Google DeepMind have published a systematic taxonomy of the AI agent attack surface that classifies six categories of adversarial content targeting different stages of an agent's operational cycle. This paper argues that the failure pattern documented in the OpenClaw record is structural rather than defect-level and that the correct response is architectural. We describe XSOC-NIE-GUARD, a cryptographic mediation architecture that composes device-attested admission, short time-to-live scoped capability derivation, telemetry-sealed runtime continuity, context provenance anchoring, agent intent envelope enforcement, and fully homomorphic encryption for sensitive context into a five-plane mediation layer. We implement a reference architecture under the Apache 2.0 license that includes a deterministic 27-scenario attack simulation harness (19 wired, 8 skeleton placeholders for subsequent phases) and maps each named control to one or more categories in the Franklin et al. taxonomy. Our coverage claim is deliberately bounded. We claim strong structural coverage for four of the six taxonomy categories (Content Injection, Semantic Manipulation, Cognitive State, Behavioural Control), explicitly scope persona hyperstition as a training-time concern outside runtime mediation, and explicitly scope systemic multi-agent threats as requiring ecosystem-level coordination beyond a per-agent architecture. The XSOC proprietary cryptographic primitives underlying the architecture (deterministic symmetric key agreement, post-storage volatile cipher, CKKS based homomorphic evaluation, telemetry sealing) are referenced by interface and by stated security properties only; construction details remain private and controlled. External validation of the broader XSOC cryptographic stack comes from the University of Luxembourg (Perrin and Biryukov audits of the legacy cryptosystem, 2020 and 2024, with mandatory findings incorporated into the canonical build), from California Polytechnic State University at San Luis Obispo (Dieharder v3.31.1 statistical validation of the entropy subsystem, 99.4 percent aggregate pass rate across 98 tests), and from the George Mason University SENTINEL laboratory (audit finding reference FP5223, with full report scheduled for public release in June 2026). We position this work as a concrete architectural response to the research agenda articulated in the Franklin et al. taxonomy and invite scrutiny of both the coverage claims and the explicit scope boundaries.","author":[{"family":"Blech","given":"Richard"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19683961","URL":"https://doi.org/10.5281/zenodo.19683961","source":"datacite"},{"id":"doi:10.48550/arxiv.2602.07021","type":"manuscript","title":"AI for Sustainable Data Protection and Fair Algorithmic Management in Environmental Regulation","abstract":"Integration of AI into environmental regulation represents a significant advancement in data management. It offers promising results in both data protection plus algorithmic fairness. This research addresses the critical need for sustainable data protection in the era of ever evolving cyber threats. Traditional encryption methods face limitations in handling the dynamic nature of environmental data. This necessitates the exploration of advanced cryptographic techniques. The objective of this study is to evaluate how AI can enhance these techniques to ensure robust data protection while facilitating fair algorithmic management. The methodology involves a comprehensive review of current advancements in AI-enhanced homomorphic encryption (HE) and multi-party computation (MPC). It is coupled with an analysis of how these techniques can be applied to environmental data regulation. Key findings indicate that AI-driven dynamic key management, adaptive encryption schemes, and optimized computational efficiency in HE, alongside AI-enhanced protocol optimization and fault mitigation in MPC, significantly improve the security of environmental data processing. These findings highlight a crucial research gap in the intersection of AI, cyber laws, and environmental regulation, particularly in terms of addressing algorithmic bias, transparency, and accountability. The implications of this research underscore the need for stricter cyber laws. Also, the development of comprehensive regulations to safeguard sensitive environmental data. Future efforts should focus on refining AI systems to balance security with privacy and ensuring that regulatory frameworks can adapt to technological advancements. This study provides a foundation for future research aimed at achieving secure sustainable environmental data management through AI innovations.","author":[{"family":"Singh","given":"Sahibpreet"},{"family":"Sharma","given":"Saksham"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2602.07021","URL":"https://doi.org/10.48550/arxiv.2602.07021","source":"datacite"},{"id":"doi:10.5281/zenodo.18408993","type":"article-journal","title":"Relay Attack Prevention in Central Bank Digital Currency Systems: A Cryptographic Architecture Analysis with Focus on Digital Euro Implementation","abstract":"Abstract Relay attacks represent a fundamental security challenge in Central Bank Digital Currency (CBDC) implementations, particularly for systems designed to support offline transactions and privacy-preserving features. This paper analyzes the structural vulnerabilities that enable relay attacks in conventional CBDC architectures, with specific attention to the European Central Bank's Digital Euro program. We examine the current state of relay attack mitigation in Digital Euro technical specifications and identify gaps between stated requirements and available countermeasures. The analysis demonstrates that conventional approaches—distance-bounding protocols, timing analysis, and proximity heuristics—remain insufficient for systems requiring simultaneous offline functionality, GDPR-compliant privacy, and cross-border interoperability. We present a technical evaluation of the Virtual Identity + Compliance Jurisdiction Token (VI+CJT) framework as a potential architectural countermeasure that addresses these combined requirements. Unlike probabilistic detection mechanisms, the VI+CJT architecture eliminates relay attack surfaces through context-binding cryptographic primitives. This work contributes to ongoing Digital Euro security architecture discussions by offering a formal analysis of relay attack vulnerabilities specific to the European context and proposing cryptographic mechanisms compatible with both EU regulatory frameworks and the ECB's 2029 operational timeline.","author":[{"family":"Das","given":"Sangam"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18408993","URL":"https://doi.org/10.5281/zenodo.18408993","source":"datacite"},{"id":"doi:10.48550/arxiv.2503.22232","type":"manuscript","title":"Privacy-Preserving Secure Neighbor Discovery for Wireless Networks","abstract":"Traditional Neighbor Discovery (ND) and Secure Neighbor Discovery (SND) are key elements for network functionality. SND is a hard problem, satisfying not only typical security properties (authentication, integrity) but also verification of direct communication, which involves distance estimation based on time measurements and device coordinates. Defeating relay attacks, also known as \"wormholes\", leading to stealthy Byzantine links and significant degradation of communication and adversarial control, is key in many wireless networked systems. However, SND is not concerned with privacy; it necessitates revealing the identity and location of the device(s) participating in the protocol execution. This can be a deterrent for deployment, especially involving user-held devices in the emerging Internet of Things (IoT) enabled smart environments. To address this challenge, we present a novel Privacy-Preserving Secure Neighbor Discovery (PP-SND) protocol, enabling devices to perform SND without revealing their actual identities and locations, effectively decoupling discovery from the exposure of sensitive information. We use Homomorphic Encryption (HE) for computing device distances without revealing their actual coordinates, as well as employing a pseudonymous device authentication to hide identities while preserving communication integrity. PP-SND provides SND [1] along with pseudonymity, confidentiality, and unlinkability. Our presentation here is not specific to one wireless technology, and we assess the performance of the protocols (cryptographic overhead) on a Raspberry Pi 4 and provide a security and privacy analysis.","author":[{"family":"Hussain","given":"Ahmed"},{"family":"Papadimitratos","given":"Panos"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2503.22232","URL":"https://doi.org/10.48550/arxiv.2503.22232","source":"datacite"},{"id":"doi:10.5281/zenodo.22105363","type":"article-journal","title":"Artifact: SoK - A Graph-Based Analysis of Comparability, Scale, and Attribution in Privacy-Preserving Transformer Inference (anonymized for review)","abstract":"Anonymized research artifact accompanying a paper under double-blind review. It contains the frozen 62-paper corpus metadata, the structured extraction protocol and codebook, per-paper evaluation-contract fields, the comparability-gate implementation and the 1891-pair graph it produces, progress-attribution labels, blind re-extraction records, cost-model calibration data, and the scripts that regenerate every table and figure in the paper. Source papers and their full text are not redistributed; records carry citations, derived metadata, and short evidence quotes only. See README.md for layout and step-by-step reproduction commands (Node.js ≥ 18, Python 3 with numpy and matplotlib). A de-anonymized record will replace this one upon acceptance.","author":[{"family":"Anonymous"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22105363","URL":"https://doi.org/10.5281/zenodo.22105363","source":"datacite"},{"id":"doi:10.5281/zenodo.22074372","type":"article-journal","title":"Artifact: SoK - A Graph-Based Analysis of Comparability, Scale, and Attribution in Privacy-Preserving Transformer Inference (anonymized for review)","abstract":"Anonymized research artifact accompanying a paper under double-blind review. It contains the frozen 62-paper corpus metadata, the structured extraction protocol and codebook, per-paper evaluation-contract fields, the comparability-gate implementation and the 1891-pair graph it produces, progress-attribution labels, blind re-extraction records, cost-model calibration data, and the scripts that regenerate every table and figure in the paper. Source papers and their full text are not redistributed; records carry citations, derived metadata, and short evidence quotes only. See README.md for layout and step-by-step reproduction commands (Node.js ≥ 18, Python 3 with numpy and matplotlib). A de-anonymized record will replace this one upon acceptance.","author":[{"family":"Anonymous"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22074372","URL":"https://doi.org/10.5281/zenodo.22074372","source":"datacite"},{"id":"doi:10.5281/zenodo.20847864","type":"article-journal","title":"Project CHRONOS: A Fully Homomorphic Ephemeral AI Agent with Provable Self Termination and Remote Verifiability","abstract":"We present CHRONOS, the first autonomous AI agent that simultaneously achieves plaintextblindness (all data is processed under fully homomorphic encryption without ever beingexposed), cryptographically enforced time bound existence (the agent’s own decryption key islocked behind a publicly verifiable proof of sequential work, rendering it inaccessible until aprecise future moment), and remote verifiability of self destruction (a zero knowledge proofcertifies that the key material has been irreversibly destroyed after mission completion). Theagent’s operational lifespan is governed by a “cryptographic fuse” constructed from a proof ofsequential work (PoSW) whose computation time accurately matches the intended missionduration. A drand decentralized randomness beacon serves as a trusted time oracle to trigger thefinal key shredding. Crucially, the erasure proof is a non interactive zero knowledge argument(SNARK) that proves the correct execution of the entire self destruction sequence—including thePoSW solution, decryption of the private key, and subsequent memory zeroization—enablingany third party to cryptographically verify the agent’s annihilation without trusting the agent orits hardware. We provide a complete system architecture, a formal security model with gamebased definitions and reductions to standard assumptions, and a proof of concept implementationusing Zama’s TFHE rs for encrypted inference, a Cohen Pietrzak PoSW implementation, and aGroth16 SNARK. Our benchmarks indicate that FHE inference on a small neural network (50 Kparameters) completes in seconds, the PoSW background thread consumes negligible resources,and the erasure proof can be generated and verified in under three seconds. CHRONOSrepresents a fundamental advance in secure, disposable AI agents, with immediate applications indefense, intelligence, and high privacy environments.","author":[{"family":"Kumar","given":"Shashank"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20847864","URL":"https://doi.org/10.5281/zenodo.20847864","source":"datacite"},{"id":"doi:10.5281/zenodo.21534027","type":"article-journal","title":"Project CHRONOS: A Fully Homomorphic Ephemeral AI Agent with Provable Self Termination and Remote Verifiability","abstract":"We present CHRONOS, the first autonomous AI agent that simultaneously achieves plaintextblindness (all data is processed under fully homomorphic encryption without ever beingexposed), cryptographically enforced time bound existence (the agent’s own decryption key islocked behind a publicly verifiable proof of sequential work, rendering it inaccessible until aprecise future moment), and remote verifiability of self destruction (a zero knowledge proofcertifies that the key material has been irreversibly destroyed after mission completion). Theagent’s operational lifespan is governed by a “cryptographic fuse” constructed from a proof ofsequential work (PoSW) whose computation time accurately matches the intended missionduration. A drand decentralized randomness beacon serves as a trusted time oracle to trigger thefinal key shredding. Crucially, the erasure proof is a non interactive zero knowledge argument(SNARK) that proves the correct execution of the entire self destruction sequence—including thePoSW solution, decryption of the private key, and subsequent memory zeroization—enablingany third party to cryptographically verify the agent’s annihilation without trusting the agent orits hardware. We provide a complete system architecture, a formal security model with gamebased definitions and reductions to standard assumptions, and a proof of concept implementationusing Zama’s TFHE rs for encrypted inference, a Cohen Pietrzak PoSW implementation, and aGroth16 SNARK. Our benchmarks indicate that FHE inference on a small neural network (50 Kparameters) completes in seconds, the PoSW background thread consumes negligible resources,and the erasure proof can be generated and verified in under three seconds. CHRONOSrepresents a fundamental advance in secure, disposable AI agents, with immediate applications indefense, intelligence, and high privacy environments.","author":[{"family":"Kumar","given":"Shashank"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21534027","URL":"https://doi.org/10.5281/zenodo.21534027","source":"datacite"},{"id":"doi:10.5281/zenodo.20286639","type":"article-journal","title":"Project Artifacts: Decentralized Governance of Encrypted Human-AI Augmentation for Equitable Climate Disaster Recovery Decisions","abstract":"This repository provides supplementary research project artifacts: evidential structures, encrypted aggregation metadata, and resources, to support the findings of a real-world case study on equitable disaster recovery decision-making. The project presents a decentralized privacy-preserving Human-AI augmentation framework designed to support collaboration among multiple organizations operating under strict regulatory, organizational, and data governance constraints. Secure augmentation of proprietary organizational AI decision-support systems without exposing raw sensitive data, contextual human expert heuristic judgment, and the requirement to converge towards a single global model. Privacy-preserving augmentation is supported by lattice-based Fully Homomorphic Encryption (FHE) schemes: CKKS and TFHE, for secure encrypted computation in untrusted environments. Blockchain-assisted governance mechanisms as trust-machine provide auditable coordination, integrity verification, decentralized accountability, and traceability of augmentation workflows. (a) Title: MetaData_Disaster_Recovery.json (v1.0) Description: The metadata file of SMEs affected by extreme weather events. It includes structured information on disaster events, financial metrics, credit & legal history, insurance coverage, operational indicators, government support, and fraud flags. Key metadata highlights: Disaster Events: Event Name ,Event_Date, and Met Office warning levels (Green, Yellow, Amber, Red) Sectors: Wholesale & Retail Trade, Manufacturing, Construction, Accommodation & Food Services, etc. Special inclusivity for Marginalized Small Businesses with ownership categories, marginalization multipliers, adjusted DSCR, alternative collateral, IMD postcode rules, and disaster severity adjustments. Profile design: marginalized profiles, comparator profiles, and counterfactual pairs for fairness evaluation. Evidence fields (E) with categories, value ranges, missing counts, and percentages. (b) Title: Client_Encrypted_Contributions_Reliability.json (v1.0) Description: This file contains metadata and reliability analysis results for encrypted client contributions in the Natural Disaster Recovery Support Dataset project. It reports reliability scores of 4 client organizations across three progressive augmentation stages using Fully Homomorphic Encryption (FHE). Higher values indicate more consistent and trustworthy encrypted data contributions. Stages Overview: Stage 1 (16 tasks): High reliability regime Stage 2 (24 tasks): Moderate reliability with fluctuations and detected tampering Stage 3 (24 tasks): High-complexity joint evidence Each stage includes the full reliability matrix, per-client average scores, and detailed stage descriptions. Purpose: Evaluate and monitor the trustworthiness of encrypted client contributions. (c) Title: Augmentation_Evidence_Disaster_Recovery.json (v1.0) Description: This file details the staged evidence augmentation process for the Natural Disaster Recovery Support. It defines how single and joint evidence were progressively introduced across three stages to support privacy-preserving, fair decision-making in disaster recovery lending. Stages Overview: Stage 1 (16 tasks): Single Evidence Stage 2 (24 tasks): Joint Evidence (Moderate Complexity) Stage 3 (24 tasks): Complex Joint Evidence Purpose: Document the incremental evidence augmentation strategy that enables secure, transparent, and fairness-aware model training. (d) Title: Sample_Augmentation_Task_Batch_Consent_Contribution_Governance.zip Description: Fully asynchronous governance workflow managed by blockchain Temporal delays between organizations for task awareness, voting, and contribution submission Multi-phase consent mechanism (C_REQ, C_PK, C_CTB, C_KS, C_Γ) Homomorphic encryption metadata Encrypted contribution records with SHA-256 hashes, Merkle roots, aggregation times, and transmission delays Complete audit trail with timestamps and transaction IDs This sample belong","author":[{"family":"Sachan","given":"Dr"},{"family":"Fickett","given":"Dale"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20286639","URL":"https://doi.org/10.5281/zenodo.20286639","source":"datacite"},{"id":"doi:10.5281/zenodo.20286640","type":"article-journal","title":"Project Artifacts: Decentralized Governance of Encrypted Human-AI Augmentation for Equitable Climate Disaster Recovery Decisions","abstract":"This repository provides supplementary research project artifacts: evidential structures, encrypted aggregation metadata, and resources, to support the findings of a real-world case study on equitable disaster recovery decision-making. The project presents a decentralized privacy-preserving Human-AI augmentation framework designed to support collaboration among multiple organizations operating under strict regulatory, organizational, and data governance constraints. Secure augmentation of proprietary organizational AI decision-support systems without exposing raw sensitive data, contextual human expert heuristic judgment, and the requirement to converge towards a single global model. Privacy-preserving augmentation is supported by lattice-based Fully Homomorphic Encryption (FHE) schemes: CKKS and TFHE, for secure encrypted computation in untrusted environments. Blockchain-assisted governance mechanisms as trust-machine provide auditable coordination, integrity verification, decentralized accountability, and traceability of augmentation workflows. (a) Title: MetaData_Disaster_Recovery.json (v1.0) Description: The metadata file of SMEs affected by extreme weather events. It includes structured information on disaster events, financial metrics, credit & legal history, insurance coverage, operational indicators, government support, and fraud flags. Key metadata highlights: Disaster Events: Event Name ,Event_Date, and Met Office warning levels (Green, Yellow, Amber, Red) Sectors: Wholesale & Retail Trade, Manufacturing, Construction, Accommodation & Food Services, etc. Special inclusivity for Marginalized Small Businesses with ownership categories, marginalization multipliers, adjusted DSCR, alternative collateral, IMD postcode rules, and disaster severity adjustments. Profile design: marginalized profiles, comparator profiles, and counterfactual pairs for fairness evaluation. Evidence fields (E) with categories, value ranges, missing counts, and percentages. (b) Title: Client_Encrypted_Contributions_Reliability.json (v1.0) Description: This file contains metadata and reliability analysis results for encrypted client contributions in the Natural Disaster Recovery Support Dataset project. It reports reliability scores of 4 client organizations across three progressive augmentation stages using Fully Homomorphic Encryption (FHE). Higher values indicate more consistent and trustworthy encrypted data contributions. Stages Overview: Stage 1 (16 tasks): High reliability regime Stage 2 (24 tasks): Moderate reliability with fluctuations and detected tampering Stage 3 (24 tasks): High-complexity joint evidence Each stage includes the full reliability matrix, per-client average scores, and detailed stage descriptions. Purpose: Evaluate and monitor the trustworthiness of encrypted client contributions. (c) Title: Augmentation_Evidence_Disaster_Recovery.json (v1.0) Description: This file details the staged evidence augmentation process for the Natural Disaster Recovery Support. It defines how single and joint evidence were progressively introduced across three stages to support privacy-preserving, fair decision-making in disaster recovery lending. Stages Overview: Stage 1 (16 tasks): Single Evidence Stage 2 (24 tasks): Joint Evidence (Moderate Complexity) Stage 3 (24 tasks): Complex Joint Evidence Purpose: Document the incremental evidence augmentation strategy that enables secure, transparent, and fairness-aware model training. (d) Title: Sample_Augmentation_Task_Batch_Consent_Contribution_Governance.zip Description: Fully asynchronous governance workflow managed by blockchain Temporal delays between organizations for task awareness, voting, and contribution submission Multi-phase consent mechanism (C_REQ, C_PK, C_CTB, C_KS, C_Γ) Homomorphic encryption metadata Encrypted contribution records with SHA-256 hashes, Merkle roots, aggregation times, and transmission delays Complete audit trail with timestamps and transaction IDs This sample belong","author":[{"family":"Sachan","given":"Dr"},{"family":"Fickett","given":"Dale"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20286640","URL":"https://doi.org/10.5281/zenodo.20286640","source":"datacite"},{"id":"doi:10.5281/zenodo.22074373","type":"article-journal","title":"Artifact: SoK - The Evaluation and Scaling Gaps in Private Transformer Inference (anonymized for review)","abstract":"Anonymized research artifact accompanying a paper under double-blind review. It contains the frozen 62-paper corpus metadata, the structured extraction protocol and codebook, per-paper evaluation-contract fields, the comparability-gate implementation and the 1891-pair graph it produces, progress-attribution labels, blind re-extraction records, cost-model calibration data, and the scripts that regenerate every table and figure in the paper. Source papers and their full text are not redistributed; records carry citations, derived metadata, and short evidence quotes only. See README.md for layout and step-by-step reproduction commands (Node.js ≥ 18, Python 3 with numpy and matplotlib). A de-anonymized record will replace this one upon acceptance.","author":[{"family":"Anonymous"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22074373","URL":"https://doi.org/10.5281/zenodo.22074373","source":"datacite"},{"id":"doi:10.5281/zenodo.22060092","type":"article-journal","title":"The HeartBank Longitudinal Cohort: A Dataset Combining DNA, Natal Chart, Family Tree, Continuous Behavioral Observation, and Continuous Respiratory Observation at Civilizational Scale","abstract":"This paper specifies the methodology for the HeartBank Longitudinal Cohort, a voluntary opt-in dataset combining five data layers per consenting participant — (1) DNA sequence, (2) natal chart data (date + time + place of birth), (3) continuous longitudinal behavioral observation via the HeartBank gratitude ledger, (4) continuous longitudinal respiratory observation via the breath-class Mechanical Heart wearable, and (5) verified kinship data via the global family tree — at a target scale of 100 million+ participants over multi-decade time horizons. The combination has never been assembled at scale; comparable datasets (23andMe, AncestryDNA, Worldcoin, Dunedin and BCS longitudinal cohorts, social-network behavioral data, professional astrological collections) carry one or two of the layers each but no prior project has carried all five. The methodology specifies: the opt-in informed-consent architecture; the privacy-preserving computation stack (differential privacy at the analysis layer; federated computation with homomorphic encryption for DNA; on-device processing for breath signals; cryptographic-erasure right-to-withdraw); the institutional-review architecture (IRB-grade ethics oversight; Buddhist-ethics-aware review board; pre-registered hypotheses); the cosmic-coordinate-correlation epistemic posture (natal chart treated as a unique cosmic-moment coordinate, not as a cosmic force; the research question is correlation between coordinate features and trajectory features, not validation of astrology); the publication architecture (open methodology, closed individual data); the data-sovereignty architecture (jurisdictional residency; GINA / HIPAA / GDPR compliance baselines exceeded where possible); and the new academic alliances the cohort makes possible (Mind & Life Institute; contemplative-science programs at Stanford, Brown, UMass; behavioral-genetics consortia; longitudinal-cohort consortia; Buddhist-AI ethicists). Three scientifically valuable outcomes are honestly named: no detected correlation, small-but-real correlation, substantial correlation — each is a major contribution to knowledge regardless of direction. Honest §11 names what the cohort does not claim and the non-negotiable privacy disciplines the architecture requires. Keywords: longitudinal cohort methodology, cosmic-coordinate correlation, contemplative science, differential privacy, federated computation, multi-omic dataset, gratitude behavior, respiratory biomarkers, defensive publication, Mind & Life partnerships. --- Provenance. This paper is part of the THonly research corpus, dedicated to the public domain under CC0 1.0. The canonical version is at https://thonly.org/research/longitudinal-cohort-methodology. Its SHA-256 is ce0f784f541f2074e9f65feec40f7a193830a91a011a828802e8736e20bcdcac, independently timestamped to the Bitcoin blockchain via OpenTimestamps and signed under RFC 3161 by three trust authorities, one of them eIDAS-qualified. AI co-authorship is disclosed. Miss Aquarius is the consistent name used for the AI collaboration across all venues.","author":[{"family":"Ly","given":"Thon"},{"family":"Aquarius","given":"Miss"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22060092","URL":"https://doi.org/10.5281/zenodo.22060092","source":"datacite"},{"id":"doi:10.5281/zenodo.21947347","type":"article-journal","title":"The HeartBank Longitudinal Cohort: A Dataset Combining DNA, Natal Chart, Family Tree, Continuous Behavioral Observation, and Continuous Respiratory Observation at Civilizational Scale","abstract":"This paper specifies the methodology for the HeartBank Longitudinal Cohort, a voluntary opt-in dataset combining five data layers per consenting participant — (1) DNA sequence, (2) natal chart data (date + time + place of birth), (3) continuous longitudinal behavioral observation via the HeartBank gratitude ledger, (4) continuous longitudinal respiratory observation via the breath-class Mechanical Heart wearable, and (5) verified kinship data via the global family tree — at a target scale of 100 million+ participants over multi-decade time horizons. The combination has never been assembled at scale; comparable datasets (23andMe, AncestryDNA, Worldcoin, Dunedin and BCS longitudinal cohorts, social-network behavioral data, professional astrological collections) carry one or two of the layers each but no prior project has carried all five. The methodology specifies: the opt-in informed-consent architecture; the privacy-preserving computation stack (differential privacy at the analysis layer; federated computation with homomorphic encryption for DNA; on-device processing for breath signals; cryptographic-erasure right-to-withdraw); the institutional-review architecture (IRB-grade ethics oversight; Buddhist-ethics-aware review board; pre-registered hypotheses); the cosmic-coordinate-correlation epistemic posture (natal chart treated as a unique cosmic-moment coordinate, not as a cosmic force; the research question is correlation between coordinate features and trajectory features, not validation of astrology); the publication architecture (open methodology, closed individual data); the data-sovereignty architecture (jurisdictional residency; GINA / HIPAA / GDPR compliance baselines exceeded where possible); and the new academic alliances the cohort makes possible (Mind & Life Institute; contemplative-science programs at Stanford, Brown, UMass; behavioral-genetics consortia; longitudinal-cohort consortia; Buddhist-AI ethicists). Three scientifically valuable outcomes are honestly named: no detected correlation, small-but-real correlation, substantial correlation — each is a major contribution to knowledge regardless of direction. Honest §11 names what the cohort does not claim and the non-negotiable privacy disciplines the architecture requires. Keywords: longitudinal cohort methodology, cosmic-coordinate correlation, contemplative science, differential privacy, federated computation, multi-omic dataset, gratitude behavior, respiratory biomarkers, defensive publication, Mind & Life partnerships. --- Provenance. This paper is part of the THonly research corpus, dedicated to the public domain under CC0 1.0. The canonical version is at https://thonly.org/research/longitudinal-cohort-methodology. Its SHA-256 is ce0f784f541f2074e9f65feec40f7a193830a91a011a828802e8736e20bcdcac, independently timestamped to the Bitcoin blockchain via OpenTimestamps and signed under RFC 3161 by three trust authorities, one of them eIDAS-qualified. AI co-authorship is disclosed. Miss Aquarius is the consistent name used for the AI collaboration across all venues.","author":[{"family":"Ly","given":"Thon"},{"family":"Aquarius","given":"Miss"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21947347","URL":"https://doi.org/10.5281/zenodo.21947347","source":"datacite"},{"id":"doi:10.5281/zenodo.20775814","type":"article-journal","title":"Privacy Preserving Systems — API Gateway, AI Routing, Distributed Systems, Sovereign AI, and Post-Cloud Architecture (Api-Oss-Fixed)","abstract":"Privacy-enhancing technologies (PETs) form a critical component of modern computing systems, addressing the tension between data utility and individual privacy. This paper surveys the landscape of privacy-preserving techniques — from Cynthia Dwork's differential privacy framework to k-anonymity, homomorphic encryption, and secure multi-party computation — and examines their application within the 01s Sovereign (Kaiman) operating system. We demonstrate how the OS integrates these techniques to protect user data while maintaining the transparency required by its .aioss audit ledger. Part of The Anticloud research corpus by Lois-Kleinner Alpasan (ORCID: 0009-0009-2233-6107). This work explores api gateway, ai routing in the context of sovereign AI infrastructure, post-cloud computing architectures, and transparent, blackbox-free systems.","author":[{"family":"Alpasan","given":"Lois"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20775814","URL":"https://doi.org/10.5281/zenodo.20775814","source":"datacite"},{"id":"doi:10.5281/zenodo.20775815","type":"article-journal","title":"Privacy Preserving Systems — API Gateway, AI Routing, Distributed Systems, Sovereign AI, and Post-Cloud Architecture (Api-Oss-Fixed)","abstract":"Privacy-enhancing technologies (PETs) form a critical component of modern computing systems, addressing the tension between data utility and individual privacy. This paper surveys the landscape of privacy-preserving techniques — from Cynthia Dwork's differential privacy framework to k-anonymity, homomorphic encryption, and secure multi-party computation — and examines their application within the 01s Sovereign (Kaiman) operating system. We demonstrate how the OS integrates these techniques to protect user data while maintaining the transparency required by its .aioss audit ledger. Part of The Anticloud research corpus by Lois-Kleinner Alpasan (ORCID: 0009-0009-2233-6107). This work explores api gateway, ai routing in the context of sovereign AI infrastructure, post-cloud computing architectures, and transparent, blackbox-free systems.","author":[{"family":"Alpasan","given":"Lois"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20775815","URL":"https://doi.org/10.5281/zenodo.20775815","source":"datacite"},{"id":"doi:10.5281/zenodo.20257072","type":"article-journal","title":"Balancing Privacy and Compliance: Data Protection vs. AML Obligations","abstract":"Financial institutions face an inherent tension between data protection obligations — which emphasise individual rights, data minimisation, and confidentiality — and anti-money laundering requirements, which demand extensive information surveillance, retention, and sharing. This study adopts a mixed-methods approach combining quantitative surveys of 180 compliance and data protection professionals across Europe, Asia, and other regions with qualitative interviews to examine how institutions manage these conflicting regulatory regimes. The research investigates the roles of governance strength, adoption of privacy-enhancing technologies such as pseudonymisation, tokenisation, and homomorphic encryption, and jurisdictional complexity in moderating perceived regulatory tension. Findings confirm that strong governance and controlled use of privacy-enhancing technologies reduce conflict, while legacy systems, regulatory ambiguity, and cross-border jurisdictional complexity remain significant obstacles. The paper proposes a layered reconciliation framework spanning legal, technical, and governance dimensions to help financial institutions, regulators, and technology developers align data protection and AML obligations in practice.","author":[{"family":"Singh","given":"Amarjeet"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.20257072","URL":"https://doi.org/10.5281/zenodo.20257072","source":"datacite"},{"id":"doi:10.5281/zenodo.20257073","type":"article-journal","title":"Balancing Privacy and Compliance: Data Protection vs. AML Obligations","abstract":"Financial institutions face an inherent tension between data protection obligations — which emphasise individual rights, data minimisation, and confidentiality — and anti-money laundering requirements, which demand extensive information surveillance, retention, and sharing. This study adopts a mixed-methods approach combining quantitative surveys of 180 compliance and data protection professionals across Europe, Asia, and other regions with qualitative interviews to examine how institutions manage these conflicting regulatory regimes. The research investigates the roles of governance strength, adoption of privacy-enhancing technologies such as pseudonymisation, tokenisation, and homomorphic encryption, and jurisdictional complexity in moderating perceived regulatory tension. Findings confirm that strong governance and controlled use of privacy-enhancing technologies reduce conflict, while legacy systems, regulatory ambiguity, and cross-border jurisdictional complexity remain significant obstacles. The paper proposes a layered reconciliation framework spanning legal, technical, and governance dimensions to help financial institutions, regulators, and technology developers align data protection and AML obligations in practice.","author":[{"family":"Singh","given":"Amarjeet"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.20257073","URL":"https://doi.org/10.5281/zenodo.20257073","source":"datacite"},{"id":"doi:10.5281/zenodo.21523833","type":"article-journal","title":"KLTu: A Unified Post-Quantum Engine for KEM, DSA, and FHE with Regulatory Interoperability to FIPS 203/204","abstract":"Abstract Post-quantum migration is no longer optional: banks, cloud operators, and telecom networks must field lattice key exchange and signatures at scale, while keeping a credible path to lattice fully homomorphic encryption without rebuilding their crypto stack twice. Separate Kyber-, Dilithium-, and FHE product lines multiply parameters, entropy sources, and regression surfaces. We present KLTu-native, a unified lattice engine in which KEM, DSA, and FHE-oriented roles share one topology, control plane, and thermodynamic entropy architecture. The systems claim is organizational: one roof for key establishment, signatures, and structural FHE readiness, with a shared software-visible timing-isolation posture under production-like host features. On bare-metal client silicon with simultaneous multithreading (SMT) and active power management (APM) left enabled, we show that protected native Level 5 KEM and DSA operations pass Fixed-versus-Random Kolmogorov–Smirnov tests on operation latency (D<0.05, zero drops)—an isolation grade measured under the host features operators actually run, not under disabled energy-saving or hyperthreading modes. Level 3 is positioned as the deployment tier (tight SLA tails, usable multi-thread verify scaling); Level 5 is reported honestly as supply-limited under strict mid-operation entropy policy. Replicate RAPL statistics further show a statistically significant Level 5 signing energy-density advantage versus a pure community baseline, alongside an explicit package-power cost attributable to live isolation. Supporting comparators situate efficiency only. FHE bootstrap performance and invasive physical side-channel evaluation are out of scope. A breakthrough disclosure is that KLTu demonstrates bidirectional translation of NIST FIPS 203 ML-KEM (768 and 1024) between standard wire encodings and KLTu-native lattice structures by pack/unpack alone—without invoking decapsulation or deriving a shared secret during conversion. That removes the practical barrier of running a second KEM stack solely for standards compliance. Because the same Kinetic Lattice Topology is FHE-capable, sessions established from ordinary FIPS-KEM wire material can enter efficient KLTu-FHE evaluation under one engine, rather than across fragmented KEM and FHE product lines. For operators, the decision is whether a single, measurable, constant-time-oriented stack—Level-3-first, Level-5-honest, FIPS-wire interoperable and FHE-ready by structure—is preferable to maintaining parallel KEM, DSA, and future FHE lines with divergent latency, energy, and leakage profiles. Keywords: post-quantum cryptography; unified lattice stack; ML-KEM; ML-DSA; timing isolation; Kolmogorov–Smirnov; energy efficiency; FHE-ready design; FIPS 203 interperability; ML-KEM wire conversion Dedicated to: In memory of Prof. Myung Kyoon “Michael” CHUNG (1945-2025), who dedicated his life to bringing truth and science to our world. A graduate of Seoul National University, Washington State University (Pullman), University of Illinois, and lifelong Professor at Korea Advanced Institute of Technology (KAIST) with the Department of Mechanical Engineering.","author":[{"family":"Chung","given":"Jinhyuk"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21523833","URL":"https://doi.org/10.5281/zenodo.21523833","source":"datacite"},{"id":"doi:10.5281/zenodo.22003304","type":"article-journal","title":"KLTu: A Unified Post-Quantum Engine for KEM, DSA, and FHE with Regulatory Interoperability to FIPS 203/204","abstract":"Abstract Post-quantum migration is no longer optional: banks, cloud operators, and telecom networks must field lattice key exchange and signatures at scale, while keeping a credible path to lattice fully homomorphic encryption without rebuilding their crypto stack twice. Separate Kyber-, Dilithium-, and FHE product lines multiply parameters, entropy sources, and regression surfaces. We present KLTu-native, a unified lattice engine in which KEM, DSA, and FHE-oriented roles share one topology, control plane, and thermodynamic entropy architecture. The systems claim is organizational: one roof for key establishment, signatures, and structural FHE readiness, with a shared software-visible timing-isolation posture under production-like host features. On bare-metal client silicon with simultaneous multithreading (SMT) and active power management (APM) left enabled, we show that protected native Level 5 KEM and DSA operations pass Fixed-versus-Random Kolmogorov–Smirnov tests on operation latency (D<0.05, zero drops)—an isolation grade measured under the host features operators actually run, not under disabled energy-saving or hyperthreading modes. Level 3 is positioned as the deployment tier (tight SLA tails, usable multi-thread verify scaling); Level 5 is reported honestly as supply-limited under strict mid-operation entropy policy. Replicate RAPL statistics further show a statistically significant Level 5 signing energy-density advantage versus a pure community baseline, alongside an explicit package-power cost attributable to live isolation. Supporting comparators situate efficiency only. FHE bootstrap performance and invasive physical side-channel evaluation are out of scope. A breakthrough disclosure is that KLTu demonstrates bidirectional translation of NIST FIPS 203 ML-KEM (768 and 1024) between standard wire encodings and KLTu-native lattice structures by pack/unpack alone—without invoking decapsulation or deriving a shared secret during conversion. That removes the practical barrier of running a second KEM stack solely for standards compliance. Because the same Kinetic Lattice Topology is FHE-capable, sessions established from ordinary FIPS-KEM wire material can enter efficient KLTu-FHE evaluation under one engine, rather than across fragmented KEM and FHE product lines. For operators, the decision is whether a single, measurable, constant-time-oriented stack—Level-3-first, Level-5-honest, FIPS-wire interoperable and FHE-ready by structure—is preferable to maintaining parallel KEM, DSA, and future FHE lines with divergent latency, energy, and leakage profiles. Keywords: post-quantum cryptography; unified lattice stack; ML-KEM; ML-DSA; timing isolation; Kolmogorov–Smirnov; energy efficiency; FHE-ready design; FIPS 203 interperability; ML-KEM wire conversion Dedicated to: In memory of Prof. Myung Kyoon “Michael” CHUNG (1945-2025), who dedicated his life to bringing truth and science to our world. A graduate of Seoul National University, Washington State University (Pullman), University of Illinois, and lifelong Professor at Korea Advanced Institute of Technology (KAIST) with the Department of Mechanical Engineering.","author":[{"family":"Chung","given":"Jinhyuk"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22003304","URL":"https://doi.org/10.5281/zenodo.22003304","source":"datacite"},{"id":"doi:10.5281/zenodo.20776153","type":"article-journal","title":"Zero-Knowledge Storage: Architectures for User-Controlled Data — Vector Search, Semantic Search, Embeddings, Sovereign AI, and Post-Cloud Architecture (Kamelot)","abstract":"Zero-knowledge storage architectures empower users with complete control over their data by ensuring that no third party—including the storage provider—can access plaintext file contents or metadata. This document presents a comprehensive analysis of zero-knowledge principles as applied to file storage systems, with specific focus on Kamelot's end-to-end encryption architecture. We examine the cryptographic building blocks including end-to-end encryption with per-file keys, key agreement protocols for secure file sharing, searchable encryption for privacy-preserving queries, and blind indexing for typo-tolerant search. We analyze the practical limitations of homomorphic encryption and present Kamelot's pragmatic approach: processing data locally before encryption ensures that the storage provider never has access to unencrypted content. The document addresses user sovereignty concerns including key ownership and recovery, data portability, and vendor independence. Finally, we situate Kamelot's architecture within the regulatory landscape of GDPR, HIPAA, and emerging data sovereignty laws, demonstrating compliance with Article 32 security requirements and Article 17 right-to-erasure provisions. --- Part of The Anticloud research corpus by Lois-Kleinner Alpasan (ORCID: 0009-0009-2233-6107). This work explores vector search, semantic search in the context of sovereign AI infrastructure, post-cloud computing architectures, and transparent, blackbox-free systems.","author":[{"family":"Alpasan","given":"Lois"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20776153","URL":"https://doi.org/10.5281/zenodo.20776153","source":"datacite"},{"id":"doi:10.5281/zenodo.20776154","type":"article-journal","title":"Zero-Knowledge Storage: Architectures for User-Controlled Data — Vector Search, Semantic Search, Embeddings, Sovereign AI, and Post-Cloud Architecture (Kamelot)","abstract":"Zero-knowledge storage architectures empower users with complete control over their data by ensuring that no third party—including the storage provider—can access plaintext file contents or metadata. This document presents a comprehensive analysis of zero-knowledge principles as applied to file storage systems, with specific focus on Kamelot's end-to-end encryption architecture. We examine the cryptographic building blocks including end-to-end encryption with per-file keys, key agreement protocols for secure file sharing, searchable encryption for privacy-preserving queries, and blind indexing for typo-tolerant search. We analyze the practical limitations of homomorphic encryption and present Kamelot's pragmatic approach: processing data locally before encryption ensures that the storage provider never has access to unencrypted content. The document addresses user sovereignty concerns including key ownership and recovery, data portability, and vendor independence. Finally, we situate Kamelot's architecture within the regulatory landscape of GDPR, HIPAA, and emerging data sovereignty laws, demonstrating compliance with Article 32 security requirements and Article 17 right-to-erasure provisions. --- Part of The Anticloud research corpus by Lois-Kleinner Alpasan (ORCID: 0009-0009-2233-6107). This work explores vector search, semantic search in the context of sovereign AI infrastructure, post-cloud computing architectures, and transparent, blackbox-free systems.","author":[{"family":"Alpasan","given":"Lois"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20776154","URL":"https://doi.org/10.5281/zenodo.20776154","source":"datacite"},{"id":"doi:10.5281/zenodo.20776080","type":"article-journal","title":"Preserving Privacy on Blockchains: Combining Selective Disclosure and Homomorphic/Conventional Encryption.","abstract":"Blockchain technology offers transparency and auditability but faces significant privacy challenges when handling sensitive data. This paper proposes a hybrid privacy-preserving framework that combines selective disclosure credentials with homomorphic and conventional encryption techniques to address the blockchain privacy paradox. The proposed framework introduces a multi-layered architecture consisting of identity and attribute management through selective disclosure, privacy-preserving computation using homomorphic encryption, and secure storage and transmission through conventional cryptographic methods. By integrating these complementary technologies, the framework enables fine-grained control over data disclosure while supporting confidential computation on encrypted information. The paper analyzes existing selective disclosure schemes, including Coconut and BBS+ credentials, alongside modern homomorphic encryption approaches and traditional encryption methods. It evaluates their performance characteristics, security properties, and suitability for blockchain environments. The study further outlines practical implementation patterns, discusses integration challenges, and identifies future research directions involving post-quantum cryptography, regulatory compliance, and performance optimization. The findings suggest that selective disclosure credentials combined with encryption-based privacy mechanisms can significantly improve confidentiality in blockchain applications while preserving the transparency and trust benefits that make blockchain systems valuable. This approach has potential applications in healthcare, finance, digital identity, supply chain management, and other privacy-sensitive domains.","author":[{"family":"Fernandes","given":"Aldrid"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20776080","URL":"https://doi.org/10.5281/zenodo.20776080","source":"datacite"},{"id":"doi:10.5281/zenodo.20776081","type":"article-journal","title":"Preserving Privacy on Blockchains: Combining Selective Disclosure and Homomorphic/Conventional Encryption.","abstract":"Blockchain technology offers transparency and auditability but faces significant privacy challenges when handling sensitive data. This paper proposes a hybrid privacy-preserving framework that combines selective disclosure credentials with homomorphic and conventional encryption techniques to address the blockchain privacy paradox. The proposed framework introduces a multi-layered architecture consisting of identity and attribute management through selective disclosure, privacy-preserving computation using homomorphic encryption, and secure storage and transmission through conventional cryptographic methods. By integrating these complementary technologies, the framework enables fine-grained control over data disclosure while supporting confidential computation on encrypted information. The paper analyzes existing selective disclosure schemes, including Coconut and BBS+ credentials, alongside modern homomorphic encryption approaches and traditional encryption methods. It evaluates their performance characteristics, security properties, and suitability for blockchain environments. The study further outlines practical implementation patterns, discusses integration challenges, and identifies future research directions involving post-quantum cryptography, regulatory compliance, and performance optimization. The findings suggest that selective disclosure credentials combined with encryption-based privacy mechanisms can significantly improve confidentiality in blockchain applications while preserving the transparency and trust benefits that make blockchain systems valuable. This approach has potential applications in healthcare, finance, digital identity, supply chain management, and other privacy-sensitive domains.","author":[{"family":"Fernandes","given":"Aldrid"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20776081","URL":"https://doi.org/10.5281/zenodo.20776081","source":"datacite"},{"id":"doi:10.5281/zenodo.21996949","type":"article-journal","title":"Project CHRONOS: A Fully Homomorphic Ephemeral AI Agent with Provable Self Termination and Remote Verifiability","abstract":"We present CHRONOS, the first autonomous AI agent that simultaneously achieves plaintextblindness (all data is processed under fully homomorphic encryption without ever beingexposed), cryptographically enforced time bound existence (the agent’s own decryption key islocked behind a publicly verifiable proof of sequential work, rendering it inaccessible until aprecise future moment), and remote verifiability of self destruction (a zero knowledge proofcertifies that the key material has been irreversibly destroyed after mission completion). Theagent’s operational lifespan is governed by a “cryptographic fuse” constructed from a proof ofsequential work (PoSW) whose computation time accurately matches the intended missionduration. A drand decentralized randomness beacon serves as a trusted time oracle to trigger thefinal key shredding. Crucially, the erasure proof is a non interactive zero knowledge argument(SNARK) that proves the correct execution of the entire self destruction sequence—including thePoSW solution, decryption of the private key, and subsequent memory zeroization—enablingany third party to cryptographically verify the agent’s annihilation without trusting the agent orits hardware. We provide a complete system architecture, a formal security model with gamebased definitions and reductions to standard assumptions, and a proof of concept implementationusing Zama’s TFHE rs for encrypted inference, a Cohen Pietrzak PoSW implementation, and aGroth16 SNARK. Our benchmarks indicate that FHE inference on a small neural network (50 Kparameters) completes in seconds, the PoSW background thread consumes negligible resources,and the erasure proof can be generated and verified in under three seconds. CHRONOSrepresents a fundamental advance in secure, disposable AI agents, with immediate applications indefense, intelligence, and high privacy environments.","author":[{"family":"Kumar","given":"Shashank"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21996949","URL":"https://doi.org/10.5281/zenodo.21996949","source":"datacite"},{"id":"doi:10.5281/zenodo.20782161","type":"article-journal","title":"Privacy Preserving Systems — API Gateway, AI Routing, Distributed Systems, Sovereign AI, and Post-Cloud Architecture (Api-Oss-Fixed)","abstract":"Privacy-enhancing technologies (PETs) form a critical component of modern computing systems, addressing the tension between data utility and individual privacy. This paper surveys the landscape of privacy-preserving techniques — from Cynthia Dwork's differential privacy framework to k-anonymity, homomorphic encryption, and secure multi-party computation — and examines their application within the 01s Sovereign (Kaiman) operating system. We demonstrate how the OS integrates these techniques to protect user data while maintaining the transparency required by its .aioss audit ledger. Part of The Anticloud research corpus by Lois-Kleinner Alpasan (ORCID: 0009-0009-2233-6107). This work explores api gateway, ai routing in the context of sovereign AI infrastructure, post-cloud computing architectures, and transparent, blackbox-free systems.","author":[{"family":"Alpasan","given":"Lois"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20782161","URL":"https://doi.org/10.5281/zenodo.20782161","source":"datacite"},{"id":"doi:10.5281/zenodo.20782162","type":"article-journal","title":"Privacy Preserving Systems — API Gateway, AI Routing, Distributed Systems, Sovereign AI, and Post-Cloud Architecture (Api-Oss-Fixed)","abstract":"Privacy-enhancing technologies (PETs) form a critical component of modern computing systems, addressing the tension between data utility and individual privacy. This paper surveys the landscape of privacy-preserving techniques — from Cynthia Dwork's differential privacy framework to k-anonymity, homomorphic encryption, and secure multi-party computation — and examines their application within the 01s Sovereign (Kaiman) operating system. We demonstrate how the OS integrates these techniques to protect user data while maintaining the transparency required by its .aioss audit ledger. Part of The Anticloud research corpus by Lois-Kleinner Alpasan (ORCID: 0009-0009-2233-6107). This work explores api gateway, ai routing in the context of sovereign AI infrastructure, post-cloud computing architectures, and transparent, blackbox-free systems.","author":[{"family":"Alpasan","given":"Lois"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20782162","URL":"https://doi.org/10.5281/zenodo.20782162","source":"datacite"},{"id":"doi:10.5281/zenodo.19635259","type":"article-journal","title":"The Zero-Knowledge Web Server (ZKWS): A Paradigm Shift in Stateless Infrastructure and Epistemic Privacy","abstract":"Abstract: This working paper introduces the Zero-Knowledge Web Server (ZKWS), a transformative infrastructure-layer protocol designed to eliminate the \"Oracle Fallacy\" inherent in modern web architectures. While current security standards like TLS/SSL protect data in transit, the host environment (RAM/CPU) remains a clear-text zone vulnerable to provider-level exfiltration and kernel exploits. ZKWS moves data \"blindness\" to the binary execution layer through a Triple-Blind Matrix: Homomorphic Routing (HR): Processing requests against encrypted routing tables. TEE-Isolation (Trusted Execution Environments): Executing logic within hardware-level enclaves (Intel SGX/AMD SEV) to prevent memory dumping. Epistemic Decoupling: Cryptographic separation of the database and web engine, where decryption occurs exclusively at the edge device. The paper further explores the integration of the Synthetic Data Contamination Index (SDCI) to mitigate recursive model collapse and addresses Cognitive Loop Burnout (CLB) by restoring absolute human agency and data sovereignty. ZKWS provides the architectural blueprint for a sovereign digital infrastructure, rendering host-level data breaches mathematically and economically obsolete in an increasingly synthetic AI era. Key Features: Attack surface reduction at the hardware-execution layer. Integration with SDCI for data provenance verification. Mitigation strategies for side-channel attacks through Constant-Time Obfuscation (CTO). Proposed \"Hybrid ZK-Sharding\" for performance optimization. Citation Note: This is a preliminary framework (Version 1.0). Part of the Bizbell Digital Ecosystem research initiative, building upon the principles of TruthSeal.pro and Vaultit.pro.","author":[{"family":"Siddiqui","given":"Jameel"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19635259","URL":"https://doi.org/10.5281/zenodo.19635259","source":"datacite"},{"id":"doi:10.5281/zenodo.19635260","type":"article-journal","title":"The Zero-Knowledge Web Server (ZKWS): A Paradigm Shift in Stateless Infrastructure and Epistemic Privacy","abstract":"Abstract: This working paper introduces the Zero-Knowledge Web Server (ZKWS), a transformative infrastructure-layer protocol designed to eliminate the \"Oracle Fallacy\" inherent in modern web architectures. While current security standards like TLS/SSL protect data in transit, the host environment (RAM/CPU) remains a clear-text zone vulnerable to provider-level exfiltration and kernel exploits. ZKWS moves data \"blindness\" to the binary execution layer through a Triple-Blind Matrix: Homomorphic Routing (HR): Processing requests against encrypted routing tables. TEE-Isolation (Trusted Execution Environments): Executing logic within hardware-level enclaves (Intel SGX/AMD SEV) to prevent memory dumping. Epistemic Decoupling: Cryptographic separation of the database and web engine, where decryption occurs exclusively at the edge device. The paper further explores the integration of the Synthetic Data Contamination Index (SDCI) to mitigate recursive model collapse and addresses Cognitive Loop Burnout (CLB) by restoring absolute human agency and data sovereignty. ZKWS provides the architectural blueprint for a sovereign digital infrastructure, rendering host-level data breaches mathematically and economically obsolete in an increasingly synthetic AI era. Key Features: Attack surface reduction at the hardware-execution layer. Integration with SDCI for data provenance verification. Mitigation strategies for side-channel attacks through Constant-Time Obfuscation (CTO). Proposed \"Hybrid ZK-Sharding\" for performance optimization. Citation Note: This is a preliminary framework (Version 1.0). Part of the Bizbell Digital Ecosystem research initiative, building upon the principles of TruthSeal.pro and Vaultit.pro.","author":[{"family":"Siddiqui","given":"Jameel"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.19635260","URL":"https://doi.org/10.5281/zenodo.19635260","source":"datacite"},{"id":"doi:10.5281/zenodo.16659484","type":"article-journal","title":"IoMT Encryption Simulation: Version 2.0.0","abstract":"Version 2.0.0 is the final archived release of the IoMT Encryption Simulation used for the doctoral dissertation: Examining Simulated Homomorphic Encryption on Data Transmissions of Always-Operating Internet of Medical Things This release contains the final Python experimental implementation, synthetic source data, measurement results, validation evidence, SPSS analysis files, and reproducibility documentation associated with the final study. Study Conditions The study evaluated 40 simulated always-operating Internet of Medical Things (IoMT) environments across three experimental conditions: - Unencrypted: ENV-01 through ENV-10- Simulated ECC: ENV-11 through ENV-25- RSA-SHE: ENV-26 through ENV-40 All nine protected transmission fields were included in the experimental workflows: - org_id- device_id- timestamp- heart_rate- bp_systolic- bp_diastolic- spo2- temperature- battery_level Unencrypted Condition The unencrypted condition served as the baseline. No encryption operation was applied to the nine protected fields. Protected values remained unchanged from the original synthetic source data. The condition produced: - Simulated encryption time: 0 seconds- Clear-text exposure: 100% Simulated ECC Implementation The simulated ECC condition uses: - ECDH with SECP384R1- HKDF-SHA256- 256-bit derived AES key- AES-256-GCM- Fresh 12-byte nonce for every protected-value encryption operation All nine protected fields are encrypted. Post-timing validation decrypts the encrypted protected values and confirms that they match the corresponding original source values. The simulated ECC condition produced 0% clear-text exposure. RSA-SHE Implementation RSA-SHE is the RSA-based simulated homomorphic encryption condition evaluated in the study. The final implementation uses: - 2048-bit RSA- RSA public exponent 65537- A new RSA key pair for each run- Run-specific encoding- RSA modular encryption of all nine protected fields- A randomized hybrid ciphertext layer- Fixed-width Base64 serialization- A predefined encrypted MAP-numerator calculation: SBP + 2(DBP) The systolic and diastolic blood-pressure operands remain encrypted during the predefined numerator calculation. Division by three to complete MAP occurs only after validation decryption. RSA-SHE is a controlled simulated-homomorphic adaptation for this predefined operation. It is not a production-grade Fully Homomorphic Encryption implementation, is not an exact reproduction of MEHE, and does not support arbitrary ciphertext computation. The RSA-SHE condition produced 0% clear-text exposure. Timing Encryption timing was measured using Python's time.perf_counter(). For each encrypted environment: - One untimed warm-up run was performed- Five timed runs were performed- The median of the five timed runs was retained No artificial timing delays or manually assigned encryption-time values were used. For RSA-SHE, the measured interval includes RSA key generation, run-specific homomorphic setup, encoding and RSA modular encryption of all nine protected fields, randomized hybrid ciphertext construction, the predefined encrypted MAP-numerator evaluation, and construction of the encrypted in-memory transmission structure. Correctness validation, validation decryption and unmasking, MAP verification, clear-text-exposure checks, audit-report generation, disk writing, file-size measurement, and average-row-length measurement occur outside the measured interval. Primary Outcomes The four primary study outcomes are: - File size (KB)- Average row length (bytes)- Simulated encryption time (seconds)- Clear-text exposure (percent) Final Descriptive Results Unencrypted — n = 10 - Mean file size: 61.64150 KB- Mean average row length: 61.02590 bytes- Mean encryption time: 0.00000000 seconds- Clear-text exposure: 100% Simulated ECC — n = 15 - Mean file size: 1825.35319 KB- Mean average row length: 428.00000 bytes- Mean encryption time: 0.06475018 seconds- Clear-text exposure: 0% RSA-SHE — n = 15 - Mean file size: 357","author":[{"family":"Anderson","given":"Devin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.16659484","URL":"https://doi.org/10.5281/zenodo.16659484","source":"datacite"},{"id":"doi:10.5281/zenodo.21858940","type":"article-journal","title":"IoMT Encryption Simulation: Version 2.0.0","abstract":"Version 2.0.0 is the final archived release of the IoMT Encryption Simulation used for the doctoral dissertation: Examining Simulated Homomorphic Encryption on Data Transmissions of Always-Operating Internet of Medical Things This release contains the final Python experimental implementation, synthetic source data, measurement results, validation evidence, SPSS analysis files, and reproducibility documentation associated with the final study. Study Conditions The study evaluated 40 simulated always-operating Internet of Medical Things (IoMT) environments across three experimental conditions: - Unencrypted: ENV-01 through ENV-10- Simulated ECC: ENV-11 through ENV-25- RSA-SHE: ENV-26 through ENV-40 All nine protected transmission fields were included in the experimental workflows: - org_id- device_id- timestamp- heart_rate- bp_systolic- bp_diastolic- spo2- temperature- battery_level Unencrypted Condition The unencrypted condition served as the baseline. No encryption operation was applied to the nine protected fields. Protected values remained unchanged from the original synthetic source data. The condition produced: - Simulated encryption time: 0 seconds- Clear-text exposure: 100% Simulated ECC Implementation The simulated ECC condition uses: - ECDH with SECP384R1- HKDF-SHA256- 256-bit derived AES key- AES-256-GCM- Fresh 12-byte nonce for every protected-value encryption operation All nine protected fields are encrypted. Post-timing validation decrypts the encrypted protected values and confirms that they match the corresponding original source values. The simulated ECC condition produced 0% clear-text exposure. RSA-SHE Implementation RSA-SHE is the RSA-based simulated homomorphic encryption condition evaluated in the study. The final implementation uses: - 2048-bit RSA- RSA public exponent 65537- A new RSA key pair for each run- Run-specific encoding- RSA modular encryption of all nine protected fields- A randomized hybrid ciphertext layer- Fixed-width Base64 serialization- A predefined encrypted MAP-numerator calculation: SBP + 2(DBP) The systolic and diastolic blood-pressure operands remain encrypted during the predefined numerator calculation. Division by three to complete MAP occurs only after validation decryption. RSA-SHE is a controlled simulated-homomorphic adaptation for this predefined operation. It is not a production-grade Fully Homomorphic Encryption implementation, is not an exact reproduction of MEHE, and does not support arbitrary ciphertext computation. The RSA-SHE condition produced 0% clear-text exposure. Timing Encryption timing was measured using Python's time.perf_counter(). For each encrypted environment: - One untimed warm-up run was performed- Five timed runs were performed- The median of the five timed runs was retained No artificial timing delays or manually assigned encryption-time values were used. For RSA-SHE, the measured interval includes RSA key generation, run-specific homomorphic setup, encoding and RSA modular encryption of all nine protected fields, randomized hybrid ciphertext construction, the predefined encrypted MAP-numerator evaluation, and construction of the encrypted in-memory transmission structure. Correctness validation, validation decryption and unmasking, MAP verification, clear-text-exposure checks, audit-report generation, disk writing, file-size measurement, and average-row-length measurement occur outside the measured interval. Primary Outcomes The four primary study outcomes are: - File size (KB)- Average row length (bytes)- Simulated encryption time (seconds)- Clear-text exposure (percent) Final Descriptive Results Unencrypted — n = 10 - Mean file size: 61.64150 KB- Mean average row length: 61.02590 bytes- Mean encryption time: 0.00000000 seconds- Clear-text exposure: 100% Simulated ECC — n = 15 - Mean file size: 1825.35319 KB- Mean average row length: 428.00000 bytes- Mean encryption time: 0.06475018 seconds- Clear-text exposure: 0% RSA-SHE — n = 15 - Mean file size: 357","author":[{"family":"Anderson","given":"Devin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21858940","URL":"https://doi.org/10.5281/zenodo.21858940","source":"datacite"},{"id":"doi:10.5281/zenodo.21617873","type":"article-journal","title":"Fully Homomorphic Behavioral Biometric Matching with Post-Quantum Lattices","abstract":"This Master’s thesis presents the design, implementation and evaluation of a lightweight privacy-preserving behavioral biometric authentication framework. The proposed system uses press-to-press keystroke timing features and the CKKS homomorphic encryption scheme implemented using OpenFHE. The enrolled biometric template and login feature vector are compared in the encrypted domain through squared Euclidean distance computation, while only the final matching score is decrypted for threshold-based authentication. A complete experimental system was developed using JavaScript, Python, Flask, NumPy and OpenFHE. The implementation demonstrates enrollment, genuine-user authentication, simulated impostor rejection and encrypted biometric matching. The system was evaluated using feature dimensionality, encrypted operation count, execution time, CPU usage and memory consumption. The research demonstrates that low-dimensional behavioral biometric features can reduce the computational burden associated with homomorphic encryption while protecting biometric templates during matching. The lattice-based foundation of CKKS also provides a post-quantum-oriented approach to protecting long-lived biometric information. This thesis was submitted to the Sri Lanka Institute of Information Technology in partial fulfilment of the requirements for the degree of Master of Science in Information Technology Specializing in Cyber Security.","author":[{"family":"Mahaarachchi","given":"Nipun"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21617873","URL":"https://doi.org/10.5281/zenodo.21617873","source":"datacite"},{"id":"doi:10.5281/zenodo.21617874","type":"article-journal","title":"Fully Homomorphic Behavioral Biometric Matching with Post-Quantum Lattices","abstract":"This Master’s thesis presents the design, implementation and evaluation of a lightweight privacy-preserving behavioral biometric authentication framework. The proposed system uses press-to-press keystroke timing features and the CKKS homomorphic encryption scheme implemented using OpenFHE. The enrolled biometric template and login feature vector are compared in the encrypted domain through squared Euclidean distance computation, while only the final matching score is decrypted for threshold-based authentication. A complete experimental system was developed using JavaScript, Python, Flask, NumPy and OpenFHE. The implementation demonstrates enrollment, genuine-user authentication, simulated impostor rejection and encrypted biometric matching. The system was evaluated using feature dimensionality, encrypted operation count, execution time, CPU usage and memory consumption. The research demonstrates that low-dimensional behavioral biometric features can reduce the computational burden associated with homomorphic encryption while protecting biometric templates during matching. The lattice-based foundation of CKKS also provides a post-quantum-oriented approach to protecting long-lived biometric information. This thesis was submitted to the Sri Lanka Institute of Information Technology in partial fulfilment of the requirements for the degree of Master of Science in Information Technology Specializing in Cyber Security.","author":[{"family":"Mahaarachchi","given":"Nipun"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21617874","URL":"https://doi.org/10.5281/zenodo.21617874","source":"datacite"},{"id":"doi:10.5281/zenodo.20593983","type":"article-journal","title":"PRIVACY-PRESERVING TECHNOLOGIES FOR CASHLESS FINANCIAL ECOSYSTEMS","abstract":"This paper provides an overview of privacy-protecting measures that can be used to secure user data and assure the safety and efficiency of digital activities in contactless financial ecosystems. Many people worry about identity theft, data breaches, and spying by unauthorised parties due to the rapid growth of digital wallets, contactless banking, and mobile payments. Modern cryptography includes safe multi-party computation, zero-knowledge proofs, and homomorphic encryption. These approaches verify transactions and safeguard sensitive data. Blockchain and other independent systems are emphasized for their ability to improve openness, reliability, and anonymity. Regulations and compliance challenges related to financial systems using privacy-enhancing technology are examined. The findings emphasize the importance of strong privacy protections to balance data security, safety, and creativity. Contactless technologies become more popular as more people believe in them.","author":[{"family":"Transformation","given":"Emerging"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20593983","URL":"https://doi.org/10.5281/zenodo.20593983","source":"datacite"},{"id":"doi:10.5281/zenodo.20593984","type":"article-journal","title":"PRIVACY-PRESERVING TECHNOLOGIES FOR CASHLESS FINANCIAL ECOSYSTEMS","abstract":"This paper provides an overview of privacy-protecting measures that can be used to secure user data and assure the safety and efficiency of digital activities in contactless financial ecosystems. Many people worry about identity theft, data breaches, and spying by unauthorised parties due to the rapid growth of digital wallets, contactless banking, and mobile payments. Modern cryptography includes safe multi-party computation, zero-knowledge proofs, and homomorphic encryption. These approaches verify transactions and safeguard sensitive data. Blockchain and other independent systems are emphasized for their ability to improve openness, reliability, and anonymity. Regulations and compliance challenges related to financial systems using privacy-enhancing technology are examined. The findings emphasize the importance of strong privacy protections to balance data security, safety, and creativity. Contactless technologies become more popular as more people believe in them.","author":[{"family":"Transformation","given":"Emerging"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20593984","URL":"https://doi.org/10.5281/zenodo.20593984","source":"datacite"},{"id":"doi:10.5281/zenodo.21947348","type":"article-journal","title":"The HeartBank Longitudinal Cohort: A Dataset Combining DNA, Natal Chart, Family Tree, Continuous Behavioral Observation, and Continuous Respiratory Observation at Civilizational Scale","abstract":"This paper specifies the methodology for the HeartBank Longitudinal Cohort, a voluntary opt-in dataset combining five data layers per consenting participant — (1) DNA sequence, (2) natal chart data (date + time + place of birth), (3) continuous longitudinal behavioral observation via the HeartBank gratitude ledger, (4) continuous longitudinal respiratory observation via the breath-class Mechanical Heart wearable, and (5) verified kinship data via the global family tree — at a target scale of 100 million+ participants over multi-decade time horizons. The combination has never been assembled at scale; comparable datasets (23andMe, AncestryDNA, Worldcoin, Dunedin and BCS longitudinal cohorts, social-network behavioral data, professional astrological collections) carry one or two of the layers each but no prior project has carried all five. The methodology specifies: the opt-in informed-consent architecture; the privacy-preserving computation stack (differential privacy at the analysis layer; federated computation with homomorphic encryption for DNA; on-device processing for breath signals; cryptographic-erasure right-to-withdraw); the institutional-review architecture (IRB-grade ethics oversight; Buddhist-ethics-aware review board; pre-registered hypotheses); the cosmic-coordinate-correlation epistemic posture (natal chart treated as a unique cosmic-moment coordinate, not as a cosmic force; the research question is correlation between coordinate features and trajectory features, not validation of astrology); the publication architecture (open methodology, closed individual data); the data-sovereignty architecture (jurisdictional residency; GINA / HIPAA / GDPR compliance baselines exceeded where possible); and the new academic alliances the cohort makes possible (Mind & Life Institute; contemplative-science programs at Stanford, Brown, UMass; behavioral-genetics consortia; longitudinal-cohort consortia; Buddhist-AI ethicists). Three scientifically valuable outcomes are honestly named: no detected correlation, small-but-real correlation, substantial correlation — each is a major contribution to knowledge regardless of direction. Honest §11 names what the cohort does not claim and the non-negotiable privacy disciplines the architecture requires. Keywords: longitudinal cohort methodology, cosmic-coordinate correlation, contemplative science, differential privacy, federated computation, multi-omic dataset, gratitude behavior, respiratory biomarkers, defensive publication, Mind & Life partnerships. --- Provenance. This paper is part of the THonly research corpus, dedicated to the public domain under CC0 1.0. The canonical version is at https://thonly.org/research/longitudinal-cohort-methodology. Its SHA-256 is b19fdcd0773da2c4f2b7ba7472b429b4bfd459e285b98c4900f15a04207fce42, independently timestamped to the Bitcoin blockchain via OpenTimestamps and signed under RFC 3161 by three trust authorities, one of them eIDAS-qualified. AI co-authorship is disclosed. Miss Aquarius is the consistent name used for the AI collaboration across all venues.","author":[{"family":"Ly","given":"Thon"},{"family":"Aquarius","given":"Miss"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21947348","URL":"https://doi.org/10.5281/zenodo.21947348","source":"datacite"},{"id":"doi:10.5281/zenodo.21821704","type":"article-journal","title":"Confidential Post-Quantum Settlement: Blind ML-DSA-65 Verification for Institutional Financial Pipelines","abstract":"Abstract Banking and mainstream finance are absorbing public-chain settlement patterns, stablecoin rails, and tokenized assets at the same time that regulators require post-quantum signatures and strict privacy on high-value flows. FIPS 204 ML-DSA-65 is the natural audit-grade signature. In custody, institutional settlement, and confidential L1 paths, that signature often sits inside an encrypted workflow. Checking it by first decrypting restores plaintext at the verifier and expands the set of systems that observe protected fields, or else requires a trusted enclave. Where institutions already keep settlement material under encryption, that forces an awkward choice between visibility and verification. Integrity proofs such as zk-STARKs excel at attesting that large batches or audit logs followed the rules. They do not, by themselves, answer whether a specific ML-DSA-65 signature is valid while the surrounding material remains under encryption. That gap is the subject of this work: a field-deployable hybrid that evaluates the linear core of ML-DSA-65 verification under leveled fully homomorphic encryption (FHE), completes verification on CPU, and keeps the protected payload out of plaintext at the verification step. Signature generation stays outside the encrypted path. Across 10 000 residual trials, infinity-norm noise stays inside a strict budget of 65 536 (observed maximum 51 489). Hybrid verification agrees with a reference FIPS 204 oracle at 100 % on 10 000 trials. End-to-end residual latency averages 2.54 ms on commodity 12th-generation Intel hardware. Because confidential paths remain attractive targets for timing and related leakage, a side-channel triad is part of the contribution: maximum absolute TVLA |t| = 1.13 (threshold 4.5), with negligible mutual information. Full verification entirely in ciphertext, and any form of signature generation under FHE, are not claimed. The contribution is a measurable building block for confidential, post-quantum, regulatory-aligned verification of ML-DSA-65—eliminating routine plaintext exposure at verify time and supporting side-channel discipline—where banking-grade privacy and blockchain-grade settlement meet. Keywords: Post-quantum cryptography, ML-DSA, FIPS 204, confidential computing, hybrid verification, leveled FHE, blockchain settlement, institutional custody, regulatory compliance, side-channel assessment. Dedicated to: In memory of Prof. Myung Kyoon “Michael” Chung (1945–2025), who dedicated his life to bringing truth and science to our world. A graduate of Seoul National University, Washington State University (Pullman), and the University of Illinois, and lifelong Professor at the Korea Advanced Institute of Science and Technology (KAIST), Department of Mechanical Engineering.","author":[{"family":"Chung","given":"Jinhyuk"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21821704","URL":"https://doi.org/10.5281/zenodo.21821704","source":"datacite"},{"id":"doi:10.5281/zenodo.21821703","type":"article-journal","title":"Confidential Post-Quantum Settlement: Blind ML-DSA-65 Verification for Institutional Financial Pipelines","abstract":"Abstract Banking and mainstream finance are absorbing public-chain settlement patterns, stablecoin rails, and tokenized assets at the same time that regulators require post-quantum signatures and strict privacy on high-value flows. FIPS 204 ML-DSA-65 is the natural audit-grade signature. In custody, institutional settlement, and confidential L1 paths, that signature often sits inside an encrypted workflow. Checking it by first decrypting restores plaintext at the verifier and expands the set of systems that observe protected fields, or else requires a trusted enclave. Where institutions already keep settlement material under encryption, that forces an awkward choice between visibility and verification. Integrity proofs such as zk-STARKs excel at attesting that large batches or audit logs followed the rules. They do not, by themselves, answer whether a specific ML-DSA-65 signature is valid while the surrounding material remains under encryption. That gap is the subject of this work: a field-deployable hybrid that evaluates the linear core of ML-DSA-65 verification under leveled fully homomorphic encryption (FHE), completes verification on CPU, and keeps the protected payload out of plaintext at the verification step. Signature generation stays outside the encrypted path. Across 10 000 residual trials, infinity-norm noise stays inside a strict budget of 65 536 (observed maximum 51 489). Hybrid verification agrees with a reference FIPS 204 oracle at 100 % on 10 000 trials. End-to-end residual latency averages 2.54 ms on commodity 12th-generation Intel hardware. Because confidential paths remain attractive targets for timing and related leakage, a side-channel triad is part of the contribution: maximum absolute TVLA |t| = 1.13 (threshold 4.5), with negligible mutual information. Full verification entirely in ciphertext, and any form of signature generation under FHE, are not claimed. The contribution is a measurable building block for confidential, post-quantum, regulatory-aligned verification of ML-DSA-65—eliminating routine plaintext exposure at verify time and supporting side-channel discipline—where banking-grade privacy and blockchain-grade settlement meet. Keywords: Post-quantum cryptography, ML-DSA, FIPS 204, confidential computing, hybrid verification, leveled FHE, blockchain settlement, institutional custody, regulatory compliance, side-channel assessment. Dedicated to: In memory of Prof. Myung Kyoon “Michael” Chung (1945–2025), who dedicated his life to bringing truth and science to our world. A graduate of Seoul National University, Washington State University (Pullman), and the University of Illinois, and lifelong Professor at the Korea Advanced Institute of Science and Technology (KAIST), Department of Mechanical Engineering.","author":[{"family":"Chung","given":"Jinhyuk"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21821703","URL":"https://doi.org/10.5281/zenodo.21821703","source":"datacite"},{"id":"doi:10.5281/zenodo.21719064","type":"article-journal","title":"KLTu-Native: Unified Kinetic Lattice Topology for KEM, DSA, and FHE-Capable Lattice Operation with Shared Timing Isolation under MAXR3","abstract":"Abstract Post-quantum migration is no longer optional: banks, cloud operators, and telecom networks must field lattice key exchange and signatures at scale, while keeping a credible path to lattice fully homomorphic encryption without rebuilding their crypto stack twice. Separate Kyber-, Dilithium-, and FHE product lines multiply parameters, entropy sources, and regression surfaces. We present KLTu-native, a unified lattice engine in which KEM, DSA, and FHE-oriented roles share one topology, control plane, and thermodynamic entropy architecture. The systems claim is organizational: one roof for key establishment, signatures, and structural FHE readiness, with a shared software-visible timing-isolation posture under production-like host features. On bare-metal client silicon with simultaneous multithreading (SMT) and active power management (APM) left enabled, we show that protected native Level 5 KEM and DSA operations pass Fixed-versus-Random Kolmogorov–Smirnov tests on operation latency (D<0.05, zero drops)—an isolation grade measured under the host features operators actually run, not under disabled energy-saving or hyperthreading modes. Level 3 is positioned as the deployment tier (tight SLA tails, usable multi-thread verify scaling); Level 5 is reported honestly as supply-limited under strict mid-operation entropy policy. Replicate RAPL statistics further show a statistically significant Level 5 signing energy-density advantage versus a pure community baseline, alongside an explicit package-power cost attributable to live isolation. Supporting comparators situate efficiency only. FHE bootstrap performance and invasive physical side-channel evaluation are out of scope. A breakthrough disclosure is that KLTu demonstrates bidirectional translation of NIST FIPS 203 ML-KEM (768 and 1024) between standard wire encodings and KLTu-native lattice structures by pack/unpack alone—without invoking decapsulation or deriving a shared secret during conversion. That removes the practical barrier of running a second KEM stack solely for standards compliance. Because the same Kinetic Lattice Topology is FHE-capable, sessions established from ordinary FIPS-KEM wire material can enter efficient KLTu-FHE evaluation under one engine, rather than across fragmented KEM and FHE product lines. For operators, the decision is whether a single, measurable, constant-time-oriented stack—Level-3-first, Level-5-honest, FIPS-wire interoperable and FHE-ready by structure—is preferable to maintaining parallel KEM, DSA, and future FHE lines with divergent latency, energy, and leakage profiles. Keywords: post-quantum cryptography; unified lattice stack; ML-KEM; ML-DSA; timing isolation; Kolmogorov–Smirnov; energy efficiency; FHE-ready design; FIPS 203 interperability; ML-KEM wire conversion Dedicated to: In memory of Prof. Myung Kyoon “Michael” CHUNG (1945-2025), who dedicated his life to bringing truth and science to our world. A graduate of Seoul National University, Washington State University (Pullman), University of Illinois, and lifelong Professor at Korea Advanced Institute of Technology (KAIST) with the Department of Mechanical Engineering.","author":[{"family":"Chung","given":"Jinhyuk"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21719064","URL":"https://doi.org/10.5281/zenodo.21719064","source":"datacite"},{"id":"doi:10.5281/zenodo.21543529","type":"article-journal","title":"Efficient Fully Homomorphic Evaluation of FIPS 203 ML-KEM Sessions via Native Wire Conversion and Measured Noise-Control Options α, β, θ, δ without Functional Bootstrapping","abstract":"Abstract Post-quantum key establishment under NIST FIPS 203 (ML-KEM) is entering production, while fully homomorphic encryption (FHE) remains constrained by the cost of functional programmable bootstrapping (PBS). Bridging the two usually forces either a heavyweight transciphering circuit that homomorphically evaluates decapsulation, or a proprietary key format that breaks compliance inventories. This paper reports two measured results that avoid both traps. First, public keys, secret keys, and ciphertexts in ML-KEM-768 (Level 3) and ML-KEM-1024 (Level 5) convert bidirectionally between FIPS 203 wire bytes and KLTu-native polynomial layouts by coefficient pack and unpack only. Conversion does not invoke decapsulation and does not produce shared secrets; when decapsulation is required, the FIPS-facing API is called explicitly and matches 45/45 official ACVP vectors per level. Second, four discrete noise-control options—denoted α, β, θ, and δ—are timed against production full PBS under Level 5 parameters (ring dimension N=32768). On a single pinned CPU core, option medians range from 42.18 µs (α) to 35.13 ms (β), while production PBS median is 15.55 s. A 50 000-step mixed schedule (70% α, 15% β, 10% θ, 5% δ) completes with zero full-PBS calls and a projected speedup of approximately 2 670× relative to an all-PBS baseline. A depth-ladder campaign further shows that α–δ alone sustain mult-depth homomorphic multiplication through depth 8 (Level 3) and depth 6 (Level 5) with decrypt match rate 1.0 and zero instrumented full-PBS calls; options are not claimed to be necessary under a strict ablation, and they are not claimed to implement programmable LUT bootstrapping. Homomorphic addition and multiplication remain exact (100% match, n=1000 per level) when the evaluator holds no secret key. Kolmogorov–Smirnov tests on conversion and evaluation surfaces are reported with FAIL gates enforced; conversion FAIL does not imply exposure of secrets or plaintext. Option algorithms are not disclosed (Level A). Together, standard FIPS 203 sessions and practical FHE evaluation coexist under one native engine without routine functional bootstrapping on the measured workloads. Dedicated to: In memory of Prof. Myung Kyoon “Michael” CHUNG (1945-2025), who dedicated his life to bringing truth and science to our world. A graduate of Seoul National University, Washington State University (Pullman), University of Illinois, and lifelong Professor at Korea Advanced Institute of Technology (KAIST) with the Department of Mechanical Engineering. 1. Introduction Post-quantum key establishment under FIPS 203 (ML-KEM) is now a deployed standard. Fully homomorphic encryption (FHE) remains limited in practice by the cost of functional bootstrapping when noise must be reset after multiplicative depth. Bridging the two worlds—evaluating on data whose session keys were established under FIPS 203—typically forces either (i) a heavyweight transciphering circuit that homomorphically evaluates decapsulation, or (ii) a redesign that abandons standard wire formats. Prior work in this series established the MAXR3 entropy substrate [1] and the KLTu-native KEM/DSA stack with FIPS 203 wire interoperability [2]. This work shows a different path, extending [2]. FIPS 203 public keys, secret keys, and ciphertexts can be mapped into a native FHE execution layout by coefficient pack/unpack alone. No shared secret is recovered at the boundary. Separately, four measured noise-control options (α, β, θ, δ) keep production multiplicative workloads free of full PBS on the sealed campaigns, while preserving exact homomorphic arithmetic. A depth-ladder evaluation (Path O) further shows mult-depth multiplication through d=8(Level 3) and d=6(Level 5) with zero instrumented full-PBS calls and decrypt match rate 1.0. The practical consequence is non-blocking noise control: multiplicative depth can advance without suspending the workload for a production full-PBS call. Together, standard PQC est","author":[{"family":"Chung","given":"Jinhyuk"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21543529","URL":"https://doi.org/10.5281/zenodo.21543529","source":"datacite"},{"id":"doi:10.5281/zenodo.21543528","type":"article-journal","title":"Efficient Fully Homomorphic Evaluation of FIPS 203 ML-KEM Sessions via Native Wire Conversion and Measured Noise-Control Options α, β, θ, δ without Functional Bootstrapping","abstract":"Abstract Post-quantum key establishment under NIST FIPS 203 (ML-KEM) is entering production, while fully homomorphic encryption (FHE) remains constrained by the cost of functional programmable bootstrapping (PBS). Bridging the two usually forces either a heavyweight transciphering circuit that homomorphically evaluates decapsulation, or a proprietary key format that breaks compliance inventories. This paper reports two measured results that avoid both traps. First, public keys, secret keys, and ciphertexts in ML-KEM-768 (Level 3) and ML-KEM-1024 (Level 5) convert bidirectionally between FIPS 203 wire bytes and KLTu-native polynomial layouts by coefficient pack and unpack only. Conversion does not invoke decapsulation and does not produce shared secrets; when decapsulation is required, the FIPS-facing API is called explicitly and matches 45/45 official ACVP vectors per level. Second, four discrete noise-control options—denoted α, β, θ, and δ—are timed against production full PBS under Level 5 parameters (ring dimension N=32768). On a single pinned CPU core, option medians range from 42.18 µs (α) to 35.13 ms (β), while production PBS median is 15.55 s. A 50 000-step mixed schedule (70% α, 15% β, 10% θ, 5% δ) completes with zero full-PBS calls and a projected speedup of approximately 2 670× relative to an all-PBS baseline. A depth-ladder campaign further shows that α–δ alone sustain mult-depth homomorphic multiplication through depth 8 (Level 3) and depth 6 (Level 5) with decrypt match rate 1.0 and zero instrumented full-PBS calls; options are not claimed to be necessary under a strict ablation, and they are not claimed to implement programmable LUT bootstrapping. Homomorphic addition and multiplication remain exact (100% match, n=1000 per level) when the evaluator holds no secret key. Kolmogorov–Smirnov tests on conversion and evaluation surfaces are reported with FAIL gates enforced; conversion FAIL does not imply exposure of secrets or plaintext. Option algorithms are not disclosed (Level A). Together, standard FIPS 203 sessions and practical FHE evaluation coexist under one native engine without routine functional bootstrapping on the measured workloads. Dedicated to: In memory of Prof. Myung Kyoon “Michael” CHUNG (1945-2025), who dedicated his life to bringing truth and science to our world. A graduate of Seoul National University, Washington State University (Pullman), University of Illinois, and lifelong Professor at Korea Advanced Institute of Technology (KAIST) with the Department of Mechanical Engineering. 1. Introduction Post-quantum key establishment under FIPS 203 (ML-KEM) is now a deployed standard. Fully homomorphic encryption (FHE) remains limited in practice by the cost of functional bootstrapping when noise must be reset after multiplicative depth. Bridging the two worlds—evaluating on data whose session keys were established under FIPS 203—typically forces either (i) a heavyweight transciphering circuit that homomorphically evaluates decapsulation, or (ii) a redesign that abandons standard wire formats. Prior work in this series established the MAXR3 entropy substrate [1] and the KLTu-native KEM/DSA stack with FIPS 203 wire interoperability [2]. This work shows a different path, extending [2]. FIPS 203 public keys, secret keys, and ciphertexts can be mapped into a native FHE execution layout by coefficient pack/unpack alone. No shared secret is recovered at the boundary. Separately, four measured noise-control options (α, β, θ, δ) keep production multiplicative workloads free of full PBS on the sealed campaigns, while preserving exact homomorphic arithmetic. A depth-ladder evaluation (Path O) further shows mult-depth multiplication through d=8(Level 3) and d=6(Level 5) with zero instrumented full-PBS calls and decrypt match rate 1.0. The practical consequence is non-blocking noise control: multiplicative depth can advance without suspending the workload for a production full-PBS call. Together, standard PQC est","author":[{"family":"Chung","given":"Jinhyuk"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21543528","URL":"https://doi.org/10.5281/zenodo.21543528","source":"datacite"},{"id":"doi:10.5281/zenodo.21523834","type":"article-journal","title":"KLTu-Native: Unified Kinetic Lattice Topology for KEM, DSA, and FHE-Capable Lattice Operation with Shared Timing Isolation under MAXR3","abstract":"Abstract Post-quantum migration is no longer optional: banks, cloud operators, and telecom networks must field lattice key exchange and signatures at scale, while keeping a credible path to lattice fully homomorphic encryption without rebuilding their crypto stack twice. Separate Kyber-, Dilithium-, and FHE product lines multiply parameters, entropy sources, and regression surfaces. We present KLTu-native, a unified lattice engine in which KEM, DSA, and FHE-oriented roles share one topology, control plane, and thermodynamic entropy architecture. The systems claim is organizational: one roof for key establishment, signatures, and structural FHE readiness, with a shared software-visible timing-isolation posture under production-like host features. On bare-metal client silicon with simultaneous multithreading (SMT) and active power management (APM) left enabled, we show that protected native Level 5 KEM and DSA operations pass Fixed-versus-Random Kolmogorov–Smirnov tests on operation latency (D<0.05, zero drops)—an isolation grade measured under the host features operators actually run, not under disabled energy-saving or hyperthreading modes. Level 3 is positioned as the deployment tier (tight SLA tails, usable multi-thread verify scaling); Level 5 is reported honestly as supply-limited under strict mid-operation entropy policy. Replicate RAPL statistics further show a statistically significant Level 5 signing energy-density advantage versus a pure community baseline, alongside an explicit package-power cost attributable to live isolation. Supporting comparators situate efficiency only. FHE bootstrap performance and invasive physical side-channel evaluation are out of scope. A breakthrough disclosure is that KLTu demonstrates bidirectional translation of NIST FIPS 203 ML-KEM (768 and 1024) between standard wire encodings and KLTu-native lattice structures by pack/unpack alone—without invoking decapsulation or deriving a shared secret during conversion. That removes the practical barrier of running a second KEM stack solely for standards compliance. Because the same Kinetic Lattice Topology is FHE-capable, sessions established from ordinary FIPS-KEM wire material can enter efficient KLTu-FHE evaluation under one engine, rather than across fragmented KEM and FHE product lines. For operators, the decision is whether a single, measurable, constant-time-oriented stack—Level-3-first, Level-5-honest, FIPS-wire interoperable and FHE-ready by structure—is preferable to maintaining parallel KEM, DSA, and future FHE lines with divergent latency, energy, and leakage profiles. Keywords: post-quantum cryptography; unified lattice stack; ML-KEM; ML-DSA; timing isolation; Kolmogorov–Smirnov; energy efficiency; FHE-ready design; FIPS 203 interperability; ML-KEM wire conversion Dedicated to: In memory of Prof. Myung Kyoon “Michael” CHUNG (1945-2025), who dedicated his life to bringing truth and science to our world. A graduate of Seoul National University, Washington State University (Pullman), University of Illinois, and lifelong Professor at Korea Advanced Institute of Technology (KAIST) with the Department of Mechanical Engineering.","author":[{"family":"Chung","given":"Jinhyuk"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21523834","URL":"https://doi.org/10.5281/zenodo.21523834","source":"datacite"},{"id":"doi:10.48550/arxiv.2512.01974","type":"manuscript","title":"The Equivalence of Fast Algorithms for Convolution, Parallel FIR Filters, Polynomial Modular Multiplication, and Pointwise Multiplication in DFT/NTT Domain","abstract":"Fast time-domain algorithms have been developed in signal processing applications to reduce the multiplication complexity. For example, fast convolution structures using Cook-Toom and Winograd algorithms are well understood. Short length fast convolutions can be iterated to obtain fast convolution structures for long lengths. In this paper, we show that well known fast convolution structures form the basis for design of fast algorithms in four other problem domains: fast parallel filters, fast polynomial modular multiplication, and fast pointwise multiplication in the DFT and NTT domains. Fast polynomial modular multiplication and fast pointwise multiplication problems are important for cryptosystem applications such as post-quantum cryptography and homomorphic encryption. By establishing the equivalence of these problems, we show that a fast structure from one domain can be used to design a fast structure for another domain. This understanding is important as there are many well known solutions for fast convolution that can be used in other signal processing and cryptosystem applications.","author":[{"family":"Parhi","given":"Keshab"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2512.01974","URL":"https://doi.org/10.48550/arxiv.2512.01974","source":"datacite"},{"id":"doi:10.5281/zenodo.18930888","type":"article-journal","title":"Y.I.N. Governance Framework: The Operating System for Cryptographically Enforceable AI Governance","abstract":"The Y.I.N. Governance Framework is a comprehensive 15-domain policy integration system that transforms fragmented AI governance requirements into a unified operational architecture. Unlike existing frameworks that organize compliance checklists, the Y.I.N. Governance Framework is specifically designed to be cryptographically enforceable through the 26-layer Y.I.N. Mazari Architecture. This framework addresses the critical gap identified by the OECD Responsible AI Due Diligence Guidance (2026): organizations face over 100 overlapping governance regimes with no systematic method to integrate and enforce them simultaneously. The Y.I.N. Governance Framework integrates the EU AI Act, ISO/IEC 42001:2023, OECD AI Principles, NIST AI Risk Management Framework, G7 Hiroshima AI Process Code of Conduct, IEEE 7000-2021, UN Guiding Principles on Business and Human Rights, GDPR, EU DORA, NIS2, HIPAA, NY Senate Bill S.7263, and over 50 additional regulatory frameworks worldwide. Key Innovation: Each policy requirement in the framework maps directly to cryptographic enforcement mechanisms in the Y.I.N. Mazari Architecture, creating the world's first governance system where compliance is mathematically provable, not procedurally documented. The framework comprises 15 integrated domains: (1) Regulatory Compliance, (2) Risk Classification & Management, (3) Privacy & Data Protection, (4) Security & Resilience, (5) Transparency & Explainability, (6) Human Oversight & Accountability, (7) Bias & Fairness, (8) Safety & Reliability, (9) Data Governance, (10) Model Governance, (11) Ethical Principles, (12) Professional Practice, (13) Incident Response & Remediation, (14) Third-Party & Supply Chain, (15) Continuous Monitoring & Improvement. Each domain maps to specific layers of the Y.I.N. Mazari Architecture for cryptographic enforcement through differential privacy, zero-knowledge proofs, homomorphic encryption, hardware-enforced finite state machines, and blockchain-anchored audit trails. This publication establishes the complete Y.I.N. governance solution: Framework (policy layer) + Architecture (cryptographic enforcement layer).","author":[{"family":"Mazari","given":"Ilyes"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18930888","URL":"https://doi.org/10.5281/zenodo.18930888","source":"datacite"},{"id":"doi:10.5281/zenodo.18930889","type":"article-journal","title":"Y.I.N. Governance Framework: The Operating System for Cryptographically Enforceable AI Governance","abstract":"The Y.I.N. Governance Framework is a comprehensive 15-domain policy integration system that transforms fragmented AI governance requirements into a unified operational architecture. Unlike existing frameworks that organize compliance checklists, the Y.I.N. Governance Framework is specifically designed to be cryptographically enforceable through the 26-layer Y.I.N. Mazari Architecture. This framework addresses the critical gap identified by the OECD Responsible AI Due Diligence Guidance (2026): organizations face over 100 overlapping governance regimes with no systematic method to integrate and enforce them simultaneously. The Y.I.N. Governance Framework integrates the EU AI Act, ISO/IEC 42001:2023, OECD AI Principles, NIST AI Risk Management Framework, G7 Hiroshima AI Process Code of Conduct, IEEE 7000-2021, UN Guiding Principles on Business and Human Rights, GDPR, EU DORA, NIS2, HIPAA, NY Senate Bill S.7263, and over 50 additional regulatory frameworks worldwide. Key Innovation: Each policy requirement in the framework maps directly to cryptographic enforcement mechanisms in the Y.I.N. Mazari Architecture, creating the world's first governance system where compliance is mathematically provable, not procedurally documented. The framework comprises 15 integrated domains: (1) Regulatory Compliance, (2) Risk Classification & Management, (3) Privacy & Data Protection, (4) Security & Resilience, (5) Transparency & Explainability, (6) Human Oversight & Accountability, (7) Bias & Fairness, (8) Safety & Reliability, (9) Data Governance, (10) Model Governance, (11) Ethical Principles, (12) Professional Practice, (13) Incident Response & Remediation, (14) Third-Party & Supply Chain, (15) Continuous Monitoring & Improvement. Each domain maps to specific layers of the Y.I.N. Mazari Architecture for cryptographic enforcement through differential privacy, zero-knowledge proofs, homomorphic encryption, hardware-enforced finite state machines, and blockchain-anchored audit trails. This publication establishes the complete Y.I.N. governance solution: Framework (policy layer) + Architecture (cryptographic enforcement layer).","author":[{"family":"Mazari","given":"Ilyes"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18930889","URL":"https://doi.org/10.5281/zenodo.18930889","source":"datacite"},{"id":"doi:10.5281/zenodo.18244672","type":"article-journal","title":"The AI Governance Crisis and Privacy-Preserving Computation: A Technical Analysis of Regulatory Compliance Solutions","abstract":"The year 2025 marked the transition from AI ethics debate to AI governance execution. Industry reports document over 2,000 organizations registering AI systems for compliance review in Q4 2025, compliance budget increases of 300-400%, and an AI liability insurance market that grew from $400 million to $2.1 billion. Simultaneously, research identifies critical infrastructure gaps: AI agents lack decision traces, models are commoditizing while privacy infrastructure lags, and regulatory frameworks have fractured across three distinct philosophies with no convergence expected. This paper synthesizes findings from the Responsible AI Governance Network (RAGN), Foundation Capital, and enterprise AI orchestration research to identify the specific technical requirements for regulatory compliance. It then presents the Y.I.N. (Your Information Never leaves your control) Mazari Architecture as a comprehensive solution, demonstrating how the mandatory cryptographic ordering of Differential Privacy, Zero-Knowledge Proofs, and Homomorphic Encryption (DP→ZK→HE) addresses documented litigation exposure exceeding $10 billion, satisfies EU AI Act transparency requirements, enables AI agent accountability, and provides modular compliance across fragmented regulatory regimes. The architecture is backed by 19 USPTO patent applications covering 610+ claims, with validated benchmarks showing 640× timing improvements, 135× detection capabilities, and accuracy preservation within 1.5 percentage points.","author":[{"family":"Mazari","given":"Ilyes"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18244672","URL":"https://doi.org/10.5281/zenodo.18244672","source":"datacite"},{"id":"doi:10.5281/zenodo.18244673","type":"article-journal","title":"The AI Governance Crisis and Privacy-Preserving Computation: A Technical Analysis of Regulatory Compliance Solutions","abstract":"The year 2025 marked the transition from AI ethics debate to AI governance execution. Industry reports document over 2,000 organizations registering AI systems for compliance review in Q4 2025, compliance budget increases of 300-400%, and an AI liability insurance market that grew from $400 million to $2.1 billion. Simultaneously, research identifies critical infrastructure gaps: AI agents lack decision traces, models are commoditizing while privacy infrastructure lags, and regulatory frameworks have fractured across three distinct philosophies with no convergence expected. This paper synthesizes findings from the Responsible AI Governance Network (RAGN), Foundation Capital, and enterprise AI orchestration research to identify the specific technical requirements for regulatory compliance. It then presents the Y.I.N. (Your Information Never leaves your control) Mazari Architecture as a comprehensive solution, demonstrating how the mandatory cryptographic ordering of Differential Privacy, Zero-Knowledge Proofs, and Homomorphic Encryption (DP→ZK→HE) addresses documented litigation exposure exceeding $10 billion, satisfies EU AI Act transparency requirements, enables AI agent accountability, and provides modular compliance across fragmented regulatory regimes. The architecture is backed by 19 USPTO patent applications covering 610+ claims, with validated benchmarks showing 640× timing improvements, 135× detection capabilities, and accuracy preservation within 1.5 percentage points.","author":[{"family":"Mazari","given":"Ilyes"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18244673","URL":"https://doi.org/10.5281/zenodo.18244673","source":"datacite"},{"id":"doi:10.5281/zenodo.18092151","type":"article-journal","title":"Universal Integration Architectures for Token-Based Authorization in Privacy-Preserving Computation: A Comprehensive Cross-Platform Technical Survey","abstract":"This comprehensive technical survey presents integration architectures for the Y.I.N. (Your Information Never leaves your control) Nine Pillars framework across 200+ commercial platforms spanning artificial intelligence (100+ LLM providers including OpenAI, Anthropic, Google DeepMind, Meta, Microsoft, Mistral AI, Baidu, Alibaba, Tencent), healthcare (50+ providers including Epic Systems, Tempus, PathAI), finance (40+ institutions including JPMorgan Chase, Goldman Sachs, BlackRock), autonomous vehicles (20+ companies including Waymo, Tesla, Cruise), telecommunications (25+ carriers including AT&T, China Mobile, Deutsche Telekom), and energy (20+ companies including Siemens Energy, NextEra) across 25+ countries. The Y.I.N. Nine Pillars architecture provides end-to-end privacy protection through: (1) Data Privacy (Differential Privacy), (2) Computation Privacy (Homomorphic Encryption), (3) Storage Privacy (Encryption at Rest), (4) Transmission Privacy (TLS 1.3), (5) Access Control (Zero-Knowledge Proofs), (6) Audit Trail (Merkle Trees), (7) Deletion Rights (Cryptographic Erasure), (8) Quantum Resistance (Lattice-based Cryptography), and (9) Token Licensing (Cryptographic Payment Enforcement). The Ninth Pillar token licensing system, covered by U.S. Patent Application 63/949,361 (filed December 28, 2025), provides cryptographic enforcement of usage rights by integrating token-derived blinding factors into homomorphic encryption operations, making computational correctness mathematically dependent on valid authorization. The system achieves 99.37% accuracy with valid tokens versus 50.7% with invalid tokens (t=147.3, p<10^-50), with security proven under CDH hardness (2^128 operations) and Ring-LWE assumptions. Integration schematics are provided for regulatory compliance with HIPAA (healthcare), SOX/DORA (finance), GDPR/EU AI Act (European Union), CCPA (California), PIPL (China), ISO 27001, NERC CIP (energy), and 15+ other frameworks. Extension directions are documented for community research including TEE hybrid architectures, MPC integration, VDF token lifetimes, key-homomorphic PRFs, flexible validation policies, hardware attestation, ABE capabilities, off-chain settlement, and DID/VC integration. Organizations seeking to implement these integration patterns may obtain licenses for individual pillars, sector packages, or the complete Nine Pillars system from the patent holder. Patent Notice: The Y.I.N. Nine Pillars architecture and Ninth Pillar token licensing system are covered by U.S. Patent Applications 63/949,361 (Ninth Pillar, filed December 28, 2025), 63/923,348 (QFED-MAZARI Quantum Extensions), 19/399,646 (Core Y.I.N. Architecture), 19/403,244 (Hardware Implementation), and 19/417,196 (SQL Database Integration), comprising 430+ claims across 15 patent applications.","author":[{"family":"Mazari","given":"Ilyes"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.18092151","URL":"https://doi.org/10.5281/zenodo.18092151","source":"datacite"},{"id":"doi:10.5281/zenodo.18092152","type":"article-journal","title":"Universal Integration Architectures for Token-Based Authorization in Privacy-Preserving Computation: A Comprehensive Cross-Platform Technical Survey","abstract":"This comprehensive technical survey presents integration architectures for the Y.I.N. (Your Information Never leaves your control) Nine Pillars framework across 200+ commercial platforms spanning artificial intelligence (100+ LLM providers including OpenAI, Anthropic, Google DeepMind, Meta, Microsoft, Mistral AI, Baidu, Alibaba, Tencent), healthcare (50+ providers including Epic Systems, Tempus, PathAI), finance (40+ institutions including JPMorgan Chase, Goldman Sachs, BlackRock), autonomous vehicles (20+ companies including Waymo, Tesla, Cruise), telecommunications (25+ carriers including AT&T, China Mobile, Deutsche Telekom), and energy (20+ companies including Siemens Energy, NextEra) across 25+ countries. The Y.I.N. Nine Pillars architecture provides end-to-end privacy protection through: (1) Data Privacy (Differential Privacy), (2) Computation Privacy (Homomorphic Encryption), (3) Storage Privacy (Encryption at Rest), (4) Transmission Privacy (TLS 1.3), (5) Access Control (Zero-Knowledge Proofs), (6) Audit Trail (Merkle Trees), (7) Deletion Rights (Cryptographic Erasure), (8) Quantum Resistance (Lattice-based Cryptography), and (9) Token Licensing (Cryptographic Payment Enforcement). The Ninth Pillar token licensing system, covered by U.S. Patent Application 63/949,361 (filed December 28, 2025), provides cryptographic enforcement of usage rights by integrating token-derived blinding factors into homomorphic encryption operations, making computational correctness mathematically dependent on valid authorization. The system achieves 99.37% accuracy with valid tokens versus 50.7% with invalid tokens (t=147.3, p<10^-50), with security proven under CDH hardness (2^128 operations) and Ring-LWE assumptions. Integration schematics are provided for regulatory compliance with HIPAA (healthcare), SOX/DORA (finance), GDPR/EU AI Act (European Union), CCPA (California), PIPL (China), ISO 27001, NERC CIP (energy), and 15+ other frameworks. Extension directions are documented for community research including TEE hybrid architectures, MPC integration, VDF token lifetimes, key-homomorphic PRFs, flexible validation policies, hardware attestation, ABE capabilities, off-chain settlement, and DID/VC integration. Organizations seeking to implement these integration patterns may obtain licenses for individual pillars, sector packages, or the complete Nine Pillars system from the patent holder. Patent Notice: The Y.I.N. Nine Pillars architecture and Ninth Pillar token licensing system are covered by U.S. Patent Applications 63/949,361 (Ninth Pillar, filed December 28, 2025), 63/923,348 (QFED-MAZARI Quantum Extensions), 19/399,646 (Core Y.I.N. Architecture), 19/403,244 (Hardware Implementation), and 19/417,196 (SQL Database Integration), comprising 430+ claims across 15 patent applications.","author":[{"family":"Mazari","given":"Ilyes"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.18092152","URL":"https://doi.org/10.5281/zenodo.18092152","source":"datacite"},{"id":"doi:10.48550/arxiv.2511.23252","type":"manuscript","title":"One-Shot Secure Aggregation: A Hybrid Cryptographic Protocol for Private Federated Learning in IoT","abstract":"Federated Learning (FL) offers a promising approach to collaboratively train machine learning models without centralizing raw data, yet its scalability is often throttled by excessive communication overhead. This challenge is magnified in Internet of Things (IoT) environments, where devices face stringent bandwidth, latency, and energy constraints. Conventional secure aggregation protocols, while essential for protecting model updates, frequently require multiple interaction rounds, large payload sizes, and per-client costs rendering them impractical for many edge deployments. In this work, we present Hyb-Agg, a lightweight and communication-efficient secure aggregation protocol that integrates Multi-Key CKKS (MK-CKKS) homomorphic encryption with Elliptic Curve Diffie-Hellman (ECDH)-based additive masking. Hyb-Agg reduces the secure aggregation process to a single, non-interactive client-to-server transmission per round, ensuring that per-client communication remains constant regardless of the number of participants. This design eliminates partial decryption exchanges, preserves strong privacy under the RLWE, CDH, and random oracle assumptions, and maintains robustness against collusion by the server and up to $N-2$ clients. We implement and evaluate Hyb-Agg on both high-performance and resource-constrained devices, including a Raspberry Pi 4, demonstrating that it delivers sub-second execution times while achieving a constant communication expansion factor of approximately 12x over plaintext size. By directly addressing the communication bottleneck, Hyb-Agg enables scalable, privacy-preserving federated learning that is practical for real-world IoT deployments.","author":[{"family":"Emmaka","given":"Imraul"},{"family":"Phuong","given":"Tran"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2511.23252","URL":"https://doi.org/10.48550/arxiv.2511.23252","source":"datacite"},{"id":"doi:10.48550/arxiv.2502.01289","type":"manuscript","title":"A Framework for Double-Blind Federated Adaptation of Foundation Models","abstract":"Foundation models (FMs) excel in zero-shot tasks but benefit from task-specific adaptation. However, privacy concerns prevent data sharing among multiple data owners, and proprietary restrictions prevent the learning service provider (LSP) from sharing the FM. In this work, we propose BlindFed, a framework enabling collaborative FM adaptation while protecting both parties: data owners do not access the FM or each other's data, and the LSP does not see sensitive task data. BlindFed relies on fully homomorphic encryption (FHE) and consists of three key innovations: (i) FHE-friendly architectural modifications via polynomial approximations and low-rank adapters, (ii) a two-stage split learning approach combining offline knowledge distillation and online encrypted inference for adapter training without backpropagation through the FM, and (iii) a privacy-boosting scheme using sample permutations and stochastic block sampling to mitigate model extraction attacks. Empirical results on four image classification datasets demonstrate the practical feasibility of the BlindFed framework, albeit at a high communication cost and large computational complexity for the LSP.","author":[{"family":"Tastan","given":"Nurbek"},{"family":"Nandakumar","given":"Karthik"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2502.01289","URL":"https://doi.org/10.48550/arxiv.2502.01289","source":"datacite"},{"id":"doi:10.5281/zenodo.20324081","type":"article-journal","title":"TEMPORAL ROTATION SECURITY PROTOCOL (TRSP) Physics-First Cryptographic Architecture: Time as the Fundamental Security Parameter","abstract":"ABSTRACT: Temporal Rotation Security Protocol (TRSP) v3 — Physics-First Cryptographic Architecture: Time as the Fundamental Quantum-Resistant Security Parameter Concept in development since at least November 2019. First complete public documentation: 2026. Version 3 adds Part 4c (Hybrid Dynamic CRATON Quorum / HDCQ) and Part 4d (CRATON Hardening Layer / CHL). This concept documents the Temporal Rotation Security Protocol — a cryptographic architecture in which security derives not from the mathematical complexity of encryption keys but from the physical irreversibility of time. Every cryptographic system currently in use — RSA, AES, elliptic curve — rests on a single foundational assumption: that breaking the encryption requires more computational time than any adversary possesses. Quantum computing, through Shor's algorithm and Grover's algorithm, is systematically dismantling this assumption. TRSP replaces it with a physically permanent alternative: a key that no longer exists cannot be recovered by any computation, quantum or classical, regardless of computational resources or future mathematical advances. The protocol operates through simultaneous multi-layer key rotation at three independent frequencies. Layer 1 (session layer) rotates every 10–100 milliseconds using hardware entropy from physical noise sources — thermal variance, clock jitter, electromagnetic fingerprint. Layer 2 (identity layer) rotates every 1–10 seconds, anchored to physically unique device characteristics that cannot be spoofed. Layer 3 (CRATON foundational layer) generates a cryptographic commitment from the unique physical state of both communicating devices at session initialisation — used once and permanently destroyed, unrepeatable at any other point in time or on any other device. Quantum resistance is structural rather than parametric. Shor's algorithm requires minutes to hours to factor key-scale integers; Layer 1 rotation windows of 10–100 milliseconds ensure the target key no longer exists when any quantum computation converges. Grover's algorithm provides quadratic speedup against static keys; against rotating keys it provides no advantage because the search target is destroyed before the search completes. As quantum hardware advances and computation accelerates, rotation windows decrease proportionally — a software parameter adjustment costing microseconds against a hardware investment requiring years. The defender's adaptation is permanently faster than the attacker's. The CRATON foundational trust layer — positioned between hardware and operating system — is not a stored value. It is a physical event: a one-time measurement of device state that generates a cryptographic commitment and is immediately destroyed. It cannot be forged by a compromised operating system, replicated on any other device, or reconstructed from any stored record. Root of trust through physical irreversibility. The inverse proposition — what breaking TRSP would prove — is documented as the second foundational contribution of this concept. A successful attack against a correctly implemented TRSP system would constitute experimental proof of one of the following physical propositions: that quantum information is globally conserved and locally accessible confirming the holographic principle; that temporal irreversibility is not absolute at quantum scale; that parallel quantum branches are accessible through computation confirming the Everett many-worlds interpretation; or that Landauer's principle is violated at computational scale. Any of these would represent the most significant scientific discovery in recorded history. TRSP is therefore simultaneously a security protocol and a physics experiment. Its security parameter is the boundary of known physical law. TRSP is an open invitation to physics. Extension 1: Spatial-Temporal Triangulation & Network Latency Mitigation A critical challenge in millisecond-scale cryptographic rotation (Δt = 10–100 ms) across standard ","author":[{"family":"Mehmetaj","given":"Ilir"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20324081","URL":"https://doi.org/10.5281/zenodo.20324081","source":"datacite"},{"id":"doi:10.5281/zenodo.18884114","type":"article-journal","title":"REVIEW OF SYNERGIZING DEEP LEARNING AND NONLINEAR MODEL PREDICTIVE CONTROL","abstract":"The operational complexity of modern industrial processes demands control frameworks that are both mathematically rigorous and computationally agile. This paper provides a systematic review of the transformative advances in Nonlinear Model Predictive Control (NMPC) integrated with Deep Learning (DL) methodologies between 2020 and 2026. We categorize the state-of-the-art into three technical pillars: neural-based system identification, computational acceleration via latent-space optimization, and robust architectures for uncertain environments. By analyzing applications across chemical engineering, fusion energy maintenance, and bionic robotics, we evaluate how hybrid frameworks-such as LSTM-based estimators and autoencoder-driven reduced-order models-mitigate the traditional trade-offs between model fidelity and real-time feasibility. The review further discusses emerging trends in cyber-secure control via homomorphic encryption and the integration of Physics-Informed Neural Networks (PINNs). Our synthesis highlights persistent gaps in formal stability proofs and model interpretability, offering a strategic roadmap for future research in autonomous industrial intelligence.","author":[{"family":"Oybek","given":"Tuyboyov"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18884114","URL":"https://doi.org/10.5281/zenodo.18884114","source":"datacite"},{"id":"doi:10.5281/zenodo.18884557","type":"article-journal","title":"REVIEW OF SYNERGIZING DEEP LEARNING AND NONLINEAR MODEL PREDICTIVE CONTROL","abstract":"The operational complexity of modern industrial processes demands control frameworks that are both mathematically rigorous and computationally agile. This paper provides a systematic review of the transformative advances in Nonlinear Model Predictive Control (NMPC) integrated with Deep Learning (DL) methodologies between 2020 and 2026. We categorize the state-of-the-art into three technical pillars: neural-based system identification, computational acceleration via latent-space optimization, and robust architectures for uncertain environments. By analyzing applications across chemical engineering, fusion energy maintenance, and bionic robotics, we evaluate how hybrid frameworks-such as LSTM-based estimators and autoencoder-driven reduced-order models-mitigate the traditional trade-offs between model fidelity and real-time feasibility. The review further discusses emerging trends in cyber-secure control via homomorphic encryption and the integration of Physics-Informed Neural Networks (PINNs). Our synthesis highlights persistent gaps in formal stability proofs and model interpretability, offering a strategic roadmap for future research in autonomous industrial intelligence.","author":[{"family":"Oybek","given":"Tuyboyov"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18884557","URL":"https://doi.org/10.5281/zenodo.18884557","source":"datacite"},{"id":"doi:10.5281/zenodo.18884115","type":"article-journal","title":"SYNERGIZING DEEP LEARNING AND NONLINEAR MODEL PREDICTIVE CONTROL: A REVIEW","abstract":"The operational complexity of modern industrial processes demands control frameworks that are both mathematically rigorous and computationally agile. This paper provides a systematic review of the transformative advances in Nonlinear Model Predictive Control (NMPC) integrated with Deep Learning (DL) methodologies between 2020 and 2026. We categorize the state-of-the-art into three technical pillars: neural-based system identification, computational acceleration via latent-space optimization, and robust architectures for uncertain environments. By analyzing applications across chemical engineering, fusion energy maintenance, and bionic robotics, we evaluate how hybrid frameworks-such as LSTM-based estimators and autoencoder-driven reduced-order models-mitigate the traditional trade-offs between model fidelity and real-time feasibility. The review further discusses emerging trends in cyber-secure control via homomorphic encryption and the integration of Physics-Informed Neural Networks (PINNs). Our synthesis highlights persistent gaps in formal stability proofs and model interpretability, offering a strategic roadmap for future research in autonomous industrial intelligence.","author":[{"family":"Oybek","given":"Tuyboyov"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18884115","URL":"https://doi.org/10.5281/zenodo.18884115","source":"datacite"},{"id":"doi:10.5281/zenodo.18828387","type":"article-journal","title":"Blueprint OS Biologique Souverain","abstract":"Résumé FRCe document, produit avec l’assistance de ChatGPT 5.2 Thinking et Gemini 3 Raisonnement, est publié sous Licence Apache 2.0. Il constitue une publication défensive (antériorité) volontaire et entre dans l'état de la technique au sens des législations applicables : art. L 611-11 CPI / art. 54(2) CBE. Le « Blueprint OS Biologique Souverain » définit une architecture complète pour la santé métabolique autonome. Il détaille 32 briques technologiques allant de la détection optique du NAD+ par capteurs multispectraux à l'IA fédérée omique, en passant par des patchs de délivrance en boucle fermée et des structures électroniques biodégradables en soie. Chaque proposition est décrite de manière « enabling », classée avec les codes IPC/CPC, et accompagnée d'une preuve d'horodatage (RFC 3161). Cet écosystème sécurise l'accès libre aux technologies fondamentales de souveraineté biologique. Abstract ENThis document, produced with the assistance of ChatGPT 5.2 Thinking and Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: EPC Art. 54(2) (European Patent Convention), French IPC Art. L 611-11 (French Intellectual Property Code), 35 U.S.C. §102(a) (United States Patent Act), Chinese Patent Law Art. 22(5) (中华人民共和国专利法), and Japanese Patent Act Art. 29(1) (特許法). The \"Sovereign Biological OS Blueprint\" outlines a full-stack architecture for autonomous metabolic health. It details 32 technological blocks including multispectral NAD+ sensing, federated omic AI, closed-loop delivery patches, and biodegradable silk electronics. Every proposal is described in an enabling manner, classified with IPC and CPC codes, and accompanied by timestamp proof (RFC 3161 / FreeTSA). This framework ensures public access to foundational biological sovereignty tools. Timestamp: 2026-03-01T22:41:54ZSHA-256: aa37064544a4d57b8b9a134d2f5993a8c5eeddccbfe37b149968e0df0e8f019c Liste des innovations & classification (IPC ; CPC) Multispectral Redox Sensor [IPC A61B 5/00 ; CPC A61B 5/1455] Federated Omic Learning [IPC G16H 50/20 ; CPC G06N 3/098] Closed-Loop Patch [IPC A61M 37/00 ; CPC A61M 37/00] Smartphone µPAD [IPC B01L 3/00 ; CPC G01N 33/52] Digital Bio-Twin [IPC G16H 50/50 ; CPC G16H 50/50] Bio-TEE Enclave [IPC G06F 21/74 ; CPC G06F 21/53] AI-Controlled Bioreactor [IPC C12M 1/36 ; CPC C12M 41/48] Bio-Chain Logistics [IPC G06Q 10/08 ; CPC G06Q 10/08] Zwitterionic Coating [IPC C08J 7/04 ; CPC C08J 7/042] AI Spectral Deconvolution [IPC G06T 5/00 ; CPC G06T 5/50] Circadian MPC Dosage [IPC A61K 31/00 ; CPC G16H 20/17] 3D DLP Microneedles [IPC B29C 64/124 ; CPC B33Y 10/00] Homomorphic Omic Compute [IPC G06F 21/62 ; CPC G06F 21/62] Redox Biometrics [IPC G06F 21/32 ; CPC G06F 21/32] Mitochondrial LNP [IPC A61K 9/127 ; CPC A61K 9/1271] TTI Smart Cap [IPC B65D 81/24 ; CPC B65D 2581/00] Proof-of-Health (PoH) [IPC G06Q 20/06 ; CPC G06Q 20/367] Bio-NAD Exosomes [IPC C12N 15/88 ; CPC A61K 9/1277] Haptic UX Interface [IPC A61M 5/42 ; CPC A61M 5/422] Enzymatic Fuel Cell [IPC H01M 8/16 ; CPC H01M 8/16] Bio-JSON Standard [IPC G16H 10/60 ; CPC G16H 10/60] Stabilized Drone Pod [IPC B64U 30/20 ; CPC B64U 2101/60] Bio-Kill-Switch [IPC G16H 40/60 ; CPC G16H 40/67] Bioprocess AI LSTM [IPC C12M 1/36 ; CPC C12M 41/48] Metabolic Tomography [IPC A61B 5/00 ; CPC G06T 7/00] Synthetic Omic Phantoms [IPC G01N 21/64 ; CPC G01N 33/50] AR-Guided Alignment [IPC G06T 19/00 ; CPC G16H 40/63] Silk Electronics [IPC H05K 1/03 ; CPC H05K 1/03] Smart Audit Contract [IPC G06F 21/64 ; CPC G06Q 20/40] Edge Knowledge Graph [IPC G16H 50/70 ; CPC G06N 5/02] Async Omic Sync [IPC G06F 17/10 ; CPC G16H 50/20] Battery-less NFC Sensor [IPC A61B 5/145 ; CPC G06K 19/07] KeywordsSovereign Biological OS, NAD+ Sensing, Federated Learning, Microneedles, Digital Twin, Bio-Cryptography, Decentralized Manufacturing, Metabolic Health, Zwitterionic Coating, Spectral Deconvolutio","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18828387","URL":"https://doi.org/10.5281/zenodo.18828387","source":"datacite"},{"id":"doi:10.5281/zenodo.18828388","type":"article-journal","title":"Blueprint OS Biologique Souverain","abstract":"Résumé FRCe document, produit avec l’assistance de ChatGPT 5.2 Thinking et Gemini 3 Raisonnement, est publié sous Licence Apache 2.0. Il constitue une publication défensive (antériorité) volontaire et entre dans l'état de la technique au sens des législations applicables : art. L 611-11 CPI / art. 54(2) CBE. Le « Blueprint OS Biologique Souverain » définit une architecture complète pour la santé métabolique autonome. Il détaille 32 briques technologiques allant de la détection optique du NAD+ par capteurs multispectraux à l'IA fédérée omique, en passant par des patchs de délivrance en boucle fermée et des structures électroniques biodégradables en soie. Chaque proposition est décrite de manière « enabling », classée avec les codes IPC/CPC, et accompagnée d'une preuve d'horodatage (RFC 3161). Cet écosystème sécurise l'accès libre aux technologies fondamentales de souveraineté biologique. Abstract ENThis document, produced with the assistance of ChatGPT 5.2 Thinking and Gemini 3 Raisonnement, is released under the Apache 2.0 licence. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: EPC Art. 54(2) (European Patent Convention), French IPC Art. L 611-11 (French Intellectual Property Code), 35 U.S.C. §102(a) (United States Patent Act), Chinese Patent Law Art. 22(5) (中华人民共和国专利法), and Japanese Patent Act Art. 29(1) (特許法). The \"Sovereign Biological OS Blueprint\" outlines a full-stack architecture for autonomous metabolic health. It details 32 technological blocks including multispectral NAD+ sensing, federated omic AI, closed-loop delivery patches, and biodegradable silk electronics. Every proposal is described in an enabling manner, classified with IPC and CPC codes, and accompanied by timestamp proof (RFC 3161 / FreeTSA). This framework ensures public access to foundational biological sovereignty tools. Timestamp: 2026-03-01T22:41:54ZSHA-256: aa37064544a4d57b8b9a134d2f5993a8c5eeddccbfe37b149968e0df0e8f019c Liste des innovations & classification (IPC ; CPC) Multispectral Redox Sensor [IPC A61B 5/00 ; CPC A61B 5/1455] Federated Omic Learning [IPC G16H 50/20 ; CPC G06N 3/098] Closed-Loop Patch [IPC A61M 37/00 ; CPC A61M 37/00] Smartphone µPAD [IPC B01L 3/00 ; CPC G01N 33/52] Digital Bio-Twin [IPC G16H 50/50 ; CPC G16H 50/50] Bio-TEE Enclave [IPC G06F 21/74 ; CPC G06F 21/53] AI-Controlled Bioreactor [IPC C12M 1/36 ; CPC C12M 41/48] Bio-Chain Logistics [IPC G06Q 10/08 ; CPC G06Q 10/08] Zwitterionic Coating [IPC C08J 7/04 ; CPC C08J 7/042] AI Spectral Deconvolution [IPC G06T 5/00 ; CPC G06T 5/50] Circadian MPC Dosage [IPC A61K 31/00 ; CPC G16H 20/17] 3D DLP Microneedles [IPC B29C 64/124 ; CPC B33Y 10/00] Homomorphic Omic Compute [IPC G06F 21/62 ; CPC G06F 21/62] Redox Biometrics [IPC G06F 21/32 ; CPC G06F 21/32] Mitochondrial LNP [IPC A61K 9/127 ; CPC A61K 9/1271] TTI Smart Cap [IPC B65D 81/24 ; CPC B65D 2581/00] Proof-of-Health (PoH) [IPC G06Q 20/06 ; CPC G06Q 20/367] Bio-NAD Exosomes [IPC C12N 15/88 ; CPC A61K 9/1277] Haptic UX Interface [IPC A61M 5/42 ; CPC A61M 5/422] Enzymatic Fuel Cell [IPC H01M 8/16 ; CPC H01M 8/16] Bio-JSON Standard [IPC G16H 10/60 ; CPC G16H 10/60] Stabilized Drone Pod [IPC B64U 30/20 ; CPC B64U 2101/60] Bio-Kill-Switch [IPC G16H 40/60 ; CPC G16H 40/67] Bioprocess AI LSTM [IPC C12M 1/36 ; CPC C12M 41/48] Metabolic Tomography [IPC A61B 5/00 ; CPC G06T 7/00] Synthetic Omic Phantoms [IPC G01N 21/64 ; CPC G01N 33/50] AR-Guided Alignment [IPC G06T 19/00 ; CPC G16H 40/63] Silk Electronics [IPC H05K 1/03 ; CPC H05K 1/03] Smart Audit Contract [IPC G06F 21/64 ; CPC G06Q 20/40] Edge Knowledge Graph [IPC G16H 50/70 ; CPC G06N 5/02] Async Omic Sync [IPC G06F 17/10 ; CPC G16H 50/20] Battery-less NFC Sensor [IPC A61B 5/145 ; CPC G06K 19/07] KeywordsSovereign Biological OS, NAD+ Sensing, Federated Learning, Microneedles, Digital Twin, Bio-Cryptography, Decentralized Manufacturing, Metabolic Health, Zwitterionic Coating, Spectral Deconvolutio","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18828388","URL":"https://doi.org/10.5281/zenodo.18828388","source":"datacite"},{"id":"doi:10.5281/zenodo.18385664","type":"article-journal","title":"Nexus Ocean: Quantum Intelligence Architecture for Deep-Sea Monitoring and Predictive Analytics","abstract":"🌊 Nexus Ocean v3.1.0 - Quantum AI Marine Intelligence Next-Generation Deep Sea Monitoring & Intelligence System Transforming oceanographic monitoring through quantum-inspired artificial intelligence 📖 Overview Nexus Ocean is a revolutionary quantum-inspired oceanographic AI system that combines six-layer neural architecture with real-time deep-sea monitoring capabilities. The system leverages cutting-edge technologies including neuromorphic computing, quantum-inspired algorithms, and autonomous decision-making to provide comprehensive ocean intelligence, predictive analytics, and emergency response capabilities. 🎯 What Makes Nexus Ocean Unique? 🧠 Six-Layer Neural Architecture: From sensor fusion to emergent intelligence ⚡ Real-Time Processing: Sub-second analysis of oceanographic data 🔮 Predictive Intelligence: Long-term forecasting with AI-powered insights 🛡️ Quantum-Resistant Security: Post-quantum cryptography and zero-trust architecture 📊 Advanced Visualization: 3D oceanographic dashboards with 10+ chart types 🌐 Edge-Cloud Hybrid: Distributed computing from deep-sea sensors to cloud analytics 🤖 Autonomous Operations: Self-learning systems with human oversight 🔗 Blockchain Verification: Immutable audit trails and decision transparency 🌟 Key Features 🛡️ Multi-Sensor Fusion Advanced integration of pressure, acoustic, thermal, magnetic, and biosignature sensors with Kalman filtering and neural preprocessing. 🧬 Real-Time Cognitive Processing Pattern recognition, anomaly detection, and contextual analysis using transformer-based models and graph neural networks. 🌌 Autonomous Decision Making AI-powered orchestration with multi-criteria optimization, Monte Carlo simulations, and risk assessment frameworks. 🔐 Quantum-Resistant Security Zero-trust architecture with lattice-based cryptography, hash-based signatures, and automated threat response. 🔮 Predictive Intelligence Long-term forecasting using LSTM/GRU networks, attention mechanisms, and causal inference models. 🌀 Emergent Intelligence Cross-layer integration with swarm orchestration, event-driven architecture, and self-learning capabilities. 📊 Advanced Analytics Dashboard 10 interactive visualizations including 3D ocean views, radar charts, and time-series analysis Location performance comparison across multiple ocean regions Real-time QII (Quantum Intelligence Index) monitoring Anomaly detection with automated alerts Export capabilities for reports and data analysis 🏗️ System Architecture The Nexus Ocean system is built on the NEXUS Architecture - a six-layer neural framework inspired by biological intelligence: ┌─────────────────────────────────────────────────────────┐ │ 🌊 SENTINEL Layer │ │ Multi-Sensor Fusion & Perception │ │ HydroSense | AcousticMatrix | ThermalGrid | GeoMag │ └────────────────────┬────────────────────────────────────┘ │ Real-time Sensor Streams ▼ ┌─────────────────────────────────────────────────────────┐ │ 🧬 CORTEX Layer │ │ Cognitive Processing & Pattern Recognition │ │ PatternWeaver | ContextEngine | KnowledgeGraph │ └────────────────────┬────────────────────────────────────┘ │ Intelligent Analysis ▼ ┌─────────────────────────────────────────────────────────┐ │ 🌌 NEXUS Layer │ │ Decision Making & Orchestration │ │ DecisionForge | ScenarioEngine | RiskCalculus │ └────────────────────┬────────────────────────────────────┘ │ Strategic Commands ▼ ┌─────────────────────────────────────────────────────────┐ │ 🛡️ AEGIS Layer │ │ Security & Emergency Response │ │ ThreatRadar | DefenseGrid | EmergencyProtocol │ └────────────────────┬────────────────────────────────────┘ │ Protected Actions ▼ ┌─────────────────────────────────────────────────────────┐ │ 🔮 ORACLE Layer │ │ Predictive Intelligence & Learning │ │ FutureSight | CognitiveLeap | TrendAnalyzer │ └────────────────────┬────────────────────────────────────┘ │ Strategic Insights ▼ ┌─────────────────────────────────────────────────────────┐ │ 🌀 SYNERGY Layer │ │ Cross-Layer Integr","author":[{"family":"Baladi","given":"Samir"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18385664","URL":"https://doi.org/10.5281/zenodo.18385664","source":"datacite"},{"id":"doi:10.5281/zenodo.18426244","type":"article-journal","title":"Nexus Ocean: Quantum Intelligence Architecture for Deep-Sea Monitoring and Predictive Analytics","abstract":"🌊 Nexus Ocean v3.1.0 - Quantum AI Marine Intelligence Next-Generation Deep Sea Monitoring & Intelligence System Transforming oceanographic monitoring through quantum-inspired artificial intelligence 📖 Overview Nexus Ocean is a revolutionary quantum-inspired oceanographic AI system that combines six-layer neural architecture with real-time deep-sea monitoring capabilities. The system leverages cutting-edge technologies including neuromorphic computing, quantum-inspired algorithms, and autonomous decision-making to provide comprehensive ocean intelligence, predictive analytics, and emergency response capabilities. 🎯 What Makes Nexus Ocean Unique? 🧠 Six-Layer Neural Architecture: From sensor fusion to emergent intelligence ⚡ Real-Time Processing: Sub-second analysis of oceanographic data 🔮 Predictive Intelligence: Long-term forecasting with AI-powered insights 🛡️ Quantum-Resistant Security: Post-quantum cryptography and zero-trust architecture 📊 Advanced Visualization: 3D oceanographic dashboards with 10+ chart types 🌐 Edge-Cloud Hybrid: Distributed computing from deep-sea sensors to cloud analytics 🤖 Autonomous Operations: Self-learning systems with human oversight 🔗 Blockchain Verification: Immutable audit trails and decision transparency 🌟 Key Features 🛡️ Multi-Sensor Fusion Advanced integration of pressure, acoustic, thermal, magnetic, and biosignature sensors with Kalman filtering and neural preprocessing. 🧬 Real-Time Cognitive Processing Pattern recognition, anomaly detection, and contextual analysis using transformer-based models and graph neural networks. 🌌 Autonomous Decision Making AI-powered orchestration with multi-criteria optimization, Monte Carlo simulations, and risk assessment frameworks. 🔐 Quantum-Resistant Security Zero-trust architecture with lattice-based cryptography, hash-based signatures, and automated threat response. 🔮 Predictive Intelligence Long-term forecasting using LSTM/GRU networks, attention mechanisms, and causal inference models. 🌀 Emergent Intelligence Cross-layer integration with swarm orchestration, event-driven architecture, and self-learning capabilities. 📊 Advanced Analytics Dashboard 10 interactive visualizations including 3D ocean views, radar charts, and time-series analysis Location performance comparison across multiple ocean regions Real-time QII (Quantum Intelligence Index) monitoring Anomaly detection with automated alerts Export capabilities for reports and data analysis 🏗️ System Architecture The Nexus Ocean system is built on the NEXUS Architecture - a six-layer neural framework inspired by biological intelligence: ┌─────────────────────────────────────────────────────────┐ │ 🌊 SENTINEL Layer │ │ Multi-Sensor Fusion & Perception │ │ HydroSense | AcousticMatrix | ThermalGrid | GeoMag │ └────────────────────┬────────────────────────────────────┘ │ Real-time Sensor Streams ▼ ┌─────────────────────────────────────────────────────────┐ │ 🧬 CORTEX Layer │ │ Cognitive Processing & Pattern Recognition │ │ PatternWeaver | ContextEngine | KnowledgeGraph │ └────────────────────┬────────────────────────────────────┘ │ Intelligent Analysis ▼ ┌─────────────────────────────────────────────────────────┐ │ 🌌 NEXUS Layer │ │ Decision Making & Orchestration │ │ DecisionForge | ScenarioEngine | RiskCalculus │ └────────────────────┬────────────────────────────────────┘ │ Strategic Commands ▼ ┌─────────────────────────────────────────────────────────┐ │ 🛡️ AEGIS Layer │ │ Security & Emergency Response │ │ ThreatRadar | DefenseGrid | EmergencyProtocol │ └────────────────────┬────────────────────────────────────┘ │ Protected Actions ▼ ┌─────────────────────────────────────────────────────────┐ │ 🔮 ORACLE Layer │ │ Predictive Intelligence & Learning │ │ FutureSight | CognitiveLeap | TrendAnalyzer │ └────────────────────┬────────────────────────────────────┘ │ Strategic Insights ▼ ┌─────────────────────────────────────────────────────────┐ │ 🌀 SYNERGY Layer │ │ Cross-Layer Integr","author":[{"family":"Baladi","given":"Samir"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18426244","URL":"https://doi.org/10.5281/zenodo.18426244","source":"datacite"},{"id":"doi:10.5281/zenodo.20692339","type":"article-journal","title":"Humans First: From Threats to PETs in Human-Machine Interaction","abstract":"Modern Human-Machine Interaction (HMI) systems increasingly rely on intimate physiological, behavioral, and cognitive data to enable adaptive interaction between humans and machines. While this opens transformative possibilities across domains such as healthcare, industrial automation, and assistive technologies, it also introduces substantial privacy, security, and safety risks that can directly translate into human harm, including ethical and societal consequences. Yet existing research tends to discuss these risks in isolation and fragmented, and the practical applicability of Privacy-Enhancing Technologies (PETs) in realistic, resource-constrained HMI environments remains insufficiently explored. This thesis addresses these gaps through three contributions. First, a systematic taxonomy-driven literature review examines risks and protections in Brain-Computer Interfaces (BCIs), a particularly high-stakes area of HMI in which direct neural interfacing makes these risks especially acute, mapping technical threats to potential human harms and examining the role of PETs as protective strategies. Second, a realistic HMI prototype integrates homomorphic encryption (HE) into a physiological stress-monitoring pipeline on constrained embedded hardware. Third, the prototype is evaluated from both technical and human-centered perspectives, demonstrating that HE preserves classification performance while introducing substantial computational overhead. Nevertheless, the results indicate that privacy-preserving machine learning using HE is feasible even on constrained embedded hardware and can realistically be integrated into practical HMI scenarios. At the same time, the evaluation highlights that HE provides a strong but narrow privacy guarantee and therefore represents only one building block within a broader, holistic approach to privacy, security, and safety in HMI systems. Overall, this thesis contributes to a holistic, human-centered analysis of risks in HMI, demonstrates the practical viability of PETs under real-world constraints, and highlights the research gaps that must be addressed to achieve meaningful protection in future HMI systems.","author":[{"family":"Bergermann","given":"Justus"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20692339","URL":"https://doi.org/10.5281/zenodo.20692339","source":"datacite"},{"id":"doi:10.5281/zenodo.20692340","type":"article-journal","title":"Humans First: From Threats to PETs in Human-Machine Interaction","abstract":"Modern Human-Machine Interaction (HMI) systems increasingly rely on intimate physiological, behavioral, and cognitive data to enable adaptive interaction between humans and machines. While this opens transformative possibilities across domains such as healthcare, industrial automation, and assistive technologies, it also introduces substantial privacy, security, and safety risks that can directly translate into human harm, including ethical and societal consequences. Yet existing research tends to discuss these risks in isolation and fragmented, and the practical applicability of Privacy-Enhancing Technologies (PETs) in realistic, resource-constrained HMI environments remains insufficiently explored. This thesis addresses these gaps through three contributions. First, a systematic taxonomy-driven literature review examines risks and protections in Brain-Computer Interfaces (BCIs), a particularly high-stakes area of HMI in which direct neural interfacing makes these risks especially acute, mapping technical threats to potential human harms and examining the role of PETs as protective strategies. Second, a realistic HMI prototype integrates homomorphic encryption (HE) into a physiological stress-monitoring pipeline on constrained embedded hardware. Third, the prototype is evaluated from both technical and human-centered perspectives, demonstrating that HE preserves classification performance while introducing substantial computational overhead. Nevertheless, the results indicate that privacy-preserving machine learning using HE is feasible even on constrained embedded hardware and can realistically be integrated into practical HMI scenarios. At the same time, the evaluation highlights that HE provides a strong but narrow privacy guarantee and therefore represents only one building block within a broader, holistic approach to privacy, security, and safety in HMI systems. Overall, this thesis contributes to a holistic, human-centered analysis of risks in HMI, demonstrates the practical viability of PETs under real-world constraints, and highlights the research gaps that must be addressed to achieve meaningful protection in future HMI systems.","author":[{"family":"Bergermann","given":"Justus"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20692340","URL":"https://doi.org/10.5281/zenodo.20692340","source":"datacite"},{"id":"doi:10.5281/zenodo.18377438","type":"article-journal","title":"Cryptographically Enforced Relay-Resistant Offline Payment System Using Algorithmic Logic Fingerprinting and Split Execution Architecture for Digital Euro and CBDC Systems","abstract":"Abstract This paper ( Patent pending concept ) presents a novel cryptographic framework for securing offline digital currency payment systems, including Central Bank Digital Currencies (CBDCs) such as the Digital Euro, against relay and replay attacks. The architecture integrates Split Execution (separating computation from authorization), Algorithmic Logic Fingerprinting (ALF) for logic-class validation, and cryptographically bound Virtual Identities (VI) and Compliance Jurisdiction Tokens (CJT). Offline payment authorization is evaluated within a Trusted Execution Environment (TEE), ensuring that cryptographic execution predicates are verified in-device before any value release. This design prevents cloned or remotely relayed tokens from being executed, as all authorization logic is session-scoped, device-specific, and context-bound. The system enforces relay resistance without relying on proximity checks, GPS, or online infrastructure—thereby supporting fully offline, privacy-preserving CBDC payments. It remains compliant with regulatory requirements through generation of Ledger-Anchored Validation Receipts (LAVR), providing auditability without user surveillance. The proposed solution is particularly suited for Digital Euro deployments, public transit systems, and unconnected retail environments, addressing a long-standing security gap in CBDC and digital cash architectures.","author":[{"family":"Das","given":"Sangam"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18377438","URL":"https://doi.org/10.5281/zenodo.18377438","source":"datacite"},{"id":"doi:10.5281/zenodo.18463973","type":"article-journal","title":"Towards Secure and Privacy-Preserving Query Processing for Encrypted Big Data in Multi-Cloud Environments: A Systematic Review","abstract":"The rapid growth of cloud computing and big data analytics has intensified concerns over privacy when sensitive data are outsourced to third-party cloud providers. Traditional encryption techniques protect data confidentiality but significantly limit the ability to perform expressive and efficient queries, particularly in distributed and multi-cloud environments. Motivated by the increasing demand for secure analytics across healthcare, finance, IoT, and collaborative cloud platforms, this review systematically examines privacy-preserving query processing techniques for encrypted data in multi-cloud settings. Following PRISMA guidelines, a systematic literature review of published peer-reviewed studies is conducted. The reviewed approaches are categorized into homomorphic encryption-based methods, searchable encryption techniques, secure multi-party computation, trusted execution environments, and hybrid architectures. The analysis highlights key trade-offs among privacy guarantees, query expressiveness, computational efficiency, and scalability. While hybrid and multi-cloud approaches improve flexibility and fault tolerance, they introduce new challenges related to leakage, communication overhead, and trust assumptions. This review identifies critical research gaps, including limited real-time support, side-channel vulnerabilities, and the absence of standardized benchmarks. Finally, future research directions are outlined, emphasizing AI-assisted encrypted querying, federated analytics, and post-quantum privacy-preserving frameworks for multi-cloud environments.","author":[{"family":"Abdullahi","given":"Ibrahim"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18463973","URL":"https://doi.org/10.5281/zenodo.18463973","source":"datacite"},{"id":"doi:10.5281/zenodo.18463974","type":"article-journal","title":"Towards Secure and Privacy-Preserving Query Processing for Encrypted Big Data in Multi-Cloud Environments: A Systematic Review","abstract":"The rapid growth of cloud computing and big data analytics has intensified concerns over privacy when sensitive data are outsourced to third-party cloud providers. Traditional encryption techniques protect data confidentiality but significantly limit the ability to perform expressive and efficient queries, particularly in distributed and multi-cloud environments. Motivated by the increasing demand for secure analytics across healthcare, finance, IoT, and collaborative cloud platforms, this review systematically examines privacy-preserving query processing techniques for encrypted data in multi-cloud settings. Following PRISMA guidelines, a systematic literature review of published peer-reviewed studies is conducted. The reviewed approaches are categorized into homomorphic encryption-based methods, searchable encryption techniques, secure multi-party computation, trusted execution environments, and hybrid architectures. The analysis highlights key trade-offs among privacy guarantees, query expressiveness, computational efficiency, and scalability. While hybrid and multi-cloud approaches improve flexibility and fault tolerance, they introduce new challenges related to leakage, communication overhead, and trust assumptions. This review identifies critical research gaps, including limited real-time support, side-channel vulnerabilities, and the absence of standardized benchmarks. Finally, future research directions are outlined, emphasizing AI-assisted encrypted querying, federated analytics, and post-quantum privacy-preserving frameworks for multi-cloud environments.","author":[{"family":"Abdullahi","given":"Ibrahim"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18463974","URL":"https://doi.org/10.5281/zenodo.18463974","source":"datacite"},{"id":"doi:10.5281/zenodo.18280900","type":"article-journal","title":"Survey of Privacy Preserving Techniques for Distributed Learning in an IoT Network","abstract":"This survey provides a comprehensive review of privacy-preserving techniques applicable to distributed learning in Internet-of-Things (IoT) environments. The paper examines classical approaches, including Differential Privacy, Homomorphic Encryption, Secure Multi-party Computation, Distributed Selective Stochastic Gradient Descent, and Anonymization, as well as more recent methods such as additive and multiplicative schemes, blockchain-based mechanisms, Bloom Filter–based preprocessing, and intrusion detection systems. Each technique is analyzed with respect to its ability to protect data privacy and its suitability for deployment on resource-constrained IoT devices. Background information on IoT architectures, device limitations, and distributed learning paradigms is provided to contextualize the discussion. The survey evaluates trade-offs among computational overhead, memory usage, communication requirements, and privacy protection, and offers guidance for selecting appropriate techniques based on application requirements and device capabilities.","author":[{"family":"Cartmell","given":"John"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18280900","URL":"https://doi.org/10.5281/zenodo.18280900","source":"datacite"},{"id":"doi:10.48550/arxiv.2601.06710","type":"manuscript","title":"Privacy-Preserving Data Processing in Cloud : From Homomorphic Encryption to Federated Analytics","abstract":"Privacy-preserving data processing refers to the methods and models that allow computing and analyzing sensitive data with a guarantee of confidentiality. As cloud computing and applications that rely on data continue to expand, there is an increasing need to protect personal, financial and healthcare information. Conventional centralized data processing methods expose sensitive data to risk of breaches, compelling the need to use decentralized and secure data methods. This paper gives a detailed review of privacy-saving mechanisms in the cloud platform, such as statistical approaches like differential privacy and cryptographic solutions like homomorphic encryption. Federated analytics and federated learning, two distributed learning frameworks, are also discussed. Their principles, applications, benefits, and limitations are reviewed, with roles of use in the fields of healthcare, finance, IoT, and industrial cases. Comparative analyses measure trade-offs in security, efficiency, scalability, and accuracy, and investigations are done of emerging hybrid frameworks to provide better privacy protection. Critical issues, including computational overhead, privacy-utility trade-offs, standardization, adversarial threats, and cloud integration are also addressed. This review examines in detail the recent privacy-protecting approaches in cloud computation and offers scholars and practitioners crucial information on secure and effective solutions to data processing.","author":[{"family":"Sarraf","given":"Gaurav"},{"family":"Pal","given":"Vibhor"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2601.06710","URL":"https://doi.org/10.48550/arxiv.2601.06710","source":"datacite"},{"id":"doi:10.5281/zenodo.17165101","type":"article-journal","title":"Effet Flynn Inverse : Analyse des Causes, Solutions et Scénarios Prospectifs","abstract":"Abstract ENThis document, produced with the assistance of ChatGPT o3 and GPT-5 Thinking, is released under the Apache 2.0 license. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: EPC Art. 54(2) (European Patent Convention), French IPC/CPI Art. L 611-11 (French Intellectual Property Code), 35 U.S.C. §102(a) (United States Patent Act), Chinese Patent Law Art. 22(5) (中华人民共和国专利法), and Japanese Patent Act Art. 29(1) (特許法). This publication addresses the inverse Flynn effect through 65 enabling innovations across devices, assays, AI, closed-loop protocols, materials, therapies, manufacturing, imaging, delivery, data standards, UX, logistics, and governance. Every invention is described in an enabling manner, classified with IPC and CPC codes, and accompanied by timestamp proof (RFC 3161 / FreeTSA). Résumé FRCe document, produit avec l’assistance de ChatGPT o3 et GPT-5 Thinking, est diffusé sous licence Apache 2.0. Il constitue une publication défensive volontaire (antériorité) et entre, dès sa diffusion publique, dans l’art antérieur au titre des régimes suivants : CBE art. 54(2), CPI art. L 611-11, 35 U.S.C. §102(a), Loi chinoise sur les brevets art. 22(5) (中华人民共和国专利法), et Loi japonaise sur les brevets art. 29(1) (特許法). Cette publication traite de l’effet Flynn inverse à travers 65 innovations habilitantes couvrant dispositifs, capteurs, tests, IA, protocoles en boucle fermée, matériaux, thérapies, procédés de fabrication, imagerie, systèmes de délivrance, normes de données, UX, logistique et gouvernance. Chaque invention est décrite de manière habilitante, classée IPC/CPC, et associée à une preuve d’horodatage (RFC 3161 / FreeTSA). Timestamp: 2025-09-13T16:43:19ZSHA-256: 55c629606ebc242c315596909e61875ab3b1653d366f1378dca5bda73e5f97b6 Liste des innovations & classification (IPC ; CPC) 1. Adaptive AI Tutor for Schools — IPC G09B 5/00 ; CPC G09B 5/022. Lightweight EEG Attention Coach — IPC A61B 5/04 ; CPC A61B 5/04083. Classroom CO2/PM Multi-sensor — IPC G01N 27/12 ; CPC F24F 11/524. PID-Driven HEPA Filtration — IPC F24F 8/80 ; CPC F24F 11/525. Salivary Cog-Biomarker Panel — IPC C12Q 1/70 ; CPC G16H 50/206. School Sleep Scheduling Trial — IPC G09B 19/00 ; CPC G16H 10/607. “Prescription Reading” for Infants — IPC A61B 10/00 ; CPC G16H 50/308. Cognitive AR Learning City — IPC G06T 19/00 ; CPC G09B 15/029. Federated “CognitCloud” Platform — IPC G16H 50/70 ; CPC G06N 20/0010. Reading Eye-Tracking Remediation — IPC A61B 5/107 ; CPC G06T 7/24611. Transdermal Melatonin Patch — IPC A61K 31/70 ; CPC A61M 37/0012. Closed-Loop tDCS for Attention — IPC A61N 1/36 ; CPC A61B 5/040813. Soil Lead Bioremediation — IPC C02F 3/28 ; CPC A62D 3/0014. Urinary Iodine Microfluidic Test — IPC G01N 21/78 ; CPC C12Q 1/2415. “Focus Mode” Network Filtering — IPC H04L 29/06 ; CPC G06F 21/6216. Low-Bandwidth Executive Games — IPC A63F 13/67 ; CPC G09B 5/1417. “Air→Score” Causal KPI — IPC G06Q 50/20 ; CPC G06F 16/245718. Standardized 10–12 min Battery — IPC G01N 33/50 ; CPC G06F 19/0019. Omega-3 School Meal Spec — IPC A23L 33/10 ; CPC A61K 36/2820. Cold-Chain for Bio-Samples — IPC B65D 81/38 ; CPC G16H 40/6321. “CognitFHIR” API Profiles — IPC G16H 10/60 ; CPC G06F 21/6222. Deep-Reading QA Module — IPC G09B 7/00 ; CPC G06F 3/04823. Classroom fNIRS Imaging — IPC A61B 5/145 ; CPC A61B 5/05524. “Offloading Index” — IPC G06Q 50/26 ; CPC G06F 16/2825. EEG Headset Manufacturing — IPC H01R 4/58 ; CPC A61B 5/040226. Wearable Sleep Multi-sensor — IPC A61B 5/11 ; CPC A61B 5/11127. Microbiome–Cognition Panel — IPC C12Q 1/6876 ; CPC G16H 50/3028. Combined Cognitive Supplement — IPC A61K 31/201 ; CPC A61P 25/2829. Population Cognitive Digital Twin — IPC G06N 10/00 ; CPC G16H 30/2030. NFC Microneedle Patch Process — IPC B29C 59/04 ; CPC A61M 37/0031. Arts & Cognition Program — IPC G09B 19/00 ; CPC A63H 2200/1232. Cognitive Data Blockchain — IPC G06Q 20/38 ; CPC G06N 20/2033. Wearable Enviro","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17165101","URL":"https://doi.org/10.5281/zenodo.17165101","source":"datacite"},{"id":"doi:10.5281/zenodo.17114008","type":"article-journal","title":"Effet Flynn Inverse : Analyse des Causes, Solutions et Scénarios Prospectifs","abstract":"Abstract ENThis document, produced with the assistance of ChatGPT o3 and GPT-5 Thinking, is released under the Apache 2.0 license. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: EPC Art. 54(2) (European Patent Convention), French IPC/CPI Art. L 611-11 (French Intellectual Property Code), 35 U.S.C. §102(a) (United States Patent Act), Chinese Patent Law Art. 22(5) (中华人民共和国专利法), and Japanese Patent Act Art. 29(1) (特許法). This publication addresses the inverse Flynn effect through 65 enabling innovations across devices, assays, AI, closed-loop protocols, materials, therapies, manufacturing, imaging, delivery, data standards, UX, logistics, and governance. Every invention is described in an enabling manner, classified with IPC and CPC codes, and accompanied by timestamp proof (RFC 3161 / FreeTSA). Résumé FRCe document, produit avec l’assistance de ChatGPT o3 et GPT-5 Thinking, est diffusé sous licence Apache 2.0. Il constitue une publication défensive volontaire (antériorité) et entre, dès sa diffusion publique, dans l’art antérieur au titre des régimes suivants : CBE art. 54(2), CPI art. L 611-11, 35 U.S.C. §102(a), Loi chinoise sur les brevets art. 22(5) (中华人民共和国专利法), et Loi japonaise sur les brevets art. 29(1) (特許法). Cette publication traite de l’effet Flynn inverse à travers 65 innovations habilitantes couvrant dispositifs, capteurs, tests, IA, protocoles en boucle fermée, matériaux, thérapies, procédés de fabrication, imagerie, systèmes de délivrance, normes de données, UX, logistique et gouvernance. Chaque invention est décrite de manière habilitante, classée IPC/CPC, et associée à une preuve d’horodatage (RFC 3161 / FreeTSA). Timestamp: 2025-09-13T16:43:19ZSHA-256: 55c629606ebc242c315596909e61875ab3b1653d366f1378dca5bda73e5f97b6 Liste des innovations & classification (IPC ; CPC) 1. Adaptive AI Tutor for Schools — IPC G09B 5/00 ; CPC G09B 5/022. Lightweight EEG Attention Coach — IPC A61B 5/04 ; CPC A61B 5/04083. Classroom CO2/PM Multi-sensor — IPC G01N 27/12 ; CPC F24F 11/524. PID-Driven HEPA Filtration — IPC F24F 8/80 ; CPC F24F 11/525. Salivary Cog-Biomarker Panel — IPC C12Q 1/70 ; CPC G16H 50/206. School Sleep Scheduling Trial — IPC G09B 19/00 ; CPC G16H 10/607. “Prescription Reading” for Infants — IPC A61B 10/00 ; CPC G16H 50/308. Cognitive AR Learning City — IPC G06T 19/00 ; CPC G09B 15/029. Federated “CognitCloud” Platform — IPC G16H 50/70 ; CPC G06N 20/0010. Reading Eye-Tracking Remediation — IPC A61B 5/107 ; CPC G06T 7/24611. Transdermal Melatonin Patch — IPC A61K 31/70 ; CPC A61M 37/0012. Closed-Loop tDCS for Attention — IPC A61N 1/36 ; CPC A61B 5/040813. Soil Lead Bioremediation — IPC C02F 3/28 ; CPC A62D 3/0014. Urinary Iodine Microfluidic Test — IPC G01N 21/78 ; CPC C12Q 1/2415. “Focus Mode” Network Filtering — IPC H04L 29/06 ; CPC G06F 21/6216. Low-Bandwidth Executive Games — IPC A63F 13/67 ; CPC G09B 5/1417. “Air→Score” Causal KPI — IPC G06Q 50/20 ; CPC G06F 16/245718. Standardized 10–12 min Battery — IPC G01N 33/50 ; CPC G06F 19/0019. Omega-3 School Meal Spec — IPC A23L 33/10 ; CPC A61K 36/2820. Cold-Chain for Bio-Samples — IPC B65D 81/38 ; CPC G16H 40/6321. “CognitFHIR” API Profiles — IPC G16H 10/60 ; CPC G06F 21/6222. Deep-Reading QA Module — IPC G09B 7/00 ; CPC G06F 3/04823. Classroom fNIRS Imaging — IPC A61B 5/145 ; CPC A61B 5/05524. “Offloading Index” — IPC G06Q 50/26 ; CPC G06F 16/2825. EEG Headset Manufacturing — IPC H01R 4/58 ; CPC A61B 5/040226. Wearable Sleep Multi-sensor — IPC A61B 5/11 ; CPC A61B 5/11127. Microbiome–Cognition Panel — IPC C12Q 1/6876 ; CPC G16H 50/3028. Combined Cognitive Supplement — IPC A61K 31/201 ; CPC A61P 25/2829. Population Cognitive Digital Twin — IPC G06N 10/00 ; CPC G16H 30/2030. NFC Microneedle Patch Process — IPC B29C 59/04 ; CPC A61M 37/0031. Arts & Cognition Program — IPC G09B 19/00 ; CPC A63H 2200/1232. Cognitive Data Blockchain — IPC G06Q 20/38 ; CPC G06N 20/2033. Wearable Enviro","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17114008","URL":"https://doi.org/10.5281/zenodo.17114008","source":"datacite"},{"id":"doi:10.5281/zenodo.17114009","type":"article-journal","title":"Effet Flynn Inverse : Analyse des Causes, Solutions et Scénarios Prospectifs","abstract":"Abstract ENThis document, produced with the assistance of ChatGPT o3 and GPT-5 Thinking, is released under the Apache 2.0 license. It is a voluntary defensive publication (prior art) and therefore enters the prior art upon release under the applicable patent statutes: EPC Art. 54(2) (European Patent Convention), French IPC/CPI Art. L 611-11 (French Intellectual Property Code), 35 U.S.C. §102(a) (United States Patent Act), Chinese Patent Law Art. 22(5) (中华人民共和国专利法), and Japanese Patent Act Art. 29(1) (特許法). This publication addresses the inverse Flynn effect through 65 enabling innovations across devices, assays, AI, closed-loop protocols, materials, therapies, manufacturing, imaging, delivery, data standards, UX, logistics, and governance. Every invention is described in an enabling manner, classified with IPC and CPC codes, and accompanied by timestamp proof (RFC 3161 / FreeTSA). Résumé FRCe document, produit avec l’assistance de ChatGPT o3 et GPT-5 Thinking, est diffusé sous licence Apache 2.0. Il constitue une publication défensive volontaire (antériorité) et entre, dès sa diffusion publique, dans l’art antérieur au titre des régimes suivants : CBE art. 54(2), CPI art. L 611-11, 35 U.S.C. §102(a), Loi chinoise sur les brevets art. 22(5) (中华人民共和国专利法), et Loi japonaise sur les brevets art. 29(1) (特許法). Cette publication traite de l’effet Flynn inverse à travers 65 innovations habilitantes couvrant dispositifs, capteurs, tests, IA, protocoles en boucle fermée, matériaux, thérapies, procédés de fabrication, imagerie, systèmes de délivrance, normes de données, UX, logistique et gouvernance. Chaque invention est décrite de manière habilitante, classée IPC/CPC, et associée à une preuve d’horodatage (RFC 3161 / FreeTSA). Timestamp: 2025-09-13T16:43:19ZSHA-256: 55c629606ebc242c315596909e61875ab3b1653d366f1378dca5bda73e5f97b6 Liste des innovations & classification (IPC ; CPC) 1. Adaptive AI Tutor for Schools — IPC G09B 5/00 ; CPC G09B 5/022. Lightweight EEG Attention Coach — IPC A61B 5/04 ; CPC A61B 5/04083. Classroom CO2/PM Multi-sensor — IPC G01N 27/12 ; CPC F24F 11/524. PID-Driven HEPA Filtration — IPC F24F 8/80 ; CPC F24F 11/525. Salivary Cog-Biomarker Panel — IPC C12Q 1/70 ; CPC G16H 50/206. School Sleep Scheduling Trial — IPC G09B 19/00 ; CPC G16H 10/607. “Prescription Reading” for Infants — IPC A61B 10/00 ; CPC G16H 50/308. Cognitive AR Learning City — IPC G06T 19/00 ; CPC G09B 15/029. Federated “CognitCloud” Platform — IPC G16H 50/70 ; CPC G06N 20/0010. Reading Eye-Tracking Remediation — IPC A61B 5/107 ; CPC G06T 7/24611. Transdermal Melatonin Patch — IPC A61K 31/70 ; CPC A61M 37/0012. Closed-Loop tDCS for Attention — IPC A61N 1/36 ; CPC A61B 5/040813. Soil Lead Bioremediation — IPC C02F 3/28 ; CPC A62D 3/0014. Urinary Iodine Microfluidic Test — IPC G01N 21/78 ; CPC C12Q 1/2415. “Focus Mode” Network Filtering — IPC H04L 29/06 ; CPC G06F 21/6216. Low-Bandwidth Executive Games — IPC A63F 13/67 ; CPC G09B 5/1417. “Air→Score” Causal KPI — IPC G06Q 50/20 ; CPC G06F 16/245718. Standardized 10–12 min Battery — IPC G01N 33/50 ; CPC G06F 19/0019. Omega-3 School Meal Spec — IPC A23L 33/10 ; CPC A61K 36/2820. Cold-Chain for Bio-Samples — IPC B65D 81/38 ; CPC G16H 40/6321. “CognitFHIR” API Profiles — IPC G16H 10/60 ; CPC G06F 21/6222. Deep-Reading QA Module — IPC G09B 7/00 ; CPC G06F 3/04823. Classroom fNIRS Imaging — IPC A61B 5/145 ; CPC A61B 5/05524. “Offloading Index” — IPC G06Q 50/26 ; CPC G06F 16/2825. EEG Headset Manufacturing — IPC H01R 4/58 ; CPC A61B 5/040226. Wearable Sleep Multi-sensor — IPC A61B 5/11 ; CPC A61B 5/11127. Microbiome–Cognition Panel — IPC C12Q 1/6876 ; CPC G16H 50/3028. Combined Cognitive Supplement — IPC A61K 31/201 ; CPC A61P 25/2829. Population Cognitive Digital Twin — IPC G06N 10/00 ; CPC G16H 30/2030. NFC Microneedle Patch Process — IPC B29C 59/04 ; CPC A61M 37/0031. Arts & Cognition Program — IPC G09B 19/00 ; CPC A63H 2200/1232. Cognitive Data Blockchain — IPC G06Q 20/38 ; CPC G06N 20/2033. Wearable Enviro","author":[{"family":"Pillet","given":"Xavier"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17114009","URL":"https://doi.org/10.5281/zenodo.17114009","source":"datacite"},{"id":"doi:10.5281/zenodo.17112704","type":"article-journal","title":"Federated Learning for Privacy-Preserving Healthcare AI Models","abstract":"Federated learning (FL) has emerged as a transformative paradigm for building collaborative healthcare AI models while safeguarding patient privacy and complying with regulations such as HIPAA and GDPR. Unlike centralized training, FL enables multiple hospitals and research centers to jointly develop a global model without exchanging raw data, thereby reducing risks of privacy breaches and promoting cross-institutional collaboration. This paper reviews recent literature (2020–2025) covering advances in privacy-preserving techniques including secure aggregation, differential privacy, and homomorphic encryption, and proposes a federated pipeline that integrates these methods for both electronic health records and medical imaging tasks. Simulated experiments with five clients illustrate that FL can achieve performance close to centralized models while substantially reducing exposure of sensitive health data, though trade-offs emerge in the form of reduced accuracy and added communication overhead. Beyond technical outcomes, the societal benefits of FL are significant: it fosters the development of AI models that generalize across diverse populations, supports early disease detection and personalized care, and enables resource-constrained institutions to contribute to and benefit from large-scale AI without compromising patient confidentiality. Ultimately, FL provides a pathway to equitable, trustworthy, and privacy-preserving healthcare innovation that can improve population health outcomes and strengthen societal trust in AI-driven medicine.","author":[{"family":"Saini","given":"Shyam"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17112704","URL":"https://doi.org/10.5281/zenodo.17112704","source":"datacite"},{"id":"doi:10.5281/zenodo.17112703","type":"article-journal","title":"Federated Learning for Privacy-Preserving Healthcare AI Models","abstract":"Federated learning (FL) has emerged as a transformative paradigm for building collaborative healthcare AI models while safeguarding patient privacy and complying with regulations such as HIPAA and GDPR. Unlike centralized training, FL enables multiple hospitals and research centers to jointly develop a global model without exchanging raw data, thereby reducing risks of privacy breaches and promoting cross-institutional collaboration. This paper reviews recent literature (2020–2025) covering advances in privacy-preserving techniques including secure aggregation, differential privacy, and homomorphic encryption, and proposes a federated pipeline that integrates these methods for both electronic health records and medical imaging tasks. Simulated experiments with five clients illustrate that FL can achieve performance close to centralized models while substantially reducing exposure of sensitive health data, though trade-offs emerge in the form of reduced accuracy and added communication overhead. Beyond technical outcomes, the societal benefits of FL are significant: it fosters the development of AI models that generalize across diverse populations, supports early disease detection and personalized care, and enables resource-constrained institutions to contribute to and benefit from large-scale AI without compromising patient confidentiality. Ultimately, FL provides a pathway to equitable, trustworthy, and privacy-preserving healthcare innovation that can improve population health outcomes and strengthen societal trust in AI-driven medicine.","author":[{"family":"Saini","given":"Shyam"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17112703","URL":"https://doi.org/10.5281/zenodo.17112703","source":"datacite"},{"id":"doi:10.48550/arxiv.2509.00332","type":"manuscript","title":"CryptoFace: End-to-End Encrypted Face Recognition","abstract":"Face recognition is central to many authentication, security, and personalized applications. Yet, it suffers from significant privacy risks, particularly arising from unauthorized access to sensitive biometric data. This paper introduces CryptoFace, the first end-to-end encrypted face recognition system with fully homomorphic encryption (FHE). It enables secure processing of facial data across all stages of a face-recognition process--feature extraction, storage, and matching--without exposing raw images or features. We introduce a mixture of shallow patch convolutional networks to support higher-dimensional tensors via patch-based processing while reducing the multiplicative depth and, thus, inference latency. Parallel FHE evaluation of these networks ensures near-resolution-independent latency. On standard face recognition benchmarks, CryptoFace significantly accelerates inference and increases verification accuracy compared to the state-of-the-art FHE neural networks adapted for face recognition. CryptoFace will facilitate secure face recognition systems requiring robust and provable security. The code is available at https://github.com/human-analysis/CryptoFace.","author":[{"family":"Ao","given":"Wei"},{"family":"Boddeti","given":"Vishnu"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2509.00332","URL":"https://doi.org/10.48550/arxiv.2509.00332","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.20000","type":"manuscript","title":"Can One Safety Loop Guard Them All? Agentic Guard Rails for Federated Computing","abstract":"We propose Guardian-FC, a novel two-layer framework for privacy preserving federated computing that unifies safety enforcement across diverse privacy preserving mechanisms, including cryptographic back-ends like fully homomorphic encryption (FHE) and multiparty computation (MPC), as well as statistical techniques such as differential privacy (DP). Guardian-FC decouples guard-rails from privacy mechanisms by executing plug-ins (modular computation units), written in a backend-neutral, domain-specific language (DSL) designed specifically for federated computing workflows and interchangeable Execution Providers (EPs), which implement DSL operations for various privacy back-ends. An Agentic-AI control plane enforces a finite-state safety loop through signed telemetry and commands, ensuring consistent risk management and auditability. The manifest-centric design supports fail-fast job admission and seamless extensibility to new privacy back-ends. We present qualitative scenarios illustrating backend-agnostic safety and a formal model foundation for verification. Finally, we outline a research agenda inviting the community to advance adaptive guard-rail tuning, multi-backend composition, DSL specification development, implementation, and compiler extensibility alongside human-override usability.","author":[{"family":"Veeraragavan","given":"Narasimha"},{"family":"Nygård","given":"Jan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.20000","URL":"https://doi.org/10.48550/arxiv.2506.20000","source":"datacite"},{"id":"doi:10.48550/arxiv.2506.11954","type":"manuscript","title":"Technical Evaluation of a Disruptive Approach in Homomorphic AI","abstract":"We present a technical evaluation of a new, disruptive cryptographic approach to data security, known as HbHAI (Hash-based Homomorphic Artificial Intelligence). HbHAI is based on a novel class of key-dependent hash functions that naturally preserve most similarity properties, most AI algorithms rely on. As a main claim, HbHAI makes now possible to analyze and process data in its cryptographically secure form while using existing native AI algorithms without modification, with unprecedented performances compared to existing homomorphic encryption schemes. We tested various HbHAI-protected datasets (non public preview) using traditional unsupervised and supervised learning techniques (clustering, classification, deep neural networks) with classical unmodified AI algorithms. This paper presents technical results from an independent analysis conducted with those different, off-the-shelf AI algorithms. The aim was to assess the security, operability and performance claims regarding HbHAI techniques. As a results, our results confirm most these claims, with only a few minor reservations.","author":[{"family":"Filiol","given":"Eric"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2506.11954","URL":"https://doi.org/10.48550/arxiv.2506.11954","source":"datacite"},{"id":"doi:10.48550/arxiv.2501.07535","type":"manuscript","title":"Code Generation for Cryptographic Kernels using Multi-word Modular Arithmetic on GPU","abstract":"Fully homomorphic encryption (FHE) and zero-knowledge proofs (ZKPs) are emerging as solutions for data security in distributed environments. However, the widespread adoption of these encryption techniques is hindered by their significant computational overhead, primarily resulting from core cryptographic operations that involve large integer arithmetic. This paper presents a formalization of multi-word modular arithmetic (MoMA), which breaks down large bit-width integer arithmetic into operations on machine words. We further develop a rewrite system that implements MoMA through recursive rewriting of data types, designed for compatibility with compiler infrastructures and code generators. We evaluate MoMA by generating cryptographic kernels, including basic linear algebra subprogram (BLAS) operations and the number theoretic transform (NTT), targeting various GPUs. Our MoMA-based BLAS operations outperform state-of-the-art multi-precision libraries by orders of magnitude, and MoMA-based NTTs achieve near-ASIC performance on commodity GPUs.","author":[{"family":"Zhang","given":"Naifeng"},{"family":"Franchetti","given":"Franz"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2501.07535","URL":"https://doi.org/10.48550/arxiv.2501.07535","source":"datacite"},{"id":"doi:10.48550/arxiv.2510.17333","type":"manuscript","title":"Comparison and performance analysis of dynamic encrypted control approaches","abstract":"Encrypted controllers using homomorphic encryption have proven to guarantee the privacy of measurement and control signals, as well as system and controller parameters, while regulating the system as intended. However, encrypting dynamic controllers has remained a challenge due to growing noise and overflow issues in the encoding. In this paper, we review recent approaches to dynamic encrypted control, such as bootstrapping, periodic resets of the controller state, integer reformulations, and FIR controllers, and equip them with a stability and performance analysis to evaluate their suitability. We complement the analysis with a numerical performance comparison on a benchmark system.","author":[{"family":"Schlor","given":"Sebastian"},{"family":"Allgöwer","given":"Frank"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2510.17333","URL":"https://doi.org/10.48550/arxiv.2510.17333","source":"datacite"},{"id":"doi:10.5281/zenodo.17385666","type":"article-journal","title":"A Systematic Literature Review on the Integration of Homomorphic Encryption and Chaotic Maps for  Secure Data Processing","abstract":"Homomorphic encryption and chaotic maps have emerged as key technologies in secure data processing, offering complementary benefits for confidentiality and computational integrity. While homomorphic encryption enables operations on encrypted data without decryption, chaotic maps provide strong pseudo-randomness and high sensitivity to initial conditions, making them highly suitable for cryptographic applications. This review examines their convergence, identifies shortcomings in prior studies, and outlines hybrid frameworks combining both methods. Results indicate that integrating chaotic maps with homomorphic encryption enhances security and performance, particularly for encrypted computation. However, computational cost and compatibility challenges remain. The review also highlights emerging trends such as lightweight adaptations for resource-constrained environments and novel chaotic map designs, providing a roadmap for more secure and scalable cryptographic systems in cloud computing and Internet of Things applications.","author":[{"family":"Abdullah Ghanim","given":"Jaber"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17385666","URL":"https://doi.org/10.5281/zenodo.17385666","source":"datacite"},{"id":"doi:10.5281/zenodo.17385667","type":"article-journal","title":"A Systematic Literature Review on the Integration of Homomorphic Encryption and Chaotic Maps for  Secure Data Processing","abstract":"Homomorphic encryption and chaotic maps have emerged as key technologies in secure data processing, offering complementary benefits for confidentiality and computational integrity. While homomorphic encryption enables operations on encrypted data without decryption, chaotic maps provide strong pseudo-randomness and high sensitivity to initial conditions, making them highly suitable for cryptographic applications. This review examines their convergence, identifies shortcomings in prior studies, and outlines hybrid frameworks combining both methods. Results indicate that integrating chaotic maps with homomorphic encryption enhances security and performance, particularly for encrypted computation. However, computational cost and compatibility challenges remain. The review also highlights emerging trends such as lightweight adaptations for resource-constrained environments and novel chaotic map designs, providing a roadmap for more secure and scalable cryptographic systems in cloud computing and Internet of Things applications.","author":[{"family":"Abdullah Ghanim","given":"Jaber"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17385667","URL":"https://doi.org/10.5281/zenodo.17385667","source":"datacite"},{"id":"doi:10.5281/zenodo.17367733","type":"article-journal","title":"The rise of federated systems in cloud-native architectures","abstract":"Federated systems represent a transformative shift in cloud computing architecture, enabling decentralized data processing while maintaining privacy and sovereignty. This technical review explores the evolution of federated approaches across multiple domains including machine learning, identity management, and cross-border collaboration. The article begins with core architectural principles that distinguish federation from traditional centralized models, including selective synchronization mechanisms and topological variations that optimize for different operational priorities. Privacy-preserving technologies like homomorphic encryption, secure multi-party computation, and differential privacy emerge as essential components for maintaining confidentiality in federated environments. Applications demonstrate particular efficacy in data-sensitive industries where regulatory considerations constrain traditional centralized approaches. Despite compelling advantages, federation introduces notable challenges in performance, resilience, and trust establishment across organizational boundaries. Future directions point toward lightweight protocols for resource-constrained environments, blockchain integration for enhanced accountability, and quantum-resistant cryptography for long-term security assurance. As federated architectures mature, standardization efforts will prove critical for widespread adoption beyond specialized use cases into mainstream enterprise deployments.","author":[{"family":"Bollam","given":"Akhilesh"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17367733","URL":"https://doi.org/10.5281/zenodo.17367733","source":"datacite"},{"id":"doi:10.5281/zenodo.17367732","type":"article-journal","title":"The rise of federated systems in cloud-native architectures","abstract":"Federated systems represent a transformative shift in cloud computing architecture, enabling decentralized data processing while maintaining privacy and sovereignty. This technical review explores the evolution of federated approaches across multiple domains including machine learning, identity management, and cross-border collaboration. The article begins with core architectural principles that distinguish federation from traditional centralized models, including selective synchronization mechanisms and topological variations that optimize for different operational priorities. Privacy-preserving technologies like homomorphic encryption, secure multi-party computation, and differential privacy emerge as essential components for maintaining confidentiality in federated environments. Applications demonstrate particular efficacy in data-sensitive industries where regulatory considerations constrain traditional centralized approaches. Despite compelling advantages, federation introduces notable challenges in performance, resilience, and trust establishment across organizational boundaries. Future directions point toward lightweight protocols for resource-constrained environments, blockchain integration for enhanced accountability, and quantum-resistant cryptography for long-term security assurance. As federated architectures mature, standardization efforts will prove critical for widespread adoption beyond specialized use cases into mainstream enterprise deployments.","author":[{"family":"Bollam","given":"Akhilesh"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17367732","URL":"https://doi.org/10.5281/zenodo.17367732","source":"datacite"},{"id":"doi:10.5281/zenodo.17164789","type":"article-journal","title":"Data Privacy and Security in AI","abstract":"The advancement of artificial intelligence (AI) technology has created both opportunities and risks in data protection and safeguarding. Given that numerous organizations are now employing AI systems in various industries, the goal of using data to drive innovation has never been urgent. This paper will review the relationship between AI, data privacy, and security and discuss the current issues and possible recommendations. Furthermore, this study introduces new approaches, including federated learning and homomorphic encryption, which preserve data integrity while still using the data. Using concrete examples from various industries, the study reveals practices and tendencies that company leaders should follow and avoid to achieve a proper balance. This paper offers an ethical approach that integrates practical recommendations for policymakers, technologists, and businesses to build user trust and progress responsibly and technically. Since most AI applications are based on big data, users’ data protection and systems’ performance and expandability are extremely important. This paper explores the issue of data privacy and security in AI and discusses promising strategies to address the problem, and guidelines for responsible AI implementation. The main priorities include the exposition of the algorithms, data anonymization methods, legal requirements, and strengthening cybersecurity. This paper presents real-life examples, and industry benchmarks to support the framework that can help organizations manage technologies in a way that addresses ethical concerns. In the future, the analysis presented in the study can help industries understand trends that help develop AI strategies that meet high privacy and security standards.","author":[{"family":"Kakarala","given":"Manikanta"},{"family":"Rongali","given":"Sateesh"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17164789","URL":"https://doi.org/10.5281/zenodo.17164789","source":"datacite"},{"id":"doi:10.5281/zenodo.17164788","type":"article-journal","title":"Data Privacy and Security in AI","abstract":"The advancement of artificial intelligence (AI) technology has created both opportunities and risks in data protection and safeguarding. Given that numerous organizations are now employing AI systems in various industries, the goal of using data to drive innovation has never been urgent. This paper will review the relationship between AI, data privacy, and security and discuss the current issues and possible recommendations. Furthermore, this study introduces new approaches, including federated learning and homomorphic encryption, which preserve data integrity while still using the data. Using concrete examples from various industries, the study reveals practices and tendencies that company leaders should follow and avoid to achieve a proper balance. This paper offers an ethical approach that integrates practical recommendations for policymakers, technologists, and businesses to build user trust and progress responsibly and technically. Since most AI applications are based on big data, users’ data protection and systems’ performance and expandability are extremely important. This paper explores the issue of data privacy and security in AI and discusses promising strategies to address the problem, and guidelines for responsible AI implementation. The main priorities include the exposition of the algorithms, data anonymization methods, legal requirements, and strengthening cybersecurity. This paper presents real-life examples, and industry benchmarks to support the framework that can help organizations manage technologies in a way that addresses ethical concerns. In the future, the analysis presented in the study can help industries understand trends that help develop AI strategies that meet high privacy and security standards.","author":[{"family":"Kakarala","given":"Manikanta"},{"family":"Rongali","given":"Sateesh"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17164788","URL":"https://doi.org/10.5281/zenodo.17164788","source":"datacite"},{"id":"doi:10.5281/zenodo.17075347","type":"article-journal","title":"Cloud Security and Privacy: A Systematic Review of Threats, Solutions, and Future Direction","abstract":"The protection of data saved, transferred, and processed in cloud settings has become a significant concern as cloud computing becomes the foundation of today's digital infrastructure. Cloud systems' multi-tenant design can be complicated, and standard security models may not take these factors into account. The current state of cloud security and privacy is thoroughly examined in this assessment, with data breaches, unsafe APIs, and violations of regulatory compliance being the most important issues. Covered are the primary methods of mitigation, such as secure key management, data encryption while in transit and at rest, and intrusion detection systems for real-time threat monitoring. The paper also speaks on the use of the latest privacy-preserving technologies like homomorphic encryption and biometric cryptosystems to preserve confidentiality and not reporting functionality. The risks and consequences of multi-tenancy and compliance are also discussed to determine the effects of shared infrastructures on data isolation and governance. With more enterprises moving workloads to the cloud, it is crucial to safeguard the data of the enterprise which includes its confidentiality, integrity, and availability. This review will provide an in-depth insight into the current security issues, research gaps, and future research trends in terms of adaptive security frameworks, standardization of security protocols, and integration of privacy-enhancing technologies towards fostering trustworthy and resilient cloud adoption.","author":[{"family":"Hussain","given":"Prof"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17075347","URL":"https://doi.org/10.5281/zenodo.17075347","source":"datacite"},{"id":"doi:10.5281/zenodo.17075348","type":"article-journal","title":"Cloud Security and Privacy: A Systematic Review of Threats, Solutions, and Future Direction","abstract":"The protection of data saved, transferred, and processed in cloud settings has become a significant concern as cloud computing becomes the foundation of today's digital infrastructure. Cloud systems' multi-tenant design can be complicated, and standard security models may not take these factors into account. The current state of cloud security and privacy is thoroughly examined in this assessment, with data breaches, unsafe APIs, and violations of regulatory compliance being the most important issues. Covered are the primary methods of mitigation, such as secure key management, data encryption while in transit and at rest, and intrusion detection systems for real-time threat monitoring. The paper also speaks on the use of the latest privacy-preserving technologies like homomorphic encryption and biometric cryptosystems to preserve confidentiality and not reporting functionality. The risks and consequences of multi-tenancy and compliance are also discussed to determine the effects of shared infrastructures on data isolation and governance. With more enterprises moving workloads to the cloud, it is crucial to safeguard the data of the enterprise which includes its confidentiality, integrity, and availability. This review will provide an in-depth insight into the current security issues, research gaps, and future research trends in terms of adaptive security frameworks, standardization of security protocols, and integration of privacy-enhancing technologies towards fostering trustworthy and resilient cloud adoption.","author":[{"family":"Hussain","given":"Prof"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17075348","URL":"https://doi.org/10.5281/zenodo.17075348","source":"datacite"},{"id":"doi:10.48550/arxiv.2507.20014","type":"manuscript","title":"Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation","abstract":"As AI-driven dataspaces become integral to data sharing and collaborative analytics, ensuring privacy, performance, and policy compliance presents significant challenges. This paper provides a comprehensive review of privacy-preserving and policy-aware AI techniques, including Federated Learning, Differential Privacy, Trusted Execution Environments, Homomorphic Encryption, and Secure Multi-Party Computation, alongside strategies for aligning AI with regulatory frameworks such as GDPR and the EU AI Act. We propose a novel taxonomy to classify these techniques based on privacy levels, performance impacts, and compliance complexity, offering a clear framework for practitioners and researchers to navigate trade-offs. Key performance metrics -- latency, throughput, cost overhead, model utility, fairness, and explainability -- are analyzed to highlight the multi-dimensional optimization required in dataspaces. The paper identifies critical research gaps, including the lack of standardized privacy-performance KPIs, challenges in explainable AI for federated ecosystems, and semantic policy enforcement amidst regulatory fragmentation. Future directions are outlined, proposing a conceptual framework for policy-driven alignment, automated compliance validation, standardized benchmarking, and integration with European initiatives like GAIA-X, IDS, and Eclipse EDC. By synthesizing technical, ethical, and regulatory perspectives, this work lays the groundwork for developing trustworthy, efficient, and compliant AI systems in dataspaces, fostering innovation in secure and responsible data-driven ecosystems.","author":[{"family":"Chandra","given":"Joydeep"},{"family":"Navneet","given":"Satyam"}],"issued":{"date-parts":[[2025]]},"DOI":"10.48550/arxiv.2507.20014","URL":"https://doi.org/10.48550/arxiv.2507.20014","source":"datacite"},{"id":"doi:10.5281/zenodo.21219655","type":"article-journal","title":"DIFFERENTIAL PRIVACY-BASED PROTECTION OF USER SHOPPING PREFERENCES IN E-COMMERCE SYSTEMS","abstract":"This study investigates the potential of differentiated privacy to safeguard the purchase preferences of online consumers while simultaneously enabling data analysis and personalized services. Protecting user privacy is a challenging endeavor due to the extensive collection of customer data by online purchasing websites. This study examines the potential of differential privacy techniques to conceal customers' private purchasing habits by introducing controlled noise to data without affecting the accuracy of the analysis. It also maintains a balance between data privacy and service quality to guarantee that enterprises can acquire critical knowledge while safeguarding customer data. The results indicate that differential privacy is a dependable and efficient method for fostering user confidence, alleviating privacy concerns, and facilitating secure data exchange in contemporary e-commerce.","author":[{"family":"Journal","given":"Advanced"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21219655","URL":"https://doi.org/10.5281/zenodo.21219655","source":"datacite"},{"id":"doi:10.5281/zenodo.20078041","type":"article-journal","title":"SHPhoneBench: A Multi-Modal Benchmark for Second-Hand Smartphone Valuation with Heterogeneous Inspection Signals","abstract":"SHPhoneBench is a multi-modal benchmark for second-hand smartphone valuation built from real recycling workflows in mainland China (January 2024 – March 2026). The full corpus contains 73,200 transaction-aligned records; a curated 15,000-sample benchmark subset carries five-tier condition grades, 18-category defect bounding boxes, cross-modal consistency labels, and differentially private transaction prices. Each record links four heterogeneous inspection signals collected during the same physical inspection:- V — multi-view exterior photographs of the device.- T — informal Chinese technician remarks (free-text).- S — structured CRM / device metadata (JSON).- IS — desktop diagnostic-tool screenshots. Tasks supported by the benchmark:- T1 — five-tier condition grading (Macro-F1, QWK).- T2 — 18-category defect detection / grounding (mAP@0.5, Grounding Acc@0.5).- T3 — binary cross-modal consistency check (F1).- T4 — calibrated price regression with 90% predictive interval (MAE in CNY, ECE). Privacy and compliance:- Faces, serial numbers, and Apple IDs are removed or masked prior to release.- List prices are released with Laplace differential privacy (epsilon = 0.1).- Collection consent is aligned with PIPL. This Zenodo record hosts the benchmark release archive, a sample manifest CSV, and a Croissant 1.1 metadata file (with MLCommons RAI 1.0 fields). Please see the accompanying paper for the full datasheet, ethics statement, and intended-use restrictions.","author":[{"family":"Authors","given":"Anonymous"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.20078041","URL":"https://doi.org/10.5281/zenodo.20078041","source":"datacite"},{"id":"doi:10.5281/zenodo.20077398","type":"article-journal","title":"SHPhoneBench: A Multi-Modal Benchmark for Second-Hand Smartphone Valuation with Heterogeneous Inspection Signals","abstract":"SHPhoneBench is a multi-modal benchmark for second-hand smartphone valuation built from real recycling workflows in mainland China (January 2024 – March 2026). The full corpus contains 73,200 transaction-aligned records; a curated 15,000-sample benchmark subset carries five-tier condition grades, 18-category defect bounding boxes, cross-modal consistency labels, and differentially private transaction prices. Each record links four heterogeneous inspection signals collected during the same physical inspection:- V — multi-view exterior photographs of the device.- T — informal Chinese technician remarks (free-text).- S — structured CRM / device metadata (JSON).- IS — desktop diagnostic-tool screenshots. Tasks supported by the benchmark:- T1 — five-tier condition grading (Macro-F1, QWK).- T2 — 18-category defect detection / grounding (mAP@0.5, Grounding Acc@0.5).- T3 — binary cross-modal consistency check (F1).- T4 — calibrated price regression with 90% predictive interval (MAE in CNY, ECE). Privacy and compliance:- Faces, serial numbers, and Apple IDs are removed or masked prior to release.- List prices are released with Laplace differential privacy (epsilon = 0.1).- Collection consent is aligned with PIPL. This Zenodo record hosts the benchmark release archive, a sample manifest CSV, and a Croissant 1.1 metadata file (with MLCommons RAI 1.0 fields). Please see the accompanying paper for the full datasheet, ethics statement, and intended-use restrictions.","author":[{"family":"Authors","given":"Anonymous"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.20077398","URL":"https://doi.org/10.5281/zenodo.20077398","source":"datacite"},{"id":"doi:10.5281/zenodo.20077399","type":"article-journal","title":"SHPhoneBench: A Multi-Modal Benchmark for Second-Hand Smartphone Valuation with Heterogeneous Inspection Signals","abstract":"SHPhoneBench is a multi-modal benchmark for second-hand smartphone valuation built from real recycling workflows in mainland China (January 2024 – March 2026). The full corpus contains 73,200 transaction-aligned records; a curated 15,000-sample benchmark subset carries five-tier condition grades, 18-category defect bounding boxes, cross-modal consistency labels, and differentially private transaction prices. Each record links four heterogeneous inspection signals collected during the same physical inspection:- V — multi-view exterior photographs of the device.- T — informal Chinese technician remarks (free-text).- S — structured CRM / device metadata (JSON).- IS — desktop diagnostic-tool screenshots. Tasks supported by the benchmark:- T1 — five-tier condition grading (Macro-F1, QWK).- T2 — 18-category defect detection / grounding (mAP@0.5, Grounding Acc@0.5).- T3 — binary cross-modal consistency check (F1).- T4 — calibrated price regression with 90% predictive interval (MAE in CNY, ECE). Privacy and compliance:- Faces, serial numbers, and Apple IDs are removed or masked prior to release.- List prices are released with Laplace differential privacy (epsilon = 0.1).- Collection consent is aligned with PIPL. This Zenodo record hosts the benchmark release archive, a sample manifest CSV, and a Croissant 1.1 metadata file (with MLCommons RAI 1.0 fields). Please see the accompanying paper for the full datasheet, ethics statement, and intended-use restrictions.","author":[{"family":"Authors","given":"Anonymous"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20077399","URL":"https://doi.org/10.5281/zenodo.20077399","source":"datacite"},{"id":"doi:10.5281/zenodo.18876341","type":"article-journal","title":"Information Physics for Safety-Critical AI","abstract":"This working paper establishes prior art for seventeen novel technical ideas derived by applying Knuth's Information Physics programme to safety-critical AI systems and cybersecurity. Background Kevin Knuth showed that physical laws -- probability theory, information theory, special relativity, quantum mechanics -- are not independent postulates. They are necessary consequences of one requirement: that partially-ordered sets (posets) be quantified consistently. Wherever a system has a natural ordering structure, that structure forces unique mathematical constraints on the system's behaviour. The author's earlier Topology of Reasoning (TOR) paper series (Zenodo DOIs: 10.5281/zenodo.18700538 and 10.5281/zenodo.18743583) established that the evidence graphs used in AI reasoning systems are themselves posets. Their topological invariants -- genus, Topological Slack, orientability -- are Knuthian valuations. Security theory becomes a sixth domain derivable from Knuth's framework. This paper extends that foundation into safety-critical AI and cybersecurity, identifying seventeen differentiators across four layers. What the Paper Contains Layer 1 -- Detection instruments These are structural and algebraic monitors derived from planarity theory and the Knuth product rule. The main near-term result is a dual-layer independence monitor that combines two mathematically orthogonal tests. The first is Topological Slack, a geometric certificate computable in O(E) time on the evidence graph. The second is a product-rule algebraic test, computable in O(1) time per sample at runtime. Together they provide two independent detection mechanisms for a correlated AI evidence failure mode called Mirror Hallucination, satisfying the IEC 61508 SIL-3 defence-in-depth requirement. Layer 2 -- Privacy-preserving forensic representation A forensic architecture in which only the topological structure of communications is stored and all content payloads are discarded before persistence. Topological invariants computed from the stored rotation-system encoding are sufficient to detect sophisticated attacks -- lateral movement, trust reversals, supply-chain depletion -- without any content ever being retained. Layer 3 -- Continuous security posture monitoring with adversarial awareness This layer addresses what happens when an adversary knows the topological certification thresholds and tries to operate within them. The countermeasure is forced-genus system design: legitimate operations are engineered so that every action necessarily increases genus, making topological silence architecturally impossible. A complementary technique inserts synthetic topological slack as honeypot subgraphs that trap an adversary who is optimising for silent attack paths. Layer 4 -- Geometric-dynamical early warning stack This layer operates before attacks execute rather than during them. A predictive staging detector identifies attack preparation by monitoring the resource allocation lattice for sum-rule violations introduced by newly staged elements. A cross-observer consistency check uses Minkowski invariant scalars derived from Knuth's causal poset construction to detect when multiple analysts or automated tools have developed inconsistent causal models of the same event stream. Relationship to Existing Work The differentiators described here are implemented through the AxoDen compositional AI safety kernel (v0.7.1, 236 automated tests), which produces mathematical certification artefacts rather than behavioural test results. The kernel is the subject of a separate publication corpus on Zenodo under ORCID 0009-0008-6435-3530. Intended Audience Researchers in AI safety, cybersecurity, topological data analysis, formal verification, and information theory. Standards bodies and certification authorities working with IEC 61508, DO-178C, ISO/IEC TR 5469:2024, and the EU AI Act. Defence and critical infrastructure procurement teams evaluating mathematically certifiable AI architectur","author":[{"family":"Yalcinkaya","given":"Erkan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18876341","URL":"https://doi.org/10.5281/zenodo.18876341","source":"datacite"},{"id":"doi:10.5281/zenodo.18876342","type":"article-journal","title":"Information Physics for Safety-Critical AI","abstract":"This working paper establishes prior art for seventeen novel technical ideas derived by applying Knuth's Information Physics programme to safety-critical AI systems and cybersecurity. Background Kevin Knuth showed that physical laws -- probability theory, information theory, special relativity, quantum mechanics -- are not independent postulates. They are necessary consequences of one requirement: that partially-ordered sets (posets) be quantified consistently. Wherever a system has a natural ordering structure, that structure forces unique mathematical constraints on the system's behaviour. The author's earlier Topology of Reasoning (TOR) paper series (Zenodo DOIs: 10.5281/zenodo.18700538 and 10.5281/zenodo.18743583) established that the evidence graphs used in AI reasoning systems are themselves posets. Their topological invariants -- genus, Topological Slack, orientability -- are Knuthian valuations. Security theory becomes a sixth domain derivable from Knuth's framework. This paper extends that foundation into safety-critical AI and cybersecurity, identifying seventeen differentiators across four layers. What the Paper Contains Layer 1 -- Detection instruments These are structural and algebraic monitors derived from planarity theory and the Knuth product rule. The main near-term result is a dual-layer independence monitor that combines two mathematically orthogonal tests. The first is Topological Slack, a geometric certificate computable in O(E) time on the evidence graph. The second is a product-rule algebraic test, computable in O(1) time per sample at runtime. Together they provide two independent detection mechanisms for a correlated AI evidence failure mode called Mirror Hallucination, satisfying the IEC 61508 SIL-3 defence-in-depth requirement. Layer 2 -- Privacy-preserving forensic representation A forensic architecture in which only the topological structure of communications is stored and all content payloads are discarded before persistence. Topological invariants computed from the stored rotation-system encoding are sufficient to detect sophisticated attacks -- lateral movement, trust reversals, supply-chain depletion -- without any content ever being retained. Layer 3 -- Continuous security posture monitoring with adversarial awareness This layer addresses what happens when an adversary knows the topological certification thresholds and tries to operate within them. The countermeasure is forced-genus system design: legitimate operations are engineered so that every action necessarily increases genus, making topological silence architecturally impossible. A complementary technique inserts synthetic topological slack as honeypot subgraphs that trap an adversary who is optimising for silent attack paths. Layer 4 -- Geometric-dynamical early warning stack This layer operates before attacks execute rather than during them. A predictive staging detector identifies attack preparation by monitoring the resource allocation lattice for sum-rule violations introduced by newly staged elements. A cross-observer consistency check uses Minkowski invariant scalars derived from Knuth's causal poset construction to detect when multiple analysts or automated tools have developed inconsistent causal models of the same event stream. Relationship to Existing Work The differentiators described here are implemented through the AxoDen compositional AI safety kernel (v0.7.1, 236 automated tests), which produces mathematical certification artefacts rather than behavioural test results. The kernel is the subject of a separate publication corpus on Zenodo under ORCID 0009-0008-6435-3530. Intended Audience Researchers in AI safety, cybersecurity, topological data analysis, formal verification, and information theory. Standards bodies and certification authorities working with IEC 61508, DO-178C, ISO/IEC TR 5469:2024, and the EU AI Act. Defence and critical infrastructure procurement teams evaluating mathematically certifiable AI architectur","author":[{"family":"Yalcinkaya","given":"Erkan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18876342","URL":"https://doi.org/10.5281/zenodo.18876342","source":"datacite"},{"id":"doi:10.4232/1.14772","type":"article-journal","title":"GESIS Panel.pop Population Sample – Standard Edition","abstract":"Das GESIS-Panel bietet eine wahrscheinlichkeitsbasierte Mixed-Mode-Access-Panel-Infrastruktur am GESIS Leibniz-Institut für Sozialwissenschaften in Mannheim. Das Projekt bietet der sozialwissenschaftlichen Community die Möglichkeit, Erhebungsdaten aus einer repräsentativen Stichprobe der deutschen Bevölkerung zu erheben. Die eingereichten Studienvorschläge werden auf der Grundlage eines wissenschaftlichen Begutachtungsverfahrens bewertet. Die Rekrutierung der Panelmitglieder erfolgte zunächst im Jahr 2013 in persönlichen Interviews, gefolgt von einer selbst durchgeführten Profilbefragung. Der Modus wurde von den Teilnehmern gewählt. Alle Teilnehmer der Profilbefragung werden als Mitglieder des Panels betrachtet und zu den alle zwei Monate stattfindenden regelmäßigen Wellen eingeladen. Die Startkohorte umfasste Anfang 2014 4900 Panelisten. Um den Panelabrieb zu kompensieren, wurde im Jahr 2016 eine Auffrischungsstichprobe mit Hilfe des German General Social Survey (ALLBUS) gezogen. Die erste Kohorte umfasst deutschsprachige Befragte im Alter zwischen 18 und 70 Jahren (zum Zeitpunkt der Einstellung) mit ständigem Wohnsitz in Deutschland, während die zweite Kohorte Befragte ab 18 Jahren ohne Obergrenze umfasst. Im Jahr 2018 wurde eine dritte Rekrutierungsstichprobe gezogen, die mit der Welle ge integriert wurde. Auch die dritte Kohorte umfasst Befragte ab 18 Jahren ohne Obergrenze.Rückwirkend wurden die Fälle bis einschließlich Welle fc (dritte Welle aus 2018) in den Daten ergänzt. Nähere Informationen finden Sie im Data Manual (ZA5664-65_sd_data-manual) und dem entsprechenden Rekrutierungsbericht (ZA5664-65_mb_recruitment2018). Die Stichproben des German General Social Survey (ALLBUS) basieren auf einer disproportionalen Stichprobe von Befragten aus West- und Ostdeutschland. Ein Designgewicht, das die Integration der beiden Rekrutierungskohorten ermöglicht, ist im Datensatz enthalten. Nähere Einzelheiten entnehmen Sie bitte den Methodenberichten der Einstellungsverfahren und dem GESIS-Panel-Referenzpapier (Bosnjak et al., 2017). Im März 2020 wurde eine Sondererhebung des GESIS-Panels zum Ausbruch des Coronavirus SARS-CoV-2 bzw. COVID-19 in Deutschland durchgeführt. Im Jahr 2021 wurde die vierte Rekrutierungsstichprobe mit Hilfe des German International Social Survey Programme (ISSP) gezogen, die mit der Welle ja integriert wurde. Die vierte Kohorte umfasst ebenfalls Befragte ab 18 Jahren ohne Obergrenze. Nähere Informationen finden Sie im entsprechenden Rekrutierungsbericht (ZA5664-65_r_i12.pdf). Im Jahr 2023 wurde die fünfte Rekrutierungsstichprobe mit Hilfe des German European Social Survey (ESS Round 11) gezogen, die mit der Welle la integriert wurde. Die fünfte Kohorte umfasst Befragte ab 18 Jahren ohne Obergrenze. Nähere Informationen finden Sie im entsprechenden Rekrutierungsbericht (ZA5664-65_r_k12.pdf). GESIS Panel Demographic Dataset Ab Version 43-0-0 ist der demografische Längsschnittdatensatz Teil des Veröffentlichungspaketes. Bei dem Datensatz handelt es sich um einen längsschnittlichen Datensatz (long format), mit harmonisierten Messungen zu demografischen Variablen: Befragten ID; Erhebungszeitpunkt; entsprechende Welle; Erhebungsjahr; Rekrutierungskohorte; Geschlecht des Befragten; Geburtsjahr; höchster Bildungsabschluss; persönliches Nettoeinkommen; Haushaltsnettoeinkommen; Familienstand; AAPOR disposition code; Einladungsmodus; Teilnahmemodus.","author":[{"family":"Gesis"}],"issued":{"date-parts":[[2026]]},"DOI":"10.4232/1.14772","URL":"https://doi.org/10.4232/1.14772","source":"datacite"},{"id":"doi:10.4232/1.14771","type":"article-journal","title":"GESIS Panel.pop Population Sample – Extended Edition","abstract":"Das GESIS-Panel bietet eine wahrscheinlichkeitsbasierte Mixed-Mode-Access-Panel-Infrastruktur am GESIS Leibniz-Institut für Sozialwissenschaften in Mannheim. Das Projekt bietet der sozialwissenschaftlichen Community die Möglichkeit, Erhebungsdaten aus einer repräsentativen Stichprobe der deutschen Bevölkerung zu erheben. Die eingereichten Studienvorschläge werden auf der Grundlage eines wissenschaftlichen Begutachtungsverfahrens bewertet. Die Rekrutierung der Panelmitglieder erfolgte zunächst im Jahr 2013 in persönlichen Interviews, gefolgt von einer selbst durchgeführten Profilbefragung. Der Modus wurde von den Teilnehmern gewählt. Alle Teilnehmer der Profilbefragung werden als Mitglieder des Panels betrachtet und zu den alle zwei Monate stattfindenden regelmäßigen Wellen eingeladen. Die Startkohorte umfasste Anfang 2014 4900 Panelisten. Um den Panelabrieb zu kompensieren, wurde im Jahr 2016 eine Auffrischungsstichprobe mit Hilfe des German General Social Survey (ALLBUS) gezogen. Die erste Kohorte umfasst deutschsprachige Befragte im Alter zwischen 18 und 70 Jahren (zum Zeitpunkt der Einstellung) mit ständigem Wohnsitz in Deutschland, während die zweite Kohorte Befragte ab 18 Jahren ohne Obergrenze umfasst. Im Jahr 2018 wurde eine dritte Rekrutierungsstichprobe gezogen, die mit der Welle ge integriert wurde. Auch die dritte Kohorte umfasst Befragte ab 18 Jahren ohne Obergrenze.Rückwirkend wurden die Fälle bis einschließlich Welle fc (dritte Welle aus 2018) in den Daten ergänzt. Nähere Informationen finden Sie im Data Manual (ZA5664-65_sd_data-manual) und dem entsprechenden Rekrutierungsbericht (ZA5664-65_mb_recruitment2018). Die Stichproben des German General Social Survey (ALLBUS) basieren auf einer disproportionalen Stichprobe von Befragten aus West- und Ostdeutschland. Ein Designgewicht, das die Integration der beiden Rekrutierungskohorten ermöglicht, ist im Datensatz enthalten. Nähere Einzelheiten entnehmen Sie bitte den Methodenberichten der Einstellungsverfahren und dem GESIS-Panel-Referenzpapier (Bosnjak et al., 2017). Im März 2020 wurde eine Sondererhebung des GESIS-Panels zum Ausbruch des Coronavirus SARS-CoV-2 bzw. COVID-19 in Deutschland durchgeführt. Im Jahr 2021 wurde die vierte Rekrutierungsstichprobe mit Hilfe des German International Social Survey Programme (ISSP) gezogen, die mit der Welle ja integriert wurde. Die vierte Kohorte umfasst ebenfalls Befragte ab 18 Jahren ohne Obergrenze. Nähere Informationen finden Sie im entsprechenden Rekrutierungsbericht (ZA5664-65_r_i12.pdf). Im Jahr 2023 wurde die fünfte Rekrutierungsstichprobe mit Hilfe des German European Social Survey (ESS Round 11) gezogen, die mit der Welle la integriert wurde. Die fünfte Kohorte umfasst Befragte ab 18 Jahren ohne Obergrenze. Nähere Informationen finden Sie im entsprechenden Rekrutierungsbericht (ZA5664-65_r_k12.pdf). GESIS Panel Demographic Dataset Ab Version 43-0-0 ist der demografische Längsschnittdatensatz Teil des Veröffentlichungspaketes. Bei dem Datensatz handelt es sich um einen längsschnittlichen Datensatz (long format), mit harmonisierten Messungen zu demografischen Variablen: Befragten ID; Erhebungszeitpunkt; entsprechende Welle; Erhebungsjahr; Rekrutierungskohorte; Geschlecht des Befragten; Geburtsjahr; Geburtsmonat; höchster Bildungsabschluss; persönliches Nettoeinkommen; Haushaltsnettoeinkommen; Familienstand; AAPOR disposition code; Einladungsmodus; Teilnahmemodus.","author":[{"family":"Gesis"}],"issued":{"date-parts":[[2026]]},"DOI":"10.4232/1.14771","URL":"https://doi.org/10.4232/1.14771","source":"datacite"},{"id":"doi:10.4232/1.14610","type":"article-journal","title":"GESIS Panel.pop Population Sample – Standard Edition","abstract":"Das GESIS-Panel bietet eine wahrscheinlichkeitsbasierte Mixed-Mode-Access-Panel-Infrastruktur am GESIS Leibniz-Institut für Sozialwissenschaften in Mannheim. Das Projekt bietet der sozialwissenschaftlichen Community die Möglichkeit, Erhebungsdaten aus einer repräsentativen Stichprobe der deutschen Bevölkerung zu erheben. Die eingereichten Studienvorschläge werden auf der Grundlage eines wissenschaftlichen Begutachtungsverfahrens bewertet. Die Rekrutierung der Panelmitglieder erfolgte zunächst im Jahr 2013 in persönlichen Interviews, gefolgt von einer selbst durchgeführten Profilbefragung. Der Modus wurde von den Teilnehmern gewählt. Alle Teilnehmer der Profilbefragung werden als Mitglieder des Panels betrachtet und zu den alle zwei Monate stattfindenden regelmäßigen Wellen eingeladen. Die Startkohorte umfasste Anfang 2014 4900 Panelisten. Um den Panelabrieb zu kompensieren, wurde im Jahr 2016 eine Auffrischungsstichprobe mit Hilfe des German General Social Survey (ALLBUS) gezogen. Die erste Kohorte umfasst deutschsprachige Befragte im Alter zwischen 18 und 70 Jahren (zum Zeitpunkt der Einstellung) mit ständigem Wohnsitz in Deutschland, während die zweite Kohorte Befragte ab 18 Jahren ohne Obergrenze umfasst. Im Jahr 2018 wurde eine dritte Rekrutierungsstichprobe gezogen, die mit der Welle ge integriert wurde. Auch die dritte Kohorte umfasst Befragte ab 18 Jahren ohne Obergrenze.Rückwirkend wurden die Fälle bis einschließlich Welle fc (dritte Welle aus 2018) in den Daten ergänzt. Nähere Informationen finden Sie im Data Manual (ZA5664-65_sd_data-manual) und dem entsprechenden Rekrutierungsbericht (ZA5664-65_mb_recruitment2018). Die Stichproben des German General Social Survey (ALLBUS) basieren auf einer disproportionalen Stichprobe von Befragten aus West- und Ostdeutschland. Ein Designgewicht, das die Integration der beiden Rekrutierungskohorten ermöglicht, ist im Datensatz enthalten. Nähere Einzelheiten entnehmen Sie bitte den Methodenberichten der Einstellungsverfahren und dem GESIS-Panel-Referenzpapier (Bosnjak et al., 2017). Im März 2020 wurde eine Sondererhebung des GESIS-Panels zum Ausbruch des Coronavirus SARS-CoV-2 bzw. COVID-19 in Deutschland durchgeführt. Im Jahr 2021 wurde die vierte Rekrutierungsstichprobe mit Hilfe des German International Social Survey Programme (ISSP) gezogen, die mit der Welle ja integriert wurde. Die vierte Kohorte umfasst ebenfalls Befragte ab 18 Jahren ohne Obergrenze. Nähere Informationen finden Sie im entsprechenden Rekrutierungsbericht (ZA5664-65_r_i12.pdf). Im Jahr 2023 wurde die fünfte Rekrutierungsstichprobe mit Hilfe des German European Social Survey (ESS Round 11) gezogen, die mit der Welle la integriert wurde. Die fünfte Kohorte umfasst Befragte ab 18 Jahren ohne Obergrenze. Nähere Informationen finden Sie im entsprechenden Rekrutierungsbericht (ZA5664-65_r_k12.pdf). GESIS Panel Demographic Dataset Ab Version 43-0-0 ist der demografische Längsschnittdatensatz Teil des Veröffentlichungspaketes. Bei dem Datensatz handelt es sich um einen längsschnittlichen Datensatz (long format), mit harmonisierten Messungen zu demografischen Variablen: Befragten ID; Erhebungszeitpunkt; entsprechende Welle; Erhebungsjahr; Rekrutierungskohorte; Geschlecht des Befragten; Geburtsjahr; höchster Bildungsabschluss; persönliches Nettoeinkommen; Haushaltsnettoeinkommen; Familienstand; AAPOR disposition code; Einladungsmodus; Teilnahmemodus.","author":[{"family":"Gesis"}],"issued":{"date-parts":[[2026]]},"DOI":"10.4232/1.14610","URL":"https://doi.org/10.4232/1.14610","source":"datacite"},{"id":"doi:10.4232/1.14609","type":"article-journal","title":"GESIS Panel.pop Population Sample – Extended Edition","abstract":"Das GESIS-Panel bietet eine wahrscheinlichkeitsbasierte Mixed-Mode-Access-Panel-Infrastruktur am GESIS Leibniz-Institut für Sozialwissenschaften in Mannheim. Das Projekt bietet der sozialwissenschaftlichen Community die Möglichkeit, Erhebungsdaten aus einer repräsentativen Stichprobe der deutschen Bevölkerung zu erheben. Die eingereichten Studienvorschläge werden auf der Grundlage eines wissenschaftlichen Begutachtungsverfahrens bewertet. Die Rekrutierung der Panelmitglieder erfolgte zunächst im Jahr 2013 in persönlichen Interviews, gefolgt von einer selbst durchgeführten Profilbefragung. Der Modus wurde von den Teilnehmern gewählt. Alle Teilnehmer der Profilbefragung werden als Mitglieder des Panels betrachtet und zu den alle zwei Monate stattfindenden regelmäßigen Wellen eingeladen. Die Startkohorte umfasste Anfang 2014 4900 Panelisten. Um den Panelabrieb zu kompensieren, wurde im Jahr 2016 eine Auffrischungsstichprobe mit Hilfe des German General Social Survey (ALLBUS) gezogen. Die erste Kohorte umfasst deutschsprachige Befragte im Alter zwischen 18 und 70 Jahren (zum Zeitpunkt der Einstellung) mit ständigem Wohnsitz in Deutschland, während die zweite Kohorte Befragte ab 18 Jahren ohne Obergrenze umfasst. Im Jahr 2018 wurde eine dritte Rekrutierungsstichprobe gezogen, die mit der Welle ge integriert wurde. Auch die dritte Kohorte umfasst Befragte ab 18 Jahren ohne Obergrenze.Rückwirkend wurden die Fälle bis einschließlich Welle fc (dritte Welle aus 2018) in den Daten ergänzt. Nähere Informationen finden Sie im Data Manual (ZA5664-65_sd_data-manual) und dem entsprechenden Rekrutierungsbericht (ZA5664-65_mb_recruitment2018). Die Stichproben des German General Social Survey (ALLBUS) basieren auf einer disproportionalen Stichprobe von Befragten aus West- und Ostdeutschland. Ein Designgewicht, das die Integration der beiden Rekrutierungskohorten ermöglicht, ist im Datensatz enthalten. Nähere Einzelheiten entnehmen Sie bitte den Methodenberichten der Einstellungsverfahren und dem GESIS-Panel-Referenzpapier (Bosnjak et al., 2017). Im März 2020 wurde eine Sondererhebung des GESIS-Panels zum Ausbruch des Coronavirus SARS-CoV-2 bzw. COVID-19 in Deutschland durchgeführt. Im Jahr 2021 wurde die vierte Rekrutierungsstichprobe mit Hilfe des German International Social Survey Programme (ISSP) gezogen, die mit der Welle ja integriert wurde. Die vierte Kohorte umfasst ebenfalls Befragte ab 18 Jahren ohne Obergrenze. Nähere Informationen finden Sie im entsprechenden Rekrutierungsbericht (ZA5664-65_r_i12.pdf). Im Jahr 2023 wurde die fünfte Rekrutierungsstichprobe mit Hilfe des German European Social Survey (ESS Round 11) gezogen, die mit der Welle la integriert wurde. Die fünfte Kohorte umfasst Befragte ab 18 Jahren ohne Obergrenze. Nähere Informationen finden Sie im entsprechenden Rekrutierungsbericht (ZA5664-65_r_k12.pdf). GESIS Panel Demographic Dataset Ab Version 43-0-0 ist der demografische Längsschnittdatensatz Teil des Veröffentlichungspaketes. Bei dem Datensatz handelt es sich um einen längsschnittlichen Datensatz (long format), mit harmonisierten Messungen zu demografischen Variablen: Befragten ID; Erhebungszeitpunkt; entsprechende Welle; Erhebungsjahr; Rekrutierungskohorte; Geschlecht des Befragten; Geburtsjahr; Geburtsmonat; höchster Bildungsabschluss; persönliches Nettoeinkommen; Haushaltsnettoeinkommen; Familienstand; AAPOR disposition code; Einladungsmodus; Teilnahmemodus.","author":[{"family":"Gesis"}],"issued":{"date-parts":[[2026]]},"DOI":"10.4232/1.14609","URL":"https://doi.org/10.4232/1.14609","source":"datacite"},{"id":"doi:10.17605/osf.io/9qdcn","type":"article-journal","title":"Determinants of Nurses’ Acceptance of a Meal Delivery Robot: The Role of Decision Making Style, Openness to Experience, Task–Technology Fit, and Trust","abstract":"1. Project Title Determinants of Nurses’ Acceptance of a Meal Delivery Robot: The Role of Decision Making Style, Openness to Experience, Task–Technology Fit, and Trust 2. Study Description This study investigates early attitudes of nursing staff/nursing students toward the potential introduction of a meal delivery robot in a clinical hospital setting. The conceptual model is adapted from Koh &amp; Yuen (2025), who propose an individual–task–technology fit framework for understanding autonomous delivery robot adoption (see also Doven et al. 2025). Early research shows that stable decision‑making styles shape how people form trust judgments (Betsch, 2004) about unfamiliar systems, including workplace technologies (Svenson et al., 2023; Svenson et al., 2024). Studies across sectors demonstrate that intuitive versus deliberative tendencies influence whether individuals rely on affective impressions or systematic evaluation when assessing automation, privacy, or expert systems, even before direct experience (Svenson et al., 2023; Svenson et al., 2024). In healthcare, where staff often evaluate innovations prospectively, these cognitive styles become especially relevant for understanding early trust in care robotics (Diab &amp; Demiris, 2025). Integrating decision‑making style and openness to experience therefore provides a theoretically grounded explanation for why some nurses expect a meal‑delivery robot to be reliable and helpful, whereas others remain cautious in the absence of hands‑on interaction. In our adaptation for healthcare robotics, we examine: • Decision making style (PID) and Openness to Experience as individual antecedents • Task–Technology Fit (TTF) as the cognitive evaluation • Trust in care robotics as the affective mediator • Intention to use as the behavioral outcome The study is conducted before nurses have direct experience with the robot. All items (in German language) are phrased as expectations. The goal is to obtain a short, low burden assessment suitable for clinical staff while maintaining theoretical rigor. Only the theoretical constructs relevant to the hypotheses are preregistered. Additional exploratory items related to workflow, safety, spatial constraints, and departmental specifics may be included in the field survey for implementation planning but are not part of the preregistered hypotheses or confirmatory analyses. 3. Hypotheses Individual Antecedents → Cognitive Evaluation H1a: Higher intuitive decision making (PID I) will be associated with lower perceived task–technology fit of the meal delivery robot. H1b: Higher deliberative decision making (PID D) will be associated with higher perceived task–technology fit. H2: Higher openness to experience will be associated with higher perceived task–technology fit. Cognitive Evaluation → Affective Evaluation H3: Task–technology fit positively predicts trust in care robotics. Affective Evaluation → Behavioral Outcome H4: Trust in care robotics positively predicts intention to use. Indirect Effects H5a: Decision making style will indirectly influence intention to use through task–technology fit and trust. H5b: Openness to experience will indirectly influence intention to use through task–technology fit and trust. H5c: Task–technology fit will indirectly influence intention to use through trust. 4. Variables and Measures Scientific Constructs (15 items total) Decision Making Style (PID short version – 4 items) Scale: 1 = trifft überhaupt nicht zu … 5 = trifft völlig zu • Ich vertraue auf meine Gefühle, wenn ich Entscheidungen treffe. • Wenn ich Entscheidungen treffe, höre ich auf mein Bauchgefühl. • Bevor ich eine Entscheidung treffe, denke ich alles sorgfältig durch. • Ich informiere mich gründlich, bevor ich eine Entscheidung treffe. Openness to Experience (2–3 items) Scale: 1 = trifft überhaupt nicht zu … 5 = trifft völlig zu • Ich bin offen für neue Erfahrungen. • Ich probiere gerne neue Dinge aus. • Ich interessiere mich für neue Ideen und Herangehe","author":[{"family":"Svenson","given":"Frithiof"}],"issued":{"date-parts":[[2026]]},"DOI":"10.17605/osf.io/9qdcn","URL":"https://doi.org/10.17605/osf.io/9qdcn","source":"datacite"},{"id":"doi:10.5281/zenodo.18512218","type":"article-journal","title":"Pre-Inference Governance — Master Specification A Canonical Regulatory Guide for Legitimate, Auditable, and Insurable AI Systems","abstract":"This master specification consolidates the canonical A7SEM / ASOSE / Pre-Inference Governance / JAQ documents into a coherent regulatory guide for oversight authorities, auditors, and policy institutions. It defines an ex-ante governance architecture in which inference and high-risk computation are treated as conditional, authorizable capabilities rather than automatic system rights. The guide establishes epistemic maturity, semantic boundary integrity, explicit authorization, and sufficient human authority as prerequisites for legitimate, auditable, and insurable AI operation. The specification formalizes glassbox governance as an ex-ante architectural property and distinguishes it from post-hoc blackbox explainability approaches. No implementation instructions, operational procedures, or technical integration guidance are provided.Commercial or operational use requires a separate written license from the author. Governance Clarification — Authorization & Execution (Canonical Note)Authorization is a prerequisite for legitimate use.Legitimate use is a prerequisite for insurability.Insurability is a prerequisite for scale, board approval, and procurement eligibility.Within the A7SEM / Pre-Inference Governance architecture, authorization is a necessary condition for legitimate system operation.Where authorization is absent or indeterminate, execution SHALL NOT proceed.Non-execution in such cases constitutes correct and intended governance behavior and does not represent malfunction, failure, or denial of service.This document is descriptive and non-operational.It grants no operational, commercial, or deployment rights.Any operational, commercial, or production use requires a separate written license from the rights holder. “Canonical semantics are protected by non-circumvention and licensing terms; absent authorization, non-use is required.” Hybrid License & Default State (Fail-Closed). Public access grants reference and citation rights only (CC BY-NC-ND 4.0 for the text). It grants no rights to implement, deploy, integrate, or rely on the governance architecture/semantics for operational, commercial, or institutional purposes (including procurement, audit, assurance, certification, or underwriting). Without explicit written authorization and a paid license, non-use / non-execution is the required posture. Interpretative Scope. This specification adopts an authorization-centric reading of applicable AI regulation, in which compliance is assessed primarily through ex-ante authorization and legitimacy conditions, not post-hoc documentation or output review. Canonical Anchor — Authorization-Centric Governance Statement This specification adopts an authorization-centric architectural reading of AI governance, including high-risk regulatory environments such as Regulation (EU) 2024/1689 (EU AI Act). In this model, inference is not treated as an automatic system right but as a conditional capability requiring prior structural authorization. Authorization precedes execution. Licensing precedes integration. Where ex-ante legitimacy cannot be established, non-execution constitutes the correct governance state. Post-hoc documentation, logging, or explainability mechanisms do not substitute for structural authorization gating. This document is canonical and non-operational; it does not provide implementation guidance, legal advice, compliance certification, or operational rights. Commercial or structural use requires a separate written license. All structured inquiries must be submitted in writing to moakarkach@hotmail.de. Default system state: fail-closed. ----------------------------------------------- Terminology Notice: “Pre-Inference Governance” (Canonical Meaning & Scope) Abstract This notice clarifies the canonical meaning of “Pre-Inference Governance” as an architectural legitimacy and authorization paradigm. It explicitly distinguishes the term from operational, infrastructural, or machine-learning usages of “pre-inference” that r","author":[{"family":"Akarkach","given":"Mounir"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18512218","URL":"https://doi.org/10.5281/zenodo.18512218","source":"datacite"},{"id":"doi:10.5281/zenodo.18512219","type":"article-journal","title":"Pre-Inference Governance — Master Specification A Canonical Regulatory Guide for Legitimate, Auditable, and Insurable AI Systems","abstract":"This master specification consolidates the canonical A7SEM / ASOSE / Pre-Inference Governance / JAQ documents into a coherent regulatory guide for oversight authorities, auditors, and policy institutions. It defines an ex-ante governance architecture in which inference and high-risk computation are treated as conditional, authorizable capabilities rather than automatic system rights. The guide establishes epistemic maturity, semantic boundary integrity, explicit authorization, and sufficient human authority as prerequisites for legitimate, auditable, and insurable AI operation. The specification formalizes glassbox governance as an ex-ante architectural property and distinguishes it from post-hoc blackbox explainability approaches. No implementation instructions, operational procedures, or technical integration guidance are provided.Commercial or operational use requires a separate written license from the author. Governance Clarification — Authorization & Execution (Canonical Note)Authorization is a prerequisite for legitimate use.Legitimate use is a prerequisite for insurability.Insurability is a prerequisite for scale, board approval, and procurement eligibility.Within the A7SEM / Pre-Inference Governance architecture, authorization is a necessary condition for legitimate system operation.Where authorization is absent or indeterminate, execution SHALL NOT proceed.Non-execution in such cases constitutes correct and intended governance behavior and does not represent malfunction, failure, or denial of service.This document is descriptive and non-operational.It grants no operational, commercial, or deployment rights.Any operational, commercial, or production use requires a separate written license from the rights holder. “Canonical semantics are protected by non-circumvention and licensing terms; absent authorization, non-use is required.” Hybrid License & Default State (Fail-Closed). Public access grants reference and citation rights only (CC BY-NC-ND 4.0 for the text). It grants no rights to implement, deploy, integrate, or rely on the governance architecture/semantics for operational, commercial, or institutional purposes (including procurement, audit, assurance, certification, or underwriting). Without explicit written authorization and a paid license, non-use / non-execution is the required posture. Interpretative Scope. This specification adopts an authorization-centric reading of applicable AI regulation, in which compliance is assessed primarily through ex-ante authorization and legitimacy conditions, not post-hoc documentation or output review. Canonical Anchor — Authorization-Centric Governance Statement This specification adopts an authorization-centric architectural reading of AI governance, including high-risk regulatory environments such as Regulation (EU) 2024/1689 (EU AI Act). In this model, inference is not treated as an automatic system right but as a conditional capability requiring prior structural authorization. Authorization precedes execution. Licensing precedes integration. Where ex-ante legitimacy cannot be established, non-execution constitutes the correct governance state. Post-hoc documentation, logging, or explainability mechanisms do not substitute for structural authorization gating. This document is canonical and non-operational; it does not provide implementation guidance, legal advice, compliance certification, or operational rights. Commercial or structural use requires a separate written license. All structured inquiries must be submitted in writing to moakarkach@hotmail.de. Default system state: fail-closed. ----------------------------------------------- Terminology Notice: “Pre-Inference Governance” (Canonical Meaning & Scope) Abstract This notice clarifies the canonical meaning of “Pre-Inference Governance” as an architectural legitimacy and authorization paradigm. It explicitly distinguishes the term from operational, infrastructural, or machine-learning usages of “pre-inference” that r","author":[{"family":"Akarkach","given":"Mounir"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18512219","URL":"https://doi.org/10.5281/zenodo.18512219","source":"datacite"},{"id":"oa:W4410771509","type":"article-journal","title":"DRACO: Decentralized Asynchronous Federated Learning Over Row-Stochastic Wireless Networks","abstract":"Emerging technologies and use cases, such as smart Internet of Things (IoT), Internet of Agents, and Edge AI, have generated significant interest in training neural networks over fully decentralized, serverless networks. A major obstacle in this context is ensuring stable convergence without imposing stringent assumptions, such as identical data distributions across devices or synchronized updates. In this paper, we introduce DRACO, a novel framework for decentralized asynchronous Stochastic Gradient Descent (SGD) over row-stochastic gossip wireless networks. Our approach leverages continuous communication, allowing edge devices to perform local training and exchange model updates along a continuous timeline, thereby eliminating the need for synchronized timing. Additionally, our algorithm decouples communication and computation schedules, enabling complete autonomy for all users while effectively addressing straggler issues. Through a thorough convergence analysis, we show that DRACO achieves high performance in decentralized optimization while maintaining low variance across users even without predefined scheduling policies. Numerical experiments further validate the effectiveness of our approach, demonstrating that controlling the maximum number of received messages per client significantly reduces redundant communication costs while maintaining robust learning performance.","author":[{"family":"Jeong","given":"Eunjeong"},{"family":"Kountouris","given":"Marios"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/ojcoms.2025.3574098","URL":"https://doi.org/10.1109/ojcoms.2025.3574098","source":"openalex"},{"id":"oa:W4416066352","type":"article-journal","title":"Advancements in Small-Object Detection (2023–2025): Approaches, Datasets, Benchmarks, Applications, and Practical Guidance","abstract":"Small-object detection (SOD) remains an important and growing challenge in computer vision and is the backbone of many applications, including autonomous vehicles, aerial surveillance, medical imaging, and industrial quality control. Small objects, in pixels, lose discriminative features during deep neural network processing, making them difficult to disentangle from background noise and other artifacts. This survey presents a comprehensive and systematic review of the SOD advancements between 2023 and 2025, a period marked by the maturation of transformer-based architectures and a return to efficient, realistic deployment. We applied the PRISMA methodology for this work, yielding 112 seminal works in the field to ensure the robustness of our foundation for this study. We present a critical taxonomy of the developments since 2023, arranged in five categories: (1) multiscale feature learning; (2) transformer-based architectures; (3) context-aware methods; (4) data augmentation enhancements; and (5) advancements to mainstream detectors (e.g., YOLO). Third, we describe and analyze the evolving SOD-centered datasets and benchmarks and establish the importance of evaluating models fairly. Fourth, we contribute a comparative assessment of state-of-the-art models, evaluating not only accuracy (e.g., the average precision for small objects (AP_S)) but also important efficiency (FPS, latency, parameters, GFLOPS) metrics across standardized hardware platforms, including edge devices. We further use data-driven case studies in the remote sensing, manufacturing, and healthcare domains to create a bridge between academic benchmarks and real-world performance. Finally, we summarize practical guidance for practitioners, the model selection decision matrix, scenario-based playbooks, and the deployment checklist. The goal of this work is to help synthesize the recent progress, identify the primary limitations in SOD, and open research directions, including the potential future role of generative AI and foundational models, to address the long-standing data and feature representation challenges that have limited SOD.","author":[{"family":"Aldubaikhi","given":"Ali"},{"family":"Patel","given":"Sarosh"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/app152211882","URL":"https://doi.org/10.3390/app152211882","source":"openalex"},{"id":"oa:W4417331005","type":"article-journal","title":"Federated Transfer Learning for Tomato Leaf Disease Detection Using Neuro-Graph Hybrid Model","abstract":"Plant diseases are currently a major threat to agricultural economies and food availability, having a negative environmental impact. Despite being a promising line of research, current approaches struggle with poor cross-site generalization, limited labels and dataset bias. Real-field complexities, such as environmental variability, heterogeneous varieties or temporal dynamics as are often overlooked. Numerous studies have been conducted to address these challenges, proposing advanced learning strategies and improved evaluation protocols. Synthetic data generation and self-supervised learning reduce dataset bias, while domain adaptation, hyperspectral, and thermal signals improve robustness across sites. However, a large portion of current methods are developed and validated mainly on clean laboratory datasets, which do not capture the variability of real-field conditions. Existing AI models often lead to imperfect detection results when dealing with field images complexities, such as dense vegetation, variable illumination or changing symptom expression. Although augmentation techniques can approximate real-world conditions, incorporating field data represents a substantial enhancement in model reliability. Federated transfer learning offers a promising approach to enhance plant disease detection, by enabling collaborative training of models across diverse agricultural environments, using in-field data but without disclosing the participants data to each others. In this study, we collaboratively trained a hybrid Graph–SNN model using federated learning (FL) to preserve data privacy, optimized for efficient use of participant resources. The model achieved an accuracy of 0.9445 on clean laboratory data and 0.6202 exclusively on field data, underscoring the considerable challenges posed by real-world conditions. Our findings demonstrate the potential of FL for privacy preserving and reliable plant disease detection under real field conditions.","author":[{"family":"Cristea","given":"Aurora"},{"family":"Dobre","given":"Ciprian"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/agriengineering7120432","URL":"https://doi.org/10.3390/agriengineering7120432","source":"openalex"},{"id":"oa:W4412711544","type":"article-journal","title":"Robust Federated Learning Against Data Poisoning Attacks: Prevention and Detection of Attacked Nodes","abstract":"Federated learning (FL) enables collaborative model building among a large number of participants without sharing sensitive data to the central server. Because of its distributed nature, FL has limited control over local data and the corresponding training process. Therefore, it is susceptible to data poisoning attacks where malicious workers use malicious training data to train the model. Furthermore, attackers on the worker side can easily manipulate local data by swapping the labels of training instances, adding noise to training instances, and adding out-of-distribution training instances in the local data to initiate data poisoning attacks. And local workers under such attacks carry incorrect information to the server, poison the global model, and cause misclassifications. So, the prevention and detection of such data poisoning attacks is crucial to build a robust federated training framework. To address this, we propose a prevention strategy in federated learning, namely confident federated learning, to protect workers from such data poisoning attacks. Our proposed prevention strategy at first validates the label quality of local training samples by characterizing and identifying label errors in the local training data, and then excludes the detected mislabeled samples from the local training. To this aim, we experiment with our proposed approach on both the image and audio domains, and our experimental results validated the robustness of our proposed confident federated learning in preventing the data poisoning attacks. Our proposed method can successfully detect the mislabeled training samples with above 85% accuracy and exclude those detected samples from the training set to prevent data poisoning attacks on the local workers. However, our prevention strategy can successfully prevent the attack locally in the presence of a certain percentage of poisonous samples. Beyond that percentage, the prevention strategy may not be effective in preventing attacks. In such cases, detection of the attacked workers is needed. So, in addition to the prevention strategy, we propose a novel detection strategy in the federated learning framework to detect the malicious workers under attack. We propose to create a class-wise cluster representation for every participating worker by utilizing the neuron activation maps of local models and analyze the resulting clusters to filter out the workers under attack before model aggregation. We experimentally demonstrated the efficacy of our proposed detection strategy in detecting workers affected by data poisoning attacks, along with the attack types, e.g., label-flipping or dirty labeling. In addition, our experimental results suggest that the global model could not converge even after a large number of training rounds in the presence of malicious workers, whereas after detecting the malicious workers with our proposed detection method and discarding them from model aggregation, we ensured that the global model achieved convergence within very few training rounds. Furthermore, our proposed approach stays robust under different data distributions and model sizes and does not require prior knowledge about the number of attackers in the system.","author":[{"family":"Ovi","given":"Pretom"},{"family":"Gangopadhyay","given":"Aryya"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/electronics14152970","URL":"https://doi.org/10.3390/electronics14152970","source":"openalex"},{"id":"oa:W4413887240","type":"article-journal","title":"Deep Learning Approaches for EEG-Motor Imagery-Based BCIs: Current Models, Generalization Challenges, and Emerging Trends","abstract":"This study critically examines the evolution of deep learning (DL) for electroencephalogram (EEG) based motor imagery (MI) decoding with a focus on real-time Brain Computer Interfaces (BCIs) development. Prior studies often prioritize accuracy in isolation, neglecting computational efficiency, interpretability, noise robustness, and neurophysiological variability across subjects and tasks, while recent DL advancements have introduced novel architectures to address these issues. This work systematically evaluates those novel architectures and emerging trends through addressing 4 research questions (RQs) based on an extensive review. Initially, over 188 papers from 3 databases were retrieved with a focus on publications from 2024 to 2025. Later, through multi-stage filtering based on strict inclusion criteria, a refined corpus of 68 high-quality studies was selected. This analysis reveals that state-of-the-art models achieve competitive accuracy, varying 85-100% on public datasets, but still face challenges in computational demands, noise resilience, generalization and BCI deployment. Additionally, preprocessing and integrated hybrid feature extraction paired with explainable AI (XAI) techniques are discussed. Emerging trends such as neuromorphic computing, federated learning (FL), and closed-loop adaptive systems offering solutions to current deployment barriers have been included in the discussion. Ethical and ecological considerations, such as data privacy, algorithmic bias, and energy efficiency, are notably represented in the literature. This review contributes a holistic framework for evaluating DL models, emphasizing the need to balance accuracy, efficiency, and adaptability. By synthesizing insights from large-scale datasets and explainability tools, this study exposes the limitations of current DL studies reliant on homogenous data, unavailability of codes to reproduce models and proposes strategies to mitigate neurophysiological variability. The finding underscores the urgency of prioritizing clinical relevance, ethical validation, and ecological robustness to bridge the lab to real-world divide, offering actionable directions for future research in low-power, generalizable, and user-centric BCI design.","author":[{"family":"Raza","given":"Aaqib"},{"family":"Yusoff","given":"Mohd"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/access.2025.3604528","URL":"https://doi.org/10.1109/access.2025.3604528","source":"openalex"},{"id":"oa:W4408791485","type":"article-journal","title":"Enhanced Privacy and Communication Efficiency in Non-IID Federated Learning With Adaptive Quantization and Differential Privacy","abstract":"Federated learning (FL) is a distributed machine learning method where multiple devices collaboratively train a model under the management of a central server without sharing underlying data. One of the key challenges of FL is the communication bottleneck caused by variations in connection speed and bandwidth across devices. Therefore, it is essential to reduce the size of transmitted data during training. Additionally, there is a potential risk of exposing sensitive information through the model or gradient analysis during training. To address both privacy and communication efficiency, we combine differential privacy (DP) and adaptive quantization methods. We use Laplacian-based DP to preserve privacy, which is relatively underexplored in FL and offers tighter privacy guarantees than Gaussian-based DP. We propose a simple and efficient global bit-length scheduler using round-based cosine annealing, along with a client-based scheduler that dynamically adapts based on client contribution estimated through dataset entropy analysis. We evaluate our approach through extensive experiments on CIFAR10, MNIST, and medical imaging datasets, using non-IID data distributions across varying client counts, bit-length schedulers, and privacy budgets. The results show that our adaptive quantization methods reduce total communicated data by up to 52.64% for MNIST, 45.06% for CIFAR10, and 31% to 37% for medical imaging datasets compared to 32-bit float training while maintaining competitive model accuracy and ensuring robust privacy through DP.","author":[{"family":"Ardıç","given":"Emre"},{"family":"Genç","given":"Yakup"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1109/access.2025.3554138","URL":"https://doi.org/10.1109/access.2025.3554138","source":"openalex"},{"id":"oa:W4413637086","type":"article-journal","title":"Identifying significant features in adversarial attack detection framework using federated learning empowered medical IoT network security","abstract":"The expansion of the Internet of Medical Things (IoHT) presents significant advantages for healthcare over improved data-driven insights and connectivity and offers critical cybersecurity challenges. Attacks are a serious risk for neural network security; recent defence mechanisms remain restricted concerning their applicability to real-world environments. The influence of adversarial attacks is essential, as they can challenge the security and reliability of Artificial Intelligence (AI) methods in crucial applications. Dealing with these vulnerabilities is vital to develop strong and reliable NNs. Therefore, the study of adversarial defence mechanisms and attack detection became an important area in the domain of AI. Machine learning (ML) and specific deep learning (DL) models have recently influenced excellent performance on challenging perceptual tasks like adversarial attack detection. Meanwhile, the federated learning (FL) method is susceptible to attacks by malicious clients. FL can complete a considerable training task effectively by attracting participants for training a DL method cooperatively, and the user privacy should be completely protected for the users only upload model parameters to the centralized server. This study presents an Adversarial Attack Detection Framework Using Federated Learning Empowered IoT Medical (AADF-FLEIoTM) model. The main intention of the AADF-FLEIoTM model is to develop adversarial attack detection using FL and an advanced hybrid model. The data normalization stage initially uses min-max normalization to scale and transform data into a consistent range. The proposed AADF-FLEIoTM employs the marine predator algorithm (MPA) model to identify and retain the most relevant features for the feature selection process. Besides, the integration of convolutional neural networks, bidirectional long short-term memory, and self-attention (SA-CNN-BiLSTM) technique is utilized for the detection and classification process. Finally, the Red-Tail Hawk (RTH)-optimizer algorithm alters the hyperparameter values of the SA-CNN-BiLSTM technique optimally and results in more excellent classification performance. The AADF-FLEIoTM approach is examined on the IoT healthcare security dataset. The performance validation of the AADF-FLEIoTM approach illustrated a superior accuracy value of 98.24% over existing models.","author":[{"family":"Sharaf","given":"Sanaa"},{"family":"Nooh","given":"Sameer"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1038/s41598-025-14913-0","URL":"https://doi.org/10.1038/s41598-025-14913-0","source":"openalex"},{"id":"oa:W4410698050","type":"article-journal","title":"Federated learning for automotive applications","abstract":"This paper presents and evaluates a new distributed learning technology, called federated learning, and its applications in automotive systems. We review and classify existing approaches to federated learning, focusing on its implementation in connected vehicles. We also evaluate the challenges associated with this application domain. Federated learning allows data to remain within vehicles, thereby avoiding costly data transfers and mitigating privacy concerns. This technology shows promise in enhancing various automotive applications. Federated learning can be applied to driver assistance systems, predictive maintenance, and personalized user experiences in connected vehicles. By keeping data local, it supports privacy and reduces communication overhead. Federated learning represents a significant advancement in the use of connected vehicle data. It offers a practical solution to privacy and data transfer issues while enhancing the performance of automotive systems through collaborative model training.","author":[{"family":"Lindskog-Münzing","given":"William"},{"family":"Prehofer","given":"Christian"}],"issued":{"date-parts":[[2025]]},"DOI":"10.26599/htrd.2025.9480055","URL":"https://doi.org/10.26599/htrd.2025.9480055","source":"openalex"},{"id":"oa:W4406163034","type":"article-journal","title":"Collaboration of IoT devices in smart home scenarios: algorithm research based on graph neural networks and federated learning","abstract":"Traditional IoT device collaboration is usually static and cannot adjust the collaboration mode between devices according to various changes, which limits work efficiency. To this end, an IoT device collaboration optimization algorithm based on graph neural network and federated learning is studied. This method abstracts various IoT device nodes and their communication relationships into graph structured data for storage, and then uses federated learning to train the graph convolutional network with graph structured data. The obtained model can be used to optimize the collaboration mode of IoT devices. During the training process, the total average MSE (mean square error) between the output and the label of the graph convolutional network model based on federated learning is 0.968; the total standard deviation of MSE is 0.0353; the total time from training to model convergence is 435.82 s, of which data transmission time accounts for 27.1% and model training time accounts for 72.9%. In a 2-h practical experiment, the graph convolutional network model based on federated learning was used to optimize the collaboration mode of smart homes, achieving a target environment residence time of 87 min and a total power consumption reduction of 0.69 kW·h. The results show that this method can effectively optimize the collaboration efficiency of IoT devices, reduce training time and network overhead, but it fails to improve the prediction accuracy of the model and may also lead to a decrease in stability.","author":[{"family":"Zhong","given":"Yuanquan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s43926-025-00096-7","URL":"https://doi.org/10.1007/s43926-025-00096-7","source":"openalex"},{"id":"oa:W4412454314","type":"article-journal","title":"Latency Analysis of UAV-Assisted Vehicular Communications Using Personalized Federated Learning with Attention Mechanism","abstract":"In this paper, unmanned aerial vehicle (UAV)-assisted vehicular communications are investigated to minimize latency and maximize the utilization of available UAV battery power. As communication and cooperation among UAV and vehicles is frequently required, a viable approach is to reduce the transmission of redundant messages. However, when the sensor data captured by the varying number of vehicles is not independent and identically distributed (non-i.i.d.), this becomes challenging. Hence, in order to group the vehicles with similar data distributions in a cluster, we utilize federated learning (FL) based on an attention mechanism. We jointly maximize the UAV’s available battery power in each transmission window and minimize communication latency. The simulation experiments reveal that the proposed personalized FL approach achieves performance improvement compared with baseline FL approaches. Our model, trained on the V2X-Sim dataset, outperforms existing methods on key performance indicators. The proposed FL approach with an attention mechanism offers a reduction in communication latency by up to 35% and a significant reduction in computational complexity without degradation in performance. Specifically, we achieve an improvement of approximately 40% in UAV energy efficiency, 20% reduction in the communication overhead, and 15% minimization in sojourn time.","author":[{"family":"Gupta","given":"Abhishek"},{"family":"Fernando","given":"Xavier"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/drones9070497","URL":"https://doi.org/10.3390/drones9070497","source":"openalex"},{"id":"oa:W4414146309","type":"article-journal","title":"A deep learning/machine learning approach for anomaly based network intrusion detection","abstract":"Introduction: The increasing complexity and frequency of cybersecurity threats necessitate the development of advanced detection systems capable of identifying both known and emerging attacks. In this study, we present a hybrid anomaly-based Network Intrusion Detection System (NIDS) that integrates multiple machine learning and deep learning algorithms, including XGBoost, Random Forest, Graph Neural Networks (GNN), Long Short-Term Memory (LSTM) networks, and Autoencoders. Methods: The proposed system was trained on a large-scale dataset comprising over 5.6 million network traffic records. Comprehensive data preprocessing and feature engineering were applied, and the Synthetic Minority Over-sampling Technique (SMOTE) was employed to address class imbalance. To enhance robustness and generalization, a weighted soft-voting ensemble strategy was used to combine predictions from the individual models. Results: The experimental evaluation demonstrated near-perfect performance, with accuracy, precision, recall, and F1-score values approaching 100% on the primary dataset. These results were validated through rigorous 5-fold cross-validation. Discussion: Evaluation on an independent benchmark dataset confirmed the strong generalizability and robustness of the proposed model across diverse intrusion scenarios. These findings highlight the effectiveness of the hybrid ensemble framework in significantly improving intrusion detection capabilities within complex and dynamic network environments.","author":[{"family":"Al-Muhanna","given":"Reem"},{"family":"Dardouri","given":"Samia"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3389/frai.2025.1625891","URL":"https://doi.org/10.3389/frai.2025.1625891","source":"openalex"},{"id":"oa:W7116750991","type":"article-journal","title":"Deep learning for sustainable development across climate, energy, agriculture and urban systems","abstract":"Deep learning (DL) has emerged as a transformative paradigm for addressing the multidimensional and data-intensive challenges of sustainable development. The purpose of this review is to examine how DL contributes to four critical domains-climate action, sustainable energy, smart agriculture, and urban development-while identifying gaps that limit large-scale deployment. Methodologically, the study synthesizes peer-reviewed research published between 2000 and 2025, covering architectures such as convolutional neural networks, recurrent networks, transformers, graph neural networks, autoencoders, and multimodal frameworks. The findings reveal key trends including the rise of physics-informed models, the integration of deep reinforcement learning in energy and transport systems, and the increasing adoption of federated and edge AI for decentralized monitoring. At the same time, recurring challenges are identified: data scarcity, limited cross-regional generalizability, deficits in explainability, and ethical concerns surrounding fairness and accountability. The review concludes that addressing these issues requires hybrid physics–AI modeling, uncertainty-aware and participatory AI frameworks, and deployment-oriented research strategies. The key contributions and implications of this work are threefold: (i) the development of a cross-domain taxonomy mapping DL methods to sustainability tasks, (ii) benchmarking insights to guide model selection and evaluation, and (iii) a forward-looking research agenda to support researchers, practitioners, and policymakers. The originality of this review lies in its cross-sectoral synthesis, which extends beyond domain-specific surveys to highlight how DL can be responsibly scaled to advance the United Nations Sustainable Development Goals (SDGs).","author":[{"family":"Sharma","given":"Harshit"},{"family":"Kaur","given":"Simran"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s43621-025-02186-6","URL":"https://doi.org/10.1007/s43621-025-02186-6","source":"openalex"},{"id":"oa:W4417118783","type":"article-journal","title":"Integrating reinforcement learning and federated meta-learning for energy-efficient wireless body sensor networks","abstract":"Abstract Intra wireless body sensor network (Intra-WBSN) is typically a short range wireless health monitoring network, Intra wireless body sensor network (Intra-WBSN) is typically a short range wireless health monitoring network, Wireless body sensor networks (WBSNs) have emerged as a transformative technology for real-time, non-invasive healthcare monitoring, enabling continuous tracking of vital physiological parameters such as heart rate, blood pressure, and glucose levels. However, their widespread deployment is hindered by critical challenges including limited energy resources, dynamic network topologies due to body movements, and stringent quality of service requirements for reliability and low latency. Conventional routing protocols, which rely on static, rule-based mechanisms, lack the adaptability and foresight needed to operate efficiently in such highly variable environments. To address these limitations, this paper proposes RELIEF-Net (reinforcement learning and federated meta-learning for energy-efficient wireless body sensor networks), a novel AI-driven framework that synergistically integrates reinforcement learning (RL) for adaptive routing, long short-term memory (LSTM) networks for predictive analytics, graph neural networks (GNNs) with attention mechanisms for energy-aware clustering, and federated meta-learning (FML) for privacy-preserving, cross-patient model personalization. At its core, RELIEF-Net employs RL to make context-aware, real-time routing decisions that optimize transmission power, prioritize critical data (e.g., cardiac alerts), and proactively reroute traffic based on LSTM-predicted energy depletion and link instability. GNN-based clustering enhances network organization and ensures efficient, topology-aware data forwarding, while FML enables decentralized learning across patients, preserving data privacy and improving scalability without centralized data aggregation. Extensive simulations demonstrate that RELIEF-Net achieves a 23% improvement in network lifetime, a packet delivery ratio (PDR) ≥ 97%, and end-to-end latency ≤ 85 ms for critical data under diverse operational scenarios. By unifying predictive intelligence, adaptive control, and scalable privacy-conscious learning, RELIEF-Net establishes a robust and sustainable solution for intelligent healthcare IoT systems, paving the way for next-generation remote patient monitoring and improved clinical outcomes.","author":[{"family":"Othman","given":"Soufiane"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s44444-025-00081-z","URL":"https://doi.org/10.1007/s44444-025-00081-z","source":"openalex"},{"id":"oa:W4415380784","type":"article-journal","title":"Artificial intelligence and students’ cognitive learning outcomes with bibliometric and content analysis for future research agenda","abstract":"This study conducted a comprehensive bibliometric and content analysis to explore the integration of artificial intelligence in students’ cognitive learning outcomes. A structured TITLE-ABS-KEY search was performed in the Scopus database using keywords such as “Artificial Intelligence,” “AI,” “Students,” and “cognitive learning outcomes,” resulting in 318 documents published between 2016 and 2025. After filtering, a final dataset of 246 research articles and conference papers was analyzed. The methodology includes bibliometric performance analysis (covering publication trends, countries, affiliations, authors, and journals) and network analysis (comprising co-word, citation, co-authorship, and bibliographic coupling). Additionally, content analysis was conducted on the ten most cited and ten focused articles addressing AI’s impact on cognitive learning outcomes. VOSviewer software was used for data analysis and visualization. Findings indicate increased research output post-2023, driven by digital transformation and global collaboration. Leading affiliations include The University of Hong Kong and Carnegie Mellon University, with the United States, China, and India as top contributing countries. Influential journals and funding bodies include the National Science Foundation and the National Natural Science Foundation of China. Notable authors include Chiu and Cukurova, while Kit Ng, Zhong, and Liu are prominent in bibliographic coupling, emphasizing AI adoption. Co-authorship analysis shows collaboration primarily among developed nations. Co-word analysis reveals Key research themes include contrastive learning, adversarial machine learning, and federated learning. Content analysis highlights AI’s transformative potential for learning, teaching, cognitive learning, and innovation. This study provides managerial and practical recommendations for students, universities, and policymakers. This study has several limitations that future studies will consider.","author":[{"family":"Ansari","given":"Shaukat"},{"family":"Qamari","given":"Ika"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s44217-025-00865-0","URL":"https://doi.org/10.1007/s44217-025-00865-0","source":"openalex"},{"id":"oa:W4407664085","type":"article-journal","title":"Framework for Addressing Imbalanced Data in Aviation with Federated Learning","abstract":"The aviation industry generates vast amounts of data across multiple stakeholders, but critical faults and anomalies occur rarely, creating inherently imbalanced datasets that complicate machine learning applications. Traditional centralized approaches are further constrained by privacy concerns and regulatory requirements that limit data sharing among stakeholders. This paper presents a novel framework for addressing imbalanced data challenges in aviation through federated learning, focusing on fault detection, predictive maintenance, and safety management. The proposed framework combines specialized techniques for handling imbalanced data with privacy-preserving federated learning to enable effective collaboration while maintaining data security. The framework incorporates local resampling methods, cost-sensitive learning, and weighted aggregation mechanisms to improve minority class detection performance. The framework is validated through extensive experiments involving multiple aviation stakeholders, demonstrating a 23% improvement in fault detection accuracy and a 17% reduction in remaining useful life prediction error compared to conventional models. Results show the enhanced detection of rare but critical faults, improved maintenance scheduling accuracy, and effective risk assessment across distributed aviation datasets. The proposed framework provides a scalable and practical solution for using distributed aviation data while addressing both class imbalance and privacy concerns, contributing to improved safety and operational efficiency in the aviation industry.","author":[{"family":"Kabashkin","given":"Igor"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/info16020147","URL":"https://doi.org/10.3390/info16020147","source":"openalex"},{"id":"oa:W4411289937","type":"article-journal","title":"Artificial Intelligence and machine learning in fraud detection for digital payments","abstract":"The financial sector has adopted artificial intelligence (AI) and machine learning (ML) for more advanced and real-time fraud detection as a result of the increased threat of fraud brought on by the global surge in digital payments. By using sophisticated algorithms like supervised learning, anomaly detection, and deep neural networks, these technologies allow systems to identify irregularities, adjust to new fraud patterns, and lower false positives. AI-driven systems are widely used by fintechs and neobanks in the US, Germany, and the EU. However, they also face issues with data quality, model transparency, and regulatory compliance, which calls for a careful balancing act between ethical oversight and technical solutions.","author":[{"family":"Davitaia","given":"Alexandre"}],"issued":{"date-parts":[[2025]]},"DOI":"10.30574/ijsra.2025.15.3.1784","URL":"https://doi.org/10.30574/ijsra.2025.15.3.1784","source":"openalex"},{"id":"oa:W4415380972","type":"article-journal","title":"Federated Incentive Learning: A Privacy-Preserving Framework for Ad Monetization and Creator Rewards in High-Concurrency Environments","abstract":"Growing regulatory pressures have upended the historical pattern of cross-site behavioral targeting and measurement, forcing ad monetization systems to reconcile personalization, creator incentives, and privacy by design. This paper introduces Federated Incentive Learning (FIL), a privacy-preserving framework that optimizes ad placement and creator rewards in high-concurrency environments without transferring raw user data. FIL combines federated learning with on-device differential privacy, integrates Google’s Privacy Sandbox primitives for interest signals, on-device auctions, and privacy-preserving attribution, and imposes explicit fairness constraints for minimum exposure and region sensitive economic weighting. A global short video case study motivates design requirements such as sub-200 ms end-to-end decision latency and transparent incentive allocation. Methodologically, the study specifies a federated objective combining click-through and conversion prediction with incentive feedback, contrasts FedAvg and FedProx for heterogeneous clients, and calibrates Gaussian mechanisms to enforce an ε-differential privacy budget. Results indicate that FIL attains approximately 92 percent of the accuracy of centralized models while yielding a 25 percent improvement in small-to-medium advertiser return on investment and a 15 percent increase in creator earnings under weighted incentives and exposure guarantee. The framework demonstrates operational feasibility for privacy-sensitive, real-time monetization markets and contributes a governance-aware and equitable approach to ad delivery and creator compensation.","author":[{"family":"Yi","given":"Xun"}],"issued":{"date-parts":[[2025]]},"DOI":"10.71465/ajbd3381","URL":"https://doi.org/10.71465/ajbd3381","source":"openalex"},{"id":"oa:W4409812826","type":"article-journal","title":"Towards Secure and Efficient Farming Using Self-Regulating Heterogeneous Federated Learning in Dynamic Network Conditions","abstract":"The advancement of precision agriculture increasingly depends on innovative technological solutions that optimize resource utilization and minimize environmental impact. This paper introduces a novel heterogeneous federated learning architecture specifically designed for intelligent agricultural systems, with a focus on combine tractors equipped with advanced nutrient and crop health sensors. Unlike conventional FL applications, our architecture uniquely addresses the challenges of communication efficiency, dynamic network conditions, and resource allocation in rural farming environments. By adopting a decentralized approach, we ensure that sensitive data remain localized, thereby enhancing security while facilitating effective collaboration among devices. The architecture promotes the formation of adaptive clusters based on operational capabilities and geographical proximity, optimizing communication between edge devices and a global server. Furthermore, we implement a robust checkpointing mechanism and a dynamic data transmission strategy, ensuring efficient model updates in the face of fluctuating network conditions. Through a comprehensive assessment of computational power, energy efficiency, and latency, our system intelligently classifies devices, significantly enhancing the overall efficiency of federated learning processes. This paper details the architecture, operational procedures, and evaluation methodologies, demonstrating how our approach has the potential to transform agricultural practices through data-driven decision-making and promote sustainable farming practices tailored to the unique challenges of the agricultural sector.","author":[{"family":"Puppala","given":"Sai"},{"family":"Sinha","given":"Koushik"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/agriculture15090934","URL":"https://doi.org/10.3390/agriculture15090934","source":"openalex"},{"id":"oa:W4408216035","type":"article-journal","title":"AnoFel: Supporting Anonymity for Privacy-Preserving Federated Learning","abstract":"Federated learning enables users to collaboratively train a machine learning model over their private datasets. Secure aggregation protocols are employed to mitigate information leakage about the local datasets from user-submitted model updates. This setup, however, still leaks the user participation in training, which can also be sensitive. Protecting user anonymity is even more challenging in dynamic environments where users may (re)join or leave the training process at any point of time. This paper introduces AnoFel, the first framework to support private and anonymous dynamic participation in federated learning (FL). AnoFel leverages several cryptographic primitives, the concept of anonymity sets, differential privacy, and a public bulletin board to support anonymous user registration, as well as unlinkable and confidential model update submission. Our system allows dynamic participation, where users can join or leave at any time without needing any recovery protocol or interaction. To assess security, we formalize a notion for privacy and anonymity in FL, and formally prove that AnoFel satisfies this notion. To the best of our knowledge, our system is the first solution with provable anonymity guarantees. To assess efficiency, we provide a concrete implementation of AnoFel, and conduct experiments showing its ability to support learning applications scaling to a large number of clients. For a TinyImageNet classification task with 512 clients, the client setup to join is less than 3 sec, and the client runtime for each training iteration takes a total of 8 sec, where the added overhead of AnoFel is 46% of the total runtime. We also compare our system with prior work and demonstrate its practicality. AnoFel client runtime is up to 5x faster than Truex et al., despite the added anonymity guarantee and dynamic user joining in AnoFel. Compared to Bonawitz et al., AnoFel is only 2x slower for added support for privacy in output, dynamic user joining, and anonymity.","author":[{"family":"Almashaqbeh","given":"Ghada"},{"family":"Ghodsi","given":"Zahra"}],"issued":{"date-parts":[[2025]]},"DOI":"10.56553/popets-2025-0051","URL":"https://doi.org/10.56553/popets-2025-0051","source":"openalex"},{"id":"oa:W4414272485","type":"article-journal","title":"AI-Driven Privacy Shield: A Secure and Privacy-Preserving Federated Learning Framework","abstract":"Emerging, highly skilled cyberattacks demand novel and robust techniques for AI-powered privacy preservation. Centralized machine learning models can be compromised with a single point of failure, a data breach, or an adversarial attack. The proposed work presents a unique AI-based Privacy Shield that enhances Privacy-Aware Hybrid Privacy-Preserving Federated Learning (HPP-FL), Blockchain-Enhanced Secure Aggregation (BESA), and Quantum-Resistant Encryption (QRE-FL). By employing an Adaptive Adversarial Training (AAT) strategy, the defense mechanism adjusts to the transforming cyber threats in real-time, thus demonstrating prevention abilities. This approach allows multiple users to collaboratively train a global deep learning model securely with minimal bandwidth and without relying on any central aggregator, similar to federated learning but built on a blockchain-based secure aggregation protocol. Additionally, quantum-resistant encryption mechanisms provide an added layer of security against emerging threats posed by quantum computing, securing the future of federated models. The framework is validated on real-world data from the healthcare, finance, and IoT domains. It shows improvements of 91.2% accuracy, 40% less data leakage, and 35% more resistance to attacks, all while using little extra computing power. This makes it possible for AI security to be scalable and future-proof, making FL a more credible privacy-protecting option for real-world uses.","author":[{"family":"Abdal","given":"Yasir"}],"issued":{"date-parts":[[2025]]},"DOI":"10.26706/ijceae.6.3.20250608","URL":"https://doi.org/10.26706/ijceae.6.3.20250608","source":"openalex"},{"id":"oa:W7132852778","type":"article-journal","title":"Privacy-Preserving Generative AI in Healthcare Systems Using Federated Learning Approaches","abstract":"The research paper focuses on the mechanism of introducing Federated Learning alongside privacy-saving strategies in Generative Artificial Intelligence to healthcare applications. The study analyses the privacy protection/model accuracy trade-off by implementing Differential Privacy and Secure Aggregation. The synthetic datasets were applied to the model to train in five rounds and included many federated clients, that improved its accuracy by 49% to 60%. The findings suggest that Federated Learning has the potential to improve the performance of AI and preserve the privacy of data at the same time. Other challenges covered in the study include mode collapse and privacy-utility trade-offs, and recommended solutions to achieve the efficient and secure healthcare AI models.","author":[{"family":"Poojari","given":"Rajesh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.64751/ijdim.2026.v5.n1.pp78-88","URL":"https://doi.org/10.64751/ijdim.2026.v5.n1.pp78-88","source":"openalex"},{"id":"oa:W4409640804","type":"article-journal","title":"Federated learning for privacy-preserving AI in human–robot collaboration for smart manufacturing","abstract":"Purpose The study aims to address privacy and security challenges in AI-driven human–robot collaboration (HRC) by developing a privacy-preserving federated learning framework. Traditional centralized AI models expose sensitive manufacturing data to cybersecurity risks, creating barriers to AI adoption in regulated industries. This research proposes a decentralized learning approach that enables robots to collaboratively train AI models without sharing raw data, ensuring compliance with privacy regulations (e.g. GDPR and CCPA). The study seeks to advance trustworthy AI-driven automation, improving robotic decision-making, scalability and real-time adaptability while safeguarding sensitive industrial information. Design/methodology/approach This study proposes a Multi-Agent Federated Reinforcement Learning (MARL-FL) framework for privacy-preserving AI in human–robot collaboration (HRC) for smart manufacturing. The framework integrates federated learning (FL), reinforcement learning (RL) and differential privacy to enhance robotic decision-making while ensuring data security. A digital twin simulation of a smart factory is used for evaluation, where collaborative robots autonomously learn and optimize tasks using decentralized AI training. Performance is assessed using model accuracy, task success rate, convergence speed and privacy leakage reduction metrics, demonstrating FL’s effectiveness in improving secure AI-driven automation. Findings Experimental results from a digital twin-based smart factory simulation demonstrate that the proposed FL-based framework achieves 91.2% model accuracy, improves task success rates by 7.6% and reduces privacy leakage risks by 41.5% compared to centralized AI models. The federated reinforcement learning approach also accelerates model convergence by 25%, enabling faster adaptation to dynamic manufacturing conditions. The study confirms that FL enhances AI-driven collaboration, operational efficiency and data security, making it a viable solution for privacy-preserving smart manufacturing. Originality/value This research is among the first to integrate federated learning, reinforcement learning and privacy-preserving AI techniques for secure human–robot collaboration in Industry 4.0. Unlike conventional AI models that rely on centralized data processing, the proposed MARL-FL framework enables secure, decentralized learning, reducing cybersecurity risks and regulatory concerns. The study provides new insights into privacy-aware AI governance in industrial automation, making it highly valuable for researchers, policymakers and manufacturers seeking trustworthy AI-driven robotics solutions.","author":[{"family":"Rahmati","given":"Milad"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1108/jimse-03-2025-0003","URL":"https://doi.org/10.1108/jimse-03-2025-0003","source":"openalex"},{"id":"oa:W4412751381","type":"article-journal","title":"Lightweight Anomaly Detection in Digit Recognition Using Federated Learning","abstract":"This study presents a lightweight autoencoder-based approach for anomaly detection in digit recognition using federated learning on resource-constrained embedded devices. We implement and evaluate compact autoencoder models on the ESP32-CAM microcontroller, enabling both training and inference directly on the device using 32-bit floating-point arithmetic. The system is trained on a reduced MNIST dataset (1000 resized samples) and evaluated using EMNIST and MNIST-C for anomaly detection. Seven fully connected autoencoder architectures are first evaluated on a PC to explore the impact of model size and batch size on training time and anomaly detection performance. Selected models are then re-implemented in the C programming language and deployed on a single ESP32 device, achieving training times as short as 12 min, inference latency as low as 9 ms, and F1 scores of up to 0.87. Autoencoders are further tested on ten devices in a real-world federated learning experiment using Wi-Fi. We explore non-IID and IID data distribution scenarios: (1) digit-specialized devices and (2) partitioned datasets with varying content and anomaly types. The results show that small unmodified autoencoder models can be effectively trained and evaluated directly on low-power hardware. The best models achieve F1 scores of up to 0.87 in the standard IID setting and 0.86 in the extreme non-IID setting. Despite some clients being trained on corrupted datasets, federated aggregation proves resilient, maintaining high overall performance. The resource analysis shows that more than half of the models and all the training-related allocations fit entirely in internal RAM. These findings confirm the feasibility of local float32 training and collaborative anomaly detection on low-cost hardware, supporting scalable and privacy-preserving edge intelligence.","author":[{"family":"Tanović","given":"Anja"},{"family":"Mezei","given":"Ivan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/fi17080343","URL":"https://doi.org/10.3390/fi17080343","source":"openalex"},{"id":"oa:W4411351109","type":"article-journal","title":"FedEmerge: An Entropy-Guided Federated Learning Method for Sensor Networks and Edge Intelligence","abstract":"Introduction: Federated Learning (FL) is a distributed machine learning paradigm where a global model is collaboratively trained across multiple decentralized clients without exchanging raw data. This is especially important in sensor networks and edge intelligence, where data privacy, bandwidth constraints, and data locality are paramount. Traditional FL methods like FedAvg struggle with highly heterogeneous (non-IID) client data, which is common in these settings. Background: Traditional FL aggregation methods, such as FedAvg, weigh client updates primarily by dataset size, potentially overlooking the informativeness or diversity of each client’s contribution. These limitations are especially pronounced in sensor networks and IoT environments, where clients may hold sparse, unbalanced, or single-modality data. Methods: We propose FedEmerge, an entropy-guided aggregation approach that adjusts each client’s impact on the global model based on the information entropy of its local data distribution. This formulation introduces a principled way to quantify and reward data diversity, enabling an emergent collective learning dynamic in which globally informative updates drive convergence. Unlike existing methods that weigh updates by sample count or heuristics, FedEmerge prioritizes clients with more representative, high-entropy data. The FedEmerge algorithm is presented with full mathematical detail, and we prove its convergence under the Polyak–Łojasiewicz (PL) condition. Results: Theoretical analysis shows that FedEmerge achieves linear convergence to the optimal model under standard assumptions (smoothness and PL condition), similar to centralized gradient descent. Empirically, FedEmerge improves global model accuracy and convergence speed on highly skewed non-IID benchmarks, and it reduces performance disparities among clients compared to FedAvg. Evaluations on CIFAR-10 (non-IID), Federated EMNIST, and Shakespeare datasets confirm its effectiveness in practical edge-learning settings. Conclusions: This entropy-guided federated strategy demonstrates that weighting client updates by data diversity enhances learning outcomes in heterogeneous networks. The approach preserves privacy like standard FL and adds minimal computation overhead, making it a practical solution for real-world federated systems.","author":[{"family":"Khan","given":"Koffka"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/s25123728","URL":"https://doi.org/10.3390/s25123728","source":"openalex"},{"id":"doi:10.48550/arxiv.2603.28434","type":"manuscript","title":"Democratizing Federated Learning with Blockchain and Multi-Task Peer Prediction","abstract":"The synergy between Federated Learning and blockchain has been considered promising; however, the computationally intensive nature of contribution measurement conflicts with the strict computation and storage limits of blockchain systems. We propose a novel concept to decentralize the AI training process using blockchain technology and Multi-task Peer Prediction. By leveraging smart contracts and cryptocurrencies to incentivize contributions to the training process, we aim to harness the mutual benefits of AI and blockchain. We discuss the advantages and limitations of our design.","author":[{"family":"Witt","given":"Leon"},{"family":"Toyoda","given":"Kentaroh"},{"family":"Samek","given":"Wojciech"},{"family":"Li","given":"Dan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.48550/arxiv.2603.28434","URL":"https://doi.org/10.48550/arxiv.2603.28434","source":"datacite"},{"id":"doi:10.5281/zenodo.21625705","type":"article-journal","title":"PrivateBoost: privacy-preserving federated gradient boosting for single-record patient devices","abstract":"Deployable implementation of the PrivateBoost protocol: histogram-based federated gradient boosting in which every client holds a single labelled record, splits its gradient and Hessian statistics into Shamir shares over a Mersenne prime field, and distributes them to independent shareholders that only ever release population-level sums. Includes the protocol crate, the shareholder and aggregator server, the client library and CLI, the Flutter mobile application, the deployment configurations, and the analysis scripts that reproduce the reported experiments and figures.","author":[{"family":"Specht","given":"Bernhard"},{"family":"Ermis","given":"Orhan"},{"family":"Garbaya","given":"Samaher"},{"family":"Schneider","given":"Reinhard"},{"family":"Chavarriaga","given":"Ricardo"},{"family":"Khadraoui","given":"Djamel"},{"family":"Tayeb","given":"Zied"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21625705","URL":"https://doi.org/10.5281/zenodo.21625705","source":"datacite"},{"id":"doi:10.5281/zenodo.21630278","type":"article-journal","title":"PrivateBoost: privacy-preserving federated gradient boosting for single-record patient devices","abstract":"Deployable implementation of the PrivateBoost protocol: histogram-based federated gradient boosting in which every client holds a single labelled record, splits its gradient and Hessian statistics into Shamir shares over a Mersenne prime field, and distributes them to independent shareholders that only ever release population-level sums. Includes the protocol crate, the shareholder and aggregator server, the client library and CLI, the Flutter mobile application, the deployment configurations, and the analysis scripts that reproduce the reported experiments and figures.","author":[{"family":"Specht","given":"Bernhard"},{"family":"Ermis","given":"Orhan"},{"family":"Garbaya","given":"Samaher"},{"family":"Schneider","given":"Reinhard"},{"family":"Chavarriaga","given":"Ricardo"},{"family":"Khadraoui","given":"Djamel"},{"family":"Tayeb","given":"Zied"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21630278","URL":"https://doi.org/10.5281/zenodo.21630278","source":"datacite"},{"id":"doi:10.5281/zenodo.20322207","type":"article-journal","title":"SecureWITT","abstract":"SecureWITT: Embedding Homomorphic Encryption into Deep Joint Source-Channel Coding against the Cliff Effect for Semantic Image Transmission. This framework integrates BFV homomorphic encryption into a Swin Transformer-based deep JSCC codec, with three synergistic mechanisms: Encrypted Gradient Tunneling (EGT), Joint Encryption Fine-tuning (JEF), and BER-Gated Inference (BGI).","author":[{"family":"Kou","given":"Guang"},{"family":"Ye","given":"Qing"},{"family":"Yuan","given":"Zhi"},{"family":"Zhu","given":"Ting"},{"family":"Fu","given":"Wei"},{"family":"Wei","given":"Guo"},{"family":"Zheng","given":"Tong"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.20322207","URL":"https://doi.org/10.5281/zenodo.20322207","source":"datacite"},{"id":"doi:10.5281/zenodo.21334273","type":"article-journal","title":"EncFormer: Secure and Efficient Transformer Inference over Encrypted Data","abstract":"Official self-contained implementation of EncFormer. The archive contains the GPU CKKS implementation, CKKS--MPC conversion, BPMax, MBNorm, BOLT GELU, real two-party EzPC/SCI execution, native dependency source, the BERT-base SST-2 checkpoint, local evaluation data, and evaluation commands.","author":[{"family":"Zhu","given":"Yufan"},{"family":"Jin","given":"Chao"},{"family":"Aung","given":"Khin"},{"family":"Xiao","given":"Xiaokui"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21334273","URL":"https://doi.org/10.5281/zenodo.21334273","source":"datacite"},{"id":"doi:10.5281/zenodo.21334274","type":"article-journal","title":"EncFormer: Secure and Efficient Transformer Inference over Encrypted Data","abstract":"Official self-contained implementation of EncFormer. The archive contains the GPU CKKS implementation, CKKS--MPC conversion, BPMax, MBNorm, BOLT GELU, real two-party EzPC/SCI execution, native dependency source, the BERT-base SST-2 checkpoint, local evaluation data, and evaluation commands.","author":[{"family":"Zhu","given":"Yufan"},{"family":"Jin","given":"Chao"},{"family":"Aung","given":"Khin"},{"family":"Xiao","given":"Xiaokui"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21334274","URL":"https://doi.org/10.5281/zenodo.21334274","source":"datacite"},{"id":"oa:W7128469018","type":"article-journal","title":"Secure and Differentially Private Edge-Cloud Federated Learning Framework for Privacy-Preserving Maritime AIS Intelligence","abstract":"Cloud computing now supports large-scale maritime analytics, yet offloading rich Automatic Identification System (AIS) data to the cloud exposes sensitive operational patterns and complicates compliance with cross-border priv... | Find, read and cite all the research you need on Tech Science Press","author":[{"family":"Khan","given":"Abuzar"},{"family":"Iqbal","given":"Abid"},{"family":"Husnain","given":"Ghassan"},{"family":"Masood","given":"Fahad"},{"family":"Al-Naeem","given":"Mohammed"},{"family":"Iqbal","given":"Dr"}],"issued":{"date-parts":[[2026]]},"DOI":"10.32604/cmc.2026.077222","URL":"https://doi.org/10.32604/cmc.2026.077222","source":"openalex"},{"id":"doi:10.5281/zenodo.21550657","type":"article-journal","title":"Enhanced Brain Tumor Detection and Privacy Preserving Using Federated Learning","abstract":"Brain cancers pose significant difficulties for both diagnosis and treatment, underscoring the necessity for precise and private-protecting detection techniques. Federated learning is used to solve this, allowing several healthcare facilities to work together to train detection models without jeopardizing patient privacy. This paper presents an approach called federated learning that may be used to improve brain tumor identification while protecting patient privacy. Brain tumors are dangerous medical disorders that need to be accurately diagnosed in order to be effectively treated. However, sharing private patient information is a common practice in traditional medical data analysis methodologies, which raises privacy issues. Federated learning helps with this by enabling cooperative training of a common model amongst several hospitals or institutions without requiring the exchange of raw data. This method protects patient privacy by having each institution train the model using its own local data and only sharing model updates. We show through trials that our method is efficient in reliably identifying brain tumors while upholding privacy norms, presenting a viable option for improving medical diagnosis without jeopardizing patient privacy.","author":[{"family":"Nandan","given":"Uday"},{"family":"Sai","given":"Chetan"},{"family":"Sai","given":"Naga"},{"family":"Viswanadapalli","given":"Anusha"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.21550657","URL":"https://doi.org/10.5281/zenodo.21550657","source":"datacite"},{"id":"doi:10.5281/zenodo.21550658","type":"article-journal","title":"Enhanced Brain Tumor Detection and Privacy Preserving Using Federated Learning","abstract":"Brain cancers pose significant difficulties for both diagnosis and treatment, underscoring the necessity for precise and private-protecting detection techniques. Federated learning is used to solve this, allowing several healthcare facilities to work together to train detection models without jeopardizing patient privacy. This paper presents an approach called federated learning that may be used to improve brain tumor identification while protecting patient privacy. Brain tumors are dangerous medical disorders that need to be accurately diagnosed in order to be effectively treated. However, sharing private patient information is a common practice in traditional medical data analysis methodologies, which raises privacy issues. Federated learning helps with this by enabling cooperative training of a common model amongst several hospitals or institutions without requiring the exchange of raw data. This method protects patient privacy by having each institution train the model using its own local data and only sharing model updates. We show through trials that our method is efficient in reliably identifying brain tumors while upholding privacy norms, presenting a viable option for improving medical diagnosis without jeopardizing patient privacy.","author":[{"family":"Nandan","given":"Uday"},{"family":"Sai","given":"Chetan"},{"family":"Sai","given":"Naga"},{"family":"Viswanadapalli","given":"Anusha"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.21550658","URL":"https://doi.org/10.5281/zenodo.21550658","source":"datacite"},{"id":"doi:10.5281/zenodo.21555490","type":"article-journal","title":"A Novel Framework for Trustworthy Privacy Preserving Machine Learning Model for Industrial IoT Systems Using Blockchain Techniques","abstract":"Industrial Internet of Things (IIoT) is changing many driving enterprises like transportation, mining, horticulture, energy and medical care. Machine Learning calculations are utilized for getting stages for IT frameworks. The IoT network unit hubs typically asset in a strange manner by making them more responsible to digital assaults. IIoT frameworks requests various situations in genuine one among them is giving security and the causes that encompass them in true viewpoints. It incorporates a system called PriModChain causes security and reliability on IIoT information by joining differential protection, Ethereum block chain and unified Machine learning. Consequently, security will be compromised and we use PriMod chain for giving protection and different compliances and created utilizing Python with attachment programming on essential PC.","author":[{"family":"Yedukondalu","given":"Dr"},{"family":"Rao","given":"Dr"},{"family":"Dugyala","given":"Raman"}],"issued":{"date-parts":[[2022]]},"DOI":"10.5281/zenodo.21555490","URL":"https://doi.org/10.5281/zenodo.21555490","source":"datacite"},{"id":"doi:10.5281/zenodo.21555491","type":"article-journal","title":"A Novel Framework for Trustworthy Privacy Preserving Machine Learning Model for Industrial IoT Systems Using Blockchain Techniques","abstract":"Industrial Internet of Things (IIoT) is changing many driving enterprises like transportation, mining, horticulture, energy and medical care. Machine Learning calculations are utilized for getting stages for IT frameworks. The IoT network unit hubs typically asset in a strange manner by making them more responsible to digital assaults. IIoT frameworks requests various situations in genuine one among them is giving security and the causes that encompass them in true viewpoints. It incorporates a system called PriModChain causes security and reliability on IIoT information by joining differential protection, Ethereum block chain and unified Machine learning. Consequently, security will be compromised and we use PriMod chain for giving protection and different compliances and created utilizing Python with attachment programming on essential PC.","author":[{"family":"Yedukondalu","given":"Dr"},{"family":"Rao","given":"Dr"},{"family":"Dugyala","given":"Raman"}],"issued":{"date-parts":[[2022]]},"DOI":"10.5281/zenodo.21555491","URL":"https://doi.org/10.5281/zenodo.21555491","source":"datacite"},{"id":"doi:10.24406/publica-4105","type":"article-journal","title":"Operational Planning Decision Support using Multi-Dimensional Data Farming","abstract":"Multi-Dimensional Data Farming (MDDF) uses Machine Learning (ML) to automate data farming to allow improved and faster decisions in highly complex multi-scale, multi-domain, and multi-level hybrid war campaigns. This has significant utility when used in support of Operational planning, allowing multiple Courses of Action (CoA) to be rapidly developed and evaluated prior to the execution of any operation. Using MDDF enables decision-makers to explore the problem space and identify multiple optimal solutions significantly faster than current techniques. MSG-186 has applied MDDF in a sand-box environment to an illustrative combined strategic campaign and tactical hybrid warfare operation resource allocation problem considering the balance between local and global optimal solutions. We have tested the technical feasibility of implementing MDDF within the Federated Mission Network operational environment at Coalition Warrior Interoperability Exercise (CWIX). Through MDDF, we aim to show it is possible to combine ML techniques exploring operations at multiple scales (Multi-Domain Operations and targeted-fidelity modelling) and optimize the strategic/operational level goal, by selecting the correct resource allocation scheme at the tactical level. This paper describes an ML-based assistant able to conduct MDDF experiments and optimization tasks on an automated basis, which was examined in detail during CWIX in 2024.","author":[{"family":"Akesson","given":"Bernt"},{"family":"Amyot-Bourgeois","given":"Maude"},{"family":"Das","given":"Sreerupa"},{"family":"Ernis","given":"Gunar"},{"family":"Gill","given":"Andrew"},{"family":"Lappi","given":"Esa"},{"family":"Nguyen","given":"Bao"},{"family":"Rolfs","given":"Chris"},{"family":"Seichter","given":"Stephan"},{"family":"Serre","given":"Lynne"},{"family":"Slyusar","given":"Vadym"},{"family":"Vaghi","given":"Alessio"},{"family":"Volbach","given":"Peter"},{"family":"Zimmermann","given":"Alexander"},{"family":"Unav"}],"issued":{"date-parts":[[2024]]},"DOI":"10.24406/publica-4105","URL":"https://doi.org/10.24406/publica-4105","source":"datacite"},{"id":"doi:10.48550/arxiv.2205.11518","type":"manuscript","title":"LIA: Privacy-Preserving Data Quality Evaluation in Federated Learning Using a Lazy Influence Approximation","abstract":"In Federated Learning, it is crucial to handle low-quality, corrupted, or malicious data. However, traditional data valuation methods are not suitable due to privacy concerns. To address this, we propose a simple yet effective approach that utilizes a new influence approximation called \"lazy influence\" to filter and score data while preserving privacy. To do this, each participant uses their own data to estimate the influence of another participant's batch and sends a differentially private obfuscated score to the central coordinator. Our method has been shown to successfully filter out biased and corrupted data in various simulated and real-world settings, achieving a recall rate of over $&gt;90\\%$ (sometimes up to $100\\%$) while maintaining strong differential privacy guarantees with $\\varepsilon \\leq 1$.","author":[{"family":"Rokvic","given":"Ljubomir"},{"family":"Danassis","given":"Panayiotis"},{"family":"Karimireddy","given":"Sai"},{"family":"Faltings","given":"Boi"}],"issued":{"date-parts":[[2022]]},"DOI":"10.48550/arxiv.2205.11518","URL":"https://doi.org/10.48550/arxiv.2205.11518","source":"datacite"},{"id":"doi:10.48550/arxiv.2305.19600","type":"manuscript","title":"Adaptive Self-Distillation for Minimizing Client Drift in Heterogeneous Federated Learning","abstract":"Federated Learning (FL) is a machine learning paradigm that enables clients to jointly train a global model by aggregating the locally trained models without sharing any local training data. In practice, there can often be substantial heterogeneity (e.g., class imbalance) across the local data distributions observed by each of these clients. Under such non-iid label distributions across clients, FL suffers from the 'client-drift' problem where every client drifts to its own local optimum. This results in slower convergence and poor performance of the aggregated model. To address this limitation, we propose a novel regularization technique based on adaptive self-distillation (ASD) for training models on the client side. Our regularization scheme adaptively adjusts to each client's training data based on the global model's prediction entropy and the client-data label distribution. We show in this paper that our proposed regularization (ASD) can be easily integrated atop existing, state-of-the-art FL algorithms, leading to a further boost in the performance of these off-the-shelf methods. We theoretically explain how incorporation of ASD regularizer leads to reduction in client-drift and empirically justify the generalization ability of the trained model. We demonstrate the efficacy of our approach through extensive experiments on multiple real-world benchmarks and show substantial gains in performance when the proposed regularizer is combined with popular FL methods.","author":[{"family":"Yashwanth","given":"M"},{"family":"Nayak","given":"Gaurav"},{"family":"Singh","given":"Arya"},{"family":"Simmhan","given":"Yogesh"},{"family":"Chakraborty","given":"Anirban"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2305.19600","URL":"https://doi.org/10.48550/arxiv.2305.19600","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.16240","type":"manuscript","title":"AFL: A Single-Round Analytic Approach for Federated Learning with Pre-trained Models","abstract":"In this paper, we introduce analytic federated learning (AFL), a new training paradigm that brings analytical (i.e., closed-form) solutions to the federated learning (FL) with pre-trained models. Our AFL draws inspiration from analytic learning -- a gradient-free technique that trains neural networks with analytical solutions in one epoch. In the local client training stage, the AFL facilitates a one-epoch training, eliminating the necessity for multi-epoch updates. In the aggregation stage, we derive an absolute aggregation (AA) law. This AA law allows a single-round aggregation, reducing heavy communication overhead and achieving fast convergence by removing the need for multiple aggregation rounds. More importantly, the AFL exhibits a property that \\textit{invariance to data partitioning}, meaning that regardless of how the full dataset is distributed among clients, the aggregated result remains identical. This could spawn various potentials, such as data heterogeneity invariance and client-number invariance. We conduct experiments across various FL settings including extremely non-IID ones, and scenarios with a large number of clients (e.g., $\\ge 1000$). In all these settings, our AFL constantly performs competitively while existing FL techniques encounter various obstacles. Our codes are available at https://github.com/ZHUANGHP/Analytic-federated-learning.","author":[{"family":"He","given":"Run"},{"family":"Tong","given":"Kai"},{"family":"Fang","given":"Di"},{"family":"Sun","given":"Han"},{"family":"Zeng","given":"Ziqian"},{"family":"Li","given":"Haoran"},{"family":"Chen","given":"Tianyi"},{"family":"Zhuang","given":"Huiping"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.16240","URL":"https://doi.org/10.48550/arxiv.2405.16240","source":"datacite"},{"id":"doi:10.48550/arxiv.2306.01176","type":"manuscript","title":"Cooperative Hardware-Prompt Learning for Snapshot Compressive Imaging","abstract":"Existing reconstruction models in snapshot compressive imaging systems (SCI) are trained with a single well-calibrated hardware instance, making their performance vulnerable to hardware shifts and limited in adapting to multiple hardware configurations. To facilitate cross-hardware learning, previous efforts attempt to directly collect multi-hardware data and perform centralized training, which is impractical due to severe user data privacy concerns and hardware heterogeneity across different platforms/institutions. In this study, we explicitly consider data privacy and heterogeneity in cooperatively optimizing SCI systems by proposing a Federated Hardware-Prompt learning (FedHP) framework. Rather than mitigating the client drift by rectifying the gradients, which only takes effect on the learning manifold but fails to solve the heterogeneity rooted in the input data space, FedHP learns a hardware-conditioned prompter to align inconsistent data distribution across clients, serving as an indicator of the data inconsistency among different hardware (e.g., coded apertures). Extensive experimental results demonstrate that the proposed FedHP coordinates the pre-trained model to multiple hardware configurations, outperforming prevalent FL frameworks for 0.35dB under challenging heterogeneous settings. Moreover, a Snapshot Spectral Heterogeneous Dataset has been built upon multiple practical SCI systems. Data and code are aveilable at https://github.com/Jiamian-Wang/FedHP-Snapshot-Compressive-Imaging","author":[{"family":"Wang","given":"Jiamian"},{"family":"Wu","given":"Zongliang"},{"family":"Zhang","given":"Yulun"},{"family":"Yuan","given":"Xin"},{"family":"Lin","given":"Tao"},{"family":"Tao","given":"Zhiqiang"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2306.01176","URL":"https://doi.org/10.48550/arxiv.2306.01176","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.17462","type":"manuscript","title":"Ferrari: Federated Feature Unlearning via Optimizing Feature Sensitivity","abstract":"The advent of Federated Learning (FL) highlights the practical necessity for the right to be forgotten for all clients, allowing them to request data deletion from the machine learning models service provider. This necessity has spurred a growing demand for Federated Unlearning (FU). Feature unlearning has gained considerable attention due to its applications in unlearning sensitive, backdoor, and biased features. Existing methods employ the influence function to achieve feature unlearning, which is impractical for FL as it necessitates the participation of other clients, if not all, in the unlearning process. Furthermore, current research lacks an evaluation of the effectiveness of feature unlearning. To address these limitations, we define feature sensitivity in evaluating feature unlearning according to Lipschitz continuity. This metric characterizes the model outputs rate of change or sensitivity to perturbations in the input feature. We then propose an effective federated feature unlearning framework called Ferrari, which minimizes feature sensitivity. Extensive experimental results and theoretical analysis demonstrate the effectiveness of Ferrari across various feature unlearning scenarios, including sensitive, backdoor, and biased features. The code is publicly available at https://github.com/OngWinKent/Federated-Feature-Unlearning","author":[{"family":"Gu","given":"Hanlin"},{"family":"Ong","given":"Win"},{"family":"Chan","given":"Chee"},{"family":"Fan","given":"Lixin"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.17462","URL":"https://doi.org/10.48550/arxiv.2405.17462","source":"datacite"},{"id":"doi:10.48550/arxiv.2210.01708","type":"manuscript","title":"Exploring Parameter-Efficient Fine-Tuning to Enable Foundation Models in Federated Learning","abstract":"Federated learning (FL) has emerged as a promising paradigm for enabling the collaborative training of models without centralized access to the raw data on local devices. In the typical FL paradigm (e.g., FedAvg), model weights are sent to and from the server each round to participating clients. Recently, the use of small pre-trained models has been shown to be effective in federated learning optimization and improving convergence. However, recent state-of-the-art pre-trained models are getting more capable but also have more parameters, known as the \"Foundation Models.\" In conventional FL, sharing the enormous model weights can quickly put a massive communication burden on the system, especially if more capable models are employed. Can we find a solution to enable those strong and readily available pre-trained models in FL to achieve excellent performance while simultaneously reducing the communication burden? To this end, we investigate the use of parameter-efficient fine-tuning in federated learning and thus introduce a new framework: FedPEFT. Specifically, we systemically evaluate the performance of FedPEFT across a variety of client stability, data distribution, and differential privacy settings. By only locally tuning and globally sharing a small portion of the model weights, significant reductions in the total communication overhead can be achieved while maintaining competitive or even better performance in a wide range of federated learning scenarios, providing insight into a new paradigm for practical and effective federated systems.","author":[{"family":"Sun","given":"Guangyu"},{"family":"Khalid","given":"Umar"},{"family":"Mendieta","given":"Matias"},{"family":"Wang","given":"Pu"},{"family":"Chen","given":"Chen"}],"issued":{"date-parts":[[2022]]},"DOI":"10.48550/arxiv.2210.01708","URL":"https://doi.org/10.48550/arxiv.2210.01708","source":"datacite"},{"id":"doi:10.48550/arxiv.2412.17373","type":"manuscript","title":"FRTP: Federating Route Search Records to Enhance Long-term Traffic Prediction","abstract":"Accurate traffic prediction, especially predicting traffic conditions several days in advance is essential for intelligent transportation systems (ITS). Such predictions enable mid- and long-term traffic optimization, which is crucial for efficient transportation planning. However, the inclusion of diverse external features, alongside the complexities of spatial relationships and temporal uncertainties, significantly increases the complexity of forecasting models. Additionally, traditional approaches have handled data preprocessing separately from the learning model, leading to inefficiencies caused by repeated trials of preprocessing and training. In this study, we propose a federated architecture capable of learning directly from raw data with varying features and time granularities or lengths. The model adopts a unified design that accommodates different feature types, time scales, and temporal periods. Our experiments focus on federating route search records and begin by processing raw data within the model framework. Unlike traditional models, this approach integrates the data federation phase into the learning process, enabling compatibility with various time frequencies and input/output configurations. The accuracy of the proposed model is demonstrated through evaluations using diverse learning patterns and parameter settings. The results show that online search log data is useful for forecasting long-term traffic, highlighting the model's adaptability and efficiency.","author":[{"family":"Ge","given":"Hangli"},{"family":"Yang","given":"Xiaojie"},{"family":"Matsunaga","given":"Itsuki"},{"family":"Huang","given":"Dizhi"},{"family":"Koshizuka","given":"Noboru"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2412.17373","URL":"https://doi.org/10.48550/arxiv.2412.17373","source":"datacite"},{"id":"doi:10.48550/arxiv.2412.18507","type":"manuscript","title":"An Empirical Analysis of Federated Learning Models Subject to Label-Flipping Adversarial Attack","abstract":"In this paper, we empirically analyze adversarial attacks on selected federated learning models. The specific learning models considered are Multinominal Logistic Regression (MLR), Support Vector Classifier (SVC), Multilayer Perceptron (MLP), Convolution Neural Network (CNN), %Recurrent Neural Network (RNN), Random Forest, XGBoost, and Long Short-Term Memory (LSTM). For each model, we simulate label-flipping attacks, experimenting extensively with 10 federated clients and 100 federated clients. We vary the percentage of adversarial clients from 10% to 100% and, simultaneously, the percentage of labels flipped by each adversarial client is also varied from 10% to 100%. Among other results, we find that models differ in their inherent robustness to the two vectors in our label-flipping attack, i.e., the percentage of adversarial clients, and the percentage of labels flipped by each adversarial client. We discuss the potential practical implications of our results.","author":[{"family":"Bhatnagar","given":"Kunal"},{"family":"Chattanathan","given":"Sagana"},{"family":"Dang","given":"Angela"},{"family":"Eranki","given":"Bhargav"},{"family":"Rana","given":"Ronnit"},{"family":"Sridhar","given":"Charan"},{"family":"Vedam","given":"Siddharth"},{"family":"Yao","given":"Angie"},{"family":"Stamp","given":"Mark"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2412.18507","URL":"https://doi.org/10.48550/arxiv.2412.18507","source":"datacite"},{"id":"doi:10.48550/arxiv.2206.02131","type":"manuscript","title":"Federated Adversarial Training with Transformers","abstract":"Federated learning (FL) has emerged to enable global model training over distributed clients' data while preserving its privacy. However, the global trained model is vulnerable to the evasion attacks especially, the adversarial examples (AEs), carefully crafted samples to yield false classification. Adversarial training (AT) is found to be the most promising approach against evasion attacks and it is widely studied for convolutional neural network (CNN). Recently, vision transformers have been found to be effective in many computer vision tasks. To the best of the authors' knowledge, there is no work that studied the feasibility of AT in a FL process for vision transformers. This paper investigates such feasibility with different federated model aggregation methods and different vision transformer models with different tokenization and classification head techniques. In order to improve the robust accuracy of the models with the not independent and identically distributed (Non-IID), we propose an extension to FedAvg aggregation method, called FedWAvg. By measuring the similarities between the last layer of the global model and the last layer of the client updates, FedWAvg calculates the weights to aggregate the local models updates. The experiments show that FedWAvg improves the robust accuracy when compared with other state-of-the-art aggregation methods.","author":[{"family":"Aldahdooh","given":"Ahmed"},{"family":"Hamidouche","given":"Wassim"},{"family":"Déforges","given":"Olivier"}],"issued":{"date-parts":[[2022]]},"DOI":"10.48550/arxiv.2206.02131","URL":"https://doi.org/10.48550/arxiv.2206.02131","source":"datacite"},{"id":"doi:10.21256/zhaw-20682","type":"article-journal","title":"Dynamic group key agreement for resource-constrained devices using blockchains","abstract":"Dynamic group key agreement (DGKA) protocols are one of the key security primitives to secure multiparty communications in decentralized and insecure environments while considering the instant changes in a communication group. However, with the ever-increasing number of connected devices, traditional DGKA protocols have performance challenges since each member in the group has to make several computationally intensive operations while verifying the keying materials to compute the resulting group key. To overcome this issue, we propose a new approach for DGKA protocols by utilizing Hyperledger Fabric framework as a blockchain platform. To this end, we migrate the communication and verification overhead of DGKA participants to the blockchain network in our developed scheme. This paradigm allows a flexible DGKA protocol that considers resource-constrained entities and trade-offs regarding distributed computation. According to our performance analysis, participants with low computing resources can efficiently utilize our protocol. Furthermore, we have demonstrated that our protocol has the same security features as other comparable protocols in the literature.","author":[{"family":"Taçyıldız","given":"Yaşar"},{"family":"Ermiş","given":"Orhan"},{"family":"Gür","given":"Gürkan"},{"family":"Alagöz","given":"Fatih"}],"issued":{"date-parts":[[2020]]},"DOI":"10.21256/zhaw-20682","URL":"https://doi.org/10.21256/zhaw-20682","source":"datacite"},{"id":"doi:10.48550/arxiv.2412.03766","type":"manuscript","title":"End to End Collaborative Synthetic Data Generation","abstract":"The success of AI is based on the availability of data to train models. While in some cases a single data custodian may have sufficient data to enable AI, often multiple custodians need to collaborate to reach a cumulative size required for meaningful AI research. The latter is, for example, often the case for rare diseases, with each clinical site having data for only a small number of patients. Recent algorithms for federated synthetic data generation are an important step towards collaborative, privacy-preserving data sharing. Existing techniques, however, focus exclusively on synthesizer training, assuming that the training data is already preprocessed and that the desired synthetic data can be delivered in one shot, without any hyperparameter tuning. In this paper, we propose an end-to-end collaborative framework for publishing of synthetic data that accounts for privacy-preserving preprocessing as well as evaluation. We instantiate this framework with Secure Multiparty Computation (MPC) protocols and evaluate it in a use case for privacy-preserving publishing of synthetic genomic data for leukemia.","author":[{"family":"Pentyala","given":"Sikha"},{"family":"Sitaraman","given":"Geetha"},{"family":"Claar","given":"Trae"},{"family":"De Cock","given":"Martine"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2412.03766","URL":"https://doi.org/10.48550/arxiv.2412.03766","source":"datacite"},{"id":"doi:10.48550/arxiv.2409.17571","type":"manuscript","title":"Incomplete quantum oblivious transfer with perfect one-sided security","abstract":"Oblivious transfer is a fundamental cryptographic primitive which is useful for secure multiparty computation. There are several variants of oblivious transfer. We consider 1 out of 2 oblivious transfer, where a sender sends two bits of information to a receiver. The receiver only receives one of the two bits, while the sender does not know which bit the receiver has received. Perfect quantum oblivious transfer with information theoretic security is known to be impossible. We aim to find the lowest possible cheating probabilities. Bounds on cheating probabilities have been investigated for complete protocols, where if both parties follow the protocol, the bit value obtained by the receiver matches the sender bit value. We instead investigate incomplete protocols, where the receiver obtains an incorrect bit value with probability pf. We present optimal non interactive protocols where Alice bit values are encoded in four symmetric pure quantum states, and where she cannot cheat better than with a random guess. We find the protocols such that for a given pf, Bob cheating probability pr is as low as possible, and vice versa. Furthermore, we show that non-interactive quantum protocols can outperform non-interactive classical protocols, and give a lower bound on Bob cheating probability in interactive quantum protocols. Importantly for optical implementations, our protocols do not require entanglement nor quantum memory.","author":[{"family":"Reichmuth","given":"David"},{"family":"Puthoor","given":"Ittoop"},{"family":"Wallden","given":"Petros"},{"family":"Andersson","given":"Erika"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2409.17571","URL":"https://doi.org/10.48550/arxiv.2409.17571","source":"datacite"},{"id":"doi:10.48550/arxiv.2307.12533","type":"manuscript","title":"PUMA: Secure Inference of LLaMA-7B in Five Minutes","abstract":"With ChatGPT as a representative, tons of companies have began to provide services based on large Transformers models. However, using such a service inevitably leak users' prompts to the model provider. Previous studies have studied secure inference for Transformer models using secure multiparty computation (MPC), where model parameters and clients' prompts are kept secret. Despite this, these frameworks are still limited in terms of model performance, efficiency, and deployment. To address these limitations, we propose framework PUMA to enable fast and secure Transformer model inference. Our framework designs high quality approximations for expensive functions such as GeLU and softmax, and significantly reduce the cost of secure inference while preserving the model performance. Additionally, we design secure Embedding and LayerNorm procedures that faithfully implement the desired functionality without undermining the Transformer architecture. PUMA is about $2\\times$ faster than the state-of-the-art framework MPCFORMER(ICLR 2023) and has similar accuracy as plaintext models without fine-tuning (which the previous works failed to achieve). PUMA can even evaluate LLaMA-7B in around 5 minutes to generate 1 token. To our best knowledge, this is the first time that a model with such a parameter size is able to be evaluated under MPC. PUMA has been open-sourced in the Github repository of SecretFlow-SPU.","author":[{"family":"Dong","given":"Ye"},{"family":"Lu","given":"Wen"},{"family":"Zheng","given":"Yancheng"},{"family":"Wu","given":"Haoqi"},{"family":"Zhao","given":"Derun"},{"family":"Tan","given":"Jin"},{"family":"Huang","given":"Zhicong"},{"family":"Hong","given":"Cheng"},{"family":"Wei","given":"Tao"},{"family":"Chen","given":"Wenguang"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2307.12533","URL":"https://doi.org/10.48550/arxiv.2307.12533","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.03136","type":"manuscript","title":"FOBNN: Fast Oblivious Inference via Binarized Neural Networks","abstract":"The remarkable performance of deep learning has sparked the rise of Deep Learning as a Service (DLaaS), allowing clients to send their personal data to service providers for model predictions. A persistent challenge in this context is safeguarding the privacy of clients' sensitive data. Oblivious inference allows the execution of neural networks on client inputs without revealing either the inputs or the outcomes to the service providers. In this paper, we propose FOBNN, a Fast Oblivious inference framework via Binarized Neural Networks. In FOBNN, through neural network binarization, we convert linear operations (e.g., convolutional and fully-connected operations) into eXclusive NORs (XNORs) and an Oblivious Bit Count (OBC) problem. For secure multiparty computation techniques, like garbled circuits or bitwise secret sharing, XNOR operations incur no communication cost, making the OBC problem the primary bottleneck for linear operations. To tackle this, we first propose the Bit Length Bounding (BLB) algorithm, which minimizes bit representation to decrease redundant computations. Subsequently, we develop the Layer-wise Bit Accumulation (LBA) algorithm, utilizing pure bit operations layer by layer to further boost performance. We also enhance the binarized neural network structure through link optimization and structure exploration. The former optimizes link connections given a network structure, while the latter explores optimal network structures under same secure computation costs. Our theoretical analysis reveals that the BLB algorithm outperforms the state-of-the-art OBC algorithm by a range of 17% to 55%, while the LBA exhibits an improvement of nearly 100%. Comprehensive proof-of-concept evaluation demonstrates that FOBNN outperforms prior art on popular benchmarks and shows effectiveness in emerging bioinformatics.","author":[{"family":"Chen","given":"Xin"},{"family":"Chen","given":"Zhili"},{"family":"Wei","given":"Shiwen"},{"family":"Gong","given":"Junqing"},{"family":"Chen","given":"Lin"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.03136","URL":"https://doi.org/10.48550/arxiv.2405.03136","source":"datacite"},{"id":"doi:10.5281/zenodo.19610547","type":"article-journal","title":"AI-Based Image Analysis in Diagnostic Devices","abstract":"AI-based image analysis has emerged as the leading application of artificial intelligence in diagnostic medical devices,with over 520 FDA-cleared AI/ML-enabled devices by 2023 -- the majority addressing radiology, pathology,ophthalmology, and dermatology image interpretation. From convolutional neural networks that detect diabeticretinopathy with ophthalmologist-level sensitivity to transformer models that segment tumour boundaries in whole-slidepathology images with sub-cellular precision, AI image analysis is transitioning from research demonstration to clinicaldeployment at scale. Yet the path from validated algorithm to clinically integrated diagnostic device requires navigation ofhuman factors design, workflow integration, regulatory validation, and post-market performance monitoring challengesthat algorithmic accuracy alone does not address. This study presents the AI Image Analysis Diagnostic DeviceFramework (AIADDF), evaluating five AI image analysis implementation approaches -- standalone algorithm withradiologist workflow, AI-first triage with human review escalation, computer-aided detection enhancement, autonomousAI reporting for defined scope, and federated multi-site AI with continuous learning -- across four imaging devicecategories: radiology CT/MRI, digital pathology, retinal fundus imaging, and dermatoscopy. Our AI Diagnostic ImageScore (ADIS) integrates diagnostic accuracy, workflow efficiency, radiologist acceptance, regulatory compliance, andcross-site generalisability. Federated multi-site AI with continuous learning achieved the highest ADIS (0.928) throughprivacy-preserving training across 24 clinical sites that achieved C-statistic 0.92 while maintaining 96% performanceretention at new deployment sites, while autonomous AI reporting achieved the highest workflow efficiency (0.955) byreducing mean report turnaround time from 48 hours to 3.2 hours for defined low-complexity imaging tasks","author":[{"family":"Popescu","given":"Helena"},{"family":"Lindberg","given":"Andreas"},{"family":"Moreau","given":"Andreas"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.19610547","URL":"https://doi.org/10.5281/zenodo.19610547","source":"datacite"},{"id":"doi:10.5281/zenodo.19610548","type":"article-journal","title":"AI-Based Image Analysis in Diagnostic Devices","abstract":"AI-based image analysis has emerged as the leading application of artificial intelligence in diagnostic medical devices,with over 520 FDA-cleared AI/ML-enabled devices by 2023 -- the majority addressing radiology, pathology,ophthalmology, and dermatology image interpretation. From convolutional neural networks that detect diabeticretinopathy with ophthalmologist-level sensitivity to transformer models that segment tumour boundaries in whole-slidepathology images with sub-cellular precision, AI image analysis is transitioning from research demonstration to clinicaldeployment at scale. Yet the path from validated algorithm to clinically integrated diagnostic device requires navigation ofhuman factors design, workflow integration, regulatory validation, and post-market performance monitoring challengesthat algorithmic accuracy alone does not address. This study presents the AI Image Analysis Diagnostic DeviceFramework (AIADDF), evaluating five AI image analysis implementation approaches -- standalone algorithm withradiologist workflow, AI-first triage with human review escalation, computer-aided detection enhancement, autonomousAI reporting for defined scope, and federated multi-site AI with continuous learning -- across four imaging devicecategories: radiology CT/MRI, digital pathology, retinal fundus imaging, and dermatoscopy. Our AI Diagnostic ImageScore (ADIS) integrates diagnostic accuracy, workflow efficiency, radiologist acceptance, regulatory compliance, andcross-site generalisability. Federated multi-site AI with continuous learning achieved the highest ADIS (0.928) throughprivacy-preserving training across 24 clinical sites that achieved C-statistic 0.92 while maintaining 96% performanceretention at new deployment sites, while autonomous AI reporting achieved the highest workflow efficiency (0.955) byreducing mean report turnaround time from 48 hours to 3.2 hours for defined low-complexity imaging tasks","author":[{"family":"Popescu","given":"Helena"},{"family":"Lindberg","given":"Andreas"},{"family":"Moreau","given":"Andreas"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.19610548","URL":"https://doi.org/10.5281/zenodo.19610548","source":"datacite"},{"id":"doi:10.5281/zenodo.21684179","type":"article-journal","title":"Supply Chain Fraud Risk Mitigation Using Federated AI Models for Continuous Transaction Integrity Verification","abstract":"Supply chain networks are increasingly vulnerable to sophisticated fraud tactics that compromise transactional integrity and threaten organizational resilience. Traditional centralized fraud detection systems struggle to scale across decentralized logistics and procurement environments due to data privacy, latency, and system heterogeneity. This review explores the integration of Federated Artificial Intelligence (AI) models as a transformative approach to mitigating fraud risks in supply chains. Federated AI enables collaborative model training across multiple stakeholders without exposing raw data, thus preserving privacy while enhancing anomaly detection capabilities. The paper examines how federated learning frameworks, combined with blockchain technology and edge intelligence, can continuously verify transaction authenticity, detect anomalies in procurement and logistics flows, and adapt to evolving fraud patterns. Furthermore, it evaluates current limitations in data standardization, model interoperability, and real-time verification under distributed conditions. Case studies of federated learning applications in financial technology, logistics automation, and smart contracts are analyzed to illustrate effectiveness and implementation strategies. The review concludes by outlining critical research directions for achieving secure, adaptive, and privacy-preserving fraud detection in future global supply chains.","author":[{"family":"Essien","given":"Iboro"},{"family":"Ajayi","given":"Joshua"},{"family":"Erigha","given":"Eseoghene"},{"family":"Obuse","given":"Ehimah"},{"family":"Ayanbode","given":"Noah"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.21684179","URL":"https://doi.org/10.5281/zenodo.21684179","source":"datacite"},{"id":"doi:10.5281/zenodo.21684180","type":"article-journal","title":"Supply Chain Fraud Risk Mitigation Using Federated AI Models for Continuous Transaction Integrity Verification","abstract":"Supply chain networks are increasingly vulnerable to sophisticated fraud tactics that compromise transactional integrity and threaten organizational resilience. Traditional centralized fraud detection systems struggle to scale across decentralized logistics and procurement environments due to data privacy, latency, and system heterogeneity. This review explores the integration of Federated Artificial Intelligence (AI) models as a transformative approach to mitigating fraud risks in supply chains. Federated AI enables collaborative model training across multiple stakeholders without exposing raw data, thus preserving privacy while enhancing anomaly detection capabilities. The paper examines how federated learning frameworks, combined with blockchain technology and edge intelligence, can continuously verify transaction authenticity, detect anomalies in procurement and logistics flows, and adapt to evolving fraud patterns. Furthermore, it evaluates current limitations in data standardization, model interoperability, and real-time verification under distributed conditions. Case studies of federated learning applications in financial technology, logistics automation, and smart contracts are analyzed to illustrate effectiveness and implementation strategies. The review concludes by outlining critical research directions for achieving secure, adaptive, and privacy-preserving fraud detection in future global supply chains.","author":[{"family":"Essien","given":"Iboro"},{"family":"Ajayi","given":"Joshua"},{"family":"Erigha","given":"Eseoghene"},{"family":"Obuse","given":"Ehimah"},{"family":"Ayanbode","given":"Noah"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.21684180","URL":"https://doi.org/10.5281/zenodo.21684180","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.16472","type":"manuscript","title":"Personalized Additive Modeling for Multi-level Federated Learning","abstract":"Contemporary AI faces the challenge of balancing generality with user-specific personalization. In federated learning (FL), this challenge is amplified by highly heterogeneous client data with complex non-IID patterns beyond standard IID assumptions. Many existing FL methods are designed for relatively restricted heterogeneity settings (e.g., a fixed number of clusters or a fixed form of personalization), limiting their robustness under complex structures. In this work, we study FL from a \\emph{multi-level non-IID} perspective, where client similarity is captured by multiple granularities of shared knowledge: global, subgroup, and client-specific components. This view captures coarse-to-fine relationships while requiring less prior knowledge of task boundaries. Building on this insight, we propose \\emph{Federated Multi-level Additive Modeling} (FeMAM), which learns multiple levels of shareable models and constructs personalized predictors via additive composition across levels. To move beyond a fixed structure, FeMAM allows models to grow and be pruned dynamically during training, adapting to diverse federated scenarios. Despite employing multiple models, FeMAM remains cost-friendly by unlocking only a small subset (one level) of models for training at a time. Extensive experiments show that FeMAM effectively approximates diverse complex non-IID structures and consistently outperforms representative clustered and personalized FL baselines.","author":[{"family":"Chen","given":"Shutong"},{"family":"Long","given":"Guodong"},{"family":"Zhou","given":"Tianyi"},{"family":"Ma","given":"Jie"},{"family":"Jiang","given":"Jing"},{"family":"Zhang","given":"Chengqi"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.16472","URL":"https://doi.org/10.48550/arxiv.2405.16472","source":"datacite"},{"id":"doi:10.48550/arxiv.2403.11166","type":"manuscript","title":"Pencil: Private and Extensible Collaborative Learning without the Non-Colluding Assumption","abstract":"The escalating focus on data privacy poses significant challenges for collaborative neural network training, where data ownership and model training/deployment responsibilities reside with distinct entities. Our community has made substantial contributions to addressing this challenge, proposing various approaches such as federated learning (FL) and privacy-preserving machine learning based on cryptographic constructs like homomorphic encryption (HE) and secure multiparty computation (MPC). However, FL completely overlooks model privacy, and HE has limited extensibility (confined to only one data provider). While the state-of-the-art MPC frameworks provide reasonable throughput and simultaneously ensure model/data privacy, they rely on a critical non-colluding assumption on the computing servers, and relaxing this assumption is still an open problem. In this paper, we present Pencil, the first private training framework for collaborative learning that simultaneously offers data privacy, model privacy, and extensibility to multiple data providers, without relying on the non-colluding assumption. Our fundamental design principle is to construct the n-party collaborative training protocol based on an efficient two-party protocol, and meanwhile ensuring that switching to different data providers during model training introduces no extra cost. We introduce several novel cryptographic protocols to realize this design principle and conduct a rigorous security and privacy analysis. Our comprehensive evaluations of Pencil demonstrate that (i) models trained in plaintext and models trained privately using Pencil exhibit nearly identical test accuracies; (ii) The training overhead of Pencil is greatly reduced: Pencil achieves 10 ~ 260x higher throughput and 2 orders of magnitude less communication than prior art; (iii) Pencil is resilient against both existing and adaptive (white-box) attacks.","author":[{"family":"Liu","given":"Xuanqi"},{"family":"Liu","given":"Zhuotao"},{"family":"Li","given":"Qi"},{"family":"Xu","given":"Ke"},{"family":"Xu","given":"Mingwei"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2403.11166","URL":"https://doi.org/10.48550/arxiv.2403.11166","source":"datacite"},{"id":"doi:10.48550/arxiv.2401.13386","type":"manuscript","title":"Privacy-Preserving Face Recognition in Hybrid Frequency-Color Domain","abstract":"Face recognition technology has been deployed in various real-life applications. The most sophisticated deep learning-based face recognition systems rely on training millions of face images through complex deep neural networks to achieve high accuracy. It is quite common for clients to upload face images to the service provider in order to access the model inference. However, the face image is a type of sensitive biometric attribute tied to the identity information of each user. Directly exposing the raw face image to the service provider poses a threat to the user's privacy. Current privacy-preserving approaches to face recognition focus on either concealing visual information on model input or protecting model output face embedding. The noticeable drop in recognition accuracy is a pitfall for most methods. This paper proposes a hybrid frequency-color fusion approach to reduce the input dimensionality of face recognition in the frequency domain. Moreover, sparse color information is also introduced to alleviate significant accuracy degradation after adding differential privacy noise. Besides, an identity-specific embedding mapping scheme is applied to protect original face embedding by enlarging the distance among identities. Lastly, secure multiparty computation is implemented for safely computing the embedding distance during model inference. The proposed method performs well on multiple widely used verification datasets. Moreover, it has around 2.6% to 4.2% higher accuracy than the state-of-the-art in the 1:N verification scenario.","author":[{"family":"Han","given":"Dong"},{"family":"Li","given":"Yong"},{"family":"Denzler","given":"Joachim"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2401.13386","URL":"https://doi.org/10.48550/arxiv.2401.13386","source":"datacite"},{"id":"doi:10.60692/he4kz-01436","type":"article-journal","title":"SviaB: Secure and verifiable multi‐instance iris remote authentication using blockchain","abstract":"IET BiometricsVolume 11, Issue 1 p. 35-50 ORIGINAL RESEARCH PAPEROpen Access SviaB: Secure and verifiable multi-instance iris remote authentication using blockchain Mahesh Kumar Morampudi, Corresponding Author Mahesh Kumar Morampudi morampudimahesh@gmail.com Department of Computer Science and Engineering, National Institute of Technology-Warangal, Telangana, India Department of Computer Science and Engineering, SRM University AP, Amaravati, Andhra Pradesh, India Correspondence Mahesh Kumar Morampudi, Department of Computer Science and Engineering, SRM University AP, Amaravati, Andhra Pradesh, India. Email: morampudimahesh@gmail.comSearch for more papers by this authorMunaga V. N. K. Prasad, Munaga V. N. K. Prasad Institute for Development and Research in Banking Technology (IDRBT), Hyderabad, IndiaSearch for more papers by this authorSurya Narayana Raju Undi, Surya Narayana Raju Undi Department of Computer Science and Engineering, National Institute of Technology-Warangal, Telangana, IndiaSearch for more papers by this author Mahesh Kumar Morampudi, Corresponding Author Mahesh Kumar Morampudi morampudimahesh@gmail.com Department of Computer Science and Engineering, National Institute of Technology-Warangal, Telangana, India Department of Computer Science and Engineering, SRM University AP, Amaravati, Andhra Pradesh, India Correspondence Mahesh Kumar Morampudi, Department of Computer Science and Engineering, SRM University AP, Amaravati, Andhra Pradesh, India. Email: morampudimahesh@gmail.comSearch for more papers by this authorMunaga V. N. K. Prasad, Munaga V. N. K. Prasad Institute for Development and Research in Banking Technology (IDRBT), Hyderabad, IndiaSearch for more papers by this authorSurya Narayana Raju Undi, Surya Narayana Raju Undi Department of Computer Science and Engineering, National Institute of Technology-Warangal, Telangana, IndiaSearch for more papers by this author First published: 12 May 2021 https://doi.org/10.1049/bme2.12042Citations: 2AboutSectionsPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinked InRedditWechat Abstract Homomorphic encryption (HE) is the most widely explored research area in the construction of privacy-preserving biometric authentication systems because of its advantages over cancellable biometrics and biometric cryptosystems. However, most of the existing privacy-preserving biometric authentication systems using HE assume that the server performs computations honestly. In a malicious server setting, the server may return an arbitrary result to save computational resources, resulting in a false accept/reject. To address this, secure and verifiable multi-instance iris authentication using blockchain (SviaB) is proposed. Paillier HE provides confidentiality for the iris templates in SviaB. The blockchain offers the integrity of the encrypted reference iris templates as well as the trust of the comparator result. The challenges of using blockchain in biometrics are also addressed in SviaB. Extensive experimental results on benchmark iris databases demonstrate that SviaB provides privacy to the iris templates with no loss of accuracy and trust in the comparator result. 1 INTRODUCTION Unlike password or token authentication systems, a biometric authentication system (BAS) has more flexibility because users do not need to carry or remember anything. Fingerprint, iris, face, etc. are the commonly used biometric modalities [1, 2]. Properties such as stability and uniqueness make the iris the most widely used of the various biometric applications in comparison ","author":[{"family":"Morampudi","given":"Mahesh"},{"family":"Prasad","given":"Munaga"},{"family":"Undi","given":"Surya"}],"issued":{"date-parts":[[2021]]},"DOI":"10.60692/he4kz-01436","URL":"https://doi.org/10.60692/he4kz-01436","source":"datacite"},{"id":"doi:10.60692/bga3e-f5h73","type":"article-journal","title":"SviaB: Secure and verifiable multi‐instance iris remote authentication using blockchain","abstract":"IET BiometricsVolume 11, Issue 1 p. 35-50 ORIGINAL RESEARCH PAPEROpen Access SviaB: Secure and verifiable multi-instance iris remote authentication using blockchain Mahesh Kumar Morampudi, Corresponding Author Mahesh Kumar Morampudi morampudimahesh@gmail.com Department of Computer Science and Engineering, National Institute of Technology-Warangal, Telangana, India Department of Computer Science and Engineering, SRM University AP, Amaravati, Andhra Pradesh, India Correspondence Mahesh Kumar Morampudi, Department of Computer Science and Engineering, SRM University AP, Amaravati, Andhra Pradesh, India. Email: morampudimahesh@gmail.comSearch for more papers by this authorMunaga V. N. K. Prasad, Munaga V. N. K. Prasad Institute for Development and Research in Banking Technology (IDRBT), Hyderabad, IndiaSearch for more papers by this authorSurya Narayana Raju Undi, Surya Narayana Raju Undi Department of Computer Science and Engineering, National Institute of Technology-Warangal, Telangana, IndiaSearch for more papers by this author Mahesh Kumar Morampudi, Corresponding Author Mahesh Kumar Morampudi morampudimahesh@gmail.com Department of Computer Science and Engineering, National Institute of Technology-Warangal, Telangana, India Department of Computer Science and Engineering, SRM University AP, Amaravati, Andhra Pradesh, India Correspondence Mahesh Kumar Morampudi, Department of Computer Science and Engineering, SRM University AP, Amaravati, Andhra Pradesh, India. Email: morampudimahesh@gmail.comSearch for more papers by this authorMunaga V. N. K. Prasad, Munaga V. N. K. Prasad Institute for Development and Research in Banking Technology (IDRBT), Hyderabad, IndiaSearch for more papers by this authorSurya Narayana Raju Undi, Surya Narayana Raju Undi Department of Computer Science and Engineering, National Institute of Technology-Warangal, Telangana, IndiaSearch for more papers by this author First published: 12 May 2021 https://doi.org/10.1049/bme2.12042Citations: 2AboutSectionsPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinked InRedditWechat Abstract Homomorphic encryption (HE) is the most widely explored research area in the construction of privacy-preserving biometric authentication systems because of its advantages over cancellable biometrics and biometric cryptosystems. However, most of the existing privacy-preserving biometric authentication systems using HE assume that the server performs computations honestly. In a malicious server setting, the server may return an arbitrary result to save computational resources, resulting in a false accept/reject. To address this, secure and verifiable multi-instance iris authentication using blockchain (SviaB) is proposed. Paillier HE provides confidentiality for the iris templates in SviaB. The blockchain offers the integrity of the encrypted reference iris templates as well as the trust of the comparator result. The challenges of using blockchain in biometrics are also addressed in SviaB. Extensive experimental results on benchmark iris databases demonstrate that SviaB provides privacy to the iris templates with no loss of accuracy and trust in the comparator result. 1 INTRODUCTION Unlike password or token authentication systems, a biometric authentication system (BAS) has more flexibility because users do not need to carry or remember anything. Fingerprint, iris, face, etc. are the commonly used biometric modalities [1, 2]. Properties such as stability and uniqueness make the iris the most widely used of the various biometric applications in comparison ","author":[{"family":"Morampudi","given":"Mahesh"},{"family":"Prasad","given":"Munaga"},{"family":"Undi","given":"Surya"}],"issued":{"date-parts":[[2021]]},"DOI":"10.60692/bga3e-f5h73","URL":"https://doi.org/10.60692/bga3e-f5h73","source":"datacite"},{"id":"doi:10.48550/arxiv.2205.09513","type":"manuscript","title":"Federated learning: Applications, challenges and future directions","abstract":"Federated learning (FL) is a system in which a central aggregator coordinates the efforts of multiple clients to solve machine learning problems. This setting allows training data to be dispersed in order to protect privacy. The purpose of this paper is to provide an overview of FL systems with a focus on healthcare. FL is evaluated here based on its frameworks, architectures, and applications. It is shown here that FL solves the preceding issues with a shared global deep learning (DL) model via a central aggregator server. This paper examines recent developments and provides a comprehensive list of unresolved issues, inspired by the rapid growth of FL research. In the context of FL, several privacy methods are described, including secure multiparty computation, homomorphic encryption, differential privacy, and stochastic gradient descent. Furthermore, a review of various FL classes, such as horizontal and vertical FL and federated transfer learning, is provided. FL has applications in wireless communication, service recommendation, intelligent medical diagnosis systems, and healthcare, all of which are discussed in this paper. We also present a thorough review of existing FL challenges, such as privacy protection, communication cost, system heterogeneity, and unreliable model upload, followed by future research directions.","author":[{"family":"Bharati","given":"Subrato"},{"family":"Mondal","given":"MRH"},{"family":"Podder","given":"Prajoy"},{"family":"Prasath","given":"VBS"}],"issued":{"date-parts":[[2022]]},"DOI":"10.48550/arxiv.2205.09513","URL":"https://doi.org/10.48550/arxiv.2205.09513","source":"datacite"},{"id":"doi:10.48550/arxiv.2412.07954","type":"manuscript","title":"MOFHEI: Model Optimizing Framework for Fast and Efficient Homomorphically Encrypted Neural Network Inference","abstract":"Due to the extensive application of machine learning (ML) in a wide range of fields and the necessity of data privacy, privacy-preserving machine learning (PPML) solutions have recently gained significant traction. One group of approaches relies on Homomorphic Encryption (HE), which enables us to perform ML tasks over encrypted data. However, even with state-of-the-art HE schemes, HE operations are still significantly slower compared to their plaintext counterparts and require a considerable amount of memory. Therefore, we propose MOFHEI, a framework that optimizes the model to make HE-based neural network inference, referred to as private inference (PI), fast and efficient. First, our proposed learning-based method automatically transforms a pre-trained ML model into its compatible version with HE operations, called the HE-friendly version. Then, our iterative block pruning method prunes the model's parameters in configurable block shapes in alignment with the data packing method. This allows us to drop a significant number of costly HE operations, thereby reducing the latency and memory consumption while maintaining the model's performance. We evaluate our framework through extensive experiments on different models using various datasets. Our method achieves up to 98% pruning ratio on LeNet, eliminating up to 93% of the required HE operations for performing PI, reducing latency and the required memory by factors of 9.63 and 4.04, respectively, with negligible accuracy loss.","author":[{"family":"Ghazvinian","given":"Parsa"},{"family":"Podschwadt","given":"Robert"},{"family":"Panzade","given":"Prajwal"},{"family":"Rafiei","given":"Mohammad"},{"family":"Takabi","given":"Daniel"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2412.07954","URL":"https://doi.org/10.48550/arxiv.2412.07954","source":"datacite"},{"id":"doi:10.48550/arxiv.2410.11184","type":"manuscript","title":"Fast and Accurate Homomorphic Softmax Evaluation","abstract":"Homomorphic encryption is one of the main solutions for building secure and privacy-preserving solutions for Machine Learning as a Service. This motivates the development of homomorphic algorithms for the main building blocks of AI, typically for the components of the various types of neural networks architectures. Among those components, we focus on the Softmax function, defined by $\\mathrm{SM}(\\mathbf{x}) = \\left(\\exp(x_i) / \\sum_{j=1}^n \\exp(x_j) \\right)_{1\\le i\\le n}$. This function is deemed to be one of the most difficult to evaluate homomorphically, because of its multivariate nature and of the very large range of values for $\\exp(x_i)$. The available homomorphic algorithms remain restricted, especially in large dimensions, while important applications such as Large Language Models (LLM) require computing Softmax over large dimensional vectors. In terms of multiplicative depth of the computation (a suitable measure of cost for homomorphic algorithms), our algorithm achieves $O(\\log n)$ complexity for a fixed range of inputs, where $n$ is the Softmax dimension. Our algorithm is especially adapted to the situation where we must compute many Softmax at the same time, for instance, in the LLM situation. In that case, assuming that all Softmax calls are packed into $m$ ciphtertexts, the asymptotic amortized multiplicative depth cost per ciphertext is, again over a fixed range, $O(1 + m/N)$ for $N$ the homomorphic ring degree. The main ingredient of our algorithms is a normalize-and-square strategy, which interlaces the exponential computation over a large range and normalization, decomposing both in stabler and cheaper smaller steps. Comparing ourselves to the state of the art, our experiments show, in practice, a good accuracy and a gain of a factor 2.5 to 8 compared to state of the art solutions.","author":[{"family":"Cho","given":"Wonhee"},{"family":"Hanrot","given":"Guillaume"},{"family":"Kim","given":"Taeseong"},{"family":"Park","given":"Minje"},{"family":"Stehlé","given":"Damien"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2410.11184","URL":"https://doi.org/10.48550/arxiv.2410.11184","source":"datacite"},{"id":"doi:10.48550/arxiv.2409.06422","type":"manuscript","title":"A Pervasive, Efficient and Private Future: Realizing Privacy-Preserving Machine Learning Through Hybrid Homomorphic Encryption","abstract":"Machine Learning (ML) has become one of the most impactful fields of data science in recent years. However, a significant concern with ML is its privacy risks due to rising attacks against ML models. Privacy-Preserving Machine Learning (PPML) methods have been proposed to mitigate the privacy and security risks of ML models. A popular approach to achieving PPML uses Homomorphic Encryption (HE). However, the highly publicized inefficiencies of HE make it unsuitable for highly scalable scenarios with resource-constrained devices. Hence, Hybrid Homomorphic Encryption (HHE) -- a modern encryption scheme that combines symmetric cryptography with HE -- has recently been introduced to overcome these challenges. HHE potentially provides a foundation to build new efficient and privacy-preserving services that transfer expensive HE operations to the cloud. This work introduces HHE to the ML field by proposing resource-friendly PPML protocols for edge devices. More precisely, we utilize HHE as the primary building block of our PPML protocols. We assess the performance of our protocols by first extensively evaluating each party's communication and computational cost on a dummy dataset and show the efficiency of our protocols by comparing them with similar protocols implemented using plain BFV. Subsequently, we demonstrate the real-world applicability of our construction by building an actual PPML application that uses HHE as its foundation to classify heart disease based on sensitive ECG data.","author":[{"family":"Nguyen","given":"Khoa"},{"family":"Budzys","given":"Mindaugas"},{"family":"Frimpong","given":"Eugene"},{"family":"Khan","given":"Tanveer"},{"family":"Michalas","given":"Antonis"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2409.06422","URL":"https://doi.org/10.48550/arxiv.2409.06422","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.03775","type":"manuscript","title":"Secure Inference for Vertically Partitioned Data Using Multiparty Homomorphic Encryption","abstract":"We propose a secure inference protocol for a distributed setting involving a single server node and multiple client nodes. We assume that the observed data vector is partitioned across multiple client nodes while the deep learning model is located at the server node. Each client node is required to encrypt its portion of the data vector and transmit the resulting ciphertext to the server node. The server node is required to collect the ciphertexts and perform inference in the encrypted domain. We demonstrate an application of multi-party homomorphic encryption (MPHE) to satisfy these requirements. We propose a packing scheme, that enables the server to form the ciphertext of the complete data by aggregating the ciphertext of data subsets encrypted using MPHE. While our proposed protocol builds upon prior horizontal federated training protocol~\\cite{sav2020poseidon}, we focus on the inference for vertically partitioned data and avoid the transmission of (encrypted) model weights from the server node to the client nodes.","author":[{"family":"Chen","given":"Shuangyi"},{"family":"Ju","given":"Yue"},{"family":"Zhu","given":"Zhongwen"},{"family":"Khisti","given":"Ashish"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.03775","URL":"https://doi.org/10.48550/arxiv.2405.03775","source":"datacite"},{"id":"doi:10.48550/arxiv.2208.08093","type":"manuscript","title":"Near Threshold Computation of Partitioned Ring Learning With Error (RLWE) Post Quantum Cryptography on Reconfigurable Architecture","abstract":"Ring Learning With Error (RLWE) algorithm is used in Post Quantum Cryptography (PQC) and Homomorphic Encryption (HE) algorithm. The existing classical crypto algorithms may be broken in quantum computers. The adversaries can store all encrypted data. While the quantum computer will be available, these encrypted data can be exposed by the quantum computer. Therefore, the PQC algorithms are an essential solution in recent applications. On the other hand, the HE allows operations on encrypted data which is appropriate for getting services from third parties without revealing confidential plain-texts. The FPGA based PQC and HE hardware accelerators like RLWE is much cost-effective than processor based platform and Application Specific Integrated Circuit (ASIC). FPGA based hardware accelerators still consume more power compare to ASIC based design. Near Threshold Computation (NTC) may be a convenient solution for FPGA based RLWE implementation. In this paper, we have implemented RLWE hardware accelerator which has 14 subcomponents. This paper creates clusters based on the critical path of all 14 subcomponents. Each cluster is implemented in an FPGA partition which has the same biasing voltage $V_{ccint}$. The clusters that have higher critical paths use higher Vccint to avoid timing failure. The clusters have lower critical paths use lower biasing voltage Vccint. This voltage scaled, partitioned RLWE can save ~6% and ~11% power in Vivado and VTR platform respectively. The resource usage and throughput of the implemented RLWE hardware accelerator is comparatively better than existing literature.","author":[{"family":"Baidya","given":"Paresh"},{"family":"Mondal","given":"Swagata"},{"family":"Paul","given":"Rourab"}],"issued":{"date-parts":[[2022]]},"DOI":"10.48550/arxiv.2208.08093","URL":"https://doi.org/10.48550/arxiv.2208.08093","source":"datacite"},{"id":"doi:10.48550/arxiv.2404.03216","type":"manuscript","title":"Accurate Low-Degree Polynomial Approximation of Non-polynomial Operators for Fast Private Inference in Homomorphic Encryption","abstract":"As machine learning (ML) permeates fields like healthcare, facial recognition, and blockchain, the need to protect sensitive data intensifies. Fully Homomorphic Encryption (FHE) allows inference on encrypted data, preserving the privacy of both data and the ML model. However, it slows down non-secure inference by up to five magnitudes, with a root cause of replacing non-polynomial operators (ReLU and MaxPooling) with high-degree Polynomial Approximated Function (PAF). We propose SmartPAF, a framework to replace non-polynomial operators with low-degree PAF and then recover the accuracy of PAF-approximated model through four techniques: (1) Coefficient Tuning (CT) -- adjust PAF coefficients based on the input distributions before training, (2) Progressive Approximation (PA) -- progressively replace one non-polynomial operator at a time followed by a fine-tuning, (3) Alternate Training (AT) -- alternate the training between PAFs and other linear operators in the decoupled manner, and (4) Dynamic Scale (DS) / Static Scale (SS) -- dynamically scale PAF input value within (-1, 1) in training, and fix the scale as the running max value in FHE deployment. The synergistic effect of CT, PA, AT, and DS/SS enables SmartPAF to enhance the accuracy of the various models approximated by PAFs with various low degrees under multiple datasets. For ResNet-18 under ImageNet-1k, the Pareto-frontier spotted by SmartPAF in latency-accuracy tradeoff space achieves 1.42x ~ 13.64x accuracy improvement and 6.79x ~ 14.9x speedup than prior works. Further, SmartPAF enables a 14-degree PAF (f1^2 g_1^2) to achieve 7.81x speedup compared to the 27-degree PAF obtained by minimax approximation with the same 69.4% post-replacement accuracy. Our code is available at https://github.com/EfficientFHE/SmartPAF.","author":[{"family":"Tong","given":"Jianming"},{"family":"Dang","given":"Jingtian"},{"family":"Golder","given":"Anupam"},{"family":"Hao","given":"Callie"},{"family":"Raychowdhury","given":"Arijit"},{"family":"Krishna","given":"Tushar"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2404.03216","URL":"https://doi.org/10.48550/arxiv.2404.03216","source":"datacite"},{"id":"doi:10.48550/arxiv.2312.11575","type":"manuscript","title":"Blind-Touch: Homomorphic Encryption-Based Distributed Neural Network Inference for Privacy-Preserving Fingerprint Authentication","abstract":"Fingerprint authentication is a popular security mechanism for smartphones and laptops. However, its adoption in web and cloud environments has been limited due to privacy concerns over storing and processing biometric data on servers. This paper introduces Blind-Touch, a novel machine learning-based fingerprint authentication system leveraging homomorphic encryption to address these privacy concerns. Homomorphic encryption allows computations on encrypted data without decrypting. Thus, Blind-Touch can keep fingerprint data encrypted on the server while performing machine learning operations. Blind-Touch combines three strategies to efficiently utilize homomorphic encryption in machine learning: (1) It optimizes the feature vector for a distributed architecture, processing the first fully connected layer (FC-16) in plaintext on the client side and the subsequent layer (FC-1) post-encryption on the server, thereby minimizing encrypted computations; (2) It employs a homomorphic encryption compatible data compression technique capable of handling 8,192 authentication results concurrently; and (3) It utilizes a clustered server architecture to simultaneously process authentication results, thereby enhancing scalability with increasing user numbers. Blind-Touch achieves high accuracy on two benchmark fingerprint datasets, with a 93.6% F1- score for the PolyU dataset and a 98.2% F1-score for the SOKOTO dataset. Moreover, Blind-Touch can match a fingerprint among 5,000 in about 0.65 seconds. With its privacy focused design, high accuracy, and efficiency, Blind-Touch is a promising alternative to conventional fingerprint authentication for web and cloud applications.","author":[{"family":"Choi","given":"Hyunmin"},{"family":"Woo","given":"Simon"},{"family":"Kim","given":"Hyoungshick"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2312.11575","URL":"https://doi.org/10.48550/arxiv.2312.11575","source":"datacite"},{"id":"doi:10.48550/arxiv.2304.11643","type":"manuscript","title":"Privacy Computing Meets Metaverse: Necessity, Taxonomy and Challenges","abstract":"Metaverse, the core of the next-generation Internet, is a computer-generated holographic digital environment that simultaneously combines spatio-temporal, immersive, real-time, sustainable, interoperable, and data-sensitive characteristics. It cleverly blends the virtual and real worlds, allowing users to create, communicate, and transact in virtual form. With the rapid development of emerging technologies including augmented reality, virtual reality and blockchain, the metaverse system is becoming more and more sophisticated and widely used in various fields such as social, tourism, industry and economy. However, the high level of interaction with the real world also means a huge risk of privacy leakage both for individuals and enterprises, which has hindered the wide deployment of metaverse. Then, it is inevitable to apply privacy computing techniques in the framework of metaverse, which is a current research hotspot. In this paper, we conduct comprehensive research on the necessity, taxonomy and challenges when privacy computing meets metaverse. Specifically, we first introduce the underlying technologies and various applications of metaverse, on which we analyze the challenges of data usage in metaverse, especially data privacy. Next, we review and summarize state-of-the-art solutions based on federated learning, differential privacy, homomorphic encryption, and zero-knowledge proofs for different privacy problems in metaverse. Finally, we show the current security and privacy challenges in the development of metaverse and provide open directions for building a well-established privacy-preserving metaverse system. For easy access and reference, we integrate the related publications and their codes into a GitHub repository: https://github.com/6lyc/Awesome-Privacy-Computing-in-Metaverse.git.","author":[{"family":"Chen","given":"Chuan"},{"family":"Li","given":"Yuecheng"},{"family":"Wu","given":"Zhenpeng"},{"family":"Mai","given":"Chengyuan"},{"family":"Liu","given":"Youming"},{"family":"Hu","given":"Yanming"},{"family":"Zheng","given":"Zibin"},{"family":"Kang","given":"Jiawen"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2304.11643","URL":"https://doi.org/10.48550/arxiv.2304.11643","source":"datacite"},{"id":"doi:10.4230/lipics.itcs.2024.14","type":"article-journal","title":"Homomorphic Indistinguishability Obfuscation and Its Applications","abstract":"In this work, we propose the notion of homomorphic indistinguishability obfuscation (HiO) and present a construction based on subexponentially-secure iO and one-way functions. An HiO scheme allows us to convert an obfuscation of circuit C to an obfuscation of C'∘C, and this can be performed obliviously (that is, without knowing the circuit C). A naïve solution would be to obfuscate C'∘iO(C). However, if we do this for k hops, then the size of the final obfuscation is exponential in k. HiO ensures that the size of the final obfuscation remains polynomial after repeated compositions. As an application, we show how to build function-hiding hierarchical multi-input functional encryption and homomorphic witness encryption using HiO.","author":[{"family":"Bhushan","given":"Kaartik"},{"family":"Koppula","given":"Venkata"},{"family":"Prabhakaran","given":"Manoj"}],"issued":{"date-parts":[[2024]]},"DOI":"10.4230/lipics.itcs.2024.14","URL":"https://doi.org/10.4230/lipics.itcs.2024.14","source":"datacite"},{"id":"doi:10.5281/zenodo.21689523","type":"article-journal","title":"Secure E - Voting System Based on Paillier Cryptography","abstract":"In the whole world the advanced security procedures are necessary to present convincing online based casting a ballot (e-voting). Trust in the voting process is therefore an important element to any voting system. Voting over the internet is not secure enough to be trusted for government elections. Choices integrated on the paper exhaust many advantages and add to the confusion of backwoods, which causes atmosphere weakening. Then web based casting a ballot come up in countries like the US, India and Brazil showed that further examination is needed to enhance the security assures for future race, to provide the characterization of votes and enable the affirmation of their reliability and legitimacy. Here, proposed the homomorphic encryption based e-voting for casting a vote, which locate these challenges. It removes every single limitation on the possible assignments of centers to different competitors as per the voters","author":[{"family":"Raut","given":"Bharati"},{"family":"Jagtap","given":"Manasi"},{"family":"Ghule","given":"Sneha"},{"family":"Jadhav","given":"Kshitija"},{"family":"Aundhakar","given":"Prof"}],"issued":{"date-parts":[[2020]]},"DOI":"10.5281/zenodo.21689523","URL":"https://doi.org/10.5281/zenodo.21689523","source":"datacite"},{"id":"doi:10.5281/zenodo.21689524","type":"article-journal","title":"Secure E - Voting System Based on Paillier Cryptography","abstract":"In the whole world the advanced security procedures are necessary to present convincing online based casting a ballot (e-voting). Trust in the voting process is therefore an important element to any voting system. Voting over the internet is not secure enough to be trusted for government elections. Choices integrated on the paper exhaust many advantages and add to the confusion of backwoods, which causes atmosphere weakening. Then web based casting a ballot come up in countries like the US, India and Brazil showed that further examination is needed to enhance the security assures for future race, to provide the characterization of votes and enable the affirmation of their reliability and legitimacy. Here, proposed the homomorphic encryption based e-voting for casting a vote, which locate these challenges. It removes every single limitation on the possible assignments of centers to different competitors as per the voters","author":[{"family":"Raut","given":"Bharati"},{"family":"Jagtap","given":"Manasi"},{"family":"Ghule","given":"Sneha"},{"family":"Jadhav","given":"Kshitija"},{"family":"Aundhakar","given":"Prof"}],"issued":{"date-parts":[[2020]]},"DOI":"10.5281/zenodo.21689524","URL":"https://doi.org/10.5281/zenodo.21689524","source":"datacite"},{"id":"doi:10.5281/zenodo.19614724","type":"article-journal","title":"Privacy-Preserving Deep Learning","abstract":"Deep learning's reliance on large datasets creates fundamental tension with privacy requirements: models trained onsensitive data can memorise and leak individual training examples through model outputs, gradients, or learnedparameters. This study presents a controlled evaluation of five privacy-preserving deep learning approaches --differential privacy SGD (DP-SGD), federated learning with secure aggregation, homomorphic encryption for inference,knowledge distillation from private models (PATE), and synthetic data generation with privacy guarantees -- across fourprivacy-sensitive tasks: medical image classification (CheXpert chest X-ray), clinical NLP (MIMIC-III dischargesummaries), financial fraud detection (credit card transactions), and recommendation systems (MovieLens-1M). Privacywas measured by formal epsilon-delta differential privacy guarantees and empirical membership inference attacksuccess rate. A total of 2,040 experiments were conducted. DP-SGD at epsilon = 8 reduced accuracy by 4.8 +- 1.2% onmedical imaging versus non-private training, while epsilon = 1 reduced accuracy by 12.4 +- 2.2%, confirming asubstantial privacy-utility trade-off. PATE achieved the best trade-off: epsilon = 2 with only 3.4 +- 0.8% accuracy loss bytransferring knowledge through noisy aggregation of teacher ensemble votes. Membership inference attack successdropped from 68.4% (non-private model) to near-random (52.8%) at epsilon = 4. Synthetic data generation preserved86.4 +- 2.8% of downstream model utility while providing formal privacy guarantees. Federated learning with secureaggregation provided communication-level privacy but remained vulnerable to model-level attacks without additional DPnoise. A practical privacy-preserving ML selection guide mapping privacy requirements, acceptable utility loss, andcomputational budget to recommended approaches is proposed.","author":[{"family":"Novak","given":"Ivan"},{"family":"Nowak","given":"Nina"},{"family":"Schmidt","given":"Ivan"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.19614724","URL":"https://doi.org/10.5281/zenodo.19614724","source":"datacite"},{"id":"doi:10.5281/zenodo.19614725","type":"article-journal","title":"Privacy-Preserving Deep Learning","abstract":"Deep learning's reliance on large datasets creates fundamental tension with privacy requirements: models trained onsensitive data can memorise and leak individual training examples through model outputs, gradients, or learnedparameters. This study presents a controlled evaluation of five privacy-preserving deep learning approaches --differential privacy SGD (DP-SGD), federated learning with secure aggregation, homomorphic encryption for inference,knowledge distillation from private models (PATE), and synthetic data generation with privacy guarantees -- across fourprivacy-sensitive tasks: medical image classification (CheXpert chest X-ray), clinical NLP (MIMIC-III dischargesummaries), financial fraud detection (credit card transactions), and recommendation systems (MovieLens-1M). Privacywas measured by formal epsilon-delta differential privacy guarantees and empirical membership inference attacksuccess rate. A total of 2,040 experiments were conducted. DP-SGD at epsilon = 8 reduced accuracy by 4.8 +- 1.2% onmedical imaging versus non-private training, while epsilon = 1 reduced accuracy by 12.4 +- 2.2%, confirming asubstantial privacy-utility trade-off. PATE achieved the best trade-off: epsilon = 2 with only 3.4 +- 0.8% accuracy loss bytransferring knowledge through noisy aggregation of teacher ensemble votes. Membership inference attack successdropped from 68.4% (non-private model) to near-random (52.8%) at epsilon = 4. Synthetic data generation preserved86.4 +- 2.8% of downstream model utility while providing formal privacy guarantees. Federated learning with secureaggregation provided communication-level privacy but remained vulnerable to model-level attacks without additional DPnoise. A practical privacy-preserving ML selection guide mapping privacy requirements, acceptable utility loss, andcomputational budget to recommended approaches is proposed.","author":[{"family":"Novak","given":"Ivan"},{"family":"Nowak","given":"Nina"},{"family":"Schmidt","given":"Ivan"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.19614725","URL":"https://doi.org/10.5281/zenodo.19614725","source":"datacite"},{"id":"doi:10.48550/arxiv.2410.21840","type":"manuscript","title":"New Permutation Decomposition Techniques for Efficient Homomorphic Permutation","abstract":"Homomorphic permutation is fundamental to privacy-preserving computations based on batch-encoding homomorphic encryption. It underpins nearly all homomorphic matrix operations and predominantly influences their complexity. Permutation decomposition as a potential approach to optimize this critical component remains underexplored. In this paper, we propose novel decomposition techniques to optimize homomorphic permutations, advancing homomorphic encryption-based privacy-preserving computations. We start by defining an ideal decomposition form for permutations and propose an algorithm searching for depth-1 ideal decompositions. Based on this, we prove the full-depth ideal decomposability of permutations used in specific homomorphic matrix transposition (HMT) and multiplication (HMM) algorithms, allowing them to achieve asymptotic improvement in speed and rotation key reduction. As a demonstration of applicability, substituting the HMM components in the best-known inference framework of encrypted neural networks with our enhanced version shows up to $3.9\\times$ reduction in latency. We further devise a new method for computing arbitrary homomorphic permutations, specifically those with weak structures that cannot be ideally decomposed. We design a network structure that deviates from the conventional scope of decomposition and outperforms the state-of-the-art technique with a speed-up of up to $1.69\\times$ under a minimal rotation key requirement.","author":[{"family":"Ma","given":"Xirong"},{"family":"Fang","given":"Junling"},{"family":"Ge","given":"Chunpeng"},{"family":"Duong","given":"Dung"},{"family":"Jiang","given":"Yali"},{"family":"Li","given":"Yanbin"},{"family":"Susilo","given":"Willy"},{"family":"Cui","given":"Lizhen"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2410.21840","URL":"https://doi.org/10.48550/arxiv.2410.21840","source":"datacite"},{"id":"doi:10.21256/zhaw-25581","type":"article-journal","title":"Trusted execution environments : applications and organizational challenges","abstract":"A lack of trust in the providers is still a major barrier to cloud computing adoption – especially when sensitive data is involved. While current privacy-enhancing technologies, such as homomorphic encryption, can increase security, they come with a considerable performance overhead. As an alternative Trusted Executing Environment (TEE) provides trust guarantees for code execution in the cloud similar to transport layer security for data transport or advanced encryption standard algorithms for data storage. Cloud infrastructure providers like Amazon, Google, and Microsoft introduced TEEs as part of their infrastructure offerings. This review will shed light on the different technological options of TEEs, as well as give insight into organizational issues regarding their usage.","author":[{"family":"Geppert","given":"Tim"},{"family":"Deml","given":"Stefan"},{"family":"Sturzenegger","given":"David"},{"family":"Ebert","given":"Nico"}],"issued":{"date-parts":[[2022]]},"DOI":"10.21256/zhaw-25581","URL":"https://doi.org/10.21256/zhaw-25581","source":"datacite"},{"id":"doi:10.48550/arxiv.2311.03470","type":"manuscript","title":"Orion: A Fully Homomorphic Encryption Framework for Deep Learning","abstract":"Fully Homomorphic Encryption (FHE) has the potential to substantially improve privacy and security by enabling computation directly on encrypted data. This is especially true with deep learning, as today, many popular user services are powered by neural networks in the cloud. Beyond its well-known high computational costs, one of the major challenges facing wide-scale deployment of FHE-secured neural inference is effectively mapping these networks to FHE primitives. FHE poses many programming challenges including packing large vectors, managing accumulated noise, and translating arbitrary and general-purpose programs to the limited instruction set provided by FHE. These challenges make building large FHE neural networks intractable using the tools available today. In this paper we address these challenges with Orion, a fully-automated framework for private neural inference using FHE. Orion accepts deep neural networks written in PyTorch and translates them into efficient FHE programs. We achieve this by proposing a novel single-shot multiplexed packing strategy for arbitrary convolutions and through a new, efficient technique to automate bootstrap placement and scale management. We evaluate Orion on common benchmarks used by the FHE deep learning community and outperform state-of-the-art by 2.38x on ResNet-20, the largest network they report. Orion's techniques enable processing much deeper and larger networks. We demonstrate this by evaluating ResNet-50 on ImageNet and present the first high-resolution FHE object detection experiments using a YOLO-v1 model with 139 million parameters. Orion is open-source for all to use at: https://github.com/baahl-nyu/orion","author":[{"family":"Ebel","given":"Austin"},{"family":"Garimella","given":"Karthik"},{"family":"Reagen","given":"Brandon"}],"issued":{"date-parts":[[2023]]},"DOI":"10.48550/arxiv.2311.03470","URL":"https://doi.org/10.48550/arxiv.2311.03470","source":"datacite"},{"id":"doi:10.48550/arxiv.2410.15215","type":"manuscript","title":"DataSeal: Ensuring the Verifiability of Private Computation on Encrypted Data","abstract":"Fully Homomorphic Encryption (FHE) allows computations to be performed directly on encrypted data without needing to decrypt it first. This \"encryption-in-use\" feature is crucial for securely outsourcing computations in privacy-sensitive areas such as healthcare and finance. Nevertheless, in the context of FHE-based cloud computing, clients often worry about the integrity and accuracy of the outcomes. This concern arises from the potential for a malicious server or server-side vulnerabilities that could result in tampering with the data, computations, and results. Ensuring integrity and verifiability with low overhead remains an open problem, as prior attempts have not yet achieved this goal. To tackle this challenge and ensure the verification of FHE's private computations on encrypted data, we introduce DataSeal, which combines the low overhead of the algorithm-based fault tolerance (ABFT) technique with the confidentiality of FHE, offering high efficiency and verification capability. Through thorough testing in diverse contexts, we demonstrate that DataSeal achieves much lower overheads for providing computation verifiability for FHE than other techniques that include MAC, ZKP, and TEE. DataSeal's space and computation overheads decrease to nearly negligible as the problem size increases.","author":[{"family":"Santriaji","given":"Muhammad"},{"family":"Xue","given":"Jiaqi"},{"family":"Lou","given":"Qian"},{"family":"Solihin","given":"Yan"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2410.15215","URL":"https://doi.org/10.48550/arxiv.2410.15215","source":"datacite"},{"id":"doi:10.48550/arxiv.2412.03924","type":"manuscript","title":"Privacy-Preserving in Medical Image Analysis: A Review of Methods and Applications","abstract":"With the rapid advancement of artificial intelligence and deep learning, medical image analysis has become a critical tool in modern healthcare, significantly improving diagnostic accuracy and efficiency. However, AI-based methods also raise serious privacy concerns, as medical images often contain highly sensitive patient information. This review offers a comprehensive overview of privacy-preserving techniques in medical image analysis, including encryption, differential privacy, homomorphic encryption, federated learning, and generative adversarial networks. We explore the application of these techniques across various medical image analysis tasks, such as diagnosis, pathology, and telemedicine. Notably, we organizes the review based on specific challenges and their corresponding solutions in different medical image analysis applications, so that technical applications are directly aligned with practical issues, addressing gaps in the current research landscape. Additionally, we discuss emerging trends, such as zero-knowledge proofs and secure multi-party computation, offering insights for future research. This review serves as a valuable resource for researchers and practitioners and can help advance privacy-preserving in medical image analysis.","author":[{"family":"Zhu","given":"Yanming"},{"family":"Yin","given":"Xuefei"},{"family":"Liew","given":"Alan"},{"family":"Tian","given":"Hui"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2412.03924","URL":"https://doi.org/10.48550/arxiv.2412.03924","source":"datacite"},{"id":"doi:10.48550/arxiv.2401.00794","type":"manuscript","title":"Privacy-Preserving Data in IoT-based Cloud Systems: A Comprehensive Survey with AI Integration","abstract":"As the integration of Internet of Things devices with cloud computing proliferates, the paramount importance of privacy preservation comes to the forefront. This survey paper meticulously explores the landscape of privacy issues in the dynamic intersection of IoT and cloud systems. The comprehensive literature review synthesizes existing research, illuminating key challenges and discerning emerging trends in privacy preserving techniques. The categorization of diverse approaches unveils a nuanced understanding of encryption techniques, anonymization strategies, access control mechanisms, and the burgeoning integration of artificial intelligence. Notable trends include the infusion of machine learning for dynamic anonymization, homomorphic encryption for secure computation, and AI-driven access control systems. The culmination of this survey contributes a holistic view, laying the groundwork for understanding the multifaceted strategies employed in securing sensitive data within IoT-based cloud environments. The insights garnered from this survey provide a valuable resource for researchers, practitioners, and policymakers navigating the complex terrain of privacy preservation in the evolving landscape of IoT and cloud computing","author":[{"family":"Dhinakaran","given":"D"},{"family":"Sankar","given":"SMU"},{"family":"Selvaraj","given":"D"},{"family":"Raja","given":"SE"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2401.00794","URL":"https://doi.org/10.48550/arxiv.2401.00794","source":"datacite"},{"id":"doi:10.48550/arxiv.2210.05560","type":"manuscript","title":"Comparison of encrypted control approaches and tutorial on dynamic systems using LWE-based homomorphic encryption","abstract":"Encrypted control has been introduced to protect controller data by encryption at the stage of computation and communication, by performing the computation directly on encrypted data. In this article, we first review and categorize recent relevant studies on encrypted control. Approaches based on homomorphic encryption, multi-party computation, and secret sharing are introduced, compared, and then discussed with respect to computational complexity, communication load, enabled operations, security, and research directions. We proceed to discuss a current challenge in the application of homomorphic encryption to dynamic systems, where arithmetic operations other than integer addition and multiplication are limited. We also introduce a homomorphic cryptosystem called ``GSW-LWE'' and discuss its benefits that allow for recursive multiplication of encrypted dynamic systems, without use of computationally expensive bootstrapping techniques.","author":[{"family":"Kim","given":"Junsoo"},{"family":"Kim","given":"Dongwoo"},{"family":"Song","given":"Yongsoo"},{"family":"Shim","given":"Hyungbo"},{"family":"Sandberg","given":"Henrik"},{"family":"Johansson","given":"Karl"}],"issued":{"date-parts":[[2022]]},"DOI":"10.48550/arxiv.2210.05560","URL":"https://doi.org/10.48550/arxiv.2210.05560","source":"datacite"},{"id":"doi:10.48550/arxiv.2203.15877","type":"manuscript","title":"Quantum Advantage from Any Non-Local Game","abstract":"We show a general method of compiling any $k$-prover non-local game into a single-prover interactive game maintaining the same (quantum) completeness and (classical) soundness guarantees (up to negligible additive factors in a security parameter). Our compiler uses any quantum homomorphic encryption scheme (Mahadev, FOCS 2018; Brakerski, CRYPTO 2018) satisfying a natural form of correctness with respect to auxiliary (quantum) input. The homomorphic encryption scheme is used as a cryptographic mechanism to simulate the effect of spatial separation, and is required to evaluate $k-1$ prover strategies (out of $k$) on encrypted queries. In conjunction with the rich literature on (entangled) multi-prover non-local games starting from the celebrated CHSH game (Clauser, Horne, Shimonyi and Holt, Physical Review Letters 1969), our compiler gives a broad framework for constructing mechanisms to classically verify quantum advantage.","author":[{"family":"Kalai","given":"Yael"},{"family":"Lombardi","given":"Alex"},{"family":"Vaikuntanathan","given":"Vinod"},{"family":"Yang","given":"Lisa"}],"issued":{"date-parts":[[2022]]},"DOI":"10.48550/arxiv.2203.15877","URL":"https://doi.org/10.48550/arxiv.2203.15877","source":"datacite"},{"id":"doi:10.48550/arxiv.2011.05296","type":"manuscript","title":"A Systematic Comparison of Encrypted Machine Learning Solutions for Image Classification","abstract":"This work provides a comprehensive review of existing frameworks based on secure computing techniques in the context of private image classification. The in-depth analysis of these approaches is followed by careful examination of their performance costs, in particular runtime and communication overhead. To further illustrate the practical considerations when using different privacy-preserving technologies, experiments were conducted using four state-of-the-art libraries implementing secure computing at the heart of the data science stack: PySyft and CrypTen supporting private inference via Secure Multi-Party Computation, TF-Trusted utilising Trusted Execution Environments and HE- Transformer relying on Homomorphic encryption. Our work aims to evaluate the suitability of these frameworks from a usability, runtime requirements and accuracy point of view. In order to better understand the gap between state-of-the-art protocols and what is currently available in practice for a data scientist, we designed three neural network architecture to obtain secure predictions via each of the four aforementioned frameworks. Two networks were evaluated on the MNIST dataset and one on the Malaria Cell image dataset. We observed satisfying performances for TF-Trusted and CrypTen and noted that all frameworks perfectly preserved the accuracy of the corresponding plaintext model.","author":[{"family":"Haralampieva","given":"Veneta"},{"family":"Rueckert","given":"Daniel"},{"family":"Passerat-Palmbach","given":"Jonathan"}],"issued":{"date-parts":[[2020]]},"DOI":"10.48550/arxiv.2011.05296","URL":"https://doi.org/10.48550/arxiv.2011.05296","source":"datacite"},{"id":"doi:10.48550/arxiv.2105.07533","type":"manuscript","title":"Private Facial Diagnosis as an Edge Service for Parkinson's DBS Treatment Valuation","abstract":"Facial phenotyping has recently been successfully exploited for medical diagnosis as a novel way to diagnose a range of diseases, where facial biometrics has been revealed to have rich links to underlying genetic or medical causes. In this paper, taking Parkinson's Diseases (PD) as a case study, we proposed an Artificial-Intelligence-of-Things (AIoT) edge-oriented privacy-preserving facial diagnosis framework to analyze the treatment of Deep Brain Stimulation (DBS) on PD patients. In the proposed framework, a new edge-based information theoretically secure framework is proposed to implement private deep facial diagnosis as a service over a privacy-preserving AIoT-oriented multi-party communication scheme, where partial homomorphic encryption (PHE) is leveraged to enable privacy-preserving deep facial diagnosis directly on encrypted facial patterns. In our experiments with a collected facial dataset from PD patients, for the first time, we demonstrated that facial patterns could be used to valuate the improvement of PD patients undergoing DBS treatment. We further implemented a privacy-preserving deep facial diagnosis framework that can achieve the same accuracy as the non-encrypted one, showing the potential of our privacy-preserving facial diagnosis as an trustworthy edge service for grading the severity of PD in patients.","author":[{"family":"Jiang","given":"Richard"},{"family":"Chazot","given":"Paul"},{"family":"Crookes","given":"Danny"},{"family":"Bouridane","given":"Ahmed"},{"family":"Celebi","given":"ME"}],"issued":{"date-parts":[[2021]]},"DOI":"10.48550/arxiv.2105.07533","URL":"https://doi.org/10.48550/arxiv.2105.07533","source":"datacite"},{"id":"doi:10.5281/zenodo.13627270","type":"article-journal","title":"Improving medical data synthesis with DP-GAN and Deep Anomaly Detection","abstract":"Ensuring the privacy of medical data in a meaningful manner is a complex task. This domain presents a plethora of unique challenges: high stakes, vast differences between possible use cases, long-established methods that limit the number of feasible solutions, and more. Consequently, an effective approach to ensuring the privacy of medical data must be easy to adopt, offer robust privacy guarantees, and minimize the reduction in data utility.The unique nature of medical data presents distinct challenges and also opportunities. We consider various types of correlations that significantly impact privacy guarantees. However, these correlations can also be used to train a model for removing anomalies and subsequently enhancing the utility of synthetic medical data.This thesis proposes a framework compatible with state-of-the-art approaches for differentially private dataset release based on the usage of Generative Adversarial Networks (GANs). Our framework uses a part of the privacy budget to train an unsupervised learning model to detect and remove anomalies. We evaluate the performance of the framework using a variety of machine-learning models and metrics. The final results show an improvement of up 13% compared to approaches not using our framework, under the same privacy budget.","author":[{"family":"Crha","given":"Vojtech"},{"family":"Hai","given":"Rihan"},{"family":"Erkin","given":"Zekeriya"},{"family":"Li","given":"Tianyu"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.13627270","URL":"https://doi.org/10.5281/zenodo.13627270","source":"datacite"},{"id":"doi:10.5281/zenodo.13627271","type":"article-journal","title":"Improving medical data synthesis with DP-GAN and Deep Anomaly Detection","abstract":"Ensuring the privacy of medical data in a meaningful manner is a complex task. This domain presents a plethora of unique challenges: high stakes, vast differences between possible use cases, long-established methods that limit the number of feasible solutions, and more. Consequently, an effective approach to ensuring the privacy of medical data must be easy to adopt, offer robust privacy guarantees, and minimize the reduction in data utility.The unique nature of medical data presents distinct challenges and also opportunities. We consider various types of correlations that significantly impact privacy guarantees. However, these correlations can also be used to train a model for removing anomalies and subsequently enhancing the utility of synthetic medical data.This thesis proposes a framework compatible with state-of-the-art approaches for differentially private dataset release based on the usage of Generative Adversarial Networks (GANs). Our framework uses a part of the privacy budget to train an unsupervised learning model to detect and remove anomalies. We evaluate the performance of the framework using a variety of machine-learning models and metrics. The final results show an improvement of up 13% compared to approaches not using our framework, under the same privacy budget.","author":[{"family":"Crha","given":"Vojtech"},{"family":"Hai","given":"Rihan"},{"family":"Erkin","given":"Zekeriya"},{"family":"Li","given":"Tianyu"}],"issued":{"date-parts":[[2024]]},"DOI":"10.5281/zenodo.13627271","URL":"https://doi.org/10.5281/zenodo.13627271","source":"datacite"},{"id":"doi:10.48550/arxiv.2410.22699","type":"manuscript","title":"Exactly Minimax-Optimal Locally Differentially Private Sampling","abstract":"The sampling problem under local differential privacy has recently been studied with potential applications to generative models, but a fundamental analysis of its privacy-utility trade-off (PUT) remains incomplete. In this work, we define the fundamental PUT of private sampling in the minimax sense, using the f-divergence between original and sampling distributions as the utility measure. We characterize the exact PUT for both finite and continuous data spaces under some mild conditions on the data distributions, and propose sampling mechanisms that are universally optimal for all f-divergences. Our numerical experiments demonstrate the superiority of our mechanisms over baselines, in terms of theoretical utilities for finite data space and of empirical utilities for continuous data space.","author":[{"family":"Park","given":"Hyun"},{"family":"Asoodeh","given":"Shahab"},{"family":"Lee","given":"Si"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2410.22699","URL":"https://doi.org/10.48550/arxiv.2410.22699","source":"datacite"},{"id":"doi:10.48550/arxiv.2405.02665","type":"manuscript","title":"Metric Differential Privacy at the User-Level Via the Earth Mover's Distance","abstract":"Metric differential privacy (DP) provides heterogeneous privacy guarantees based on a distance between the pair of inputs. It is a widely popular notion of privacy since it captures the natural privacy semantics for many applications (such as, for location data) and results in better utility than standard DP. However, prior work in metric DP has primarily focused on the item-level setting where every user only reports a single data item. A more realistic setting is that of user-level DP where each user contributes multiple items and privacy is then desired at the granularity of the user's entire contribution. In this paper, we initiate the study of one natural definition of metric DP at the user-level. Specifically, we use the earth-mover's distance ($d_\\textsf{EM}$) as our metric to obtain a notion of privacy as it captures both the magnitude and spatial aspects of changes in a user's data. We make three main technical contributions. First, we design two novel mechanisms under $d_\\textsf{EM}$-DP to answer linear queries and item-wise queries. Specifically, our analysis for the latter involves a generalization of the privacy amplification by shuffling result which may be of independent interest. Second, we provide a black-box reduction from the general unbounded to bounded $d_\\textsf{EM}$-DP (size of the dataset is fixed and public) with a novel sampling based mechanism. Third, we show that our proposed mechanisms can provably provide improved utility over user-level DP, for certain types of linear queries and frequency estimation.","author":[{"family":"Imola","given":"Jacob"},{"family":"Chowdhury","given":"Amrita"},{"family":"Chaudhuri","given":"Kamalika"}],"issued":{"date-parts":[[2024]]},"DOI":"10.48550/arxiv.2405.02665","URL":"https://doi.org/10.48550/arxiv.2405.02665","source":"datacite"},{"id":"oa:W4303453208","type":"article-journal","title":"Trusted Data Sharing Mechanism Based on Blockchain and Federated Learning in Space‐Air‐Ground Integrated Networks","abstract":"Network data is distributed data on electricity, with the explosive growth of network data, and it has become an inevitable trend of network development to synergized and shared crossdomain scattered data and enhances the value transmission of network data. Federated learning, as a technology that combines data value delivery and data privacy security, is widely concerned in the process of data sharing. However, currently federated learning is used within a single business system. In the process of crossdomain data sharing, how to ensure the data trust, model trust, and result trust of federated learning is still an urgent problem to be solved. To this end, we designed to use blockchain structure to record each behavior of data sharing. Based on its tamper‐proof and traceability, combined with cryptography technology, we constructed an endogenous trusted architecture for crossdomain data sharing. In addition, a reverse auction node incentive mechanism based on high credit preference is designed to solve the common problems in data sharing, such as low enthusiasm of users in sharing, unstable data quality of contributions, and unreasonable distribution of data sharing benefits. Through theoretical analysis and experimental verification, it can be seen that the incentive mechanism designed in this paper can meet the authenticity, user rationality, and budget feasibility. On this basis, it can motivate users to participate in data sharing, improve the average quality of data shared by users, and ensure security and trustworthiness and resist malicious attacks to a certain extent.","author":[{"family":"Li","given":"Da"},{"family":"Guo","given":"Qinglei"},{"family":"Yang","given":"Chao"},{"family":"Yan","given":"Han"}],"issued":{"date-parts":[[2022]]},"DOI":"10.1155/2022/5338876","URL":"https://doi.org/10.1155/2022/5338876","source":"openalex"},{"id":"oa:W3088635668","type":"article-journal","title":"Advancing Fusion with Machine Learning Research Needs Workshop Report","abstract":"Abstract Machine learning and artificial intelligence (ML/AI) methods have been used successfully in recent years to solve problems in many areas, including image recognition, unsupervised and supervised classification, game-playing, system identification and prediction, and autonomous vehicle control. Data-driven machine learning methods have also been applied to fusion energy research for over 2 decades, including significant advances in the areas of disruption prediction, surrogate model generation, and experimental planning. The advent of powerful and dedicated computers specialized for large-scale parallel computation, as well as advances in statistical inference algorithms, have greatly enhanced the capabilities of these computational approaches to extract scientific knowledge and bridge gaps between theoretical models and practical implementations. Large-scale commercial success of various ML/AI applications in recent years, including robotics, industrial processes, online image recognition, financial system prediction, and autonomous vehicles, have further demonstrated the potential for data-driven methods to produce dramatic transformations in many fields. These advances, along with the urgency of need to bridge key gaps in knowledge for design and operation of reactors such as ITER, have driven planned expansion of efforts in ML/AI within the US government and around the world. The Department of Energy (DOE) Office of Science programs in Fusion Energy Sciences (FES) and Advanced Scientific Computing Research (ASCR) have organized several activities to identify best strategies and approaches for applying ML/AI methods to fusion energy research. This paper describes the results of a joint FES/ASCR DOE-sponsored Research Needs Workshop on Advancing Fusion with Machine Learning, held April 30–May 2, 2019, in Gaithersburg, MD (full report available at https://science.osti.gov/-/media/fes/pdf/workshop-reports/FES_ASCR_Machine_Learning_Report.pdf ). The workshop drew on broad representation from both FES and ASCR scientific communities, and identified seven Priority Research Opportunities (PRO’s) with high potential for advancing fusion energy. In addition to the PRO topics themselves, the workshop identified research guidelines to maximize the effectiveness of ML/AI methods in fusion energy science, which include focusing on uncertainty quantification, methods for quantifying regions of validity of models and algorithms, and applying highly integrated teams of ML/AI mathematicians, computer scientists, and fusion energy scientists with domain expertise in the relevant areas.","author":[{"family":"Humphreys","given":"DA"},{"family":"Kupresanin","given":"Ana"},{"family":"Boyer","given":"Mark"},{"family":"Canik","given":"JM"},{"family":"Chang","given":"CS"},{"family":"Cyr","given":"Eric"},{"family":"Granetz","given":"R"},{"family":"Hittinger","given":"J"},{"family":"Kolemen","given":"Egemen"},{"family":"Lawrence","given":"Earl"},{"family":"Pascucci","given":"Valerio"},{"family":"Patra","given":"Aditya"},{"family":"Schissel","given":"DP"}],"issued":{"date-parts":[[2020]]},"DOI":"10.1007/s10894-020-00258-1","URL":"https://doi.org/10.1007/s10894-020-00258-1","source":"openalex"},{"id":"oa:W4308671132","type":"manuscript","title":"Enhancing Efficiency in Multidevice Federated Learning through Data Selection","abstract":"Ubiquitous wearable and mobile devices provide access to a diverse set of data. However, the mobility demand for our devices naturally imposes constraints on their computational and communication capabilities. A solution is to locally learn knowledge from data captured by ubiquitous devices, rather than to store and transmit the data in its original form. In this paper, we develop a federated learning framework, called Centaur, to incorporate on-device data selection at the edge, which allows partition-based training of a deep neural nets through collaboration between constrained and resourceful devices within the multidevice ecosystem of the same user. We benchmark on five neural net architecture and six datasets that include image data and wearable sensor time series. On average, Centaur achieves ~19% higher classification accuracy and ~58% lower federated training latency, compared to the baseline. We also evaluate Centaur when dealing with imbalanced non-iid data, client participation heterogeneity, and different mobility patterns. To encourage further research in this area, we release our code at https://github.com/nokia-bell-labs/data-centric-federated-learning","author":[{"family":"Mo","given":"Fan"},{"family":"Malekzadeh","given":"Mohammad"},{"family":"Chatterjee","given":"Soumyajit"},{"family":"Kawsar","given":"Fahim"},{"family":"Mathur","given":"Akhil"}],"issued":{"date-parts":[[2022]]},"DOI":"10.48550/arxiv.2211.04175","URL":"https://doi.org/10.48550/arxiv.2211.04175","source":"openalex"},{"id":"oa:W4309953431","type":"manuscript","title":"Resource-Constrained Decentralized Federated Learning via Personalized Event-Triggering","abstract":"Federated learning (FL) is a popular technique for distributing machine learning (ML) across a set of edge devices. In this paper, we study fully decentralized FL, where in addition to devices conducting training locally, they carry out model aggregations via cooperative consensus formation over device-to-device (D2D) networks. We introduce asynchronous, event-triggered communications among the devices to handle settings where access to a central server is not feasible. To account for the inherent resource heterogeneity and statistical diversity challenges in FL, we define personalized communication triggering conditions at each device that weigh the change in local model parameters against the available local network resources. We theoretically recover the $O(\\ln{k} / \\sqrt{k})$ convergence rate to the globally optimal model of decentralized gradient descent (DGD) methods in the setup of our methodology. We provide our convergence guarantees for the last iterates of models, under relaxed graph connectivity and data heterogeneity assumptions compared with the existing literature. To do so, we demonstrate a $B$-connected information flow guarantee in the presence of sporadic communications over the time-varying D2D graph. Our subsequent numerical evaluations demonstrate that our methodology obtains substantial improvements in convergence speed and/or communication savings compared to existing decentralized FL baselines.","author":[{"family":"Zehtabi","given":"Shahryar"},{"family":"Hosseinalipour","given":"Seyyedali"},{"family":"Brinton","given":"Christopher"}],"issued":{"date-parts":[[2022]]},"DOI":"10.48550/arxiv.2211.12640","URL":"https://doi.org/10.48550/arxiv.2211.12640","source":"openalex"},{"id":"oa:W4404295022","type":"article-journal","title":"Advancements and Challenges in Federated Learning : A Survey","abstract":"Machine Learning (ML) has significantly impacted daily life by automating tasks and enhancing decision-making across diverse sectors such as healthcare, finance, and transportation. However, concerns regarding data privacy, particularly in sensitive domains like healthcare and finance, have impeded its widespread adoption. Federated Learning (FL), pioneered by Google in 2016, presents a promising solution by enabling devices to collaborate on model training without sharing raw data. In FL, each device retains its data locally, and only model updates are exchanged with a central server for aggregation. This decentralized approach preserves data privacy while still improving model performance through collaborative learning. FL has garnered interest across industries, with approximately 32% of companies intending to integrate it into their systems soon. Furthermore, investment in FL is projected to rise substantially from $107 million in 2020 to $538 million by 2025, as indicated by forecasts from KPMG. This paper provides an overview of FL, its applications across various sectors, current adoption trends, and future growth prospects, highlighting its significance in addressing data privacy concerns while advancing machine learning capabilities.","author":[{"family":"Prajapat","given":"Vivek"},{"family":"Rathore","given":"Narendra"},{"family":"Sethi","given":"Kamal"},{"family":"Rajput","given":"Shiv"}],"issued":{"date-parts":[[2024]]},"DOI":"10.1109/acroset62108.2024.10743243","URL":"https://doi.org/10.1109/acroset62108.2024.10743243","source":"openalex"},{"id":"oa:W4297006047","type":"article-journal","title":"HealthGuard: An Intelligent Healthcare System Security Framework Based on Machine Learning","abstract":"Utilization of the Internet of Things and ubiquitous computing in medical apparatuses have “smartified” the current healthcare system. These days, healthcare is used for more than simply curing patients. A Smart Healthcare System (SHS) is a network of implanted medical devices and wearables that monitors patients in real-time to detect and avert potentially fatal illnesses. With its expanding capabilities comes a slew of security threats, and there are many ways in which a SHS might be exploited by malicious actors. These include, but are not limited to, interfering with regular SHS functioning, inserting bogus data to modify vital signs, and meddling with medical devices. This study presents HealthGuard, an innovative security architecture for SHSs that uses machine learning to identify potentially harmful actions taken by users. HealthGuard monitors the vitals of many SHS-connected devices and compares the vitals to distinguish normal from abnormal activity. For the purpose of locating potentially dangerous actions inside a SHS, HealthGuard employs four distinct machine learning-based detection approaches (Artificial Neural Network, Decision Tree, Random Forest, and k-Nearest Neighbor). Eight different smart medical devices were used to train HealthGuard for a total of twelve harmless occurrences, seven of which are common user activities and five of which are disease-related occurrences. HealthGuard was also tested for its ability to defend against three distinct forms of harmful attack. Our comprehensive analysis demonstrates that HealthGuard is a reliable security architecture for SHSs, with a 91% success rate and in F1-score of 90% success.","author":[{"family":"Sundas","given":"Amit"},{"family":"Badotra","given":"Sumit"},{"family":"Bharany","given":"Salil"},{"family":"Almogren","given":"Ahmad"},{"family":"Eldin","given":"Sayed"},{"family":"Rehman","given":"Ateeq"}],"issued":{"date-parts":[[2022]]},"DOI":"10.3390/su141911934","URL":"https://doi.org/10.3390/su141911934","source":"openalex"},{"id":"oa:W4286888504","type":"manuscript","title":"Deep Learning in Human Activity Recognition with Wearable Sensors: A Review on Advances","abstract":"Mobile and wearable devices have enabled numerous applications, including activity tracking, wellness monitoring, and human--computer interaction, that measure and improve our daily lives. Many of these applications are made possible by leveraging the rich collection of low-power sensors found in many mobile and wearable devices to perform human activity recognition (HAR). Recently, deep learning has greatly pushed the boundaries of HAR on mobile and wearable devices. This paper systematically categorizes and summarizes existing work that introduces deep learning methods for wearables-based HAR and provides a comprehensive analysis of the current advancements, developing trends, and major challenges. We also present cutting-edge frontiers and future directions for deep learning-based HAR.","author":[{"family":"Zhang","given":"Shibo"},{"family":"Li","given":"Yaxuan"},{"family":"Zhang","given":"Shen"},{"family":"Shahabi","given":"Farzad"},{"family":"Xia","given":"Stephen"},{"family":"Deng","given":"Yu"},{"family":"Alshurafa","given":"Nabil"}],"issued":{"date-parts":[[2021]]},"DOI":"10.48550/arxiv.2111.00418","URL":"https://doi.org/10.48550/arxiv.2111.00418","source":"openalex"},{"id":"oa:W3097471692","type":"article-journal","title":"Digital Twins: State of the art theory and practice, challenges, and open research questions","abstract":"Digital Twin was introduced over a decade ago, as an innovative all-encompassing tool, with perceived benefits including real-time monitoring, simulation, optimisation and accurate forecasting. However, the theoretical framework and practical implementations of digital twin (DT) are yet to fully achieve this vision at scale. Although an increasing number of successful implementations exist in research and industrial works, sufficient implementation details are not publicly available, making it difficult to fully assess their components and effectiveness, to draw comparisons, identify successful solutions, share lessons, and thus to jointly advance and benefit from the DT methodology. This work first presents a review of relevant DT research and industrial works, focusing on the key DT features, current approaches in different domains, and successful DT implementations, to infer the key DT components and properties, and to identify current limitations and reasons behind the delay in the widespread implementation and adoption of digital twin. This work identifies that the major reasons for this delay are: the fact the DT is still a fast evolving concept; the lack of a universal DT reference framework, e.g. DT standards are scarce and still evolving; problem- and domain-dependence; security concerns over shared data; lack of DT performance metrics; and reliance of digital twin on other fast-evolving technologies. Advancements in machine learning, Internet of Things (IoT) and big data have led to significant improvements in DT features such as real-time monitoring and accurate forecasting. Despite this progress and individual company-based efforts, certain research and implementation gaps exist in the field, which have so far prevented the widespread adoption of the DT concept and technology; these gaps are also discussed in this work. Based on reviews of past work and the identified gaps, this work then defines a conceptualisation of DT which includes its components and properties; these also validate the uniqueness of DT as a concept, when compared to similar concepts such as simulation, autonomous systems and optimisation. Real-life case studies are used to showcase the application of the conceptualisation. This work discusses the state-of-the-art in DT, addresses relevant and timely DT questions, and identifies novel research questions, thus contributing to a better understanding of the DT paradigm and advancing the theory and practice of DT and its allied technologies.","author":[{"family":"Sharma","given":"Angira"},{"family":"Kosasih","given":"Edward"},{"family":"Zhang","given":"Jie"},{"family":"Brintrup","given":"Alexandra"},{"family":"Calinescu","given":"Anisoara"}],"issued":{"date-parts":[[2022]]},"DOI":"10.1016/j.jii.2022.100383","URL":"https://doi.org/10.1016/j.jii.2022.100383","source":"openalex"},{"id":"oa:W4408287608","type":"article-journal","title":"DataSHIELD: mitigating disclosure risk in a multi-site federated analysis platform","abstract":"Motivation: The validity of epidemiologic findings can be increased using triangulation, i.e. comparison of findings across contexts, and by having sufficiently large amounts of relevant data to analyse. However, access to data is often constrained by practical considerations and by ethico-legal and data governance restrictions. Gaining access to such data can be time-consuming due to the governance requirements associated with data access requests to institutions in different jurisdictions. Results: DataSHIELD is a software solution that enables remote analysis without the need for data transfer (federated analysis). DataSHIELD is a scientifically mature, open-source data access and analysis platform aligned with the 'Five Safes' framework, the international framework governing safe research access to data. It allows real-time analysis while mitigating disclosure risk through an active multi-layer system of disclosure-preventing mechanisms. This combination of real-time remote statistical analysis, disclosure prevention mechanisms, and federation capabilities makes DataSHIELD a solution for addressing many of the technical and regulatory challenges in performing the large-scale statistical analysis of health and biomedical data. This paper describes the key components that comprise the disclosure protection system of DataSHIELD. These broadly fall into three classes: (i) system protection elements, (ii) analysis protection elements, and (iii) governance protection elements. Availability and implementation: Information about the DataSHIELD software is available in https://datashield.org/ and https://github.com/datashield.","author":[{"family":"Avraam","given":"Demetris"},{"family":"Wilson","given":"Rebecca"},{"family":"Chan","given":"Noemi"},{"family":"Banerjee","given":"Soumya"},{"family":"Bishop","given":"Tom"},{"family":"Butters","given":"OW"},{"family":"Cadman","given":"Tim"},{"family":"Cederkvist","given":"Luise"},{"family":"Duijts","given":"Liesbeth"},{"family":"Escribà-Montagut","given":"Xavier"},{"family":"Garner","given":"Hugh"},{"family":"Gonçalves","given":"Gonçalo"},{"family":"González","given":"Juan"},{"family":"Haakma","given":"Sido"},{"family":"Hartlev","given":"Mette"},{"family":"Hasenauer","given":"Jan"},{"family":"Huth","given":"Manuel"},{"family":"Hyde","given":"Eleanor"},{"family":"Jaddoe","given":"Vincent"},{"family":"Marcon","given":"Yannick"},{"family":"Mayrhofer","given":"Michaela"},{"family":"Molnárgábor","given":"Fruzsina"},{"family":"Morgan","given":"Andreï"},{"family":"Murtagh","given":"Madeleine"},{"family":"Nestor","given":"Marc"},{"family":"Andersen","given":"Anne‐marie"},{"family":"Parker","given":"Simon"},{"family":"Moira","given":"Angela"},{"family":"Schwarz","given":"Florian"},{"family":"Strandberglarsen","given":"Katrine"},{"family":"Swertz","given":"Morris"},{"family":"Welten","given":"Marieke"},{"family":"Wheater","given":"Stuart"},{"family":"Burton","given":"Paul"}],"issued":{"date-parts":[[2024]]},"DOI":"10.1093/bioadv/vbaf046","URL":"https://doi.org/10.1093/bioadv/vbaf046","source":"openalex"},{"id":"doi:10.21203/rs.3.rs-9747984/v1","type":"article-journal","title":"On a Symmetric Homomorphic Encryption Scheme for Multidimensional Data in IoT","abstract":"Abstract Homomorphic encryption (HE) schemes enable one to perform certain operations on the encrypted data without decrypting. In literature, HE schemes that allow simple computations on encrypted data have been known for a long time and many efficient HE encryption schemes have been developed. But, most of them are designed for single-dimensional data where encryption is performed on individual elements sequentially, rather than on multidimensional structures as a whole, such as on a vector or an array of elements simultaneously. This makes them inefficient HE for multidimensional data. Due to the emergence of various Internet of Things (IoTs) applications such as smart health, smart grid, smart city, and other domains where IoT devices are required to transmit their sensed data for further processing, it is desirable to have an encryption scheme that can deal with multidimensional data. Because of its wide range of applicability in various IoT applications, in this paper, we propose a Symmetric Homomorphic Encryption for Multidimensional data ( symhem ). Our scheme is easily implementable because of its simple definition. At the same time, it is secure because of its two-step encryption. We have also conducted various simulations and compared the symhem with the other state-of-the-art HE schemes.","author":[{"family":"Upadhyay","given":"Manvi"},{"family":"Choudhuri","given":"Manoj"},{"family":"Yadav","given":"Ram"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9747984/v1","URL":"https://doi.org/10.21203/rs.3.rs-9747984/v1","source":"europepmc"},{"id":"doi:10.21203/rs.3.rs-10097660/v1","type":"article-journal","title":"Federated Learning Parameter Protection Based on Homomorphic Encryption and Selective User Decryption","abstract":"Abstract Existing security schemes of Federated learning mostly rely on encryption mechanism and participant reliability, but it is difficult to effectively defend against attack threats such as member reasoning, attribute reasoning, model reversal, etc. Meanwhile, the low-quality data users in the system will directly slow down the training speed and undermine the stability of the system due to their unreliable infrastructure and potential evil motives. To solve these problems, this paper proposes a federated learning model parameter protection scheme based on threshold homomorphic encryption. This scheme is based on distributed Paillier homomorphic encryption mechanism. Its core is to split the private key through secret sharing technology, so as to eliminate the risk brought by the single private key holder and resist the collusion attack of malicious participants and semi trusted servers. On this basis, top-t high-quality trusted nodes (active contributors) are selected through the data quality evaluation mechanism to participate in the decryption task. At the same time, combined with ECDSA digital signature technology, it ensures the integrity of encryption parameters in the transmission process and the authentication of end-to-end communication. Experiments demonstrate that, compared to existing homomorphic encryption schemes, our approach achieves faster model learning convergence with comparable accuracy. Specifically, fewer iterations are required for model parameters to stabilize, yielding approximately 10% higher learning efficiency. This resolves security threats posed by malicious users' inference attacks and server-side aggregation tampering, while safeguarding the reliability of parameter transmission.","author":[{"family":"Li","given":"Zhangbing"},{"family":"Xiao","given":"Mingyu"},{"family":"Xiao","given":"Jiantian"},{"family":"Li","given":"Jinsheng"},{"family":"Zhang","given":"Shaobo"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-10097660/v1","URL":"https://doi.org/10.21203/rs.3.rs-10097660/v1","source":"europepmc"},{"id":"doi:10.1038/s41746-026-02790-4","type":"article-journal","title":"A multiparty homomorphic encryption approach to confidential federated Kaplan-Meier survival analysis.","abstract":"Abstract The proliferation of real-world health data enables multi-institutional survival studies, yet privacy constraints preclude centralizing sensitive records. We present a privacy-preserving federated Kaplan–Meier framework based on threshold CKKS (Cheon-Kim-Kim-Song) homomorphic encryption that supports approximate floating-point computation and encrypted aggregation of per-time-point counts while exposing only public outputs. Sites compute aligned at-risk and event tallies on a shared time grid and encrypt compact vectors; a coordinator aggregates ciphertexts; and a decryptor committee produces partial shares fused per block to recover aggregated plaintexts without releasing per-time-point tables. We prove correctness, stability, and slot-optimal vector packing, and derive scaling laws showing that communication grows linearly with the number of sites and predictably with the number of time points. Empirically, using synthetic breast-cancer data ( N = 60,000) distributed across 500 sites, encrypted federated curves match the pooled oracle to numerical precision. In contrast, plaintext protocols permit trivial reconstruction by subtraction; our threshold-gated design precludes this attack under the stated threat model, enabling high-fidelity survival estimation with predictable overhead and substantially reduced privacy risk.","author":[{"family":"Veeraragavan","given":"Narasimha"},{"family":"Boudko","given":"Svetlana"},{"family":"Nygård","given":"Jan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41746-026-02790-4","URL":"https://doi.org/10.1038/s41746-026-02790-4","source":"europepmc"},{"id":"doi:10.1371/journal.pone.0349432","type":"article-journal","title":"Threshold-adaptive pruning with multi-key homomorphic encryption for communication-efficient secure federated learning.","abstract":"Under the federated learning framework, frequent parameter interactions between edge devices and servers result in communication inefficiency, while conventional encryption methods fail to resist multi-node collusion attacks. To address these challenges, this paper proposes an optimized federated learning scheme integrating adaptive channel pruning with multi-key homomorphic encryption. First, we construct a dynamic threshold determination mechanism that automatically calibrates channel pruning rates through precision feedback during the pre-pruning phase, achieving the optimal balance between model compression and accuracy, while significantly reducing communication bandwidth consumption compared to traditional algorithms. Second, based on the Brakerski-Gentry-Vaikuntanathan (BGV) multi-key fully homomorphic encryption architecture, we design a distributed public-key encryption protocol that enables aggregation servers to securely fuse multi-source model parameters without decryption, resisting collusion attacks from up to C&#x2009;-&#x2009;1 nodes (where C denotes the total number of devices). Experiments on MNIST and CIFAR-10 datasets demonstrate that our scheme significantly reduces communication overhead through two complementary mechanisms: adaptive pruning reduces both the computational burden of local training and the volume of parameters transmitted per round, while multi-key BGV encryption ensures privacy-preserving aggregation without decryption. This work provides a novel technical pathway for privacy-preserving federated learning in resource-constrained scenarios.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1371/journal.pone.0349432","URL":"https://doi.org/10.1371/journal.pone.0349432","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-9660473/v1","type":"article-journal","title":"HE-CloudML: A Privacy-Preserving Framework for Secure Machine Learning Inference over Encrypted Cloud Data Using Homomorphic Encryption","abstract":"Abstract The widespread adoption of cloud-based Machine Learning as a Service (MLaaS) exposes sensitive user data to critical privacy risks during inference, as plaintext data must typically be processed by untrusted cloud servers. This paper presents HE-CloudML, a unified privacy-preserving framework for secure deep neural network (DNN) inference over encrypted cloud data using Homomorphic Encryption (HE). HE-CloudML is architected as a three-tier system comprising a client-side CKKS encryption module, a cloud-side HE inference engine, and a distributed key management layer, ensuring that raw input data is never exposed to the server at any stage of computation. The framework introduces HE-compatible polynomial activation function approximations via degree-5 Chebyshev minimax polynomials, an optimized SIMD ciphertext batching strategy exploiting Ring Learning With Errors (RLWE) slot packing, and an adaptive lazy bootstrapping pipeline to substantially reduce homomorphic evaluation depth and inference latency. A formal security analysis under the IND-CPA model grounded in the RLWE hardness assumption demonstrates resistance to inference, model inversion, and membership inference attacks. Comprehensive experiments across three domains benchmark image classification (MNIST: 99.28%, CIFAR-10: 90.37%), medical imaging (93.61%), and financial fraud detection (96.44%) demonstrate that HE-CloudML achieves near-plaintext accuracy with a maximum accuracy drop of 1.81%, while delivering up to 26.9× latency improvements over CryptoNets.","author":[{"family":"Abdirahman","given":"Abdullahi"},{"family":"Hashi","given":"Abdirahman"},{"family":"Dahir","given":"Ubaid"},{"family":"Elmi","given":"Mohamed"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9660473/v1","URL":"https://doi.org/10.21203/rs.3.rs-9660473/v1","source":"europepmc"},{"id":"doi:10.20944/preprints202601.0638.v1","type":"manuscript","title":"Multi-Group Fully Homomorphic Encryption Scheme Based on LWE and NTRU","abstract":"Multi-Group Homomorphic Encryption (MGHE) is a pivotal advance in secure multi-party computation, integrating merits of Multi-Party Homomorphic Encryption (MPHE) and Multi-Key Homomorphic Encryption (MKHE) to eliminate MPHE’s fixed-party limitation and mitigate MKHE’s ciphertext expansion from dynamic enrollment. However, the efficient single-key FINAL scheme cannot extend to multi-party scenarios, due to the challenge of defining valid multiplication for vector NTRU ciphertexts, which hinders its use in multi-group bootstrapping and curbs efficiency. To address this, additive secret sharing is adopted to convert vector NTRU ciphertext multiplication into secret share multiplication, enabling shared bootstrapping key generation within groups. For the first time, a multi-group ciphertext bootstrapping algorithm based on LWE and NTRU is proposed. Bootstrapping tasks are decomposed for parallel processing, and a hybrid product algorithm is designed to aggregate subtask outputs, boosting multi-group bootstrapping speed to match that of single-key ciphertexts. Noise accumulation is analyzed, with 100-bit and 128-bit security parameter sets selected for validation. Experiments show that 30/50-party multi-group bootstrapping takes only 1.87/2.58 seconds respectively.","author":[{"family":"Li","given":"Yongheng"},{"family":"Wen","given":"Jing"},{"family":"Liang","given":"Shaoling"},{"family":"Kong","given":"Fanqi"},{"family":"Huang","given":"Baohua"}],"issued":{"date-parts":[[2026]]},"DOI":"10.20944/preprints202601.0638.v1","URL":"https://doi.org/10.20944/preprints202601.0638.v1","source":"europepmc"},{"id":"doi:10.1093/bib/bbaf648","type":"article-journal","title":"KmerCrypt: private k-mer search with homomorphic encryption.","abstract":"Abstract Outsourcing the storage and analysis of genomic data to third-party servers is often necessary due to the scale of modern datasets, but it introduces significant privacy challenges that must be addressed to ensure secure handling. K-mer-based analyses offer broad applications across genomics research, clinical diagnostics, pathogen surveillance, and metagenomic classification, though implementation requires careful ethical and technical considerations, particularly when processing human genomic data in clinical settings. We present a novel protocol utilizing homomorphic encryption that enables a client to store a fully encrypted version of a genome on an untrusted server and perform private k-mer searches. The protocol ensures the server never gains access to the client’s non-encrypted genome sequence, nor does it learn the content of any k-mer query. After a one-time client-side encryption of the genome, the server performs all computations on ciphertext, returning only encrypted results that can be decrypted solely by the data owner. This framework transforms an honest but curious cloud server into a secure storage and computation system, enabling practical and confidential querying of encrypted, client-owned genomic data. The system supports exact k-mer searches on genomic data, as well as position weight matrix searches. Finally, we provide KmerCrypt, a private k-mer search toolkit that implements this protocol, offering researchers an efficient and secure solution for querying encrypted genomic datasets without compromising privacy.","author":[{"family":"Provatas","given":"Kimonas"},{"family":"Mouratidis","given":"Ioannis"},{"family":"Georgakopoulos-Soares","given":"Ilias"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1093/bib/bbaf648","URL":"https://doi.org/10.1093/bib/bbaf648","source":"europepmc"},{"id":"doi:10.20944/preprints202601.1157.v1","type":"manuscript","title":"Secure and Verifiable Edge-Federated Learning with Homomorphic Encryption and a Trusted Execution Environment for UAV Communication","abstract":"Edge drones continuously collect sensitive information such as telemetry data during missions, making it difficult to apply centralized model training directly due to privacy protection, security compliance, and regulatory constraints. Although federated learning (FL) can avoid sharing raw data, existing federated learning schemes based solely on homomorphic encryption (HE) still face security risks in drone scenarios, such as gradient inversion, member inference, and malicious update injection. To address this, we propose a secure and verifiable edge federated learning framework for parameter-efficient model adaptation in drone scenarios. The framework introduces homomorphic encryption for model updates on the device side to protect the privacy of updates before transmission and aggregation. Simultaneously, on the server side, decryption, aggregation, and verification are performed through a remotely authenticated Trusted Execution Environment (TEE), thereby limiting the server's access to plaintext updates and reducing the feasibility of gradient inversion and member inference attacks at the system level. Furthermore, an aggregation signature mechanism is introduced to batch verify the identity and update integrity of participating nodes, effectively preventing malicious or tampered updates from participating in aggregation, thus overcoming the shortcomings of existing HE-FL schemes in terms of poisoning resistance and verifiability. Experimental results show that, while ensuring safety and verifiability, the proposed method improves model accuracy by 3% compared to the comparative scheme, while maintaining better performance in terms of computation and communication overhead, thus verifying the practicality and deployability of the framework in resource-constrained UAV edge environments.","author":[{"family":"Su","given":"Huachang"},{"family":"Zhao","given":"Yekang"},{"family":"Zhang","given":"Wenrui"},{"family":"Zhang","given":"Hongling"},{"family":"Huang","given":"Shitao"},{"family":"Zhong","given":"Sheng"},{"family":"Zhou","given":"Xiaoyang"}],"issued":{"date-parts":[[2026]]},"DOI":"10.20944/preprints202601.1157.v1","URL":"https://doi.org/10.20944/preprints202601.1157.v1","source":"europepmc"},{"id":"doi:10.22541/au.176232803.35206392/v1","type":"article-journal","title":"Privacy-Preserving Clinical Analytics with Threshold Homomorphic Encryption: Insights from Hematologic Toxicity During Craniospinal Irradiation","abstract":"The increasing need for safeguarding medical data privacy, particularly in research involving sensitive patient information, requires innovative solutions. This work proposes a conceptual architecture that uses (threshold) homomorphic encryption for statistical analysis on encrypted medical data shared between different institutions. By utilizing this type of encryption, sensitive patient data is kept secure throughout the analysis process, minimizing the risk of re-identification. The data is encrypted locally before being processed on secure computation servers, ensuring privacy while enabling several statistical analysis. Performing computation on encrypted data is expensive and this has led to widespread skepticism regarding its practicality. Our work shows that, with recent advances and careful design and engineering the technology can indeed be harnessed to facilitate medical research. This method aligns with key data protection regulations and lays the groundwork for more privacy-preserving collaborative research in the medical field. This paper presents a conceptual model rather than a full platform implementation. We securely replicate, on encrypted patient-level data from pediatric craniospinal irradiation, three routine statistics—Pearson’s correlation, Wilcoxon rank-sum, and χ 2 —in a three-party threshold-HE workflow. Across tested sizes, χ 2 ≤ 0 . 5 % error, Pearson’s r ≈0–5.7% (≈5% at n =512), Wilcoxon’s z ≈13.5–21.9%, with millisecond-scale runtimes ( ≈ 5 0 – 8 5 0 m s ). We also outline a concise systems blueprint for cross-institution analytics (Fig. [1](#fig-cap-0001)).","author":[{"family":"Ţurcaş","given":"George"},{"family":"Gugulea","given":"George"},{"family":"Lupaşcu","given":"Cristian"},{"family":"Togan","given":"Mihai"},{"family":"Ţurcaş","given":"Andrada"}],"issued":{"date-parts":[[2025]]},"DOI":"10.22541/au.176232803.35206392/v1","URL":"https://doi.org/10.22541/au.176232803.35206392/v1","source":"europepmc"},{"id":"doi:10.20944/preprints202505.1442.v1","type":"manuscript","title":"Outsourced Privacy-Preserving Feature Selection Based on Fully Homomorphic Encryption","abstract":"Feature selection is a technique that extracts a meaningful subset from a set of features in training data. When the training data is large-scale, appropriate feature selection enables the removal of redundant features, which can improve generalization performance, accelerate the training process, and enhance the interpretability of the model. This study proposes a privacy-preserving computation model for feature selection. Generally, when the data owner and analyst are the same, there is no need to conceal the private information. However, when they are different parties or when multiple owners exist, an appropriate privacy-preserving framework is required. Although various private feature selection algorithms, they all require two or more computing parties and do not guarantee security in environments where no external party can be fully trusted. To address this issue, we propose the first outsourcing algorithm for feature selection using fully homomorphic encryption. Compared to a prior two-party algorithm, our result improves the time and space complexity O(kn2)) to O(knlog3n) and O(kn), where k and n denote the number of features and data samples, respectively. We also implemented the proposed algorithm and conducted comparative experiments with the naive one. The experimental result shows the efficiency of our method even with small datasets.","author":[{"family":"Wakiyama","given":"Koki"},{"family":"Sakamoto","given":"Hiroshi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.20944/preprints202505.1442.v1","URL":"https://doi.org/10.20944/preprints202505.1442.v1","source":"europepmc"},{"id":"doi:10.21203/rs.3.rs-6964665/v1","type":"article-journal","title":"Benchmarking Homomorphic Encryption on Low-Power Devices: Trade-offs Between PHE and FHE","abstract":"Abstract Homomorphic Encryption (HE) allows computation on ciphertexts, ensuring strong privacy for applications like smart metering, healthcare, and financial analytics. However, instantiating HE on power-constrained embedded devices is challenging as the computation and memory footprints are excessively high—particularly for Fully Homomorphic Encryption (FHE) schemes like BFV and CKKS. Partly Homomorphic Encryption (PHE) like Paillier is lightweight but less functional. This work evaluates the trade-offs among PHE and FHE with an investigation of Paillier, BFV, and CKKS on three representative platforms: ESP32, Raspberry Pi 4, and Arduino Uno. Performance measures like encryption and decryption time, ciphertext size, memory, and energy are compared. Experimentation demonstrates FHE to be impractical on 8-bit microcontrollers but efficient on 32- and 64-bit platforms. Of interest on Raspberry Pi 4, BFV and CKKS demonstrate sub-10 ms encryption times and consume below 5 J per 100 operations, both outpacing Paillier on speed and energy efficiency. Our work refutes the argument on the impracticability of FHE on embedded devices and provides practical advice on selecting among HE schemes according to platform capability. Our research fills the gap between theoretical cryptography and realistic deployment and promotes the use of HE as an enabling solution to trusted edge computing.","author":[{"family":"Khatusuriya","given":"Het"},{"family":"Patel","given":"Dhvani"},{"family":"Parmar","given":"Martin"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-6964665/v1","URL":"https://doi.org/10.21203/rs.3.rs-6964665/v1","source":"europepmc"},{"id":"doi:10.3390/s25123700","type":"article-journal","title":"Smart Grid IoT Framework for Predicting Energy Consumption Using Federated Learning Homomorphic Encryption.","abstract":"Homomorphic Encryption (HE) introduces new dimensions of security and privacy within federated learning (FL) and internet of things (IoT) frameworks that allow preservation of user privacy when handling data for FL occurring in Smart Grid (SG) technologies. In this paper, we propose a novel SG IoT framework to provide a solution for predicting energy consumption while preserving user privacy in a smart grid system. The proposed framework is based on the integration of FL, edge computing, and HE principles to provide a robust and secure framework to conduct machine learning workloads end-to-end. In the proposed framework, edge devices are connected to each other using P2P networking, and the data exchanged between peers is encrypted using Cheon–Kim–Kim–Song (CKKS) fully HE. The results obtained show that the system can predict energy consumption as well as preserve user privacy in SG scenarios. The findings provide an insight into the SG IoT framework that can help network researchers and engineers contribute further towards developing a next-generation SG IoT system.","author":[{"family":"Jerkovic","given":"Filip"},{"family":"Sarkar","given":"Nurul"},{"family":"Ali","given":"Jahan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/s25123700","URL":"https://doi.org/10.3390/s25123700","source":"europepmc"},{"id":"doi:10.1093/bioinformatics/btae754","type":"article-journal","title":"Privacy-preserving framework for genomic computations via multi-key homomorphic encryption.","abstract":"Abstract Motivation The affordability of genome sequencing and the widespread availability of genomic data have opened up new medical possibilities. Nevertheless, they also raise significant concerns regarding privacy due to the sensitive information they encompass. These privacy implications act as barriers to medical research and data availability. Researchers have proposed privacy-preserving techniques to address this, with cryptography-based methods showing the most promise. However, existing cryptography-based designs lack (i) interoperability, (ii) scalability, (iii) a high degree of privacy (i.e. compromise one to have the other), or (iv) multiparty analyses support (as most existing schemes process genomic information of each party individually). Overcoming these limitations is essential to unlocking the full potential of genomic data while ensuring privacy and data utility. Further research and development are needed to advance privacy-preserving techniques in genomics, focusing on achieving interoperability and scalability, preserving data utility, and enabling secure multiparty computation. Results This study aims to overcome the limitations of current cryptography-based techniques by employing a multi-key homomorphic encryption scheme. By utilizing this scheme, we have developed a comprehensive protocol capable of conducting diverse genomic analyses. Our protocol facilitates interoperability among individual genome processing and enables multiparty tests, analyses of genomic databases, and operations involving multiple databases. Consequently, our approach represents an innovative advancement in secure genomic data processing, offering enhanced protection and privacy measures. Availability and implementation All associated code and documentation are available at https://github.com/farahpoor/smkhe.","author":[{"family":"Namazi","given":"Mina"},{"family":"Farahpoor","given":"Mohammadali"},{"family":"Ayday","given":"Erman"},{"family":"Pérez-González","given":"Fernando"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1093/bioinformatics/btae754","URL":"https://doi.org/10.1093/bioinformatics/btae754","source":"europepmc"},{"id":"doi:10.20944/preprints202502.0927.v1","type":"manuscript","title":"Smart Grid IoT Framework Integrating Peer-to-Peer Federated Learning with Homomorphic Encryption","abstract":"Homomorphic Encryption (HE) introduces new dimensions of security and privacy within federated learning (FL) and Internet of Things (IoT) frameworks that allow preservation of user privacy when handling data for FL occurring Smart Grid (SG) technologies. In this paper, we propose a novel SG IoT framework to provide a solution of predicting energy consumption while preserving user-privacy in a smart grid system. The proposed framework is based on the integration of FL, edge computing, and HE principles to provide a robust and secure framework to conduct machine learning workloads end-to-end. In the proposed framework, edge devices are connected to each other using P2P networking and the data exchanged between peers is encrypted using CKKS fully HE. The results obtained show that the system can predict energy consumption as well as preserve user privacy in SG scenarios. The findings provide an insight into the SG IoT framework that can help network researchers and engineers to contribute further towards developing a next generation SG IoT system.","author":[{"family":"Jerkovic","given":"Filip"},{"family":"Sarkar","given":"Nurul"},{"family":"Ali","given":"Jahan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.20944/preprints202502.0927.v1","URL":"https://doi.org/10.20944/preprints202502.0927.v1","source":"europepmc"},{"id":"doi:10.20944/preprints202504.0302.v1","type":"manuscript","title":"<span style=\"color: black; mso-themecolor: text1;\">Lattice-Based Multi-Key Homomorphic Encryption Scheme Without CRS","abstract":"Multi-key homomorphic encryption is widely applied into outsourced computing and privacy-preserving applications in multi-user scenarios. However, the existence of CRS weakens the ability of users to independently generate public keys, and it is difficult to implement in decentralized systems or scenarios with low trust requirements. In order to reduce excessive reliance on public parameters, a multi-key homomorphic encryption scheme without pre-setting CRS is proposed based on a distributed key generation protocol. The proposed scheme does not require the pre-generation and distribution of CRS, which enhances the security and decentralization of the scheme. Furthermore, in order to further protect the plaintext privacy from each user, by embedding the specified target user into the ciphertext, this paper proposes an enhanced multi-key homomorphic encryption scheme that only allows only the target user to decrypt. Finally, this paper applies the proposed lattice-based multi-key homomorphic encryption scheme into the data submission stage of the perceived users, and thereby proposes a crowd-sensing scheme with privacy preservation.","author":[{"family":"Zhang","given":"Hongyi"},{"family":"Shang","given":"Mengxue"},{"family":"Liu","given":"Hanzhuo"},{"family":"Zhang","given":"Dandan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.20944/preprints202504.0302.v1","URL":"https://doi.org/10.20944/preprints202504.0302.v1","source":"europepmc"},{"id":"doi:10.3233/978-1-61499-532-6-58","type":"article-journal","title":"Statistical Analysis Methods Using Secure Multiparty Computation","abstract":"This chapter gives an overview of privacy-preserving versions of the analysis methods and algorithms that are most commonly used in statistical analysis. We discuss methods for data collection and sharing, and describe privacy-preserving database joins and sorting. From simple statistical measures, we discuss the count and sum of elements, quantiles, the five-number summary, frequency tables, mean, variance, covariance and standard deviation. We look into outlier detection and explore privacy-preserving versions of several different statistical tests, such as Student's t-test, Wilcoxon rank-sum and signed-rank tests and the &amp;chi;2-test. We discuss how to evaluate the significance of the test statistic in the privacy-preserving environment. We give several options for linear regression and conclude the chapter with a privacy-preserving method for data classification.","author":[{"family":"Liina","given":"Kamm"},{"family":"Dan","given":"Bogdanov"},{"family":"Alisa","given":"Pankova"},{"family":"Riivo","given":"Talviste"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3233/978-1-61499-532-6-58","URL":"https://doi.org/10.3233/978-1-61499-532-6-58","source":"crossref"},{"id":"doi:10.1136/bmjhci-2024-101384","type":"article-journal","title":"Enabling health data analyses across multiple private datasets with no information sharing using secure multiparty computation.","abstract":"The UK’s health datasets are among the most comprehensive and inclusive globally, enabling groundbreaking research during the COVID-19 pandemic. However, restrictions on data sharing between secure data environments (SDEs) imposed limitations on the ability to carry out joint analyses across multiple separate datasets. There are currently significant efforts underway to enable such analyses using methods such as federated analytics (FA) and virtual SDEs. FA involves distributed data analysis without sharing raw data but does require sharing summary statistics. Virtual SDEs in principle allow researchers to access data across multiple SDEs, but in practice, data transfers may be restricted by information governance concerns. Secure multiparty computation (SMPC) is a cryptographic approach that allows multiple parties to perform joint analyses over private datasets with zero information sharing. SMPC may eliminate the need for data-sharing agreements and statistical disclosure control, offering a compelling alternative to FA and virtual SDEs. SMPC comes with a higher computational burden than traditional pooled analysis. However, efficient implementations of SMPC can enable a wide range of practical, secure analyses to be carried out. This perspective reviews the strengths and limitations of FA, virtual SDEs and SMPC as approaches to joint analyses across SDEs. We argue that while efforts to implement FA and virtual SDEs are ongoing in the UK, SMPC remains underexplored. Given its unique advantages, we propose that SMPC deserves greater attention as a transformative solution for enabling secure, cross-SDE analyses of private health data.","author":[{"family":"Kerr","given":"Steven"},{"family":"Robertson","given":"Chris"},{"family":"Sudlow","given":"Cathie"},{"family":"Sheikh","given":"Aziz"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1136/bmjhci-2024-101384","URL":"https://doi.org/10.1136/bmjhci-2024-101384","source":"europepmc"},{"id":"doi:10.1007/s42979-025-04503-2","type":"article-journal","title":"Fast and Secure Multiparty Querying over Federated Graph Databases.","abstract":"Abstract We have developed a framework for efficient privacy preserving multi-party querying (PPMQ) over federated graph databases, leveraging Secure Multi-Party Computation (SMPC) protocols to enhance data security. The system offers two distinct security protocols: a client-based protocol and a server-based protocol. In the client-based protocol, standard SMPC techniques are employed, allowing computations to be performed on data without exposing the data itself. The server-based protocol employs SMPC to facilitate secure data processing and is further enhanced by encrypted hashing, which adds an additional layer of security to prevent data exposure. We conducted experiments comparing PPMQ with Neo4j Fabric and two previous systems, SMPQ and Conclave. The results indicate that PPMQ’s execution times and overheads are comparable to those of Neo4j Fabric, while outperforming both SMPQ and Conclave, demonstrating its superior efficiency. Additionally, PPMQ, like SMPQ and Conclave, utilises an honest but curious security model. However, it enhances the security of the server protocol, making it more robust against brute force attacks and providing stronger privacy guarantees than previous solutions.","author":[{"family":"Aljuaid","given":"Nouf"},{"family":"Lisitsa","given":"Alexei"},{"family":"Schewe","given":"Sven"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s42979-025-04503-2","URL":"https://doi.org/10.1007/s42979-025-04503-2","source":"europepmc"},{"id":"doi:10.2196/80178","type":"article-journal","title":"Integration of Federated Learning and Blockchain in Health Care: Tutorial on Medical Data, Architectures, Privacy, Security, and Regulatory Compliance.","abstract":"The convergence of artificial intelligence (AI), blockchain technology, and health care represents one of the most transformative yet technically challenging frontiers in computational medicine. As health care systems adopt data-driven paradigms for precision medicine and clinical decision support, the need for secure, privacy-preserving, and collaborative learning frameworks has become critical. This tutorial introduces a comprehensive, clinically oriented, and compliance-aware framework integrating federated learning (FL) and blockchain for secure and privacy-preserving health care analytics. FL enables collaborative training across distributed institutions without raw data sharing, in alignment with privacy regulations such as the Health Insurance Portability and Accountability Act (HIPAA) and the General Data Protection Regulation (GDPR). However, FL remains vulnerable to model poisoning and gradient leakage. To address these risks, we introduce blockchain-based FL (BCFL), which leverages blockchain's immutable ledger and decentralized consensus to enhance trust, verifiability, and auditability. The tutorial's main contributions include (1) a taxonomy of diverse medical data types and their FL requirements; (2) three integration architectures (fully coupled, semicoupled, and loosely coupled) analyzed for security, scalability, and regulatory compliance; (3) a security analysis of health care-specific vulnerabilities and mitigation strategies using advanced cryptography, such as zero-knowledge proofs, homomorphic encryption, and differential privacy; and (4) a regulatory compliance framework addressing HIPAA, GDPR, and United States Food and Drug Administration guidelines for AI-enabled medical devices. We demonstrate BCFL's relevance across major health care applications, including disease prediction, medical imaging, patient monitoring, and drug discovery, and highlight emerging research directions such as quantum-resilient cryptography, scalable interoperability, and automated compliance. This tutorial serves as a foundational resource for advancing secure, compliant, and collaborative AI in health care; fostering privacy-preserving analytics; and improving patient outcomes.","author":[{"family":"Oa","given":"Dambri"}],"issued":{"date-parts":[[2026]]},"DOI":"10.2196/80178","URL":"https://doi.org/10.2196/80178","source":"pubmed"},{"id":"doi:10.1007/s10791-025-09627-w","type":"article-journal","title":"Genomic privacy and security in the era of artificial intelligence and quantum computing.","abstract":"The rapid advancements in sequencing technologies have greatly increased access to genomic data stored in public databases. This has raised significant privacy and security concerns. This review emphasizes the importance of protecting genomic data by analyzing vulnerabilities in current storage and sharing practices. It examines the risks genetic databases face from cyber-attacks and internal breaches, focusing especially on advanced AI-driven threats and quantum computing vulnerabilities. The review explores machine learning methods designed to secure data. It highlights algorithms that prioritize privacy while maintaining data confidentiality, such as differential privacy, federated learning, and synthetic data generation using Generative Adversarial Networks (GANs). Findings demonstrate progress in mitigating common privacy breaches like re-identification and inference attacks. However, persistent vulnerabilities remain, particularly to emerging threats such as model inversion and membership inference attacks. The review advocates an integrated approach combining robust legislative frameworks with advanced technology to address genomic privacy challenges. It calls for intensified research efforts to safeguard genomic information. In particular, there is an urgent need to adopt quantum-resistant cryptographic methods, including lattice-based encryption and blockchain-integrated security frameworks. The paper emphasizes the necessity for genomics researchers to prioritize data privacy and security. This ensures responsible handling of genomic information in research.","author":[],"issued":{"date-parts":[[2025]]},"DOI":"10.1007/s10791-025-09627-w","URL":"https://doi.org/10.1007/s10791-025-09627-w","source":"pubmed"},{"id":"doi:10.1016/j.csbj.2025.06.009","type":"article-journal","title":"Revolutionizing healthcare data analytics with federated learning: A comprehensive survey of applications, systems, and future directions.","abstract":"Federated learning (FL)-a distributed machine learning that offers collaborative training of global models across multiple clients. FL has been considered for the design and development of many FL systems in various domains. Hence, we present a comprehensive survey and analysis of existing FL systems, drawing insights from more than 250 articles published in 2019-2024. Our review elucidates the functioning of FL systems, particularly in comparison with alternative distributed learning approaches. Considering the healthcare domain as an example, we define the building blocks of a typical FL healthcare system, including system architecture, federation scale, data partitioning, open-source frameworks, ML models, and aggregation algorithms. Furthermore, we identify and discuss key challenges associated with the design and implementation of FL systems within the healthcare sector while outlining the directions of future research. In general, through systematic categorization and analysis of existing FL systems, we offer insights to design efficient, accurate, and privacy-preserving healthcare applications using cutting-edge FL techniques.","author":[{"family":"Nt","given":"Madathil"},{"family":"Fk","given":"Dankar"},{"family":"An","given":"Belkacem"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1016/j.csbj.2025.06.009","URL":"https://doi.org/10.1016/j.csbj.2025.06.009","source":"pubmed"},{"id":"doi:10.1038/s41598-026-61276-1","type":"article-journal","title":"A lightweight anonymous authentication scheme for federated learning.","abstract":"Federated learning enables collaborative model training between central servers and distributed clients without collecting users' raw sensitive data, which effectively promotes the large-scale deployment of intelligent collaborative services. Considering the high sensitivity of local training data and model gradient parameters in federated learning, protecting identity privacy and interaction security has become extremely critical. Therefore, mutual identity authentication is indispensable to restrict illegal client access and prevent malicious parameter transmission and data tampering. In this paper, we propose a lightweight anonymous authentication scheme for federated learning (FedLAS), which realizes secure mutual authentication between servers and clients and establishes a shared session key for subsequent encrypted interaction. In particular, the proposed scheme eliminates the reliance on high-cost cryptographic operations such as bilinear pairing, thus minimizing computational and communication overhead. Furthermore, informal security analysis demonstrates that our FedLAS scheme can resist multiple common attacks and meet predefined security requirements. Extensive comparative experiments show that the FedLAS scheme achieves excellent performance in computational and communication cost. It is well suitable for resource-constrained federated learning scenarios.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-61276-1","URL":"https://doi.org/10.1038/s41598-026-61276-1","source":"pubmed"},{"id":"doi:10.3233/shti260281","type":"article-journal","title":"Traceability in Federated Learning in Healthcare.","abstract":"Federated learning (FL) enables privacy-preserving analytics on distributed healthcare data, but achieving transparency remains a critical challenge for trust and accountability. This scoping review focuses on traceability as a core component of transparency and systematically assesses traceability in FL within healthcare. Following the PRISMA-ScR methodology, we screened 125 articles across four databases, of which 53 met our inclusion criteria. Preliminary results show that 77.36% of articles rely on blockchain to support traceability, while only a small subset addresses healthcare applications, and comprehensive evaluation frameworks regarding traceability are largely lacking. Future research should explore the integration of blockchain in federated learning platforms in healthcare for traceability to enhance trust, auditability and accountability in clinical practice.","author":[{"family":"Fk","given":"Tang"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3233/shti260281","URL":"https://doi.org/10.3233/shti260281","source":"pubmed"},{"id":"doi:10.1038/s41598-026-46689-2","type":"article-journal","title":"Optimized IoT clustering and assignment in semi-synchronous federated learning.","abstract":"This study focuses on the important task of optimizing device clustering and assigning them to edge servers, while also implementing data redistribution in hierarchical semi-synchronous federated learning within the realm of advancing edge computing. Our research goal is to increase the performance and scalability of federated learning systems by improving resource allocation and data processing efficiency, which will in turn enhance edge computing frameworks. The current literature does not have thorough methods that can effectively combine model accuracy with optimal device clustering algorithms in hierarchical semi-synchronous federated learning, leading to below-par performance and inefficient use of resources. This difference highlights the need for creative measures that enhance not only model training accuracy but also the grouping of devices as opposed to current methods. The study utilizes a Graph Neural Network (GNN) to group IoT devices according to their hardware features and local datasets, then applies the K-means algorithm to create efficient device clusters. After that, Hybrid Data Redistribution is used to equalize local datasets in each cluster, and Proximal Policy resource allocation optimization algorithm is implemented to allocate devices to edge servers according to bandwidth usage, and energy consumption based on real-time updates, ultimately enabling hierarchical semi-synchronous federated learning to improve model training. The results show a 15% increase in clustering metrics compared to current algorithms, showcasing how our method improves device assignment and data redistribution in hierarchical semi-synchronous federated learning, addressing issues in model accuracy and resource optimization.","author":[{"family":"Aaph","given":"Kazem"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-46689-2","URL":"https://doi.org/10.1038/s41598-026-46689-2","source":"pubmed"},{"id":"doi:10.3390/s26144460","type":"article-journal","title":"Digital Twin-Enabled Dynamic Aggregation for Efficient Federated Learning.","abstract":"Federated learning (FL) enables collaborative model training without sharing raw data, but it faces challenges due to client heterogeneity, leading to inefficiency and reduced accuracy. This paper proposes a digital twin (DT)-based dynamic FL aggregation method to address these issues. The framework integrates a DT layer on the server side to perform preaggregation evaluations, simulating various aggregation strategies to select the optimal approach before actual global aggregation. An adaptive clustering method based on K-means is employed to group clients with similar characteristics, and a hierarchical aggregation evaluation strategy is designed to optimize both intra-cluster and inter-cluster aggregation, with the goal of minimizing latency and energy consumption while maximizing model accuracy. Simulation results on the MNIST and CIFAR-10 datasets demonstrate that the proposed method not only accelerates model convergence and improves accuracy but also significantly reduces training latency and energy consumption costs compared with baseline FL algorithms. This DT-assisted approach delivers a practical and effective optimization solution for federated learning deployment over large-scale heterogeneous IoT sensor networks.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26144460","URL":"https://doi.org/10.3390/s26144460","source":"pubmed"},{"id":"doi:10.1109/tpami.2026.3712213","type":"article-journal","title":"Factor-Assisted Federated Learning for Personalized Optimization with Heterogeneous Data.","abstract":"Federated learning is an emerging distributed machine learning framework aimed at protecting data privacy. Data heterogeneity is one of the core challenges in federated learning, which could severely degrade the convergence rate and prediction performance of deep neural networks. To address this issue, we develop a novel personalized federated learning framework for heterogeneous data, which we refer to as FedSplit. This modeling framework is motivated by the finding that data on different clients contain both common knowledge and personalized knowledge. Then the hidden elements in each neural layer can be split into shared and personalized groups. With this decomposition, a novel objective function is established and optimized. We demonstrate that FedSplit enjoys a faster convergence speed than the standard federated learning method both theoretically and empirically. The generalization bound of the FedSplit method is also studied. To practically implement the proposed method on real datasets, factor analysis is introduced to facilitate the decoupling of hidden elements. This leads to a practically implemented model for FedSplit, which we further refer to as FedFac. We demonstrate by simulation studies that using factor analysis can well recover the underlying shared/personalized decomposition. The superior prediction performance of FedFac is further verified empirically by comparison with various state-of-the-art federated learning methods on several real datasets.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tpami.2026.3712213","URL":"https://doi.org/10.1109/tpami.2026.3712213","source":"pubmed"},{"id":"doi:10.1038/s41598-026-58571-2","type":"article-journal","title":"Privacy-aware vaccine recommendation using federated learning and blockchain.","abstract":"Suitable vaccines for individuals are suggested by the vaccine recommendation system regarding certain criteria. Nevertheless, the existing studies didn't augment the vaccine recommendation system centered on users' symptoms and medical history among several geographical locations in the hyperledger fabric blockchain. Thus, in this paper, Krichevsky Dirichlet Trofimov-based latent Dirichlet allocation (KDT-LDA) and federated learning-Expcos bidirectional distillation long short-term memory (FL-EBiDLSTM)-based vaccine recommendation systems using symptoms and medical history are presented. Primarily, the vaccine symptoms dataset is taken. Then, the pre-processing is done based on named entity recognition, tokenization, and stemming. Later, by employing NSR-KMeans, the pre-processed data is grouped. Later, KDT-LDA-based symptoms and medical history modeling and adversarial debiasing-ClinicalBERT-based word embedding are carried out. Simultaneously, from the pre-processed data, the polarity score is identified. By utilizing the Cauchy Cubic-based fuzzy inference system, the labelling is performed regarding the polarity score. After that, the labelling outcomes are trained by the EBiDLSTM-based sentiment nature identification. Then, the natural language processing features are extracted from the symptoms and medical history modeling outcomes. After that, by using Spearman rank correlation, feature correlation is performed. Lastly, the appropriate vaccine is predicted based on FL-EBiDLSTM. Here, to solve the issue of training the patient data among various locations, FL is included. In real-time, vaccine demand users register with the hyperledger fabric blockchain and upload their medical history. Later, the vaccine recommendation system suggests the vaccines concerning the history. As per the outcomes, the proposed model achieved a high accuracy of 99% and outperformed prevailing techniques.","author":[{"family":"Ck","given":"Shinzeer"},{"family":"As","given":"Kushwaha"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-58571-2","URL":"https://doi.org/10.1038/s41598-026-58571-2","source":"pubmed"},{"id":"doi:10.1038/s41598-026-56332-9","type":"article-journal","title":"CASCADENCE: a layered cascade defense mechanism for federated learning.","abstract":"This paper aims to enhance the security and robustness of Federated Learning (FL) systems through a multi-layered defense. We address the critical challenge of protecting distributed learning environments from adversarial attacks while maintaining high model performance during both the training and operational phases. The proposed framework is based on integrated approaches that utilize a Gaussian filter with Discrete Fourier Transform (DFT), adversarial training with differential privacy, JPEG compression, randomized smoothing, and adversarial logit pairing. It integrates multiple defense mechanisms based on system requirements, focusing on preserving model performance while ensuring robust protection during both training and testing phases. Our approach extends beyond existing solutions by introducing various staged defense implementations and analyzing their synergistic effects. Experimental results demonstrate that the proposed ensemble defense mechanism achieves the highest performance, maintaining 98.21% accuracy and an F1 score of 0.98 under attack conditions, compared to a baseline accuracy of 90.87%.","author":[{"family":"Sw","given":"Hashmi"},{"family":"Rm","given":"Shukla"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-56332-9","URL":"https://doi.org/10.1038/s41598-026-56332-9","source":"pubmed"},{"id":"doi:10.1109/jbhi.2026.3715197","type":"article-journal","title":"Federated Learning with Global Model Hint for Medical Image Object Detection.","abstract":"Building an ideal medical image object detection model often requires sufficient training data, which can be challenging to obtain in practical scenarios. Manual annotation is labor-intensive, and sharing datasets may raise data privacy concerns. Although federated learning can partially address these issues, we find an amplified feature drift problem when it is directly applied to medical image object detection. Motivated by the observation that the global model's parameters tend to align more closely with those of the oracle model than with those of the client models, we propose FedMHDet: Model Hint Federated Learning Detection Model, a novel federated learning detection framework. During the training phase, FedMHDet leverages multi-scale feature consistency as a global model hint to guide client models, thus mitigating the feature drift problem. Extensive experiments on pulmonary lesion and brain tumor detection tasks show that FedMHDet achieves favorable overall performance. Compared to the strongest baseline under each corresponding metric, it improves average AP by 1.05 and 0.19, and average sensitivity by 1.10 and 0.43 on the two tasks, respectively. We also provide in-depth analyses to support the practical use of our method. The code is available at https://github.com/bbamai/FedMHDet.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/jbhi.2026.3715197","URL":"https://doi.org/10.1109/jbhi.2026.3715197","source":"pubmed"},{"id":"doi:10.3233/shti260452","type":"article-journal","title":"Recovering Sensitive Medical Text in Federated Learning.","abstract":"Federated Learning (FL) allows institutions to train shared models without exchanging raw data, making it a promising approach for healthcare applications that involve sensitive electronic health records (EHRs). However, despite this distributed design, the gradients exchanged during training can still reveal private information. In this study, we analyze how vulnerable transformer-based language models are to gradient inversion attacks, focusing on the Decepticons method, which can reconstruct original training text from shared gradients. We simulate a cross-silo FL setup with three types of French clinical reports (genetic, anesthesia, and birth records) to evaluate how batch size and sequence length affect reconstruction quality. Our experiments show that a malicious server can recover clinical text with high accuracy: token-level recovery exceeded 95% when training with batch size 1 and remained above 60% for sequences of up to 512 tokens. Reconstructed examples contained identifying elements (names, dates, genetic markers), revealing serious privacy risks for real-world use. These results emphasize that FL alone is insufficient for sensitive clinical text and that privacy-preserving defenses must be integrated before real-world deployment.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3233/shti260452","URL":"https://doi.org/10.3233/shti260452","source":"pubmed"},{"id":"doi:10.1007/s10916-026-02436-8","type":"article-journal","title":"Security Analysis of a Federated Learning Framework for Medical Image-to-Image Translation.","abstract":"Federated Learning (FL) emerged as a privacy-preserving paradigm for collaborative training of deep learning models across institutions without sharing patient data. This approach has been applied to complex tasks such as medical image-to-image (I2I) translation, including MRI-to-synthetic CT (sCT) generation. However, existing federated I2I frameworks often assume privacy preservation as an inherent property of FL rather than a requirement to be explicitly validated, leaving their robustness to representative adversarial threat scenarios largely unexplored. In this study, we evaluated the vulnerability of a federated MRI-to-sCT translation framework (FedSynthCT-Brain) to three representative attack classes: Deep Leakage from Gradients (DLG), Federated Membership Inference Attack (FedMIA), and data poisoning. The efficacy of corresponding defense mechanisms, such as Secure Aggregation (SecAgg) and Byzantine-robust median aggregation (FedMedian), were assessed. DLG enabled only the recovery of coarse anatomical structures, with no clinically identifiable details (SSIM &#x2264; 0.16, PSNR &#x2264; 11&#xa0;dB) across clients, suggesting limited vulnerability under the evaluated DLG setting. In contrast, FedMIA achieved high membership discrimination, with AUC scores between 0.92 and 0.99, revealing a critical privacy vulnerability. The introduction of SecAgg reduced AUC values to near-random levels (0.23-0.56) across all centers without impacting synthesis quality. Under high-noise poisoning, the standard federated averaging (FedAvg) aggregation rendered the federation inoperative, while FedMedian restored performance close to the no-poisoning baseline in most scenarios, with significant residual degradation in specific center configurations. At low noise levels, the advantage of FedMedian was less consistent, as low-level noise injection may be indistinguishable from natural heterogeneity across centers, potentially enabling stealthy degradation. These findings demonstrate that federated I2I translation frameworks are not inherently secure and require explicit, multi-layered evaluation. As FL is increasingly adopted in clinical workflows, our results underscore the necessity of integrating cryptographic, algorithmic, and infrastructural safeguards for secure deployment.","author":[{"family":"Cb","given":"Raggio"},{"family":"Mf","given":"Spadea"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1007/s10916-026-02436-8","URL":"https://doi.org/10.1007/s10916-026-02436-8","source":"pubmed"},{"id":"doi:10.1038/s41598-026-60725-1","type":"article-journal","title":"Differentially private federated learning for localized control of infectious disease dynamics.","abstract":"In times of epidemics, swift reaction is necessary to mitigate epidemic spreading. For this reaction, localized approaches have several advantages, limiting necessary resources and reducing the impact of interventions on a larger scale. However, training a separate machine learning (ML) model on a local scale is often not feasible due to limited available data. Centralizing the data is also challenging because of its high sensitivity and privacy constraints. In this study, we consider a localized strategy based on the German counties and communities managed by the related local health authorities (LHA). For the preservation of privacy to not oppose the availability of detailed situational data, we propose a privacy-preserving forecasting method that can assist public health experts and decision makers. ML methods with federated learning (FL) train a shared model without centralizing raw data. Considering the counties, communities or LHAs as clients and finding a balance between utility and privacy, we study a FL framework with client-level differential privacy (DP). We train a shared multilayer perceptron on sliding windows of recent case counts to forecast the number of cases in the future, while clients exchange only norm-clipped updates and the server aggregates updates with DP noise. We evaluate the approach on COVID-19 data on county-level during two phases: November 2020 and March 2022 (Omicron). As expected, very strict privacy ([Formula: see text]) yields unstable, unusable forecasts. At a moderately strong but still privacy-preserving level ([Formula: see text]), the DP model closely approaches the non-DP model: [Formula: see text] (vs. 0.96) and mean absolute percentage error (MAPE) [Formula: see text] in November 2020; [Formula: see text] (vs. 0.90) and MAPE [Formula: see text] in March 2022. Overall, our results support the feasibility of privacy-preserving collaboration among health authorities for local forecasting. In the evaluated COVID-19 phases, client-level DP-FL delivered useful county-level predictions with formal privacy guarantees under the stated threat model. The appropriate privacy budget should nevertheless be re-evaluated for other epidemic phases and applications.","author":[{"family":"Mj","given":"Kühn"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-60725-1","URL":"https://doi.org/10.1038/s41598-026-60725-1","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-9499405/v1","type":"article-journal","title":"Defending Patient Privacy in Federated Learning with Shadow Model","abstract":"Abstract Federated Learning (FL) has become a approach for training models together in privacy-sensitive domains like healthcare, where sharing of raw data is frequently restricted. But recent research has shown that Gradient Inversion Attacks (GIAs) can use shared model updates to recreate sensitive training data, which puts patient privacy at risk. Current defense mechanisms, like differential privacy and gradient perturbation techniques, often use uniform protection strategies that either make the model work less well or don't protect important data areas well enough.This study presents an enhanced privacy-preserving framework based on an augmented Shadow Defense (Shadow Def) mechanism to address these limitations.The suggested method includes a multi-phase, region-aware defense strategy that uses Fast Fourier Transform for frequency-domain sensitivity analysis, Mean Squared Error (MSE) for vulnerability mapping, and multi-component noise injection. To make the system more resistant to reconstruction attacks while keeping training stable, gradient direction perturbation, temporal smoothing, and adaptive noise scaling are also used. The suggested method has undergone comprehensive testing in a federated learning context, utilizing simulated gradient inversion attacks on two prominent medical imaging datasets: Chest X-Ray and Eye PACS. The adversarial training process significantly reduces reconstruction quality (i.e., reconstruction errors), as shown by a higher RMS error (MSE), a lower structural similarity index (SSIM), and a lower peak signal-to-noise ratio (PSNR). Conversely, the model's classification accuracy remains elevated, exhibiting only a slight decline in performance. These results show that there is a good balance between protecting patient privacy and making the model available in a healthcare setting.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9499405/v1","URL":"https://doi.org/10.21203/rs.3.rs-9499405/v1","source":"europepmc"},{"id":"doi:10.1038/s41746-026-02958-y","type":"article-journal","title":"Nationwide federated learning for histopathology: secure deployment across Germany behind firewalls.","abstract":"Federated Learning (FL) enables collaborative training across institutions without sharing sensitive data, a solution for privacy-preserving AI in medical imaging. However, hospital deployment remains challenging due to strict data protection regulations, heterogeneous infrastructures, and limited network accessibility behind firewalls. We introduce TheODen, an open-source framework for Federated training on histopathology Whole Slide Imaging (WSI). It requires no open client-side ports, enabling training through firewalls via a secure reverse-proxy architecture. We conducted, to our knowledge, the first nationwide FL study for histopathology segmentation of colorectal cancer across three German university hospitals, using breast and colorectal cancer datasets without opening firewall ports. TheODen achieves robust segmentation, with global average dice scores of 0.764 on BCSS and 0.754 on SemiCOL despite data heterogeneity and network constraints. These findings underline TheODen's potential to facilitate secure, large-scale collaborations between medical institutions and to accelerate clinical translation of AI models under real-world infrastructure constraints, providing a privacy-preserving-by-design architecture for future collaborations.","author":[{"family":"Ml","given":"Eich"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41746-026-02958-y","URL":"https://doi.org/10.1038/s41746-026-02958-y","source":"pubmed"},{"id":"doi:10.1038/s41598-026-58414-0","type":"article-journal","title":"HCFL: hybrid contribution-driven federated learning for fair and efficient optimization.","abstract":"Federated Learning (FL) enables collaborative model training across decentralized data silos while preserving data privacy. However, client selection strategies in conventional FL processes typically rely on single-dimensional evaluation metrics, which fail to capture data diversity and overlook the dynamic nature of client contributions, particularly in domains characterized by sparse and heterogeneous data, such as healthcare and drug discovery. These limitations ultimately hinder the global model's generalization ability and reduce training efficiency. To address these challenges, this paper proposes an adaptive FL framework that employs a hybrid contribution evaluation mechanism as the core principle for client selection and resource management. The proposed approach quantifies each client's effectiveness by integrating two complementary dimensions: (i) a performance-based evaluation that measures the immediate impact of a client's update on the global optimization trajectory, and (ii) a coverage-based evaluation that estimates data diversity in the latent embedding space without exposing raw data. By combining these two criteria, the hybrid mechanism ensures that highly contributive clients are preferentially selected while preventing the permanent exclusion of any participant, thereby maintaining a balanced trade-off between efficiency and fairness. Experimental results demonstrate that the proposed framework outperforms existing FL baselines in terms of training efficiency, data utilization, and fairness.","author":[{"family":"Wg","given":"Choi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-58414-0","URL":"https://doi.org/10.1038/s41598-026-58414-0","source":"pubmed"},{"id":"doi:10.1109/isbi61048.2026.11515381","type":"article-journal","title":"HYPERBOLIC MODEL AGGREGATION FOR FEDERATED LEARNING IN FMRI.","abstract":"The privacy-sensitive nature of clinical data often limits the use of machine learning in medical imaging applications, particularly for modalities with high acquisition costs such as functional MRI (fMRI). Federated learning mitigates data-sharing barriers by training site-specific models locally and aggregating weights centrally into a server model. However, small and heterogeneous per-site samples in medical imaging heighten the need for robust model-aggregation strategies. In this work, we introduce a federated aggregation scheme based on hyperbolic geometry to provide a robust and flexible approach to federated model weight integration. The proposed scheme is plug-and-play for standard federated learning loops. Empirically, our method improves stability and accuracy across multi-site fMRI data from ABIDE I, yielding more consistent convergence versus methods based on Euclidean mean and median. Codes are publicly available at https://github.com/Jiyao96/FedHAvg.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/isbi61048.2026.11515381","URL":"https://doi.org/10.1109/isbi61048.2026.11515381","source":"pubmed"},{"id":"doi:10.1038/s41598-026-57744-3","type":"article-journal","title":"Federated learning and digital twins for lifecycle optimization in Urban building renewal.","abstract":"The restructuring of aging Chinese city infrastructure requires new approaches based on computational intelligence and optimization lifecycle structures. Current building renovation methods are limited by the lack of seamless linkage between real-time operation data and predictive lifecycle management, particularly regarding privacy and multi-building management in heterogeneous environments. This paper introduces a simulation framework that combines federated deep reinforcement learning and behavioural digital twins to restore Chinese buildings into smart cities. The architecture involves a three-layer system: a physical layer with IoT-enabled sensing networks in distributed building clusters, a digital twin layer with real-time BIM-to-operational model synchronization and LiDAR-enhanced geometric precision, and an intelligent layer with privacy-reflecting federated proximal policy optimization for distributed decision-making. The framework addresses the gap between fixed 3D representations and variable behavioural modelling by integrating continuous learning processes that respond to changing occupancy, energy consumption, and structural decay. Simulation studies of Chinese urban residential communities show better performance: 27.3% lower lifecycle operational costs, 34.6% improved energy efficiency with the same thermal comfort, and 39.7% better structural integrity prediction with CFRP-optimized improvements. The federated learning architecture results in 5.8% cost savings and 6.2% emission reduction, offering scalable, privacy-sensitive urban renewal decision support for China's modernization efforts.","author":[{"family":"Asm","given":"Metwally"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-57744-3","URL":"https://doi.org/10.1038/s41598-026-57744-3","source":"pubmed"},{"id":"doi:10.1109/tpami.2026.3697332","type":"article-journal","title":"Privacy-preserving Online Federated Learning for Massive Infinite Streams.","abstract":"Online federated learning (OFL) is essential for privacy-preserving collaborative online analytics over decentralized streams. Different from batch-based FL, OFL faces new challenges including longitudinal privacy leakage, and accumulated utility loss and communication costs, caused by the infinite data streams. This paper first extends the definition of traditional differential privacy (DP) to OFL, to provide window-based privacy protection with a tunable granularity for infinite streams. By analyzing baseline methods, a generic sampling-based solution framework is then proposed for designing a DP-enhanced OFL algorithm. We prove that despite the DP constraint, the sampling solution framework can achieve an asymptotic optimality when time tends to infinity. Finally, we present Sampling$^{3}$-OFL, an adaptive triple-sampling strategy driven by deep reinforcement learning, which can dynamically determine a near-optimal sampling strategy with significant gains in both utility and efficiency. Extensive experiments on six real-world datasets demonstrate that Sampling$^{3}$-OFL can scale to millions of streams, and achieves utility improvements of 0.74%-15.84% and communication cost reductions of 33.33%-95.24% across these datasets compared to state-of-the-art methods.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tpami.2026.3697332","URL":"https://doi.org/10.1109/tpami.2026.3697332","source":"pubmed"},{"id":"doi:10.1038/s41598-026-56310-1","type":"article-journal","title":"Resource management for blockchain enhanced federated learning in wireless edge networks.","abstract":"Though machine learning is widely used in wireless edge networks, the transmission of raw data still suffers from security and privacy leakage. Federated learning (FL) addresses these privacy concerns by enabling model training without sharing raw data. However, traditional centralized FL is vulnerable to a single point of failure. Blockchain-based federated learning (BFL) technology can provide FL with a more reliable and secure environment. In wireless edge networks with limited resources, BFL systems encounter challenges related to computing demands and network transmission overhead. To address these issues, we propose a BFL framework for wireless edge networks, which includes local client training, a consensus process, and edge server aggregation. A client selection policy is designed to exclude low-quality clients that could degrade training efficiency and accuracy. Additionally, a joint client selection and resource allocation scheme is implemented to optimize the allocation of computing and bandwidth resources necessary for BFL training and consensus. Simulation results demonstrate that the proposed approach improves BFL system accuracy while reducing delay.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-56310-1","URL":"https://doi.org/10.1038/s41598-026-56310-1","source":"pubmed"},{"id":"doi:10.1038/s41598-026-63242-3","type":"article-journal","title":"Privacy-preserving clustered federated learning via differential privacy and homomorphically encrypted prototypes.","abstract":"Clustered federated learning (CFL) is an effective paradigm for handling statistical heterogeneity by grouping clients with similar data characteristics and learning cluster-specific models. However, existing CFL methods often expose sensitive clustering signals or cluster-specific updates to the server, which may reveal latent client similarity relations and weaken privacy protection. To address this issue, we propose Privacy-Preserving Clustered Federated Learning (PPCFL), a split-stream framework that integrates adaptive Gaussian perturbation with threshold Paillier encrypted aggregation. In PPCFL, backbone updates are protected by adaptive Gaussian perturbation before plaintext aggregation, while clustering signatures and cluster-head updates are first perturbed by stream-specific adaptive Gaussian mechanisms and then uploaded under threshold Paillier encryption. The server performs ciphertext-domain aggregation for clustering prototypes and cluster-head updates, whereas plaintext prototypes and cluster-level decrypted aggregates are recovered by a qualified threshold-decryption client subset without giving the server decryption capability. In addition, PPCFL adopts round-wise budget growth, utility-aware refinement, and adaptive clipping-threshold updates to improve the privacy-utility trade-off under dynamic Non-IID settings. Experiments on MNIST, Fashion-MNIST, and CIFAR-10 show that PPCFL achieves the highest final-round accuracy among the evaluated methods in the reported settings while providing enhanced protection for clustering-related information and cluster-specific updates. Under the representative Dirichlet setting [Formula: see text], PPCFL improves the final accuracy over DP-FedAvg by 0.33, 1.73, and 2.62 percentage points on MNIST, Fashion-MNIST, and CIFAR-10, respectively, and over IFCA by 0.98, 8.28, and 10.24 percentage points.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-63242-3","URL":"https://doi.org/10.1038/s41598-026-63242-3","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-9888411/v1","type":"article-journal","title":"On the Interaction Between Personalization and Optimization in Federated Learning for Medical Image Classification","abstract":"Abstract Federated learning (FL) has emerged as a promising paradigm for privacy-preserving medical image analysis, enabling collaborative model training across distributed institutions without sharing sensitive patient data. However, two key challenges remain: instability of model optimization under non-independent and identically distributed(non-IID)data, and the need for client-specific adaptation. In this paper, we propose a hybrid federated learning framework that integrates proximal regularization (FedProx) with personalized model adaptation(FedPer) to address these challenges jointly. Unlike prior work that treats these mechanisms independently, we investigate their interaction and demonstrate how their combination improves both convergence stability and generalization. We evaluate the proposed framework on the MURA dataset under simulated heterogeneous client distributions. Experimental results show that the hybrid approach consistently outperforms standard federated baselines, achieving up to 96. These findings highlight the importance of jointly considering optimization stability and personalization in federated learning systems, particularly for real-world medical imaging applications.","author":[{"family":"Ansar","given":"Rimsha"},{"family":"Salma","given":"Zainab"},{"family":"Neira","given":"Raquel"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9888411/v1","URL":"https://doi.org/10.21203/rs.3.rs-9888411/v1","source":"europepmc"},{"id":"doi:10.1109/tpami.2026.3699626","type":"article-journal","title":"Federated Learning via Variational Bayesian Inference: Personalization, Sparsity and Clustering.","abstract":"Federated learning (FL) is a promising framework that models distributed machine learning while protecting the privacy of clients. However, FL suffers performance degradation from heterogeneous and limited data. To alleviate the degradation, we present a novel personalized Bayesian FL approach named pFedBayes. By using the trained global distribution from the server as the prior distribution of each client, each client adjusts its own distribution by minimizing the sum of the reconstruction error over its personalized data and the KL divergence with the downloaded global distribution. Then, we propose a sparse personalized Bayesian FL approach named sFedBayes to enhance the inference efficiency. To overcome the extreme heterogeneity in non-i.i.d. data, we propose a clustered Bayesian FL model named cFedbayes by learning different prior distributions for different clients. Theoretical analysis gives the generalization error bound of three approaches and shows that the generalization error rates of the proposed approaches achieve minimax optimality up to a logarithmic factor. Moreover, cFedBayes achieves a cluster-level generalization error bound, rather than a single uniform bound in pFedBayes. Numerous experiments demonstrate that the proposed approaches have better performance than other advanced personalized methods on private models in the presence of heterogeneous and limited data.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tpami.2026.3699626","URL":"https://doi.org/10.1109/tpami.2026.3699626","source":"pubmed"},{"id":"doi:10.1038/s41598-026-58988-9","type":"article-journal","title":"Mitigating request flooding attack in named data networking using federated learning.","abstract":"Named Data Networking (NDN) represents a paradigm shift toward content-centric architectures but remains critically vulnerable to Interest Flooding Attacks (IFAs), where malicious actors overwhelm router Pending Interest Tables with spurious requests, causing service degradation and denial-of-service. To address the limitations of existing approaches, including high false positives in threshold-based methods and substantial overhead in centralized learning, we propose FL-IFAshield, a novel federated learning framework for adaptive IFA mitigation. Our solution integrates dynamic Poisson-EMA thresholding for accurate flood detection, entropy-aware federated aggregation to handle non-IID traffic distributions across edge routers, and Byzantine-robust mechanisms with differential privacy guarantees. Comprehensive evaluation on the FIT/IoT-LAB testbed with 100 routers demonstrates exceptional performance: 93.1% F1-score in attack detection, only 5% false positives, 28 ms average end-to-end latency ([Formula: see text]), and over 90% legitimate Interest Satisfaction Ratio under sophisticated collusive attacks, while maintaining minimal computational overhead (&lt;9% CPU utilization on ARMv8 routers). FL-IFAshield significantly improves security performance, offering 35% higher accuracy than static thresholding and 60% lower communication overhead than centralized approaches. While simpler heuristic baselines naturally incur marginally lower computational footprints, our solution delivers the optimal overall operational balance among high precision, low end-to-end latency ([Formula: see text]), and resource efficiency in constrained edge computing environments.","author":[{"family":"Ml","given":"Benmaidi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-58988-9","URL":"https://doi.org/10.1038/s41598-026-58988-9","source":"pubmed"},{"id":"doi:10.3233/shti260241","type":"article-journal","title":"A Durable Backdoor Attack on Medical Imaging via Federated Learning.","abstract":"Federated Learning (FL) enables multiple healthcare institutions to jointly train models without sharing raw patient data, making it a natural fit for privacy-sensitive medical applications. However, its distributed and partially trusted nature exposes it to backdoor attacks, in which malicious clients inject hidden behaviors that activate at inference. In this paper, we propose a durable backdoor attack designed for FL. The attack leverages a Generative Adversarial Network guided by the global model to generate synthetic data that closely matches the distribution of benign clients. Furthermore, we introduce a two-step strategy that enhances the robustness of the injected backdoor. We evaluate our method on the MedMNIST benchmark under a non-IID data distribution to simulate realistic medical scenarios. Experimental results demonstrate that our approach achieves a durable backdoor effect, persisting even under limited attacker participation. Therefore, there is a need for more resilient defense mechanisms to ensure the trustworthiness of Federated Learning in medical applications.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3233/shti260241","URL":"https://doi.org/10.3233/shti260241","source":"pubmed"},{"id":"doi:10.3390/s26134322","type":"article-journal","title":"FedHSFV: Federated Learning for Finger Vein Recognition via Hierarchical Decoupling and Subspace Metric.","abstract":"Finger vein recognition (FVR) has significant potential in biometrics due to its high accuracy and intrinsic liveness detection capabilities. However, the increasingly stringent privacy regulations have presented severe data security challenges for traditional centralized training. While federated learning (FL) mitigates these privacy concerns through a decentralized training paradigm, conventional FL algorithms that seek a single global model experience significant performance degradation on non-independent and identically distributed (Non-IID) data in real-world cross-institutional deployments. This degradation stems primarily from a dual-heterogeneity issue that involves domain shift caused by hardware discrepancies across acquisition devices, and label skew resulting from nonoverlapping user identities. To address this dual-heterogeneity challenge, we propose a personalized federated learning framework driven by hierarchical parameter decoupling and subspace metric. First, we designed a hierarchical parameter decoupling architecture. Macroscopically, the architecture retains the classifier locally to isolate label heterogeneity; microscopically, it introduces an additive parameter decomposition that decouples the feature extractor on a global full-rank basis (to capture domain-invariant semantics, namely, the shared physiological vein topologies) and a local low-rank adapter (that accommodates device-specific characteristics, such as hardware-induced noise and illumination discrepancies). Furthermore, we propose a subspace similarity matching strategy based on principal angles on the Grassmann manifold. By exploiting the geometric properties of low-rank projection matrices, this strategy accurately quantifies the underlying distribution discrepancies among clients to guide personalized weighted aggregation. Extensive experiments on six public finger vein datasets demonstrate that the proposed framework significantly improves the overall recognition performance and mitigates performance degradation caused by data heterogeneity.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26134322","URL":"https://doi.org/10.3390/s26134322","source":"pubmed"},{"id":"doi:10.1038/s41598-026-56609-z","type":"article-journal","title":"Privacy-preserving federated learning for interpretable student at-risk prediction across schools.","abstract":"Early-warning systems for at-risk students increasingly rely on predictive models trained on sensitive educational records. However, centralized learning pipelines raise concerns about privacy, institutional data sovereignty, and auditability, particularly when student-level data are shared across institutions. This study presents Federated Learning for At-Risk Student Prediction with Differential Privacy and Proof-Before-Train (FL-AtRisk-DP-PBT), a federated learning framework for multi-school at-risk prediction that integrates Federated Averaging (FedAvg)-based training, client-side DP, and a PBT protocol for verifiable logging of client participation and model states. The framework uses a single interpretable global logistic-regression classifier and is evaluated under centralized, standard federated, and FL&#x2009;+&#x2009;DP+PBT regimes on three educational datasets: a primary merged cohort of 14,003 students partitioned into 10 simulated schools, the xAPI-Edu-Data click-stream corpus, and the Students Performance in Exams dataset. On the primary dataset, the centralized model achieves 99.14% accuracy, F1&#x2009;=&#x2009;0.9915, and area under the curve (AUC)&#x2009;=&#x2009;0.9998, while the FedAvg and FL&#x2009;+&#x2009;DP+PBT variants achieve 98.61%/0.9863/0.9993 and 98.00%/0.9802/0.9992, respectively. On xAPI and Exams, FL&#x2009;+&#x2009;DP+PBT reaches approximately 93-94% accuracy, F1&#x2009;&#x2248;&#x2009;0.92-0.93, and AUC&#x2009;&#x2248;&#x2009;0.97-0.98. Coefficient-based feature-importance analysis indicates that FL&#x2009;+&#x2009;DP+PBT preserves broadly similar interpretation patterns to the centralized and non-private federated baselines. The PBT ablation introduces only small metric changes relative to DP-only federated training. Overall, the results suggest that interpretable federated at-risk prediction can retain competitive utility while keeping student records local and adding privacy-preserving and verifiable training mechanisms. These findings should be interpreted within the evaluated datasets, simulated school partitions, and label definitions.","author":[{"family":"Ak","given":"Ghafi"},{"family":"Mh","given":"Shafiabadi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-56609-z","URL":"https://doi.org/10.1038/s41598-026-56609-z","source":"pubmed"},{"id":"doi:10.1038/s41598-026-49544-6","type":"article-journal","title":"BLEND blockchain and federated learning enabled data sharing network.","abstract":"Combining blockchain and federated learning has emerged as a promising solution for secure, privacy-preserving data sharing and collaborative training of machine learning models in decentralised settings. However, their methods suffer from scalability issues, computational overhead, and challenges in preserving privacy in dynamic environments such as IoT, healthcare, and smart cities. Although blockchain-based solutions ensure data confidence and provenance, they come at the cost of the high computational overhead involved in transaction validation and consensus. In a similar spirit to federated learning, where local models are trained on sensitive data, and only aggregation is performed without centralising the data, traditional approaches fail to provide efficient aggregation and do not allow for secure transmission. This paper proposes BLEND, a Blockchain- and federated-learning-enabled network for Data sharing, as a new framework to tackle this issue. It will integrate a new consensus protocol, adaptive encryption schemes, and smart contract-based aggregation to deliver outstanding security, scalability, and operational efficiency. The proposed framework automatically adjusts its encryption strength in real time based on threat intelligence, enabling secure data while enhancing model performance. We show through experiments that BLEND outperforms existing blockchain-based federated learning methods in terms of latency, computation cost, and model accuracy by several orders of magnitude. Exploiting this commonality, the frame can reduce the existing framework's computational overhead by as much as 20%, while maintaining around 90-95% of the original model's accuracy and achieving lower latency than traditional methods. The framework enables privacy-friendly practical scenarios that enable large-scale, distributed data sharing and model training. BLEND offers industries such as healthcare and IoT a performant, robust, and scalable instruction-level collaborative solution that does not compromise performance or privacy.","author":[{"family":"Mi","given":"Reddy"},{"family":"Kr","given":"Pradeep"},{"family":"Gb","given":"Madhavi"},{"family":"Ys","given":"Reddy"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-49544-6","URL":"https://doi.org/10.1038/s41598-026-49544-6","source":"pubmed"},{"id":"doi:10.1109/tpami.2026.3688672","type":"article-journal","title":"Boosting the Performance of Decentralized Federated Learning via Catalyst Acceleration.","abstract":"Decentralized Federated Learning has emerged as an alternative to centralized architectures due to its faster training, privacy preservation, and reduced communication overhead. In decentralized communication, the server aggregation phase in Centralized Federated Learning shifts to the client side, which means that clients connect with each other in a peer-to-peer manner. However, compared to the centralized mode, data heterogeneity in Decentralized Federated Learning will cause larger variances between aggregated models, which leads to slow convergence in training and poor generalization performance in tests. To address these issues, we introduce Catalyst Acceleration and propose an acceleration Decentralized Federated Learning algorithm called DFedCata. It consists of two main components: the Moreau envelope function, which primarily addresses parameter inconsistencies among clients caused by data heterogeneity, and Nesterov's extrapolation step, which accelerates the aggregation phase. Theoretically, we prove the optimization error bound and generalization error bound of the algorithm, providing a further understanding of the nature of the algorithm and the theoretical perspectives on the hyperparameter choice. Empirically, we demonstrate the advantages of the proposed algorithm in both convergence speed, computational cost, and generalization performance on CIFAR10/100 and Tiny-ImageNet with various non-iid data distributions. Moreover, extensive experiments are conducted to validate the theoretical properties of DFedCata, showing strong consistency between theory and empirical observations.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tpami.2026.3688672","URL":"https://doi.org/10.1109/tpami.2026.3688672","source":"pubmed"},{"id":"doi:10.3390/s26113325","type":"article-journal","title":"Stability-Controlled Continual Federated Learning for Energy-Harvesting AIoT Systems.","abstract":"Energy-harvesting (EH) AIoT systems enable long-term autonomous operation but suffer from time-varying energy availability, which makes stable learning difficult. In such environments, federated learning (FL) is prone to energy depletion (blackout), while continual learning is required to handle evolving data distributions, leading to a trade-off between energy stability and catastrophic forgetting. In this paper, we propose a stability-controlled continual federated learning framework that jointly regulates local training intensity and rehearsal usage based on the residual energy state. The proposed method is derived from a Lyapunov drift-plus-penalty formulation and implemented as a lightweight mode-based control policy. Simulation results using real solar energy traces show that the proposed method significantly reduces blackout while improving accuracy and mitigating forgetting compared to existing approaches. These results demonstrate the effectiveness of energy-aware joint control for stable continual federated learning in EH-AIoT systems.","author":[{"family":"Dk","given":"Noh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26113325","URL":"https://doi.org/10.3390/s26113325","source":"pubmed"},{"id":"doi:10.1016/j.neunet.2026.109193","type":"article-journal","title":"FedEBM: Robust graph federated learning via energy-based model.","abstract":"Graph Federated Learning (GFL), as a vital component of graph neural networks, has found extensive applications in real-world scenarios. However, real-world graph data often suffers from label noise due to factors such as mislabeled data or malicious attacks. Existing methods for handling noisy labels primarily focus on centralized approaches, which perform poorly when directly applied to distributed settings and struggle to operate effectively on large-scale datasets. To address noisy labels in GFL, we propose a novel method called FedEBM. First, recognizing that clients in GFL often exhibit suboptimal learning capabilities under adverse conditions such as sample sparsity or label imbalance, we innovatively apply an Energy-Based Model (EBM) to tackle the noisy label problem. The EBM discriminates between clean and noisy samples based on their energy score. Even when clean samples are scarce, it implicitly delineates the energy region boundary by elevating the energy score of noisy samples, thereby separating clean and noisy samples. Furthermore, the energy score across categories in the EBM does not necessitate changes in other categories' energy score, avoiding probability competition on minority classes and enhancing sensitivity to minority class features. Comparative experiments across multiple public datasets demonstrate that FedEBM outperforms six baseline methods under various noise rates, noise types, and client numbers. Specifically, FedEBM outperforms the second-best method by an average margin of 5.43% on small-scale datasets and 12.58% on large-scale datasets.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.neunet.2026.109193","URL":"https://doi.org/10.1016/j.neunet.2026.109193","source":"pubmed"},{"id":"doi:10.1038/s41598-026-50882-8","type":"article-journal","title":"Federated learning with swarm intelligence for efficient and secure medical image analysis.","abstract":"Collaborative learning in healthcare faces challenges, including strict regulations and fragmented data. This research introduces a federated learning framework that employs swarm intelligence to augment communication and enhance the analysis of medical images. The method optimizes hyperparameters, selects features, and assigns aggregation weights to federated clients simultaneously by combining Particle Swarm Optimization (PSO) and the Firefly Algorithm (FA) with deep Convolutional Neural Networks (CNNs). The framework was tested on three medical datasets: COVID-19 chest X-rays (5,856 images), monkeypox skin images (569 images), and breast cancer mammograms (320 images). These datasets were shared among four fake healthcare institutions. It strives to strike a balance between privacy, communication costs, and classification accuracy. The results showed that the test was 96.71% accurate in detecting COVID-19, 96.06% accurate in classifying monkeypox, and 97.0% accurate in diagnosing breast cancer. The framework was able to handle noise and attacks from individuals who sought to disrupt it, which reduced communication rounds by 25-30%. A privacy-utility analysis revealed that there were acceptable trade-offs, with accuracy remaining above 94%. This study employs robust privacy measures and statistical validation. It also shows how to use medical AI in smaller healthcare settings without putting patients' privacy at risk.","author":[{"family":"Ma","given":"Sayedelahl"},{"family":"Rm","given":"Farouk"},{"family":"Ae","given":"Ali"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-50882-8","URL":"https://doi.org/10.1038/s41598-026-50882-8","source":"pubmed"},{"id":"doi:10.1038/s41746-026-02710-6","type":"article-journal","title":"Flexible and scalable federated learning with deep feature prompts for digital pathology.","abstract":"Collaborative learning across medical institutions is essential for building robust and generalisable digital pathology models. Federated learning (FL) enables collaboration without centralising data, yet its adoption is limited by high communication costs, model heterogeneity, and privacy concerns. We propose Federated Deep Feature Prompting (FedDFP), an efficient FL framework tailored for heterogeneous clinical environments. FedDFP introduces lightweight, client-specific learnable prompts applied to patch-level embeddings from whole-slide images. By sharing only these compact prompts, FedDFP reduces communication overhead by over 99.9% compared to standard FL while improving classification accuracy. Extensive experiments on TCGA-IDH, CAMELYON16 and CAMELYON17 show that FedDFP consistently outperforms standard and personalised FL baselines, achieving mean AUC gains of 0.11-0.13 over local-only training and up to 0.10 over the strongest federated methods. FedDFP also converges 2-4&#xd7; faster and remains effective across diverse feature extractors and multiple-instance learning architectures, demonstrating scalability, flexibility and privacy-aware collaboration.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41746-026-02710-6","URL":"https://doi.org/10.1038/s41746-026-02710-6","source":"pubmed"},{"id":"doi:10.1016/j.neunet.2026.109080","type":"article-journal","title":"Fed-DiTTab: Diffusion transformer for tabular data generation in federated learning.","abstract":"Imbalanced data is prevalent in real-world classification tasks, and traditional machine learning methods often struggle to effectively learn from minority class samples. This implies that predictive models can achieve better performance only when sufficient and balanced training data are available. However, due to privacy protection and the limitations of data silos, data cannot be directly shared. Federated learning offers a feasible solution by enabling multiple clients to collaboratively train a shared model without exposing local data. Nevertheless, the performance of federated learning can still be significantly degraded under imbalanced data distributions. To address this issue, we propose Fed-DiTTab, which first employs DiTTab to oversample minority class samples on each client, thereby mitigating class imbalance in federated learning while preserving data privacy to a certain extent. Extensive experiments on public datasets validate the effectiveness of Fed-DiTTab, showing that it significantly outperforms other methods discussed in this paper on imbalanced datasets. Furthermore, the ablation study validates the necessity of the synthetic data mechanism, demonstrating that training solely on raw data fails to effectively capture minority class features in severely imbalanced data environments.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.neunet.2026.109080","URL":"https://doi.org/10.1016/j.neunet.2026.109080","source":"pubmed"},{"id":"doi:10.1038/s41598-026-60588-6","type":"article-journal","title":"BlockFedZTA: a trust-aware federated learning framework for secure multi-organizational intrusion detection.","abstract":"The design of a privacy-preserved intrusion detection system for supply chain networks is challenging because of strict data privacy requirements, heterogeneous data distributions, and unreliable participating nodes. This study proposes BlockFedZTA, a framework that integrates federated learning, XGBoost, trust-aware aggregation, and a lightweight commitment-based integrity verification mechanism. In the proposed approach, each participant trains a local model and shares only a salted SHA-256 commitment without exposing model parameters. The aggregation mechanism assigns weights according to validation performance, reducing the influence of low-quality or potentially malicious updates. Experiments were conducted using a unified dataset containing 100,000 instances and 130 features representing five classes (Normal, DoS, Probe, R2L, and U2R), distributed among three organizations under non-IID conditions. The framework was evaluated under no-drift, moderate-drift, and severe-drift scenarios. Five-fold cross-validation produced average accuracies of 0.966, 0.964, and 0.963, respectively. Statistical analysis confirmed that the trust-aware aggregation strategy significantly outperformed FedAvg under drift conditions (p&#x2009;&lt;&#x2009;0.01). Additional comparison with FedAvg, Krum, Multi-Krum, Median Aggregation, Trimmed Mean, FLTrust, and FoolsGold revealed higher performance with an accuracy of 0.96470 in both mild and severe situations of the data drift problem. Moreover, our model was highly resistant to label poisonings, ensuring an accuracy of more than 0.962 even with a high level of 60%. Scaling analysis with up to 50 clients again confirmed high performance and superiority over FedAvg with moderate communication overhead. For example, communication costs went up from 1640.74 KB to 20451.89 KB per round; meanwhile, the number of audit log bytes needed rose only from 3.40 KB to 174.02 KB. Repeated runs of the algorithm ensured a stable average accuracy of 0.9647 with a standard deviation of 0.0002. Thus, BlockFedZTA ensures a robust federated IDS approach in a supply chain environment.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-60588-6","URL":"https://doi.org/10.1038/s41598-026-60588-6","source":"pubmed"},{"id":"doi:10.1016/j.neunet.2026.109298","type":"article-journal","title":"FedCAD: Cross-modal semantic alignment and distillation for cross-domain heterogeneous federated learning.","abstract":"Recent advances in the IoT and edge intelligence have made the deployment of CLIP-style image-text models in edge-cloud architectures increasingly common for collaborative sensing. However, resource heterogeneity at the edge limits the feasibility of using a unified backbone across devices with different computational budgets. In addition, non-IID data and domain shifts can disrupt image-text alignment and cause semantic drift during federated aggregation. To address these challenges, we propose FedCAD, a federated learning framework for effective knowledge collaboration under cross-domain distribution shifts and heterogeneous edge environments through cross-modal semantic alignment and multi-level distillation. FedCAD consists of three main components: (i) a Latency-Distribution Co-aware Clustering (LDCC) strategy with heterogeneous model orchestration to alleviate resource disparities and straggler effects; (ii) a decouple-then-align dual-stage feature alignment mechanism that suppresses domain-specific noise and preserves the geometric consistency of the joint image-text embedding space, thereby enhancing cross-domain generalization; and (iii) a multi-stage collaborative distillation protocol spanning intra-cluster, inter-cluster, and cloud levels, which promotes cross-cluster semantic complementarity and cross-architecture knowledge fusion to mitigate knowledge fragmentation. Experimental results show that FedCAD consistently improves classification accuracy and target-domain generalization across multiple cross-domain classification benchmarks. In addition, results on Flickr30K and MSCOCO further demonstrate its effectiveness in image-text matching and cross-modal retrieval. Under the frozen-backbone and lightweight-adapter setting, only adapter parameters are transmitted during communication rounds, which to some extent supports its deployment applicability in bandwidth-constrained networks.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.neunet.2026.109298","URL":"https://doi.org/10.1016/j.neunet.2026.109298","source":"pubmed"},{"id":"doi:10.1109/tnnls.2026.3691817","type":"article-journal","title":"A Method for Data Augmentation in Vertical Federated Learning Addressing Data Heterogeneity.","abstract":"Vertical federated learning (VFL) can aggregate data features from participating parties and is applicable to data collaboration in various fields. To address data heterogeneity in VFL, this article proposes a framework tailored for heterogeneous environments. First, to mitigate performance degradation caused by imbalanced local data across clients, we exploit conditional generative adversarial networks (CGANs) for targeted data augmentation, and propose a data balancing model named FeCWGAN-GP based on CGAN. This model pretrains a local CGAN for each client to perform local private data compensation, thereby alleviating the problem of decreased model performance. Second, to handle local model parameter distribution shifts induced by heterogeneity, leveraging both the sample size proportions and the Wasserstein distance in the model parameter space to capture parameter distribution shifts due to data heterogeneity, a parameter aggregation algorithm named WFedDA based on sample size and parameter distribution is proposed. This method calculates weights based on the sample size proportions of participants and the distribution difference between local model parameters and global model parameters, thus optimizing the global model. Finally, to address the instability of local model parameters caused by data heterogeneity, a stochastic gradient descent (SGD) method with a dual smoothing mechanism named SGD-MA is proposed. This method uses an exponential moving average (EMA) to process gradients and parameters sequentially, which reduces the fluctuation of gradients and the instability of parameter updates, thereby improving the stability of the training process. Experiments on the public datasets MNIST, CIFAR-10, Fashion-MNIST, and MIMIC-III demonstrate that the methods proposed in this article can effectively address the issues caused by data heterogeneity in multidata-source environments, significantly improving the generalization capability and stability of the global model.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tnnls.2026.3691817","URL":"https://doi.org/10.1109/tnnls.2026.3691817","source":"pubmed"},{"id":"doi:10.1038/s41598-026-58383-4","type":"article-journal","title":"SecureTrust-FL: trust-aware privacy-preserving federated learning for network intrusion detection.","abstract":"The proliferation of distributed network environments and the Internet of Things (IoT) has increased the need for privacy-preserving intrusion detection systems capable of operating effectively under heterogeneous and non-independent and identically distributed (non-IID) data conditions. This paper proposes SecureTrust-FL, a trust-aware federated learning framework for privacy-preserving intrusion detection. The framework integrates Federated Learning, Blockchain-based Trust Management, Differential Privacy, FGSM-based Adversarial Learning, and Zero-Trust Security principles to support secure collaborative learning without requiring raw data sharing among participating entities. The framework is evaluated using three benchmark intrusion detection datasets, namely CICIDS2017, UNSW-NB15, and BoT-IoT, which are treated as independent federated clients. Experimental results demonstrate that the proposed framework achieves an overall Accuracy of 92.91% &#xb1; 0.45%, Balanced Accuracy of 93.25% &#xb1; 0.43%, Macro F1-Score of 92.89% &#xb1; 0.45%, and AUC-ROC of 95.50% &#xb1; 0.40% across heterogeneous datasets. The results indicate that the federated model can effectively learn from distributed and heterogeneous data while preserving data privacy. Further analysis reveals the impact of class imbalance on intrusion detection performance, particularly in datasets containing skewed attack distributions, highlighting the importance of Balanced Accuracy and F1-Score in addition to overall Accuracy. Differential privacy experiments demonstrate the privacy-utility trade-off, where stronger privacy protection leads to a reduction in model performance. Adversarial robustness evaluation using FGSM perturbations also shows a noticeable decline in detection performance, indicating the need for stronger defense mechanisms against adversarial attacks. In addition, the trust ledger enhances transparency and accountability by monitoring client participation and recording the trust scores used during trust-weighted aggregation and maintaining trust records throughout the collaborative learning process. The results demonstrate that SecureTrust-FL provides an effective framework for privacy-preserving collaborative intrusion detection while integrating trust management, privacy protection, and secure federated learning within a unified architecture.","author":[{"family":"Ns","given":"Alshammari"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-58383-4","URL":"https://doi.org/10.1038/s41598-026-58383-4","source":"pubmed"},{"id":"doi:10.1007/s12539-026-00825-8","type":"article-journal","title":"VFLING: Vertical Federated Learning for Multi-Omics Data Integration with Graphs.","abstract":"Modern machine learning models leveraging multi-omics data face significant privacy challenges due to the sensitive nature of patient information. Communication overhead and missing features in each institution can lead to a substantial decline in federated learning performance. In response to these concerns, we propose VFLING-Federated Learning for Multi-Omics Data Integration with Graphs-a secure one-shot communication federated learning framework. We note that medical data reflects disease characteristics from different omics, while the relationships between samples exhibit relative stability across these omics. To minimize data transmission while maximizing the utilization of each participant's feature information we develop a strategy that transmits not only the local features but also the relationships or topology in one-shot communication. By fusing the omics based on the locally learned graph structure instead of features, VFLING enables improved performance even when some features are missing from individual parties. Extensive experiments demonstrate that VFLING outperforms existing frameworks, paving the way for applications in the medical field. Local features and graph topology are shared to the trainable server in a single communication, enhancing model accuracy through integrated data. This approach improves robustness despite incomplete information. Local Parties and Server Integration: Local parties learn and transmit both local features and graph topology to the server in a single communication, maximizing effective information transfer. The trainable server then integrates data across parties using graph relationships, enhancing model robustness and accuracy despite incomplete feature sets.","author":[{"family":"Rs","given":"Huang"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1007/s12539-026-00825-8","URL":"https://doi.org/10.1007/s12539-026-00825-8","source":"pubmed"},{"id":"doi:10.1186/s12911-026-03553-7","type":"article-journal","title":"TrainTracks - federated learning for reproducible research on sensitive medical data.","abstract":"Reproducibility of computational algorithms is a challenging but crucial requirement for medical research and an important component of trustworthy training and application of AI algorithms. Federated Learning (FL) is commonly used to enable privacy-preserving AI in medical research. One prerequisite of reproducibility is traceability. A majority of publications on traceable FL platforms leverage blockchain technology to achieve traceable FL. In healthcare settings, resource-efficient alternatives to blockchains are possible; however, their traceability features require separate design considerations.","author":[{"family":"Fk","given":"Tang"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1186/s12911-026-03553-7","URL":"https://doi.org/10.1186/s12911-026-03553-7","source":"pubmed"},{"id":"doi:10.1038/s41598-026-59680-8","type":"article-journal","title":"BRIDGE-T: addressing temporal unreliability in federated learning for edge-enabled IoT networks.","abstract":"Federated Learning (FL) in edge-enabled Internet of Things (IoT) networks faces considerable challenges owing to intermittent client participation and distributional drift, and which undermines the stability of a global model's optimization. This coupled impact introduces temporal unreliability, in turn, impairing the training stability. State-of-the-art FL frameworks typically address these challenges in isolation and overlook their coupled impact particularly during the client reintegration process. In order to address this limitation, we propose BRIDGE-T, i.e., a reliability-aware FL framework that addresses temporal unreliability in edge-enabled IoT networks. BRIDGE-T encompasses three components, i.e., (i) Prototype Contrastive Drift Alignment (PCDA) to constrain cross-client representation divergence under evolving non-Independent and Identically Distributed (non-IID) data, (ii) Prototype Query Agreement (PQA) to estimate round-wise clients reliability via cross-client prediction consistency on shared prototypes, and (iii) Reliability-Weighted Asynchronous-aware Aggregation (RWAA) to regulate clients' influence and attenuate stale or misaligned clients' updates. Extensive experiments under varying intermittency and distributional drift on CIFAR-10, CIFAR-100, MNIST, and TON-IoT suggest that BRIDGE-T achieves smoother convergence and greater robustness to client reintegration vis-&#xe0;-vis the state-of-the-art FL frameworks.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-59680-8","URL":"https://doi.org/10.1038/s41598-026-59680-8","source":"pubmed"},{"id":"doi:10.1016/j.neunet.2026.109302","type":"article-journal","title":"Reliability-aware modality completion with cross-modal distillation for federated learning with missing modalities.","abstract":"Multimodal federated learning (MFL) enables multiple clients to collaboratively train a global model from decentralized data without sharing local privacy-sensitive information. However, practical MFL is challenged by cross-client heterogeneity and missing modalities, which jointly induce client drift, incomplete semantic representations, and degraded global generalization. To address these issues, we propose FCKMD, a robust multimodal federated learning framework for heterogeneous and incomplete-modality settings. Specifically, FCKMD introduces a heterogeneity-adaptive modality expert encoding mechanism, in which a sample-wise router dynamically selects suitable expert paths for different client data and adopts a bypass strategy for missing modalities. To compensate for incomplete observations, FCKMD further employs a cross-modal reconstruction module together with a reliability-aware constraint strategy that adjusts the supervision strength. To further improve prediction robustness, a cross-view consistency transfer scheme is developed to distill discriminative knowledge from the fused multimodal branch into unimodal branches. These components are jointly optimized under a unified objective that integrates classification, reconstruction, and distillation losses. Experimental results on CREMA-D, Crisis-MMD, and UCI-HAR demonstrate that FCKMD outperforms representative baselines and achieves strong robustness and generalization under different missing rates, heterogeneity levels, and client participation ratios.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.neunet.2026.109302","URL":"https://doi.org/10.1016/j.neunet.2026.109302","source":"pubmed"},{"id":"doi:10.1109/jbhi.2026.3695064","type":"article-journal","title":"Privacy-Enhanced Vertical Federated Learning for Healthcare via Directional Noise and Subset Representations.","abstract":"Vertical federated learning (VFL) allows healthcare institutions to train models on complementary patient features without sharing raw data, but strong differential privacy often causes severe utility loss and labeled medical data are limited.We propose HEAL, a privacy-enhanced VFL framework that jointly learns subset representations and optimizes the direction of privacy-preserving noise. HEAL first constructs importance-aware feature subsets and performs multi-level contrastive pre-training to exploit unlabeled data and unify heterogeneous feature spaces. It then applies direction-optimized differential privacy to preserve formal $(\\epsilon, \\delta)$-privacy while reducing gradient distortion, followed by collaborative task learning for healthcare prediction. Across four healthcare datasets, HEAL improves accuracy by 2.6-4.7% over state-of-the-art baselines, reaches 96.2% of centralized performance at $\\epsilon =1.0$, and degrades gradient-inversion reconstruction quality by 20-35%. These results show that privacy protection and representation learning can reinforce each other, rather than treating privacy only as a performance cost.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/jbhi.2026.3695064","URL":"https://doi.org/10.1109/jbhi.2026.3695064","source":"pubmed"},{"id":"doi:10.3791/70175","type":"article-journal","title":"Intelligent Federated Learning Framework for Non-colocated and Heterogeneous Datasets.","abstract":"Federated learning has significant potential for distributed model training while preserving privacy, but it faces challenges related to convergence, fairness, and interpretability due to the heterogeneity of non-colocated datasets. This study proposes an Intelligent Federated Learning Framework (IFLF) to address these challenges through an adaptive approach using a multi-layer architecture consisting of Data, Client, Aggregation, Adaptation, and Optimization, and Interpretability layers. The framework demonstrates stable convergence under non-IID data distributions, with aggregation strategies supporting balanced optimization. The learning-rate modulation approach contributes to stable training by integrating heterogeneous client updates and reducing divergence during optimization. Explainable AI techniques, including SHAP and LIME, are incorporated to improve transparency at both client and global levels. The IFLF is evaluated on four benchmark datasets (FEMNIST (vision), FLamby (healthcare imaging), FedGraphNN (graph learning), and CICIDS2017 (cybersecurity)). The framework achieved an average accuracy of 92.8%, with faster convergence and reduced performance variability across clients.","author":[{"family":"Nd","given":"Navghare"},{"family":"Lm","given":"Gladence"},{"family":"Aa","given":"Bhosle"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3791/70175","URL":"https://doi.org/10.3791/70175","source":"pubmed"},{"id":"doi:10.14293/pr2199.003337.v1","type":"article-journal","title":"Secure Aggregation Techniques in Federated Learning for Vehicle Data Analytics","abstract":"The rapid evolution of Intelligent Transportation Systems (ITS) and Autonomous Vehicles (AVs) has generated a massive influx of vehicular data. While this data is pivotal for enhancing traffic safety and predictive maintenance, privacy concerns regarding location history and driving patterns remain a significant barrier. Federated Learning (FL) offers a decentralized alternative to traditional machine learning by training models locally on vehicles; however, FL is still susceptible to poisoning attacks and inference-based privacy leaks. This study investigates the efficacy of secure aggregation techniques—specifically Homomorphic Encryption (HE) and Multi-Party Computation (MPC)—in maintaining model accuracy while ensuring robust privacy. Using a quantitative experimental design and the Bosch Vehicle Motion Dataset, we demonstrate that while secure aggregation introduces a latency overhead of 12-18%, it successfully mitigates reconstruction attacks without compromising convergence rates. Our findings suggest that a hybrid approach is optimal for real-time vehicular analytics.","author":[{"family":"Brown","given":"Emily"},{"family":"Jones","given":"Sarah"},{"family":"Miller","given":"David"},{"family":"Elvis","given":"Grace"}],"issued":{"date-parts":[[2026]]},"DOI":"10.14293/pr2199.003337.v1","URL":"https://doi.org/10.14293/pr2199.003337.v1","source":"europepmc"},{"id":"doi:10.3233/shti260367","type":"article-journal","title":"Good for All, Not Good Enough for One: Reuse Dilemma in Federated Learning.","abstract":"Federated learning (FL) promises privacy-aware collaboration in healthcare, but real-world adoption remains limited by infrastructural and organizational hurdles. In this paper, we reflect on our experience developing and later bypassing our own general-purpose FL framework, in favor of a task-specific pipeline. This case exposed five core barriers to reuse, ranging from workflow misalignment to governance constraints, that often go unaddressed in technical design. Rather than prescribing one approach over another, we argue for a shift toward modular, interoperable tools that can better accommodate the diversity of research contexts. Our findings highlight the need for realistic infrastructure thinking: one that acknowledges both the promise and the limits of reuse in practice.","author":[{"family":"Lm","given":"Peeters"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3233/shti260367","URL":"https://doi.org/10.3233/shti260367","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-9991236/v1","type":"article-journal","title":"Embodied brain-cerebellum federated learning for satellite-assisted low-altitude wireless networks","abstract":"Abstract Satellite-assisted low-altitude wireless networks need distributed intelligence that handles scarce labels, limited onboard resources and short UAV contact windows. We propose EBC-FL, an embodied brain--cerebellum federated learning framework in which ground stations maintain semantic memory, MEO satellites coordinate regional aggregation and low-altitude agents execute lightweight local policies. EBC-FL combines descriptor-based semantic inference, adapter updates and a control-entropy-aware sensing--communication--computing--control objective. Against conventional and strong baselines, EBC-FL is not always the highest-success method in benign traffic, but provides a stronger safety-critical reliability--semantic-awareness--upload tradeoff. In the safety scenario, it improves service success from 85.26% to 88.52%, reduces entropy-induced failures from 14.09% to 7.89%, and lowers upload from 209.42 MB to 75.79 MB relative to risk-aware scheduling. Controlled mobility and contact-window sweeps confirm robustness under high-mobility, short-window operation.","author":[{"family":"Jing","given":"Yi"},{"family":"Jiang","given":"Chunxiao"},{"family":"Wang","given":"Jiawei"},{"family":"Sun","given":"Jiachen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9991236/v1","URL":"https://doi.org/10.21203/rs.3.rs-9991236/v1","source":"europepmc"},{"id":"doi:10.1016/j.neunet.2026.109378","type":"article-journal","title":"A unified basis decomposition framework for addressing the federated learning trilemma: Communication-efficiency, personalization, and privacy.","abstract":"Federated learning (FL) is confronted with a fundamental trilemma: simultaneously achieving communication efficiency, personalized adaptation, and privacy protection. Current approaches typically optimize one objective at the expense of the others, failing to provide a unified solution. To address this challenge, we propose Federated Basis Decomposition (FedBD), a unified framework that leverages layer-wise basis decomposition to achieve balanced optimization across all three dimensions. The key innovation lies in reformulating the learning process as a subspace optimization problem: each network layer is represented as a linear combination of globally shared basis models, enabling clients to exchange only low-dimensional scalar weights rather than full model parameters. FedBD directly resolves the trilemma by drastically reducing communication overhead through compact parameter transmission, while simultaneously enabling effective personalization by decoupling globally shared scalar weights from private local parameters including Batch Normalization statistics. Furthermore, the framework strengthens privacy protection by projecting raw gradients into a random basis subspace. We validate our framework on three medical imaging datasets. The results demonstrate that FedBD achieves an accuracy density gain of up to 9.73&#x202f;&#xd7;&#x202f;, effectively maintaining competitive accuracy with an order-of-magnitude reduction in communication overhead compared to conventional FL methods. Ultimately, FedBD offers a unified solution to the FL trilemma, paving the way for practical deployment in complex real-world scenarios.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.neunet.2026.109378","URL":"https://doi.org/10.1016/j.neunet.2026.109378","source":"pubmed"},{"id":"doi:10.3389/fdata.2026.1769948","type":"article-journal","title":"Quantifying energy and accuracy trade-offs of federated learning on wearable health devices.","abstract":"The rapid development of wearable health tools has made it possible to continuously monitor physiological conditions for preventive care. However, stringent privacy laws, including HIPAA and GDPR, require decentralized methods such as federated learning (FL) to safeguard personal patient information. Nonetheless, empirical profiling in this paper finds that typical FL implementations are plagued by a serious performance trilemma; a naive federated model attains a 35.3 percent energy savings (3.84 vs. 5.93 kJ in the centralized models), but at the cost of a disastrous performance penalty of 13.87 percentage points (84.94 vs. 98.81 percent in centralized models). The failure in research is largely due to the on-device computational load of 4.24 MFLOPs per training sample, resulting in a \"straggler\" bottleneck that increases the total training duration to 1,066.26 s, almost 70 times longer than centralized training. As a result, the introduction of the hybrid hierarchical federated split learning (H-FedSL) architecture helps in strategically splitting the neural network at a cut layer to divide the workload between wearable and nearby edge servers. The methodology provides a new framework that offloads the heavy and deep-layer computations to the edge server, leaving the shallow feature extraction to the point of operation, and sends only privacy-sensitive abstractions of the smashed data, rather than raw signals. The integration of asynchronous protocols will help manage device heterogeneity and resource-aware client selection, thereby achieving the aim of H-FedSL to restore the gold-standard accuracy of 98.81% with the state-of-the-art 35.3% energy efficiency of the federated model. Thus, a technically and economically feasible pathway will be provided for deploying medical-grade AI on resource-constrained Internet of Medical Things (IoMT) devices.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/fdata.2026.1769948","URL":"https://doi.org/10.3389/fdata.2026.1769948","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-6729707/v1","type":"article-journal","title":"Federated Learning for Privacy-Preserving Smart Cities: A Secure and Scalable Machine Learning Framework","abstract":"Abstract The exponential growth of data in smart city infrastructures—from traffic systems to health monitoring and surveillance—has created unprecedented opportunities for machine learning applications. However, centralizing such diverse and sensitive data introduces serious challenges related to data privacy, regulatory compliance, and system scalability. In this paper, we propose a secure and scalable federated learning (FL) framework tailored for smart city environments, enabling decentralized model training while preserving data locality and privacy. The framework integrates key technologies including differential privacy, secure aggregation, and edge device optimization to ensure robust model performance and security under real-world conditions. The framework is implemented and simulated using TensorFlow with synthetic smart city data streams, evaluating the system across key metrics such as training accuracy, communication cost, latency, and model convergence. Our experimental results show that the proposed FL framework achieves high prediction accuracy (94.3%) with significantly reduced bandwidth consumption and strong privacy guarantees. This work contributes a deployable architecture for future smart cities, offering an effective balance between intelligent data use and citizen data rights.","author":[{"family":"Juneja","given":"Deepak"},{"family":"Singh","given":"Arvinder"},{"family":"Singh","given":"Jagvinder"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-6729707/v1","URL":"https://doi.org/10.21203/rs.3.rs-6729707/v1","source":"preprints"},{"id":"doi:10.2139/ssrn.5420991","type":"manuscript","title":"Verifiable and Privacy-Preserving Decentralized Collaboration for Machine Learning Model Improvement","abstract":"Abstract Decentralized collaboration, particularly in complex domains like machine learning (ML) model development, faces hurdles regarding trust, intellectual property (IP) protection, and the verification of contributions. Traditional centralized platforms introduce bottlenecks, while existing decentralized approaches often lack mechanisms to privately verify complex, computationally intensive work. This paper introduces a framework integrating Zero-Knowledge Proofs (ZKPs) and smart contracts to enable trustless, verifiable, and privacy-preserving collaborative ML model improvement process. Our system allows \"Improvers\" to cryptographically prove superior model performance compared to a baseline, without revealing proprietary model parameters prematurely. The framework features on-chain job management, client-side ZKP generation, on-chain verification, automated selection of the best contributor based on verified proofs, secure solution submission via cryptographic commitments, and an arbiter-based dispute resolution mechanism. We analyze different ZKP workflow variations, evaluating performance and suitability. Performance analysis demonstrates the feasibility of client-side proof generation for individual models, while highlighting the resource demands of proof aggregation, suggesting its suitability for server-side work. On-chain gas cost evaluation indicates the system's economic viability for Layer 2 (L2) scaling solution deployments under specific cost assumptions. This work provides a first step for secure and verifiable collaboration in ML and potentially other software development tasks in decentralized environments.","author":[{"family":"Burgos","given":"Jay"},{"family":"Sedlar","given":"Urban"},{"family":"Pustišek","given":"Matevž"}],"issued":{"date-parts":[[2025]]},"DOI":"10.2139/ssrn.5420991","URL":"https://doi.org/10.2139/ssrn.5420991","source":"crossref"},{"id":"doi:10.21203/rs.3.rs-9637381/v1","type":"article-journal","title":"Federated Learning, Temporal Convolutional Networks, Electric Vehicle Charging Infrastructure, Intrusion Detection, Cybersecurity, Privacy-Preserving Machine Learning","abstract":"Abstract The rapid growth of electric vehicle (EV) charging infrastructure has increased exposure to cyberattacks, while conventional centralized intrusion detection remains poorly suited to distributed and privacy-sensitive EVSE environments. This study proposes a federated learning framework based on a Dual-Attention Temporal Convolutional Network (DA-TCN) for collaborative cyberattack detection without centralized data sharing. The model combines dilated causal convolutions with channel and temporal attention to capture complex dependencies in multimodal EVSE telemetry. To improve federated optimization, we introduce Federated Stochastic Weight Averaging (FedSWA) and client-local mixup augmentation, which together enhance generalization and robustness across distributed clients. Evaluated on the CICEVSE2024 dataset, the proposed framework achieved 99.18\\% accuracy and 98.95\\% macro F1-score, outperforming a centralized baseline (98.96\\% accuracy) while maintaining perfect recall for benign and cryptojacking traffic. With an inference latency of 7.21~ms per batch and a model size of 4.07~MB, the framework is well suited for real-time edge deployment. A complementary Isolation Forest module provides unsupervised zero-day anomaly detection. These results demonstrate that federated learning can deliver accurate, lightweight, and deployment-ready intrusion detection for privacy-preserving EV charging infrastructure.","author":[{"family":"Ragab","given":"Mohammed"},{"family":"Alhussian","given":"Hitham"},{"family":"Abdulkadir","given":"Said"},{"family":"Eltahir","given":"Majdy"},{"family":"Alwadain","given":"Ayed"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9637381/v1","URL":"https://doi.org/10.21203/rs.3.rs-9637381/v1","source":"europepmc"},{"id":"doi:10.1109/tpami.2026.3715766","type":"article-journal","title":"FedAdOb: Privacy-Preserving Federated Deep Learning with Adaptive Obfuscation. ","abstract":"Federated learning (FL) has emerged as a collaborative approach that allows multiple clients to jointly learn a machine learning model without sharing their private data. The concern about privacy leakage, albeit demonstrated under specific conditions [1], has triggered numerous follow-up research in designing powerful attacking methods and effective defending mechanisms aiming to thwart these attacking methods. Nevertheless, privacy-preserving mechanisms employed in these defending methods invariably lead to compromised model performances due to a fixed obfuscation applied to private data or gradients. In this article, we, therefore, propose a novel adaptive obfuscation mechanism, coined FedAdOb, to protect private data without yielding original model performances. Technically, FedAdOb utilizes passport-based adaptive obfuscation to ensure data privacy in both horizontal and vertical federated learning settings. The privacy-preserving capabilities of FedAdOb, specifically with regard to private features and labels, are theoretically proven through Theorems 1 and 2. Furthermore, extensive experimental evaluations conducted on various datasets and network architectures demonstrate the effectiveness of FedAdOb by manifesting its superior trade-off between privacy preservation and model performance, surpassing existing methods.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tpami.2026.3715766","URL":"https://doi.org/10.1109/tpami.2026.3715766","source":"pubmed"},{"id":"doi:10.3390/diagnostics16132029","type":"article-journal","title":"FLAME: Federated Learning and Aggregated Multi-Model Ensemble for Multi-Class Alzheimer's Disease Stage Classification from Structured Clinical Data.","abstract":"Background/Objectives : The precise identification of Alzheimer's disease (AD) stages through clinical data is crucial for early diagnosis and suitable therapy. This classification remains troublesome due to overlap in cognitive profiles across different phases of illness progression. This study presents a comprehensive and advanced diagnostic system, termed FLAME, featuring an enhanced federated learning architecture for privacy-preserving multi-institutional implementation. It provides a systematic review of machine learning (ML) and deep learning (DL) models for the classification of five stages of Alzheimer's disease (AD). The models include cognitively normal (CN), subjective memory complaints (SMC), early mild cognitive impairment (EMCI), late mild cognitive impairment (LMCI), and Alzheimer's disease (AD). Methods : Sixteen traditional machine learning models and eleven deep learning architectures-including FT-Transformer and NODE-were evaluated using a structured clinical dataset comprising 362 features. A hybrid ensemble was created at the probability level by combining the two top-performing models, LightGBM and a five-layer DNN. The weights of this ensemble were automatically optimised using a Genetic Algorithm (GA) with Macro-F1 as the fitness criterion, confirmed stable across 30 independent runs (w&#x2605;=0.5024&#xb1;0.0001). A federated learning architecture was then established, deploying the DNN across non-IID clients while keeping LightGBM centralised. We examine four distinct aggregation algorithms: FedAvg, FedProx, FedNova, and SCAFFOLD. Results : Among all deep learning architectures, FT-Transformer achieved the highest standalone performance (accuracy = 0.7810, &#x3ba; = 0.7081). The five-layer deep neural network (DNN) was selected as the DL representative for the hybrid ensemble. LightGBM attained superior machine learning performance (accuracy = 0.8156, &#x3ba; = 0.7537), confirmed deterministic across 10 seeds. The LightGBM vs. XGBoost difference is not statistically significant (McNemar p=0.4227). The GA-optimised hybrid ensemble (w = 0.685) surpassed both individual baselines across all evaluation metrics. The FedNova hybrid design achieved superior overall performance in federated configurations, surpassing all centralised arrangements in accuracy (accuracy = 0.8213, &#x3ba; 0.7614). Conclusions : Evolutionary ensemble optimisation combined with federated learning provides a robust, scalable, and privacy-preserving solution for AD stage classification, offering a clinically viable framework for real-world multi-institutional decision-support systems. However, the AD class remains severely under-recalled across all configurations (F1 &#x2264; 0.21), identifying this as the primary open challenge for clinical translation.","author":[{"family":"Lb","given":"Ammar"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/diagnostics16132029","URL":"https://doi.org/10.3390/diagnostics16132029","source":"pubmed"},{"id":"doi:10.1038/s41598-026-63991-1","type":"article-journal","title":"Manual federated simulation for multiple sclerosis integrating XGBoost algorithm with SHAP explanation.","abstract":"Multiple sclerosis (MS) is a chronic autoimmune disorder of the central nervous system, underscoring the importance of early and accurate diagnosis. In this study investigates the predictive modelling of MS progression in patients with Clinically Isolated Syndrome (CIS), privacy-preserving for a federated and explainable Machine Learning (ML) framework. To address missing data while preserving inter-feature dependencies, Multivariate Imputation by Chained Equations (MICE) with iterative imputers was employed. Classification was performed using the Extreme Gradient Boosting (XGBoost) algorithm. Model interpretability was developed through Explainable Artificial Intelligence (XAI) techniques, specifically Shapley Additive Explanations (SHAP). To ensure data confidentiality and simulate decentralized clinical environments, an in silico federated learning framework was applied. Experimental results demonstrated strong predictive performance, achieving 96.7% accuracy and 99% ROC-AUC during training, 92.5% accuracy in validation, and 81.8% accuracy with an AUC of 88% on the test set. For the Federated Learning (FL) simulation, the model maintained competitive performance, yielding an accuracy of 76.3% and an AUC of 83.9%. The proposed approach supports early diagnosis, enhances clinical trust through interpretability, and promotes secure data collaboration, thereby contributing to more informed and transparent clinical decision-making and improved patient care.","author":[{"family":"He","given":"Ghazy"},{"family":"Zh","given":"Ali"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-63991-1","URL":"https://doi.org/10.1038/s41598-026-63991-1","source":"pubmed"},{"id":"doi:10.1038/s41598-026-55847-5","type":"article-journal","title":"Federated MobileNetV2 with ensemble meta-learning for privacy-preserving brain tumor classification.","abstract":"The identification of brain tumors from MRI images is very crucial for the selection of an appropriate treatment. However, the existing solution has issues with privacy and data sharing. To address this challenge, this paper proposes the use of federated learning. The proposed solution employs a light convolutional backbone and some adaptive local meta-learners. The proposed solution employs MobileNetV2 as the feature extractor. This is fine-tuned for many clients using a combination of FedAvg and FedProx regularization. Each client also trains a few meta-learners (MLP, SVM, and ELM) using the local feature embeddings, enabling people to obtain personalized predictions without sharing their private information. For inference, the framework supports both probability-level averaging across client ensembles and deployable single-client prediction using only the local meta-learners of one client. On the Brain Tumor MRI Dataset, containing 7023 image slices across glioma, meningioma, no tumor, and pituitary classes, the proposed framework achieved a maximum observed accuracy of 99.57%. Across four repeated runs, it achieved 99.29% +/- 0.20% accuracy with a 95% confidence interval of 98.97% to 99.61%, while maintaining strong macro-F1 performance and a macro-average ROC-AUC of 0.998690. Under the same preprocessing and split protocol, it outperformed internally reimplemented CNN+FedAvg and CNN+FedAvg+FedProx baselines and preserved near-centralized ROC-AUC performance. Communication analysis showed that exchanging the MobileNetV2 backbone required 149.89&#xa0;MB per round for five clients, corresponding to an 83.34% reduction relative to a ResNet-50 backbone.","author":[{"family":"Vs","given":"Panwar"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-55847-5","URL":"https://doi.org/10.1038/s41598-026-55847-5","source":"pubmed"},{"id":"doi:10.1038/s41598-026-57660-6","type":"article-journal","title":"Federated deep reinforcement learning enabled hierarchical Edge-Fog-Cloud architecture for intelligent task offloading in 6G networks.","abstract":"Latency sensitive, computation intensive and mobility aware applications in Edge Fog Cloud environments have increased the demand to develop intelligent offloading task mechanism that can dynamically scale to changing network conditions whilst remaining scalable, energy efficient, and preserving data privacy. Traditional, heuristic and centralized based learning offloading methods frequently have problems in accommodating heterogeneous workloads, non- stationary environments and privacy limitations associated with the next generation distributed computer system. In order to overcome these drawbacks, the present paper will suggest a Federated Deep Q-Learning (FDQL)-based task offloading framework, which incorporates deep reinforcement learning and federated learning to support adaptive, decentralized and privacy-conscious decision-making across hierarchical Edge Fog Cloud architectures. The framework proposed solves task offloading as a Markov Decision Process, with the execution decisions being trained based on the joint consideration of the latency, bandwidth availability, queue length, computational load, and energy state, as well as user mobility, without sharing raw data during federated model aggregation. In comparison to the current CNN-, LSTM-, SVM-, and rule-based methods, which use fixed threshold values or rely on centralized training, the FDQL architecture allows collaborative learning between distributed edge nodes, enhancing generalization as well as resilience as network conditions evolve. Large-scale experimental analysis is performed using a trace-driven simulation based on a publicly available task offloading dataset of tasks and the performance is evaluated based on the latency, energy consumption, task success rate, robustness analysis, and computational efficiency. Experimental findings indicate that the proposed FDQL framework demonstrates improved performance under distributed and resource-constrained environments compared to baseline approaches since shorter latency, increased energy efficiency, and more predictable execution-layer selection are achieved. The significance of federated learning, mobility awareness, and bandwidth-aware optimization in the stability of the performance is also confirmed by ablation studies. In order to achieve a better level of transparency and trustworthiness, SLA-based confusion matrix analysis and ROC analysis are performed as well as SHAP-based explainability analysis, which proves that the decisions made by FDQL are based on physically interesting, as well as SLA-relevant features, like latency, bandwidth, and resource use. All in all, the designed FDQL framework is a successful, interpretable, and scalable approach to intelligent task offloading, so it would fit perfectly into the implementation of the 6G-enabled application, such as smart cities, industrial internet of things, and autonomous systems in the future.","author":[{"family":"Smv","given":"Pandian"},{"family":"Es","given":"Vinothkumar"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-57660-6","URL":"https://doi.org/10.1038/s41598-026-57660-6","source":"pubmed"},{"id":"doi:10.1038/s41598-026-50865-9","type":"article-journal","title":"A federated learning-benchmarking framework for privacy-preserving UAV intrusion detection using adaptive aggregation algorithms.","abstract":"The recent explosive growth of Unmanned Aerial Vehicles (UAVs) has contributed to their high vulnerability to cyber-attacks including Denial of Service (DoS), identity impersonation and unauthorized access to data. The UAV networks have the inherent risks of centralized Intrusion Detection Systems (IDS) which pose critical privacy risks and points of failure, and therefore the decentralized and privacy-preserving learning paradigms are required. The paper presents federated learning architecture called FedDrone-Shield( Federated Learning Framework for Drone Security and Shield against Intrusions), which is used in the task of detecting UAV intrusions in the scenarios of Independent and Identically Distributed (IID) data and the assessment of several aggregation algorithms: FedAvg, FedProx, FedAdam, FedMedian, and ClusterAvg. A significant amount of experiments that were carried out on a dataset on anomaly detection of UAVs prove that FedAdam and ClusterAvg outperform other aggregation strategies by achieving test accuracies of 99.98, F1-scores of 0.9999, and impressively low loss values of 0.0009-0.0014. FedMedian has also closely competitive performance, whereas FedAvg and FedProx are slightly less accurate and slower converging. Client-level assessments also show a consistent high precision, recall and F1-score across all attack types, with weighted F1-scores between 0.9997 and 0.9999, which again shows that there is reliable detection performance amongst distributed UAV clients. These findings make FedDrone-Shield a strong and feasible bench-marking model of federated intrusion detection in UAV networks proving that adaptive aggregation approaches contribute to a substantial improvement of detection accuracy, training, and data privacy. This means that the proposed structure offers a robust basis of intrusion detection that is safe and ensures privacy in distributed UAVs.","author":[{"family":"Me","given":"Masud"},{"family":"Ma","given":"Hossain"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-50865-9","URL":"https://doi.org/10.1038/s41598-026-50865-9","source":"pubmed"},{"id":"doi:10.1038/s41598-026-50003-5","type":"article-journal","title":"Federated learning-enabled privacy-preserving framework for seizure forecasting and affective state analysis using multi-modal EEG-ECG data.","abstract":"Seizure forecasting and affective state analysis using EEG-ECG data play a pivotal role in advancing neurological and mental health monitoring. However, existing methods such as Fed-Transformer, Res-1D CNN, and Fed-ESD suffer from privacy risks, inefficient feature extraction, and high computational overhead, limiting their effectiveness in real-world applications. To overcome these challenges, this study proposes NeuroFedSense, a novel Federated Learning-enabled Privacy-Preserving Framework that integrates a Temporal Convolutional Network (TCN) with an Attention Mechanism for accurate seizure forecasting and affective state analysis using EEG-ECG data, ensuring enhanced feature selection, interpretability and efficient decentralized training. The model leverages adaptive attention-based optimization and weighted feature selection to improve classification performance while ensuring data privacy. Implemented using TensorFlow, NeuroFedSense achieves 99.54% accuracy, 99.62% precision, 99.34% recall, and a 99.46% F1-score, outperforming Fed-Transformer (97.10% accuracy), Res-1D CNN (81.62% accuracy), and FML (99.10% accuracy). The ROC-AUC score of 0.99 further establishes its superiority over competing models. Additionally, the federated approach reduces energy consumption per node by 30% and optimizes communication efficiency by minimizing data transmission by 15% over 100 rounds. By ensuring high accuracy, improved privacy, reduced computational overhead, and enhanced energy efficiency, NeuroFedSense sets a new benchmark for decentralized, real-time seizure prediction and affective state monitoring. These findings underscore its potential for deployment in intelligent, privacy-preserving healthcare applications, addressing critical challenges in remote neurological monitoring.","author":[{"family":"Vs","given":"Arulmurugan"},{"family":"Ss","given":"Vidhya"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-50003-5","URL":"https://doi.org/10.1038/s41598-026-50003-5","source":"pubmed"},{"id":"doi:10.1371/journal.pone.0343669","type":"article-journal","title":"Privacy-preserving multimodal federated learning pipeline for cyber-resilient healthcare systems.","abstract":"The integration of Internet of Things (IoT) devices and electronic medical records (EMRs) has transformed healthcare delivery but has also created new vulnerabilities to cyberattacks that threaten both data confidentiality and patient safety. Conventional centralized machine learning approaches for intrusion detection are impractical in this domain due to strict privacy regulations, heterogeneous data sources, and the risk of single points of failure. To address these challenges, we propose a secure distributed machine learning pipeline for cyber-resilient healthcare systems. The framework combines federated optimization with split learning for sensitive EMR data, robust aggregation to mitigate poisoned updates, and differential privacy with secure aggregation to protect against inference attacks. Multimodal fusion is enabled through temporal consistency regularization for IoT traffic and cross-layer contrastive alignment to link EMR representations, ensuring improved anomaly detection across diverse healthcare environments. Experiments conducted on representative IoT and EMR datasets demonstrate that the proposed pipeline achieves accuracy of 0.942 on IoT data, 0.931 on EMR data, and 0.953&#xa0;in the combined setting, with corresponding F1-scores of 0.921, 0.908, and 0.932. Ranking metrics further confirm superiority with AUROC up to 0.961 and AUPRC up to 0.947, outperforming deep baselines by margins of +0.025 to +0.033. Robustness analysis shows graceful degradation under client poisoning ([Formula: see text] at 30% malicious clients) and resilience under severe communication constraints (accuracy [Formula: see text] at 90% update sparsification). Detection latency is reduced to an average of 5.9 time steps, compared to 7.8 for the strongest deep baseline. These results highlight that secure distributed pipelines can deliver both strong detection capabilities and regulatory compliance, providing a practical path toward safeguarding next-generation healthcare infrastructures against evolving cyber threats.","author":[{"family":"Mim","given":"Tanvir"},{"family":"Hr","given":"Rabby"},{"family":"Mh","given":"Arif"},{"family":"Ny","given":"Nadia"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1371/journal.pone.0343669","URL":"https://doi.org/10.1371/journal.pone.0343669","source":"pubmed"},{"id":"doi:10.1038/s41598-026-49902-4","type":"article-journal","title":"Hierarchical proof of trust a Byzantine fault tolerant federated learning framework for industrial IoT applications.","abstract":"The Industrial Internet of Things (IIoT) presents significant challenges for training machine learning models due to data privacy concerns, heterogeneous data distributions, and limited bandwidth. This paper proposes HPoT (Hierarchical Proof-of-Trust), a novel federated learning-based consensus algorithm for IIoT infrastructure that addresses these limitations while ensuring data integrity and model fidelity. HPoT integrates blockchain technology with federated learning to create a decentralized, privacy-preserving framework enabling collaborative model training without raw data sharing. The algorithm features: (1) robust aggregation with weighted averaging and anomaly detection for heterogeneous datasets, (2) adaptive weighting prioritizing updates from trustworthy devices, (3) reputation scoring to detect malicious nodes, and (4) hierarchical three-tier architecture optimized for IIoT constraints. Using Byzantine fault tolerance principles, HPoT tolerates up to [Formula: see text] malicious participants while maintaining convergence. Experimental evaluation demonstrates [Formula: see text] reduction in communication rounds, [Formula: see text] decrease in per-node energy consumption, and [Formula: see text] lower consensus latency compared to traditional federated learning, while maintaining comparable accuracy ([Formula: see text] vs [Formula: see text]). The algorithm effectively handles data heterogeneity, communication constraints, and adversarial attacks common in IIoT environments. HPoT provides a scalable, secure, and energy-efficient solution for privacy-preserving machine learning on resource-constrained IIoT devices, advancing practical deployment of collaborative AI in industrial settings.","author":[{"family":"Sk","given":"Sharma"},{"family":"Ps","given":"Rathore"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-49902-4","URL":"https://doi.org/10.1038/s41598-026-49902-4","source":"pubmed"},{"id":"doi:10.1038/s41598-026-51460-8","type":"article-journal","title":"Resource-efficient federated machine unlearning via evolutionary synaptic pruning for cloud-based distributed learning systems.","abstract":"With an increasing emphasis on user data control and privacy regulations, such as the General Data Protection Regulation (GDPR), machine unlearning (MU) has emerged as a crucial mechanism for managing data in AI systems. MU enables models to remove the influence of specific user data upon request. This problem becomes more complex in collaborative settings such as Federated Learning (FL), where data remains distributed across multiple clients, giving rise to Federated Unlearning (FU). In large-scale deployments, particularly those supported by cloud infrastructure, retraining models to satisfy data deletion requests can be computationally expensive, energy-intensive, and disruptive to ongoing services. This highlights the importance of having effective methods to unlearn outdated practices that can hinder the growth of a system and compromise data privacy awareness. We propose PRUNE-FL (Privacy-preserving Retention-focused Unlearning with Neuro-Evolution in Federated Learning), a framework that uses relevance-guided pruning and evolutionary optimization to delete the influence of the targeted data. PRUNE-FL is different from methods that rely on retraining or coarse parameter updates because it focuses on finding and changing the parameters that are most closely related to the data that needs to be forgotten. The approach is based on a synaptic relevance scoring system to figure out how each model parameter relates to certain client or class-level data. This makes it easier to find parameters that are connected to the target data. Then, the unlearning task is set up as a multi-objective problem to find a balance between the overall performance and the forgetting of unnecessary information. Finally, a genetic algorithm is used to implement an evolutionary pruning strategy. It seeks the optimal pruning settings that operate within the constraints of federated learning. Thus, PRUNE-FL helps in unlearning specific data without having to retrain the whole system. Tests on the CIFAR-10 dataset in both IID (Independent and Identically Distributed) and non-IID settings show that PRUNE-FL has higher accuracy. It also effectively removes the influence of the targeted data. The results also show it to be strong against bad patterns like backdoor triggers. Overall, PRUNE-FL enhances privacy by selectively unlearning the data and using fewer resources in federated environments.","author":[{"family":"Dkjb","given":"Saini"},{"family":"Bk","given":"Rai"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-51460-8","URL":"https://doi.org/10.1038/s41598-026-51460-8","source":"pubmed"},{"id":"doi:10.1038/s41598-026-52361-6","type":"article-journal","title":"Blockchain-enhanced federated learning for IoT security and privacy using the GSR-C2N model.","abstract":"The rapid proliferation of Internet of Things devices has intensified the demand for security, privacy-preserving, and scalable machine learning solutions. Federated Learning (FL) supports the decentralized training of models across distributed devices without transporting raw data, whereas blockchain provides a trusted, transparent mechanism for integrity and reaching consensus. The paper explains an FL framework that incorporates blockchain technology, based on the GSR-C2N model for identifying crypto-mining malware. It is demonstrated that the system's feature extraction and optimization processes are optimized to address security problems in IoT by utilizing blockchain to verify model updates and build trust and privacy. The given structure has demonstrated superiority to the current practices. This model was 96.85% accurate and 97.51% specific on the crypto-mining malware data set using 10-fold cross-validation, making it applicable to smart healthcare and smart city applications with IoT-based systems. In addition, when combined with regulatory compliance, homomorphic encryption can strengthen data and privacy management, underscoring the model's effectiveness in advanced IoT systems.","author":[{"family":"Ma","given":"Ansari"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-52361-6","URL":"https://doi.org/10.1038/s41598-026-52361-6","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-9252646/v1","type":"article-journal","title":"Poison-Resilient and Privacy-Preserving Federated Learning Scheme in Mobile Systems","abstract":"Abstract Mobile systems including smartphones, IoT devices generate massive high value data, while conventional centralized data collection and analysis suffer from security and privacy vulnerabilities. Federating learning, an emerging paradigm in machine learning, collaboratively train s high performance model via participants sharing local training model updates rather than raw data, thereby preserving the privacy of their local datasets, provides a new approach for the secure extraction of data value in mobile system. However, malicious participants may inject carefully crafted poisoned samples into their local datasets, with the intent of disrupting the convergence of the global model or inducing targeted misclassification. Consequently, the identification of such malicious participants are of critical importance in federated learning. To address these challenges, this paper proposes poison-resilient and privacy-preserving federated learning scheme in mobile system. It not only detects poisoning attacks in both vertical federated learning and horizontal federated learning, but also removes prior assumptions regarding client data distributions and restrictions on the proportion of adversarial participants. In addition, a convergence control parameter is introduced to regulate the model’s convergence rate. The security, correctness, fairness, and robustness of the proposed scheme are formally analyzed and rigorously proven. Experimental results demonstrate that our scheme detects poisoning attacks effectively while maintaining high accuracy and model training efficiency.","author":[{"family":"Zhao","given":"Quanyu"},{"family":"Gu","given":"Chenrui"},{"family":"Jiang","given":"Bingbing"},{"family":"Zhou","given":"Yuanjian"},{"family":"Jing","given":"Zhengjun"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9252646/v1","URL":"https://doi.org/10.21203/rs.3.rs-9252646/v1","source":"europepmc"},{"id":"doi:10.1371/journal.pone.0351957","type":"article-journal","title":"FEDI-CODE: A federated and causally-informed framework for dementia risk prediction using multi-site patient data.","abstract":"Early detection of dementia is critical for timely intervention and disease management, yet it remains a challenging task due to the fragmented nature of healthcare data and the need for privacy-preserving solutions. This paper proposes FEDI-CODE, a Federated and Causally Informed Dementia Estimation framework that integrates deep learning, federated learning, and counterfactual inference to predict dementia risk across distributed patient data sources. FEDI-CODE is designed to operate without centralizing sensitive medical data, enabling collaborative training across institutions while preserving privacy. It combines temporal modeling of longitudinal imaging and clinical data with individualized treatment effect estimation for modifiable risk factors such as alcohol consumption, weight, and cardiovascular indicators. A fusion module aggregates representations from each site to form a global prediction head. Extensive experiments on simulated multi-site dementia datasets demonstrate that FEDI-CODE achieves an accuracy of 83.7%, a precision of 83%, a recall of 81%, an F1-score of 82%, and an AUC-ROC of 0.86, outperforming standard federated models and deep learning baselines by notable margins. The model also generalizes well to external datasets, achieving 79.2% accuracy and 0.80 AUC-ROC, confirming its robustness. Furthermore, FEDI-CODE produces interpretable causal insights by estimating individual treatment effects, offering actionable clinical value. These results highlight FEDI-CODE as a scalable, interpretable, and privacy-aware solution for early dementia screening and personalized risk assessment.","author":[{"family":"Ms","given":"Uddin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1371/journal.pone.0351957","URL":"https://doi.org/10.1371/journal.pone.0351957","source":"pubmed"},{"id":"doi:10.1109/tpami.2026.3711847","type":"article-journal","title":"Aligning Condensed Graph via Hashing: A New Insight for Federated Graph Learning.","abstract":"Federated Graph Learning (FGL) aims to maximize the benefits of each graph owner, which is a common form of distributed graph learning under privacy-preserving conditions. As the landscape of local clients becomes increasingly diverse in terms of both model architectures and topological complexities, graph heterogeneity turns out to be one of the significant challenges to efficient collaboration among clients. Beyond existing paradigms, we delve a fresh insight into revisiting FGL as a semantic condensed graph alignment problem in this work. From this perspective, HashFGL is proposed for heterogeneous FGL through aligning condensed graphs via hashing in a symbiotic space. Specifically, the core of HashFGL lies in that it introduces a cross-client symbiotic space to facilitate effective collaboration. Within this space, an efficient hash-based semantic encoding strategy is proposed to model each local client while balancing coordinated resilience and semantic consistency. Furthermore, we derive an elaborate graph condenser based on the above strategy, which condenses original graphs with semantics and structure-preserving property, to maintain the effectiveness of condensed graph alignment for FGL. Formal theoretical analysis further reveals that HashFGL can effectively alleviate the problem of graph heterogeneity. Experimental results on three large-scale graphs, employing standard partitioning strategies and a pioneering, more realistic partitioning that we introduced, demonstrate the efficacy and scalability of HashFGL.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tpami.2026.3711847","URL":"https://doi.org/10.1109/tpami.2026.3711847","source":"pubmed"},{"id":"doi:10.3390/biomedicines14051010","type":"article-journal","title":"Privacy-Preserving Hybrid GA-LSTM Ensemble for Typhoid Detection Using Optimised Clinical Feature Selection.","abstract":"Background/Objectives: Typhoid fever remains a major public health challenge in many low-income countries, where overlapping clinical symptoms and the limited reliability of conventional diagnostic procedures hinder accurate diagnosis. This study aims to develop a reliable and efficient diagnostic framework that automates typhoid fever detection from clinical data while preserving patient privacy. Methods: To achieve this objective, we propose a hybrid framework combining genetic algorithm (GA)-based feature selection, a Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) deep learning classifier, and federated learning. The GA identifies the most informative clinical features, reducing redundancy and computational complexity. The selected features are then used to train a CNN-LSTM model in a federated learning setup using the Federated Averaging (FedAvg) algorithm, enabling collaborative model training across multiple clients without sharing raw patient data. Results: Experimental results show that the proposed framework achieves 92% accuracy, with a strong F1-score and satisfactory sensitivity. Compared to models trained on the full feature set, the proposed approach requires less memory and shorter training time, while maintaining balanced performance under class imbalance. Conclusions: These results demonstrate that integrating evolutionary feature selection, deep sequential learning, and federated training provides an effective and privacy-aware solution for multi-class typhoid fever diagnosis. The proposed framework is particularly suitable for clinical environments with limited data access and constrained resources.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/biomedicines14051010","URL":"https://doi.org/10.3390/biomedicines14051010","source":"pubmed"},{"id":"doi:10.1002/advs.75906","type":"article-journal","title":"Machine Learning-Driven Prediction of Microplastic Aging Processes and Environmental Risk Assessment Across Multi-Media Systems.","abstract":"Machine learning (ML) holds promise for reconstructing microplastic (MP) aging and assessing risks, but current studies rely on small-scale, accelerated laboratory datasets and single environmental medium models that miss cross-media transport and environmental interactions in real-world MP lifecycles. To realize its potential for reconstructing spatiotemporal aging trajectories and toxicological assessment of MPs, this perspective provides a paradigm shift in ML application from fragmented data-fitting to a holistic, privacy-preserving, physics-aware strategy. A novel probabilistic framework reconstructs the environmental history of field-sampled MPs through mechanistic fingerprinting, using Bayesian inference to reconcile multi-evidence signals and improve trajectory models for source attribution and risk assessment. Furthermore, we propose the TRACE framework (TRansport, Aging, Corona, Ecotoxicity), which moves beyond the isolated modeling of aging processes and toxicity endpoints. By integrating physics-informed models with causal discovery, TRACE captures the reciprocal feedback loops between physicochemical evolution and eco-corona formation, thereby mechanistically linking surface transformations to biological risks. To support this data-intensive architecture, we advocate for federated learning (FL) to dismantle privacy barriers. This approach facilitates secure, multi-institutional collaborative modeling without raw data exchange, harmonizing heterogeneous datasets. Ultimately, this cohesive strategy bridges laboratory-field disparities, moving toward predictive, evidence-based, and targeted mitigation efforts in global plastic pollution governance.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1002/advs.75906","URL":"https://doi.org/10.1002/advs.75906","source":"pubmed"},{"id":"doi:10.1080/10255842.2026.2690183","type":"article-journal","title":"A systematic review of machine learning approaches for phonocardiogram classification.","abstract":"Recent advances in machine learning (ML) have accelerated automated analysis of phonocardiogram (PCG) signals, yet prior surveys often narrow their scope (e.g. omitting segmentation, focusing only on classical ML or overlooking recent deep learning [DL] trends) and provide limited methodological transparency. We present a comprehensive, Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA)-guided synthesis of PCG classification research published predominantly between 2021 and 2025, resulting in 151 studies included in the final synthesis. The systematic search was conducted across IEEE Xplore, PubMed/MEDLINE, Scopus, SpringerLink, ScienceDirect and Google Scholar, using Boolean combinations of heart sound/PCG-related terms and ML keywords. The review maps the full pipeline - acquisition, preprocessing, segmentation, feature extraction and classification - and distinguishes feature representations as single-independent-variable (SIV) and double-independent-variable (DIV). We compare classical classifiers with modern DL and hybrid architectures, and summarize public/proprietary datasets and evaluation practices. Our analysis shows that DL - especially convolutional neural network (CNN)-based approaches-dominates recent work, with growing interest in hybrid CNN-recurrent neural networks (RNNs) models and attention/transformer architectures. Segmentation is handled either explicitly (standalone or embedded) or bypassed in end-to-end designs. Despite high benchmark performance, comparability is hindered by heterogeneous datasets, non-uniform splits and metric choices; generalization under noise, device variability and pediatric vs. adult domain shift remains a key challenge. Bridging technical progress to clinical utility requires robustness, interpretability, efficient on-device inference and privacy-preserving training. We conclude with practical recommendations for standardized evaluation protocols, interpretable modeling, multimodal fusion, edge deployment and federated/multi-center learning. By integrating methodological and clinical perspectives, this review connects state-of-the-art results with the requirements of scalable, clinically viable PCG-based screening systems.","author":[{"family":"Er","given":"Sykes"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1080/10255842.2026.2690183","URL":"https://doi.org/10.1080/10255842.2026.2690183","source":"pubmed"},{"id":"doi:10.1109/tpami.2026.3713184","type":"article-journal","title":"BRFedTD: A Novel Framework for Federated Reinforcement Learning With Byzantine-Resilient Policy Evaluation.","abstract":"Federated reinforcement learning (FRL) enables distributed agents to collaboratively evaluate policies without sharing raw data, making it a promising approach for privacy-preserving and scalable decision-making. However, the presence of Byzantine agents, which may behave arbitrarily or maliciously, poses significant challenges to the robustness and reliability of FRL. To address this issue, in this paper, we propose a novel trimmed mean-based robust federated policy evaluation framework called Byzantine-resilient federated temporal difference learning (BRFedTD), and establish a finite-time convergence theory of BRFedTD. This framework effectively addresses the combined challenges of linear function approximation, heterogeneous Markov decision processes (MDPs), multiple local updates, and robust aggregation. To support more accurate confidence interval estimation and policy uncertainty analysis, we further derive the asymptotic distribution of the estimation error, showing that BRFedTD achieves asymptotic normality and efficiency in the case of identical MDPs without Byzantine attacks. This represents, to the best of our knowledge, the first asymptotic normality result established in FRL. Extensive numerical experiments demonstrate the robustness and effectiveness of the proposed algorithm, and corroborate that it generalizes to deep reinforcement learning and performs well on complex control tasks.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tpami.2026.3713184","URL":"https://doi.org/10.1109/tpami.2026.3713184","source":"pubmed"},{"id":"doi:10.3390/electronics14081674","type":"article-journal","title":"TeeDFuzzer: Fuzzing Trusted Execution Environment","abstract":"The Trusted Execution Environment (TEE) is crucial for safeguarding the ecosystem of embedded systems. It uses isolation to minimize the TCB (Trusted Computing Base) and protect sensitive software. It is vital because devices handle vast, potentially sensitive data. Leveraging ARM TrustZone, widely used in mobile and IoT for TEEs, it ensures hardware protection via security extensions, though needing firmware and software stack support. Despite the reputation of TEEs for high security, TrustZone-aided ones have vulnerabilities. Fuzzing, as a practical bug-finding technique, has seen limited research in the context of TEE. The unique software architecture of TrustZone-assisted TEE complicates the direct application of traditional fuzzing methods. Moreover, simplistic approaches, such as feeding random input values into TEE through the API functions of the rich operating system, fail to uncover deeper, latent bugs within the TEE code. In this paper, we present a fuzzing strategy for TrustZone-assisted TEE that utilizes inferred dependencies between Trusted Kernel system calls to uncover deep-seated TEE bugs. We implemented our approach on OP-TEE, where it successfully identified 17 crashes, including one previously undetected kernel bug.","author":[{"family":"Wen","given":"Sheng"},{"family":"Xu","given":"Liam"},{"family":"Tian","given":"Liwei"},{"family":"Liu","given":"Suping"},{"family":"Ding","given":"Yong"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3390/electronics14081674","URL":"https://doi.org/10.3390/electronics14081674","source":"crossref"},{"id":"doi:10.1038/s41598-026-48141-x","type":"article-journal","title":"SPHTRLM: secure and privacy-preserving hyperparameter-tuned reinforcement learning method for robot path finding in dynamic environments.","abstract":"Autonomous robot navigation within a dynamic environment is a complicated issue since environmental factors keep on changing, safety remains a factor, and issues of data privacy concern are also on the increase. The existing reinforcement learning (RL) navigation systems mainly focus on path performance and avoidance of collisions but do not focus on privacy protection, adaptation learning stability, and real deployment. This research aims to overcome these constraints by suggesting a novel framework Secure and Privacy-Preserving Hyperparameter-Tuned RL Model (SPHTRLM) to the efficient generation of path plans in grid ecosystems with dynamic environments. The framework incorporates adjusted Q-learning with federated learning (FL) based distributed updates, refined differentiated privacy, minimal encrypted parameter exchange, adaptive reward shaping and automatic hyperparameter optimization. In a further attempt to enhance practicability, the proposed architecture also embraces mobility conscious aggregation and heterogeneous model support of resource-limited robotic platforms. The suggested SPHTRLM has a success rate of (95% &#xb1; 2%), and it is better than the comparable one Q-learning (87% &#xb1; 4%) and Deep RL (DRL) baselines (88%) when these methods were evaluated under the same condition. The framework minimizes distances to the average path with a reduction of 20&#x2013;25% and convergence is speeded up by around 35% compared to normal Q-learning. When the obstacles are very thick then the collision rate becomes and the obstacle reduces to 0.08, and the safety of the navigation process improves. Although there are additional privatization mechanisms, the computational costs are minimal (8&#x2013;12%), and the average decision time is 110&#x2013;125 ms, which meets the real-time operational capabilities. Privacy analysis with formally stated membership inference and reconstruction attacks provide status of attack rate less than 5% attack success with both white and black box adversary. These findings underscore that SPHTRLM is a feasible way of achieving the goals of ensuring navigation, learning consistency, safety as well as privacy protection to give credible acceptance to using autonomous robotic systems in dynamic and data-sensitive environment.","author":[{"family":"Rr","given":"Dewangan"},{"family":"Bk","given":"Dewangan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-48141-x","URL":"https://doi.org/10.1038/s41598-026-48141-x","source":"pubmed"},{"id":"doi:10.1038/s41598-026-51804-4","type":"article-journal","title":"Multi-modal federated learning with differential privacy for privacy-preserving healthcare AI.","abstract":"The growing adoption of artificial intelligence in healthcare highlights the need for models that can leverage heterogeneous patient data while preserving strict privacy requirements. This paper proposes a novel multi-modal federated learning framework with differential privacy for decentralized healthcare AI. The model integrates electronic health records and ECG time-series using modality-specific encoders and a shared latent fusion network, enabling comprehensive representation learning without centralizing sensitive data. Differential privacy is incorporated into local updates to provide formal guarantees against information leakage in federated aggregation. Extensive experiments on real-world healthcare datasets show that the proposed method achieves [Formula: see text] accuracy, [Formula: see text] precision, [Formula: see text] recall, [Formula: see text] F1-score, and [Formula: see text] AUC, outperforming centralized, single-modality, and non-private baselines. The framework also converges [Formula: see text] faster than single-modality federated learning, reaching [Formula: see text] accuracy in 35 rounds. An ablation study confirms the contribution of multi-modal fusion and class balancing, while client variance analysis shows the lowest performance deviation ([Formula: see text]) under heterogeneous distributions. These results indicate that combining federated optimization, differential privacy, and multi-modal learning provides an effective framework for privacy-preserving clinical AI, with potential for deployment in distributed healthcare settings.","author":[{"family":"Mr","given":"Hasan"},{"family":"Mi","given":"Ahmed"},{"family":"Tk","given":"Ishika"},{"family":"Ha","given":"Shoaib"},{"family":"Mj","given":"Hossen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-51804-4","URL":"https://doi.org/10.1038/s41598-026-51804-4","source":"pubmed"},{"id":"doi:10.3390/jimaging12050205","type":"article-journal","title":"Federated Learning with Differential Privacy for Ultrasound Breast Cancer Classification: An Empirical Study.","abstract":"Breast cancer is a critical global health challenge, and deep learning shows transformative potential for medical image classification. However, privacy regulations such as HIPAA and GDPR create barriers to centralized data aggregation across institutions. This paper presents an empirical evaluation of federated learning (FL) for breast cancer classification in ultrasound images, systematically comparing seven deep learning architectures (ResNet-50, VGG16, VGG19, DenseNet-121, MobileNetV2, Vision Transformer, CoAtNet) across three FL algorithms (FedAvg, FedProx, FedOpt) with client-side differential privacy (DP). Using a simulated federation of eight institutions, we evaluate three clinically relevant classification scenarios. Federated models achieve performance comparable to centralized baselines-98.52% accuracy for normal/abnormal screening, 89.53% for three-class classification-with ViT-small and DenseNet-121 exceeding their centralized counterparts in several configurations. Under strong DP constraints (noise multiplier &#x3b7;=2.0, yielding conservative privacy budget estimates of &#x3b5;&lt;1.0 with &#x3b4;=10-5), screening accuracy remains above 82%, though diagnostic tasks incur substantial degradation (best 68.42%). Our findings provide empirical guidance on architecture selection, FL algorithm choice, and privacy-utility trade-offs for privacy-preserving breast cancer diagnosis, while identifying key challenges for clinical deployment.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/jimaging12050205","URL":"https://doi.org/10.3390/jimaging12050205","source":"pubmed"},{"id":"doi:10.1038/s41598-026-51535-6","type":"article-journal","title":"FedDriftGuard adaptive federated learning with differential privacy for concept drift in edge environments.","abstract":"Federated learning (FL) has become a highly promising paradigm for privacy-preserving distributed model training by enabling edge devices to train without sharing raw data. But in practice, edge environments are both non-stationary and asymmetric, with varying data distributions due to shifts in user behaviour, sensing conditions, and overall environmental dynamics. This causes concept drift (sudden, gradual, and recurrent), leading to poor model performance, slower convergence, and predictive bias. Current approaches to FL are not combined to tackle problems of drift adaptation, differential privacy (DP) and resource efficiency (FedAvg, DP-FedAvg). To address these constraints, we present FedDriftGuard. This Federated learning layer unifies client-level drift detection, drift-adaptive aggregation, and adaptable differential privacy into a single, FLE architecture-compatible system. The proposed DP-DriftNet model implements attention-based time encoding to capture changing data patterns and drift-directed feature weighting to allow greater flexibility in the presence of distributional changes. A drift-optimal privacy scheduler allocates noise probabilistically, subject to a limited privacy budget, thereby enforcing an appropriate privacy-utility trade-off without cancelling formal DP guarantees. Also, update sparsification, compression and periodic transmission techniques are used to reduce communication overhead. Decades of experimentation on real-world and synthetic drift datasets have shown that FedDriftGuard outperforms baseline FL techniques, achieving accuracy and F1-score gains of 9-14% and 11-17%, respectively, with adaptation latency 28% shorter and communication cost 20-35% lower. Such findings are statistically significant and confirm the soundness of the suggested method. FedDriftGuard offers effective, scalable privacy-preserving learning in adaptable, edge-drifting environments.","author":[{"family":"Hn","given":"Bhusarapu"},{"family":"Ts","given":"Sreenivas"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-51535-6","URL":"https://doi.org/10.1038/s41598-026-51535-6","source":"pubmed"},{"id":"doi:10.1038/s41598-026-45277-8","type":"article-journal","title":"TrustFed-RHIO: an optimization-driven differential privacy federated learning framework for secure and explainable IIoT attack detection.","abstract":"A rapid proliferation of industrial internet of things (IIoT) systems has increased the vulnerability of interconnected devices for sophisticated cyberattacks, which necessitates intelligent and privacy-preserving solution for security. This paper presents TrustFed-RHIO, a novel hybrid model which integrates rock hyrax intelligence optimization (RHIO) for optimal selection of feature with a trustworthy differential privacy-enhanced federated learning (TrustFed) scheme for the collaborative detection of attack. The algorithm of RHIO mimics collective intelligence of rock hyrax colonies for identifying most discriminative features, thus reducing dimensionality and enhancing efficiency of classifier. The proposed TrustFed-RHIO scheme ensures, secure, distributed learning over multiple IIoT nodes on embedding differential privacy mechanisms thus mitigating the data leakage risks and adversarial inference. A suggested scheme is thus empowered with explainable artificial intelligence (XAI) scheme termed SHAP and LIME for enhancing interpretability and trust in model predictions. Experimental estimation on benchmark IIoT dataset shows that TrustFed-RHIO attains superior performance on detection accuracy, robustness against adversarial attacks, and privacy preservation on comparing existing schemes. At last, this framework supports secure storage of cloud on detection outcomes, thus enabling scalable deployment in the real-world IIoT framework. The performance evaluation is carried on benchmark dataset CCIoT2024-DIAD and performance is estimated for various metrics like latency, accuracy, precision, recall, F1-score, specificity, training time, and so on. Overall analysis shows that the proposed model is effective in detecting IIoT attacks.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-45277-8","URL":"https://doi.org/10.1038/s41598-026-45277-8","source":"pubmed"},{"id":"doi:10.1049/htl2.70080","type":"article-journal","title":"Federated Learning for Thoracic Disease Classification Using Convolutional Neural Networks and Differential Privacy.","abstract":"Early diagnosis of thoracic diseases using chest x-ray imaging remains a critical challenge, particularly in resource-constrained healthcare environments where data sharing is restricted due to privacy concerns. Federated learning (FL) offers a decentralized solution by enabling collaborative model training without sharing sensitive patient data. However, integrating privacy-preserving mechanisms such as differential privacy (DP) introduces additional challenges related to performance degradation and computational overhead. In this study, we present a unified FL framework for multi-label thoracic disease classification using multiple convolutional neural network (CNN) architectures, including ResNet50, DenseNet169, EfficientNet variants and MobileNetV3. Unlike prior studies focusing on single-model evaluation, this work provides a controlled comparative analysis under identical FL settings and investigates the impact of client scalability (5-10 clients) on model performance. Furthermore, we conduct a comprehensive empirical analysis of the privacy utility trade-off by integrating DP with varying privacy budgets ( &#x3b5; &#xa0;=&#xa0;1, 15 and 30). Experimental results on the CheXpert and NIH Chest x-ray14 datasets demonstrate that the proposed EfficientNet-B3-based federated model achieves a mean AUC of 0.8027, while maintaining robustness across decentralized settings. The integration of DP leads to a predictable reduction in performance, with mean AUC ranging from 0.60 to 0.64, highlighting the inherent trade-off between privacy and diagnostic accuracy. The findings emphasize the practical viability of FL for privacy-sensitive medical imaging applications and provide insights into model selection, scalability and privacy configuration for real-world deployment. The source code for this study is publicly accessible at https://github.com/Zulqarnain8-8/FEDERATED_LEARNING_FOR_THORACIC_DISEASE_CLASSIFICATION.","author":[{"family":"Sj","given":"Hussain"},{"family":"Mz","given":"Aslam"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1049/htl2.70080","URL":"https://doi.org/10.1049/htl2.70080","source":"pubmed"},{"id":"doi:10.1049/htl2.70079","type":"article-journal","title":"APB-FLDPA: Adaptive Personalized Blockchain-Federated Learning With Differential Privacy and Attention for Privacy-Preserving Healthcare Analytics.","abstract":"Developing robust medical artificial intelligence (AI) requires collaboration across multiple institutions, but strict data protection regulations such as HIPAA and GDPR prevent centralized patient data sharing. Existing federated learning (FL) methods often exhibit 15%-30% performance degradation in real-world clinical settings due to data heterogeneity, security threats, and privacy constraints. We present APB-FLDPA, a privacy-preserving federated learning framework for secure multi-hospital disease prediction. APB-FLDPA integrates five key innovations: (i) adaptive Byzantine-resilient aggregation using dynamic client trust scoring, (ii) self-attention for automated clinical feature importance, (iii) selective differential privacy applied at the final aggregation stage, (iv) cluster-aware personalization to handle cross-institutional heterogeneity, and (v) a lightweight blockchain module to ensure model integrity. Evaluated across five institutions using large-scale Diabetes (183,000 patients) and Thyroid (6840 patients) datasets, APB-FLDPA achieved 90.8% accuracy for diabetes and 83.8% accuracy for thyroid disease, with minimal performance loss (&lt;0.2%) compared to centralized learning. Statistical tests confirmed significant improvements, and selective differential privacy outperformed conventional methods by 5.6% in accuracy. These results show that APB-FLDPA provides a scalable, high-performance and privacy-compliant solution for real-world federated medical&#xa0;AI.","author":[{"family":"Mkh","given":"Chowdhury"},{"family":"Pk","given":"Mondal"},{"family":"Mai","given":"Mozumder"},{"family":"Hc","given":"Kim"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1049/htl2.70079","URL":"https://doi.org/10.1049/htl2.70079","source":"pubmed"},{"id":"doi:10.2139/ssrn.6218028","type":"manuscript","title":"Federated-survival: A Federated Learning Framework for Privacy-Preserving Survival Analysis","abstract":"Federated Learning (FL) has emerged as a promising paradigm for collaborative, privacy-conscious model training; however, its application to survival analysis remains at an early stage of development. A significant barrier to progress in this domain is the absence of standardized benchmarking tools, making it difficult to compare methods, validate innovations, and establish best practices. To address this gap, we introduce \\texttt{federated-survival}, a comprehensive, open-source Python framework designed for privacy-preserving survival analysis. This framework offers a systematic integration of seven mainstream survival methodologies, encompassing both classical statistical approaches, such as the Cox Proportional Hazards model, and state-of-the-art deep learning architectures, including DeepHit and PC-Hazard. Recognizing the unique complexities of survival data in distributed settings, \\texttt{federated-survival} implements tailored data augmentation techniques to enhance model robustness and generalizability. Furthermore, to ensure rigorous privacy protection, the proposed framework integrates multiple differential privacy mechanisms, allowing users to navigate the critical trade-off between privacy guarantees and model utility with fine-grained control.The \\texttt{federated-survival} package is freely available at\\url{https://pypi.org/project/federated-survival/} along with installation instructions.","author":[{"family":"Wang","given":"Wenjun"},{"family":"Hong","given":"Wang"},{"family":"Zhang","given":"Zhuan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.2139/ssrn.6218028","URL":"https://doi.org/10.2139/ssrn.6218028","source":"crossref"},{"id":"doi:10.1371/journal.pcbi.1014530","type":"article-journal","title":"Eleven quick tips for Biomedical Federated Learning.","abstract":"Modern statistical and machine learning techniques are effective at describing, testing hypotheses and making predictions from complex data. This effectiveness is strongly influenced by the volume and heterogeneity of available data. In many fields, including much of biomedicine, large centralized datasets are not available because of cost, privacy, regulatory or other restrictions. In these cases, smaller datasets are distributed across a large number of independent sites. Medical record data is a classic example of this challenge: the total number of patients may be large, but their records are distributed across many health systems and cannot easily be centralized. Federated learning (FL) is a machine learning paradigm that enables training and validation of a shared model in settings of decentralized data. FL can improve model accuracy and generalizability by increasing sample size, but has trade-offs ranging from operational complexity to data-privacy risks to the potential to introduce unexpected imbalances in model accuracy. We outline ten tips for successfully and sustainably implementing FL for Biomedical applications, ensuring both ethical data governance and improved model performance in sensitive domains.","author":[{"family":"Vs","given":"Malladi"},{"family":"Jc","given":"Bélisle"},{"family":"Aat","given":"Bui"},{"family":"Pc","given":"Boutros"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1371/journal.pcbi.1014530","URL":"https://doi.org/10.1371/journal.pcbi.1014530","source":"pubmed"},{"id":"doi:10.1109/tpami.2026.3722165","type":"article-journal","title":"Decentralized Federated Learning by Partial Message Exchange.","abstract":"Decentralized federated learning (DFL) has emerged as a transformative server-free paradigm that enables collaborative learning over large-scale heterogeneous networks. However, it continues to face fundamental challenges, including data heterogeneity, restrictive assumptions for theoretical analysis, and de graded convergence when standard communication- or privacy enhancing techniques are applied. To overcome these drawbacks, this paper develops a novel algorithm, PaME (DFL by Partial Message Exchange). The central principle is to allow only randomly selected sparse coordinates to be exchanged between two neighbor nodes. As a result, PaME significantly reduces communication costs while simultaneously limiting the exposure of data-sensitive information during transmission. The latter property is rigorously characterized by a formal reconstruction risk theory under partial observation. Moreover, the algorithm is proven to converge in expectation to a stationary point at a linear rate, provided that the gradient is locally Lipschitz continuous and the communication matrix is doubly stochastic. These two mild assumptions not only dispenses with many restrictive conditions commonly imposed by existing DFL methods but also enables PaME to effectively address data heterogeneity. Furthermore, comprehensive numerical experiments demonstrate its superior performance compared with several representative decentralized learning algorithms.","author":[{"family":"Gy","given":"Li"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tpami.2026.3722165","URL":"https://doi.org/10.1109/tpami.2026.3722165","source":"pubmed"},{"id":"doi:10.20944/preprints202608.1914.v1","type":"manuscript","title":"Differential Privacy Synthetic Tabular Data Generation for Federated Learning","abstract":"Machine learning in healthcare struggles for one main reason: good data is hard to come by. Medical records are sensitive, and the rules on sharing them between hospitals are strict. A common workaround is synthetic data, that is artificial records that copy the statistics of real ones; however, on its own, it offers no real privacy guarantee, and earlier studies show it can still leak details about the patients behind it. We introduce an approach tackling both problems at once. It builds on a previous UMAP-based generator methodology and extends it so several hospitals can work together without sharing raw records, in a federated learning form. Then, the proposed methodology adds differential privacy so every shared quantity carries a formal (εDP,δ) guarantee, tracked by a custom privacy accountant. Two versions of the methodology are tested: a partially synthetic one, where a reference hospital validates the others’ rows, and a fully synthetic one, where each centre builds its own data from aggregated cluster statistics. Both are evaluated on four publicly available healthcare tabular datasets spanning prostate cancer (PI-CAI, CIA), breast cancer (BC-MLR), and cardiovascular disease (fiv-CardioDB), to test whether the observed trends generalize beyond a single clinical domain. As a key takeaway, it is demonstrated that going federated barely affects quality, but the privacy cost depends strongly on the moment in the procedure when the noise is injected, and this pattern holds consistently across all four datasets.","author":[{"family":"Almató-Baucells","given":"Mariona"},{"family":"Lázaro","given":"Carla"},{"family":"Angulo","given":"Cecilio"}],"issued":{"date-parts":[[2026]]},"DOI":"10.20944/preprints202608.1914.v1","URL":"https://doi.org/10.20944/preprints202608.1914.v1","source":"europepmc"},{"id":"doi:10.3390/s26165297","type":"article-journal","title":"Privacy-Preserving and Poisoning-Robust Federated Learning for Industrial IoT.","abstract":"With the rapid development of Industrial Internet of Things (IIoT), large amounts of sensor data generated by industrial devices and edge nodes have become the basis of intelligent manufacturing applications. Collaborative modeling on these distributed data is important for tasks such as anomaly detection, equipment monitoring, and predictive maintenance. Federated learning offers a practical way to train models without exposing raw sensor data, but it still faces privacy leakage and malicious poisoning attacks. To address these issues, this paper proposes a hierarchical privacy protection and poisoning-robust defense framework for industrial federated learning. Starting from the sensitivity differences among parameters at different model layers, the proposed method designs a hierarchical privacy-budget allocation strategy that enhances protection for sensitive information while minimizing the performance impact of perturbation. Meanwhile, a multi-layer, multi-feature anomaly-detection mechanism is adopted to identify malicious updates by jointly exploiting directional consistency, scale stability, and inter-layer similarity, and majority voting together with update clipping is used to further improve system robustness. Experiments on Fashion-MNIST, MVTec AD, and C-MAPSS demonstrate that the proposed method can effectively suppress global-model degradation under multiple poisoning attacks and achieves a favorable balance among privacy protection strength, robustness, and training efficiency.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26165297","URL":"https://doi.org/10.3390/s26165297","source":"pubmed"},{"id":"doi:10.1016/j.neunet.2026.109519","type":"article-journal","title":"Multi-FedLinks: Highly reliable decentralized multibiometric federated learning links.","abstract":"In recent years, the recognition accuracies of deep learning-based biometric recognition methods, which rely on large amounts of biometric data for training, have significantly increased. However, in practical applications, biometric data are often distributed in small and fragmented amounts among various local clients. Implementing distributed biometric recognition is therefore greatly important. Most existing distributed biometric methods are implemented by federated learning, and these methods suffer from two problems. (1) The current methods are overwhelmingly limited to addressing distributed single-biometric recognition problems and are not applicable to distributed multibiometric recognition. (2) The conversion from traditional local learning to distributed learning with multiterminal cooperation poses a series of security hazards that have not been addressed. To address these issues, a decentralized multibiometric federated learning links (Multi-FedLinks) model for distributed multibiometric recognition is proposed in this paper. The model consists of multiple FedLink structures, which are resistant to Byzantine attacks. Collaboration among the multiple FedLink structures is implemented with a third-party server to achieve multibiometric federated learning. Experimental results on the NUPT-FPV dataset demonstrate that the superiority of Multi-FedLinks model. Our code can be found in https://github.com/HYMu99/Multi-FedLinks.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.neunet.2026.109519","URL":"https://doi.org/10.1016/j.neunet.2026.109519","source":"pubmed"},{"id":"doi:10.3390/s26165059","type":"article-journal","title":"Optimal Transport-Based Heterogeneous Federated Learning for Chest X-Rays.","abstract":"In some federated learning (FL) scenarios, discrepancies in local client devices result in inconsistent image resolutions, which motivates clients to adopt models with different depths and widths. Existing heterogeneous federated learning methods struggle to maintain model accuracy while preserving computational efficiency. To tackle this issue, this paper proposes a heterogeneous federated learning framework based on optimal transport (OT) and cross-layer alignment. The framework addresses the inconsistency of model depth via cross-layer alignment, fuses parameters of layers with different widths using optimal transport, and develops an aggregation strategy for multiple heterogeneous models. Experiments demonstrate that our method can improve model accuracy by up to 1.65% while maintaining satisfactory efficiency.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26165059","URL":"https://doi.org/10.3390/s26165059","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-9613468/v1","type":"article-journal","title":"A Lightweight Anonymous Authentication Scheme for Federated Learning","abstract":"Abstract Federated learning enables collaborative model training between central servers and distributed clients without collecting users’ raw sensitive data, which effectively promotes the large-scale deployment of intelligent collaborative services. Considering the high sensitivity of local training data and model gradient parameters in federated learning, protecting identity privacy and interaction security has become extremely critical. Therefore, mutual identity authentication is indispensable to restrict illegal client access and prevent malicious parameter transmission and data tampering. In this paper, we propose a lightweight anonymous authentication scheme for federated learning (FedLAS), which realizes secure mutual authentication between servers and clients and establishes a shared session key for subsequent encrypted interaction. In particular, the proposed scheme eliminates the reliance on high-cost cryptographic operations such as bilinear pairing, thus minimizing computational and communication overhead. Furthermore, informal security analysis demonstrates that our FedLAS scheme can resist multiple common attacks and meet predefined security requirements. Extensive comparative experiments show that the FedLAS scheme achieves excellent performance in computational and communication cost. It is well suitable for resource-constrained federated learning scenarios.","author":[{"family":"Wu","given":"Shu"},{"family":"Meng","given":"Guoqiang"},{"family":"Lu","given":"Linlin"},{"family":"Dong","given":"Xiaojuan"},{"family":"Tian","given":"Sai"},{"family":"Chen","given":"Jindou"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9613468/v1","URL":"https://doi.org/10.21203/rs.3.rs-9613468/v1","source":"europepmc"},{"id":"doi:10.1109/tmi.2026.3725153","type":"article-journal","title":"ToPPFed: Topological Prototype-Enhanced Personalized Federated Learning for Neuropsychiatric Disorders Identification.","abstract":"Functional connectivity networks (FCNs) derived from functional magnetic resonance imaging (fMRI) have been widely used to characterize topological alterations of brain networks in neuropsychiatric disorders (NDs). Given the frequent restrictions on direct multi-site fMRI data sharing, federated learning (FL) offers a collaborative modeling paradigm without exchanging raw neuroimaging data. However, conventional parameter-averaging FL approaches struggle under cross-site non-IID distributions. Prototype-based FL provides a promising alternative, yet existing designs implicitly rely on spatially structured image data and fail to capture the topology-centric semantics of FCNs. To bridge this gap, we propose ToPPFed, a Topological Prototype-Enhanced Personalized Federated Learning framework for multi-site classification between subjects with each studied disorder and normal controls (NCs). ToPPFed introduces a Graph Topological Prototype Learning module to extract discriminative topology-aware prototypes from FCNs and a Contrastive Mask-Induced Residual Scaling mechanism to adaptively integrate group-level priors into individual representations. By exchanging topology prototypes instead of raw data or full model parameters, ToPPFed supports cross-site collaboration while reducing direct data exposure. Experiments on multi-site fMRI datasets of three representative NDs show that ToPPFed improves accuracy (ACC) by 1.7-7.9 percentage points over the best-performing federated baseline on each dataset. Interpretability analyses indicate that ToPPFed highlights model-derived discriminative brain regions and functional connections. The topology-aware exchange of node and edge prototypes offers an effective framework for collaborative FCN modeling across imaging sites without centralizing neuroimaging data.","author":[{"family":"Km","given":"Kendrick"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tmi.2026.3725153","URL":"https://doi.org/10.1109/tmi.2026.3725153","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-10801503/v1","type":"article-journal","title":"Federated Learning Framework with Differential Privacy over Homomorphic Vector Encryption for Data-Sensitive Applications","abstract":"Abstract Federated Learning (FL) has become a popular paradigm in recent years, attributed to its concept of insight sharing for ensuring data privacy. It has been found to be of extensive utility in applications that regularly deal with confiden-tial data and has been widely used in the fields of healthcare and the Internet of Medical Things (IoMT). However, an equal amount of research is conducted to reverse-engineer the used datasets from the shared insights, and a native FL implementation alone is not suitable for these sensitive applications. We pro-pose a lightweight secure FL framework incorporating both Differential Privacy and Homomorphic Encryption at the vector level to minimize the efficiency of reverse-engineering algorithms over the transmitted insights in IoMT and other resource-constrained networks. Through experiments on multiple image classi-fication datasets and comparison with secure federated learning baselines, the proposed framework demonstrates the feasibility of combining differential pri-vacy with encrypted aggregation while quantifying the resulting predictive and cryptographic overhead.","author":[{"family":"Narula","given":"Manu"},{"family":"Meena","given":"Jasraj"},{"family":"Vishwakarma","given":"Dinesh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-10801503/v1","URL":"https://doi.org/10.21203/rs.3.rs-10801503/v1","source":"europepmc"},{"id":"doi:10.3390/bioengineering13060603","type":"article-journal","title":"A Method for Workout Video Classification via Explainable and Federated Learning.","abstract":"In recent years, the widespread availability of wearable devices and smartphones has enabled the large-scale collection of human activity data, fostering new opportunities for automatic workout recognition and personalized fitness monitoring. However, the centralized storage of video recordings raises critical privacy concerns, particularly when raw data contain identifiable individuals. Federated Machine Learning provides a paradigm designed with the aim of reducing privacy risks; here, models are collaboratively trained across distributed clients without sharing their sensitive data. In this paper, we propose an approach for workout video classification with Federated Machine Learning, enhanced by explainability through Gradient-weighted Class-Activation Mapping. The proposed method is evaluated on a real-world multi-class exercise video dataset, organized into eight biomechanically coherent macro-classes. In the experimental analysis, we consider several federated configurations in terms of the number of clients, the chosen aggregation strategy, and global communication rounds. The obtained results demonstrate that different aggregation strategies achieve comparable overall accuracy, while explainability effectively highlights the discriminative regions associated with exercise execution, revealing meaningful differences in model behavior between aggregation strategies and uncovering misclassifications driven by contextual biases, demonstrating the trustworthiness of the proposed approach for explainable workout video classification.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/bioengineering13060603","URL":"https://doi.org/10.3390/bioengineering13060603","source":"pubmed"},{"id":"doi:10.3389/fmed.2026.1904004","type":"article-journal","title":"Federated learning for privacy-preserving ophthalmic artificial intelligence: clinical applications and translational challenges.","abstract":"Federated learning (FL) is increasingly relevant to ophthalmology because retinal photographs, optical coherence tomography (OCT), OCT angiography, visual fields, and linked clinical records are clinically valuable but difficult to pool across institutions. In this narrative review, we synthesize ophthalmology-focused FL literature across diabetic retinopathy (DR), glaucoma, age-related macular degeneration (AMD), pediatric retinal disease, multi-disease retinal diagnostics, and emerging ophthalmic platforms. Current evidence suggests that FL can support collaborative AI development without centralizing raw patient data, and selected studies show performance close to centralized training under controlled retrospective or multicenter experimental conditions. For example, multicenter glaucoma detection from volumetric OCT achieved an AUC of 0.92 with FL compared with 0.94 for centralized training. However, FL is privacy-enhancing rather than privacy-complete, and most ophthalmic FL systems have not yet undergone prospective clinical validation. Model updates may remain vulnerable to gradient inversion, membership inference, poisoning, site-level bias, and latent identity or attribute leakage. For eye-care networks, the main value of FL is therefore not simply algorithmic performance but a governance model for privacy-conscious collaboration. Prospective validation, interoperability, explainability, workflow integration, privacy auditing, and clear responsibility for monitoring are needed before FL-enabled ophthalmic AI can be deployed routinely.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/fmed.2026.1904004","URL":"https://doi.org/10.3389/fmed.2026.1904004","source":"pubmed"},{"id":"doi:10.1016/j.neunet.2026.109470","type":"article-journal","title":"MoFedAGR: Mitigating client drift with adaptive gradient regularization and global momentum in federated learning.","abstract":"Federated learning is a novel distributed machine learning framework with privacy-protection, yet it is vulnerable to the effects of heterogeneous data. Heterogeneous data drive client models that overfit local datasets and depart from the global optimum during local training, which is termed client drift. To address the impact of client drift, we approach this issue from the perspectives of optimization and generalization. We comprehensively considering the effects of client drift during the training process, and quantifying it as the aggregation error. We first propose adaptive gradient regularization, which is based on gradient regularization and further and applies different regularization strengths to each parameter based on the magnitude of the parameter variance between the local model and the global model, thereby mitigating the performance degradation caused by aggregation error and helping model converge to a flatter minimum. In order to obtain the variance between local and global models to compute our adaptive gradient regularization term, we introduce global momentum from the server side as the approximation of global gradient and further utilize it as a gradient correction term. Next, we propose MoFedAGR, which combines gradient correction term and adaptive gradient regularization term, helping client models converge to a consistent flat minimum. We have provided the theoretical convergence bounds of the algorithm we proposed. Furthermore, experiments on several image classification datasets demonstrate that our algorithm significantly improves model performance while exhibiting strong generalization capabilities.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.neunet.2026.109470","URL":"https://doi.org/10.1016/j.neunet.2026.109470","source":"pubmed"},{"id":"doi:10.3390/s26154712","type":"article-journal","title":"DBST-FL: Dynamic Behavioural and Semantic Trust for Robust Federated Learning in Industrial IoT.","abstract":"Federated learning (FL) has emerged as an effective paradigm for collaborative model training in Industrial Internet of Things (IIoT) environments by enabling distributed devices to learn shared models without exchanging raw data. However, existing FL defence mechanisms predominantly rely on either behavioural analysis of client updates or semantic validation of model performance, limiting their ability to detect sophisticated poisoning and stealthy backdoor attacks that evade single-dimensional trust assessment. This paper proposes DBST-FL, a dynamic behavioural and semantic trust framework for robust federated learning in the Industrial IoT. The proposed framework evaluates each client through two complementary trust dimensions: a behavioural trust layer that measures gradient alignment, historical consistency, and collective deviation and a semantic trust layer that assesses benign utility and template-free semantic stress validation using server-side data. The two trust scores are integrated through a non-compensatory multiplicative trust fusion mechanism, ensuring that weaknesses in one trust dimension cannot be masked by strengths in the other. The resulting trust score guides a trust-aware aggregation strategy that reduces the influence of malicious participants while preserving the contributions of reliable clients. Extensive experiments are conducted on the Edge-IIoTset and UNSW-NB15 datasets using ANN, 1D-CNN, and LSTM models under multiple poisoning and backdoor attack scenarios. The proposed framework achieves overall classification performance competitive with the strongest robust aggregation baselines while consistently delivering stronger resilience against adversarial attacks and lower backdoor attack success rates than representative trust-based and Byzantine-robust aggregation methods, all while maintaining linear per-round computational complexity suitable for large-scale IIoT deployments. The results demonstrate that integrating behavioural and semantic trust within a unified aggregation framework provides an effective and scalable defence against advanced adversarial threats in federated learning.","author":[{"family":"Ak","given":"Tom"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26154712","URL":"https://doi.org/10.3390/s26154712","source":"pubmed"},{"id":"doi:10.3390/e28060630","type":"article-journal","title":"SCKM: Symmetric Co-Skew Moment for User Selection in Federated Learning.","abstract":"We introduce the symmetric co-skewness moment (SCKM)-a third-order informational dissimilarity metric that consistently outperforms state-of-the-art client-selection heuristics in federated learning (FL) under heterogeneous data. Unlike similarity-driven schemes, SCKM minimizes redundancy by favoring clients with complementary gradients, delivering faster and more stable convergence even at high heterogeneity levels. Operating on highly compressed 0.5% gradient summaries, our framework provides two operating modes for different deployment scales: (i) SCKM-Select directly ranks and schedules a small candidate pool, whereas (ii) SCKM-Cluster adds a fast, elbow-guided clustering step to scalably choose from thousands of users. We evaluate both variants on a VGG-16 model across multiple non-IID partition schemes and initializations, observing consistent gains over leading cosine-similarity, loss-sketch, and max-diversity baselines-without increasing the communication budget.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/e28060630","URL":"https://doi.org/10.3390/e28060630","source":"pubmed"},{"id":"doi:10.1109/tnnls.2026.3718132","type":"article-journal","title":"Enhanced Spectral Clustering Robust Aggregation for Lens Detection in Federated Learning Against Byzantine Attacks.","abstract":"Although federated learning (FL) addresses the issues of centralized data storage and privacy leakage, its distributed nature makes it vulnerable to malicious clients. These malicious participants introduce malicious parameters during the training procedure, which can significantly impact the model's accuracy. Existing algorithms such as Krum and median defend against attacks by capturing low-order features of data and requiring prior data. However, these approaches struggle to counteract gradually evolving Byzantine attacks. Therefore, we propose an unsupervised approach based on an enhanced spectral clustering algorithm (SCA) to identify malicious updates. First, we design a method for constructing an undirected weighted graph using the Gaussian kernel function. This approach maps features among data into an infinite-dimensional Hilbert space, enabling better capture of high-order data features. Second, due to the similarity between Byzantine updates and benign updates in Euclidean space and cosine similarity scenarios, traditional robust aggregation algorithms fail to recognize them, causing models to fail to converge. To address this, a new lens detection method is designed. We calculate the Laplacian matrix through the undirected weighted graph and employ normalized cut (NCut) to partition the Laplacian matrix. This transforms the problem of identifying malicious clients into a graph partitioning problem. Furthermore, we conduct a convergence analysis of the proposed SCA. Experiments on four datasets under independent and identically distributed (IID) and non-IID (Non-IID) settings show this method has strong robustness to all tested attacks, while other defense methods cannot resist all of them.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/tnnls.2026.3718132","URL":"https://doi.org/10.1109/tnnls.2026.3718132","source":"pubmed"},{"id":"doi:10.3389/frai.2026.1895239","type":"article-journal","title":"GA-AFedOD: gradient-aligned active federated learning for resource-aware object detection in edge industrial IoT.","abstract":"Visual object detection is essential for defect inspection and process monitoring in edge-deployed Industrial Internet of Things (IIoT). Yet, training accurate detectors across distributed factories faces stringent constraints on data privacy, annotation budgets, and uplink communication. Standard federated learning (FL) preserves locality but often wastes labeling resources on redundant frames and overlooks detection-specific gradient alignment when scheduling clients. To bridge this gap, we propose Gradient-Aligned Active Federated Object Detection (GA-AFedOD), a unified framework that jointly optimizes annotation selection, client participation, and model aggregation as a constrained stochastic program. A novel utility metric integrates box-level uncertainty, prototype diversity, gradient alignment, and resource pricing, enabling edge clients to perform locally guided active querying while the server solves a lightweight primal-dual problem for budget-aware client scheduling. We prove a submodular approximation guarantee for the greedy sampling rule and establish a non-convex convergence bound that explicitly captures the impact of label budgets, client drift, and compression noise. This article further clarifies the relationship with recent federated active learning and industrial detection studies, adds parameter and theory-diagnostic analyses, and distinguishes controlled simulation evidence from real-world deployment validation on industrial datasets such as RasPiDets, Electric Power Fitting Dataset (EPFD), and Diverse Insulator Dataset (DINS). Controlled simulation results show that GA-AFedOD achieves considerably higher mean average precision (mAP) while reducing both annotation costs and uplink consumption by over 40% compared with competitive baselines.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/frai.2026.1895239","URL":"https://doi.org/10.3389/frai.2026.1895239","source":"pubmed"},{"id":"doi:10.20944/preprints202605.1674.v1","type":"manuscript","title":"Poisoning Attacks in Federated Learning: An Accountability-Oriented Survey with Centralized Learning as Baseline","abstract":"Artificial intelligence (AI) systems are increasingly deployed in critical domains such as healthcare, finance, defense, and transportation. These deployments, however, face grow- ing risks from poisoning attacks that corrupt training data, manipulate model updates, or implant covert backdoors. Such attacks undermine trust, reduce transparency, and challenge the safe and accountable use of AI in high-stakes settings. This survey examines poisoning attacks in federated learning (FL), using centralized learning as a comparative baseline to clarify how the threat landscape changes when data, updates, and control are distributed. Rather than reintroducing a generic poisoning taxonomy as a standalone contri- bution, we position the paper relative to prior surveys, identify what remains insufficiently covered, and synthesize representative primary studies through an accountability-oriented lens. Because FL introduces additional vulnerabilities, including untrusted servers, non-IID heterogeneity, and limited observability of client behavior, we analyze how these properties expand the poisoning threat surface. We review state-of-the-art countermeasures, including Byzantine-robust aggregation, anomaly detection, validation-based defenses, and cryp- tographic prevention mechanisms such as malicious-secure aggregation, authenticated update handling, and verifiable aggregation protocols. Particular emphasis is placed on accountability-enabling mechanisms such as auditability, traceability, and forensic readi- ness, which are essential to responsible AI and regulatory compliance. Our analysis identi- fies persistent research gaps, including the lack of unified privacy-robustness-accountability frameworks, insufficient defenses against server-side attacks, limited verifiability tools, and scalability challenges in real-world FL systems. The paper’s contribution is therefore not to claim that accountability-oriented FL defenses are new, but to clarify the survey scope, position prior surveys directly, and integrate aggregation-based, cryptographic, and governance-oriented strands within an Accountability-Integrated Taxonomy (AIT) and an evidence-oriented discussion of trustworthy federated learning.","author":[{"family":"Mohammed","given":"Safiia"},{"family":"Alhadidi","given":"Dima"},{"family":"Ngom","given":"Alioune"}],"issued":{"date-parts":[[2026]]},"DOI":"10.20944/preprints202605.1674.v1","URL":"https://doi.org/10.20944/preprints202605.1674.v1","source":"europepmc"},{"id":"doi:10.1007/s41666-025-00226-4","type":"article-journal","title":"Multimodal Federated Learning in Healthcare: A Review.","abstract":"Recent advancements in multimodal machine learning have empowered the development of accurate and robust AI systems in the medical domain, especially within centralized database systems. Simultaneously, Federated Learning (FL) has progressed, providing a decentralized mechanism where data need not be consolidated, thereby enhancing the privacy and security of sensitive healthcare data. The integration of these two concepts supports the ongoing progress of multimodal learning in healthcare while ensuring the security and privacy of patient records within local data-holding agencies. This paper offers a concise overview of the significance of FL in healthcare and outlines the current state-of-the-art approaches to Multimodal Federated Learning (MMFL) within the healthcare domain. It comprehensively examines the existing challenges in the field, shedding light on the limitations of present models. Finally, the paper outlines potential directions for future advancements in the field, aiming to bridge the gap between cutting-edge AI technology and the imperative need for patient data privacy in healthcare applications.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1007/s41666-025-00226-4","URL":"https://doi.org/10.1007/s41666-025-00226-4","source":"pubmed"},{"id":"doi:10.1016/j.cmpb.2026.109454","type":"article-journal","title":"Federated learning: A new frontier in the exploration of multi-institutional medical imaging data.","abstract":"Artificial intelligence has transformed the perspective of medical imaging, leading to a genuine technological revolution in modern computer-assisted healthcare systems. However, ubiquitously featured deep learning (DL) systems require access to a considerable amount of data, facilitating proper knowledge extraction and generalization. Access to such extensive resources may be hindered due to the time and effort required to convey ethical agreements, set up and carry the acquisition procedures through, and manage the datasets adequately with a particular emphasis on proper anonymization. One of the pivotal challenges in the DL field is data integration from various sources acquired using different hardware vendors, diverse acquisition protocols, experimental setups, and even inter-operator variabilities. In this paper, we review the federated learning (FL) concept that fosters the integration of large-scale heterogeneous datasets from multiple institutions in training DL models. In contrast to a centralized approach, the decentralized FL procedure promotes training DL models while preserving data privacy at each institution involved. We formulate the FL principle and comprehensively review general and specialized medical imaging aggregation and learning algorithms, enabling the generation of a globally generalized model. We meticulously go through the challenges in constructing FL-based systems, such as data and model heterogeneities across the institutions, resilience to potential attacks on data privacy, and the variability in computational and communication resources among the entangled sites that might induce efficiency issues of the entire system. Finally, we explore the up-to-date open frameworks for rapid FL-based algorithm prototyping, comprehensively present real-world implementations of FL systems and shed light on future directions in this intensively growing field.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.cmpb.2026.109454","URL":"https://doi.org/10.1016/j.cmpb.2026.109454","source":"pubmed"},{"id":"doi:10.1109/jbhi.2026.3702502","type":"article-journal","title":"FedDPI-SH: A Quality and Similarity Aware Federated Learning Framework for Medical Image Analysis.","abstract":"Federated learning (FL) enables decentralized medical image analysis while preserving data privacy. However, conventional methods overlook client data heterogeneity and inter-client feature similarity, resulting in suboptimal performance. In this paper, our proposed FedDPI-SH framework addresses these limitations through quality and similarity-aware weighted aggregation. The framework introduces a Data Performance Index (DPI) quantifying client reliability through dataset size, image resolution, label distribution balance, duplication levels, and cross-client generalization accuracy. Client Representation Similarity Matrix (CRSM) measures inter-client feature alignment via cosine similarity. FedDPI-SH combines these components to compute aggregation weights, prioritizing high-quality clients during feature extractor updates while maintaining personalized classifiers. Evaluation across three medical imaging datasets (Medical MNIST, PathMNIST, HAM10000) under severe non-IID conditions demonstrates improvements, with HAM10000 achieving 81.24% balanced accuracy, outperforming MOON (58.85%), FedAvg (45.22%), FedAvgM (17.66%), and FedProx (12.64%).The framework addresses data quality heterogeneity through explicit assessment of duplication, label balance, and resolution in federated medical imaging.","author":[{"family":"Gl","given":"P"},{"family":"Ab","given":"George"},{"family":"Vmr","given":"Tummala"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/jbhi.2026.3702502","URL":"https://doi.org/10.1109/jbhi.2026.3702502","source":"pubmed"},{"id":"doi:10.1109/jbhi.2026.3693747","type":"article-journal","title":"GTAFL: Addressing Test-Agnostic Long-Tailed Federated Learning in Medical Image.","abstract":"In healthcare, protecting patient privacy is crucial due to the sensitivity of medical data and its extensive accessibility. Federated Learning (FL) offers a decentralized and privacy-preserving training paradigm, making it an ideal solution for healthcare applications. A critical challenge in healthcare is that real-world medical data often exhibits long-tailed distributions in local and global views. Existing methods addressing long-tailed FL problem typically assume that the model will be evaluated on uniform test data distribution. However, Practical test data in healthcare systems is often agnostic and unpredictable, leading to potential model failures in realworld scenarios. In this paper, we introduce a novel task termed Test-Agnostic Long-Tailed Federated Learning and propose GTAFL, a comprehensive framework to address this challenge. During the training stage, GTAFL employs adaptive re-sampling, expert classifier retraining, and selfsupervised learning to correct biased classifiers and distorted feature spaces caused by long-tailed training distributions. During the inference stage, an ensemble mechanism combines retrained expert classifiers to handle test data with unknown distributions. Extensive experiments on CIFAR10 and two medical datasets manifest that our framework outperforms other state-of-the art methods.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1109/jbhi.2026.3693747","URL":"https://doi.org/10.1109/jbhi.2026.3693747","source":"pubmed"},{"id":"doi:10.1016/j.neunet.2026.109095","type":"article-journal","title":"DDFL: dual defense against poisoning attacks in privacy-preserving federated learning.","abstract":"Federated learning (FL) offers a solution to data silos by enabling collaborative training of a global model across decentralized environments. However, when operating with semi-honest servers or malicious clients, traditional FL faces critical privacy and security challenges. Existing defense strategies often struggle to address both privacy and poisoning attacks effectively, as enhanced privacy protections can increase parameter similarity across clients, unintentionally complicating the detection of malicious behavior. Moreover, most poisoning defenses are primarily server-side, resulting in a one-sided approach that is insufficient to handle increasingly sophisticated attack patterns. Therefore, we introduce a dual defense framework against poisoning attacks in privacy-preserving federated learning (DDFL), which effectively tackles both privacy and security challenges in FL. To enhance privacy, we have clients randomly slice and reassemble model parameters before uploading them to the server, thereby safeguarding client privacy without compromising the server's ability to detect potential malicious behaviors in the system. For stronger security, we incorporate meta-learning and knowledge distillation techniques on the client side, alongside Byzantine-robust methods on the server side, effectively mitigating the impact of malicious clients. Extensive evaluations on three benchmark datasets demonstrate that DDFL not only protects clients' sensitive information but also outperforms existing defense strategies in resisting poisoning attacks, achieving higher model accuracy and faster convergence.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.neunet.2026.109095","URL":"https://doi.org/10.1016/j.neunet.2026.109095","source":"pubmed"},{"id":"doi:10.1016/j.mex.2026.104066","type":"article-journal","title":"JS-Drift: A reproducible Jensen-Shannon divergence procedure for drift-aware client weighting in federated learning.","abstract":"Federated learning lets institutions train a shared model without exchanging raw data, but standard aggregation (FedAvg) assumes client data distributions are stationary. In practice they drift over time, and aggregation that ignores this lets unstable clients degrade the global model. This article describes JS-Drift, a reproducible, model-agnostic procedure that quantifies per-client temporal drift using Jensen-Shannon (JS) divergence between a client's label distribution in consecutive communication rounds, converts that divergence into a stability weight via a single sensitivity parameter &#x3b3;, and folds the weight into the aggregation step. The procedure requires no architectural changes, adds negligible overhead, and drops into any FedAvg-style training loop. We give the full algorithm, exact computation steps, parameter-selection guidance, and an open implementation, and we validate that the procedure behaves as intended on three structurally different tabular-classification settings.&#x2022;Computes a per-client, per-round drift coefficient from JS divergence between consecutive local label distributions; only a small class-proportion summary is shared, so raw data never leave the client.&#x2022;Maps the drift coefficient to an aggregation weight through one interpretable parameter &#x3b3;, down-weighting clients with high distributional shift.&#x2022;Is model-agnostic and integrates into any FedAvg-style round in a few lines of code; a public repository reproduces every step.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.mex.2026.104066","URL":"https://doi.org/10.1016/j.mex.2026.104066","source":"pubmed"},{"id":"doi:10.1038/s41598-026-59796-x","type":"article-journal","title":"PureChain web-based energy predictor with federated learning Dirichlet for real-time energy consumption forecasting.","abstract":"Accurate energy consumption forecasting in smart grids requires privacy-preserving learning mechanisms that remain effective under heterogeneous data distributions and support real-time operation. Existing federated learning approaches remain limited by poor performance under data heterogeneity, unvalidated architectural assumptions, and blockchain consensus mechanisms that are too slow for real-time grid operations. This paper presents PureChain, a blockchain-integrated federated learning framework that combines federated averaging, Dirichlet partitioning, LSTM-based forecasting, and a permissioned blockchain for secure client isolation and model rollback. A partitioning strategy is introduced to improve training stability under extreme non-IID conditions ([Formula: see text]), revealing that the distributional impact of a given Dirichlet parameter is dataset-dependent. To support low-latency smart grid applications, a permissioned blockchain employing proof-of-authority and association (PoA[Formula: see text]) consensus achieves 2.0&#xa0;s transaction latency and 20.88 TPS, outperforming Hyperledger Fabric and Quorum in the evaluated setting. Experimental results on two energy-consumption datasets show that LSTM consistently outperforms BiLSTM under high data heterogeneity, achieving an average R[Formula: see text] of 0.9184 across clients at [Formula: see text]. Smart contract security assessment further yields a threat score of 98.5/100, demonstrating the framework's suitability for privacy-sensitive smart grid deployments. The contribution lies in the integration and systematic validation of established federated learning, forecasting, and blockchain technologies within a unified smart grid architecture.","author":[{"family":"Al","given":"Nanteza"},{"family":"Lac","given":"Ahakonye"},{"family":"Ds","given":"Kim"},{"family":"Jm","given":"Lee"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-59796-x","URL":"https://doi.org/10.1038/s41598-026-59796-x","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-6933179/v1","type":"article-journal","title":"MerkleFL: A Secure Decentralized Federated Learning Framework for Healthcare with Model Integrity Verification","abstract":"Abstract Decentralized federated learning (DFL) offers a privacy-preserving approach for collaboratively training models across distributed healthcare entities without sharing raw patient data. However, ensuring the integrity and reliability of model updates in such decentralized settings remains a significant challenge. This paper introduces \\textit{MerkleFL}, a secure and efficient DFL framework that integrates cluster-based aggregation with Merkle Tree-based verification to detect and reject tampered or unauthorized model contributions. The proposed system employs a lightweight integrity-checking mechanism where each model update is associated with a cryptographic Merkle Root, enabling trustless verification at the cluster level. A dynamic leader election protocol facilitates intra-cluster coordination without relying on central servers or blockchain consensus. Experimental evaluations conducted on the NIH ChestX-ray14 dataset demonstrate that MerkleFL achieves faster convergence, higher classification accuracy, and lower training loss compared to existing DFL schemes such as Gossip-DFL, Ring-DFL, and Blockchain-DFL. The results confirm that MerkleFL effectively balances security, scalability, and performance, making it a practical solution for federated healthcare AI applications.","author":[{"family":"Verma","given":"Ashwin"},{"family":"Pathak","given":"Sunil"},{"family":"Bhattacharya","given":"Pronaya"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-6933179/v1","URL":"https://doi.org/10.21203/rs.3.rs-6933179/v1","source":"europepmc"},{"id":"doi:10.1038/s41598-026-53053-x","type":"article-journal","title":"A blockchain-assisted secure federated learning architecture for intrusion detection in internet of things networks.","abstract":"The fast growth of Internet of Things (IoT) systems has made them very susceptible to advanced cyber-attacks, and an intelligent and privacy-sustainable intrusion detection system is required. Conventional centralized intrusion detection frameworks have the drawbacks of data privacy threats, scale constraints, and a single point of failure, and conventional federated learning still faces the threat of malicious client membership and a lack of trust during model aggregation. To overcome these obstacles, this paper suggests a federated learning (B-FL) system based on a Blockchain to ensure safe and reliable intrusion detection in the distributed Internet of Things. The framework proposed is a combination of federated and blockchain-based trust management to guarantee decentralized collaborative model training and maintain data confidentiality. The use of smart contract-based verification tools and trust-weighted aggregation counteracts the adversarial threats, such as model poisoning, data manipulation, free-rider behavior, and Sybil attacks. Testing is performed on the CICIoT2023 dataset, which consists of traffic produced by 105 IoT devices in 33 different types of attacks, and it allows testing all the aspects of its work in terms of a real and heterogeneous network. Findings reveal that the proposed B-FL model has high detection rates, high convergence stability, and enhanced robustness as compared to traditional methods of centralized and federated intrusion detection. Another study, Receiver Operating Characteristic (ROC) analysis, supports the presence of excellent discriminative ability with respect to several classes of intrusion. Though the integration of the blockchain has a marginal increase in computing overhead, it benefits the system in terms of transparency, reliability, and aggregation security significantly. In general, the suggested framework offers a scalable, privacy-aware, and trust-conscious IoT intrusion detection system in the next generation to enable secure collaborative intelligence in dynamic and adversarial IoT environments such as mining and mineral-processing environments.","author":[{"family":"Sm","given":"Akhtar"},{"family":"Aa","given":"Alhashmi"},{"family":"Aa","given":"Darem"},{"family":"Aa","given":"Alofairi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-53053-x","URL":"https://doi.org/10.1038/s41598-026-53053-x","source":"pubmed"},{"id":"doi:10.1038/s41598-026-59830-y","type":"article-journal","title":"FireSmoke-FL: a privacy-preserving federated learning framework for real-time fire and smoke detection.","abstract":"The early detection of fire and smoke is a significant aspect in the prevention of disasters in smart cities and the development of extensive monitoring systems. In such scenarios, the early response is critical in ensuring the safety and prevention of further damage. The traditional sensor-based approach in the detection of fire and smoke using heat sensors and smoke detectors is prone to several limitations such as delayed response times, increased false alarm rates, and the inability of the system to adjust to changing environmental conditions. To improve the detection of fire and smoke in the context of a smart-city and the development of extensive monitoring systems, a privacy-preserving vision-based framework called FireSmoke-FL is proposed. FireSmoke-FL is a Federated Learning-based framework that integrates an enhanced YOLOv11 model. The model sensitivity is enhanced by incorporating a high-resolution P2 detection head and an attention-refined C2PSA-iEMA module. These improvements aim to enhance the model's robustness in detecting fire and smoke in the presence of background interference, such as fog, clouds, and varying illumination conditions. To ensure the privacy of data collected from edge devices, such as IoT devices and drones, FireSmoke-FL is designed to enable devices to learn locally without sharing data. The Dynamic Average Fusion Algorithm (DAFA) is adapted to improve the performance of the model through the adaptive selection of the clients based on the quality of the models developed locally. The FireSmoke-FL framework is validated through extensive experiments on the Fire and Smoke dataset and the Indoor Fire Smoke dataset. The results show that the framework achieves 96.7% mAP at 84.7 FPS and 96.5% mAP at 82.4 FPS on the Fire and Smoke dataset and the Indoor Fire Smoke dataset, respectively. These findings indicate that the proposed model possesses accuracy, efficiency, scalability, and privacy protection.","author":[{"family":"Ai","given":"Alzahrani"},{"family":"Ah","given":"Al"},{"family":"Aa","given":"Alhabshy"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-59830-y","URL":"https://doi.org/10.1038/s41598-026-59830-y","source":"pubmed"},{"id":"doi:10.3390/s26113545","type":"article-journal","title":"Federated Learning Based on Fuzzy Fusion Rules for Chemical Production Process Fault Diagnosis.","abstract":"Process data plays a vital role in diagnosing fault sources in chemical production. However, such data contain rich process information and are often sensitive, making direct analysis infeasible due to privacy concerns. Although federated learning mitigates data leakage risks, the conventional averaging strategy falls short in achieving high fault identification accuracy, especially under non-independent and identically distributed (non-IID) client data. To overcome this challenge, we propose a personalized federated learning framework, in which a Takagi-Sugeno (T-S) fuzzy fusion rule is designed. Then, the personalized model is constructed through a structured procedure: fuzzification of model parameter distances, definition of fuzzy rules, fuzzy inference, and defuzzification. Moreover, layer-wise fusion is employed to enhance the precision of aggregation. Evaluations on the Tennessee Eastman (TE) process demonstrate that our method achieves superior fault identification accuracy. The results validate the efficacy of the proposed Fuzzy Rule-Based Federated Layer-wise Fusion (FedFZ) framework in industrial fault diagnosis under heterogeneous data distributions.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26113545","URL":"https://doi.org/10.3390/s26113545","source":"pubmed"},{"id":"doi:10.14293/pr2199.003338.v1","type":"article-journal","title":"Design of Privacy-Preserving Federated Learning Models for Auto Insurance Telematics","abstract":"The integration of telematics in the automobile insurance industry has facilitated the transition toward Usage-Based Insurance (UBI). However, the centralized collection of granular GPS and accelerometer data poses significant privacy risks to policyholders. This study proposes a PrivacyPreserving Federated Learning (PPFL) framework designed to train risk-prediction models without necessitating the transfer of raw sensor data to a central cloud server. By utilizing Federated Averaging (FedAvg) integrated with Differential Privacy (DP), the model enables insurance providers to collaboratively learn driving patterns while maintaining local data residency on user devices. Experimental results using the UAH-DriveSet demonstrate that the proposed PPFL model achieves predictive accuracy comparable to centralized models ($AUC \\approx 0.89$) while strictly adhering to privacy guarantees. The study concludes that PPFL addresses the critical \"privacy-utility\" trade-off, offering a viable path for the ethical adoption of telematics in highly regulated insurance markets.","author":[{"family":"Davis","given":"John"},{"family":"Smith","given":"Jennifer"},{"family":"Williams","given":"Robert"}],"issued":{"date-parts":[[2026]]},"DOI":"10.14293/pr2199.003338.v1","URL":"https://doi.org/10.14293/pr2199.003338.v1","source":"europepmc"},{"id":"doi:10.34133/research.1299","type":"article-journal","title":"Experimentally Validated Quantum-Secure Federated Learning over a Multi-user Quantum Network.","abstract":"Federated learning enables decentralized, privacy-preserving training but remains vulnerable to privacy leakage in the quantum era. Quantum federated learning (QFL) offers a promising path toward enhanced security and efficiency. However, a practical and experimentally validated QFL protocol utilizing near-term quantum techniques to address data privacy has been lacking. Here, we present QuNetQFL, a QFL protocol implemented on quantum networks, in which local model updates are masked with distributed quantum secret keys, offering information-theoretic security during aggregation. We experimentally validate the protocol on a 4-client quantum network and benchmark its performance using the generated keys on quantum and real-world datasets. Adding a single quantum client substantially improves global accuracy for classifying multipartite entangled and nonstabilizer quantum datasets. For language tasks, we apply QuNetQFL to sentiment analysis by federated fine-tuning of a hybrid classical-quantum language model, achieving comparable and robust performance in simulation and on real quantum hardware. Large-scale simulations further demonstrate scalability to 200 clients for handwritten-digit recognition, with rapid convergence and a 75% reduction in communication cost via model compression. Our work establishes a practical and scalable route to quantum-secure federated learning for the emerging quantum internet.","author":[{"family":"Zp","given":"Liu"},{"family":"Xy","given":"Cao"},{"family":"Hw","given":"Liu"},{"family":"Xr","given":"Sun"},{"family":"Jy","given":"Shen"},{"family":"Ys","given":"Lu"},{"family":"Hl","given":"Yin"},{"family":"Zb","given":"Chen"}],"issued":{"date-parts":[[2026]]},"DOI":"10.34133/research.1299","URL":"https://doi.org/10.34133/research.1299","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-9746460/v1","type":"article-journal","title":"Securing IoMT with Federated Learning: A Hybrid Deep Learning and Ensemble-Based Intrusion Detection Framework","abstract":"Abstract The rapid proliferation of Internet of Medical Things (IoMT) devices has significantly enhanced healthcare connectivity while simultaneously increasing exposure to sophisticated cyber threats such as distributed denial-of-service attacks, malware propagation, and data exfiltration. Traditional centralized intrusion detection systems (IDS) are increasingly unsuitable for IoMT environments due to strict privacy regulations, decentralized data ownership, device heterogeneity, and communication constraints. To address these challenges, this paper proposes a privacy-preserving federated intrusion detection framework that integrates a hybrid Deep Neural Network (DNN) and XGBoost model within a federated learning (FL) architecture. Unlike conventional FedAvg-based approaches, the proposed method leverages the complementary strengths of deep feature extraction and gradient-boosting classification, enhanced by differential privacy, secure aggregation, and robust aggregation techniques to mitigate adversarial poisoning attacks. The problem is formally modeled as distributed empirical risk minimization under non-independent and identically distributed (non-IID) client data, with convergence behavior analyzed under bounded gradient variance and client drift. Extensive experiments conducted on WUSTL-EHMS-2020, UNSW-NB15, and BoT-IoT datasets demonstrate that the proposed framework achieves detection accuracy of up to 91.4% with F1-scores exceeding 90%, closely approaching centralized baselines such as XGBoost and LightGBM while preserving strict data locality. Furthermore, communication overhead and privacy-utility trade-offs are quantitatively evaluated, showing that differential privacy introduces less than 1% performance degradation under realistic noise budgets. The framework also demonstrates robustness against non-IID data heterogeneity, straggler effects, and adversarial model poisoning. These results establish the proposed approach as a scalable, privacy-compliant, and efficient solution for securing next-generation IoMT systems, enabling trustworthy AI-driven healthcare cybersecurity.","author":[{"family":"Nwokoro","given":"Ifeanyi"},{"family":"Osaghae","given":"Edgar"},{"family":"Kayode","given":"Saheed"},{"family":"Sibe","given":"Tombari"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9746460/v1","URL":"https://doi.org/10.21203/rs.3.rs-9746460/v1","source":"europepmc"},{"id":"doi:10.1038/s41598-026-54935-w","type":"article-journal","title":"Improving IoT security through an explainable hybrid CNN-transformer model and federated learning.","abstract":"The rapid proliferation of Internet of Things (IoT) devices has intensified cybersecurity threats, exposing critical infrastructure to sophisticated intrusion attacks. Existing intrusion detection systems (IDS) typically rely on centralized architectures that compromise data privacy, or employ single-architecture models that fail to capture both local spatial patterns and long-range temporal dependencies in network traffic simultaneously. To address these limitations, this paper proposes a novel explainable hybrid CNN-Transformer model integrated with federated learning (FL) for privacy-preserving intrusion detection in IoT environments. The proposed framework uniquely combines four key components not previously integrated in this context: a dual-block CNN-Transformer architecture, federated learning with FedAvg aggregation, multi-class attack classification, and Local Interpretable Model-Agnostic Explanations (LIME) for decision transparency. Evaluated on the IoT-23 dataset across both federated and non-federated scenarios, the proposed model achieves 94.89% accuracy in federated binary classification and 92.17% in federated multi-class classification, outperforming standalone CNN and ensemble baselines by significant margins. Generalizability is further validated through an ablation study on the CIC IoT-DIAD 2024 dataset. The integration of LIME provides actionable feature-level explanations that support real-time decision-making for network security analysts, advancing both the interpretability and trustworthiness of automated IoT intrusion detection systems.","author":[{"family":"Am","given":"Al"},{"family":"Rm","given":"Al"},{"family":"Ah","given":"Sable"},{"family":"Ss","given":"Alshamrani"},{"family":"Km","given":"Alshmrany"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-54935-w","URL":"https://doi.org/10.1038/s41598-026-54935-w","source":"pubmed"},{"id":"doi:10.2196/69985","type":"article-journal","title":"Explainable AI Approaches in Federated Learning: Systematic Review.","abstract":"Background Artificial intelligence (AI) has, in the recent past, experienced a rebirth with the growth of generative AI systems such as ChatGPT and Bard. These systems are trained with billions of parameters and have enabled widespread accessibility and understanding of AI among different user groups. Widespread adoption of AI has led to the need for understanding how machine learning (ML) models operate to build trust in them. An understanding of how these models generate their results remains a huge challenge that explainable AI seeks to solve. Federated learning (FL) grew out of the need to have privacy-preserving AI by having ML models that are decentralized but still share model parameters with a global model. Objective This study sought to examine the extent of development of the explainable AI field within the FL environment in relation to the main contributions made, the types of FL, the sectors it is applied to, the models used, the methods applied by each study, and the databases from which sources are obtained. Methods A systematic search in 8 electronic databases, namely, Web of Science Core Collection, Scopus, PubMed, ACM Digital Library, IEEE Xplore, Mendeley, BASE, and Google Scholar, was undertaken. Results A review of 26 studies revealed that research on explainable FL is steadily growing despite being concentrated in Europe and Asia. The key determinants of FL use were data privacy and limited training data. Horizontal FL remains the preferred approach for federated ML, whereas post hoc explainability techniques were preferred. Conclusions There is potential for development of novel approaches and improvement of existing approaches in the explainable FL field, especially for critical areas. Trial Registration OSF Registries 10.17605/OSF.IO/Y85WA; https://osf.io/y85wa","author":[{"family":"Tunduny","given":"Titus"},{"family":"Shibwabo","given":"Bernard"}],"issued":{"date-parts":[[2025]]},"DOI":"10.2196/69985","URL":"https://doi.org/10.2196/69985","source":"pubmed"},{"id":"doi:10.1021/acs.jmedchem.5c03681","type":"article-journal","title":"Developing Predictive Models by Sharing Predictions - An Investigation of a Federated Learning Approach for ADMET Predictions.","abstract":"Machine learning models for ADMET prediction benefit from large, diverse data sets, yet such data are typically siloed across organizations. Federated learning (FL) enables collaborative modeling while preserving data privacy. Here, we investigate a student-teacher model (STM) framework in which organizations train internal models on proprietary data and share predictions on a public data set to generate pseudolabels for a centralized student model. As a proof of concept, 11 pharmaceutical companies contributed predictions for rat steady-state volume of distribution, yielding a pseudolabeled data set of &#x223c;133,000 compounds. The resulting student model achieved performance comparable to individual teacher models on an external test set (RMSE &#x2248; 0.51 vs 0.47-0.61). Compared with FL approaches such as MELLODY and Effiris, STM offers a simpler workflow that avoids direct data sharing or iterative collaboration, providing a practical and scalable framework for secure cross-company model development.","author":[{"family":"Dvs","given":"Green"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1021/acs.jmedchem.5c03681","URL":"https://doi.org/10.1021/acs.jmedchem.5c03681","source":"pubmed"},{"id":"doi:10.3233/shti260485","type":"article-journal","title":"Enabling Privacy-Preserving Federated Learning in Healthcare: The FLAME Architecture and Policy Framework.","abstract":"Federated Learning enables collaborative AI development in healthcare without sharing patient data, addressing privacy and regulatory constraints like GDPR and HIPAA. We present FLAME, an open-source platform developed within the German PrivateAIM project, designed to ensure privacy-compliant and auditable federated analytics. FLAME uses a hub-and-node architecture, integrating privacy-enhancing technologies with a dynamic policy framework that specifies permissions and conditions for data access and algorithm execution. This framework allows distributed policy evaluation and enforcement across institutions. Initial deployments at German university hospitals demonstrated FLAME's capability to conduct federated analyses on clinical and genomic data with model performance comparable to centralized approaches. The system offers fine-grained access control, audit logging, and minimal overhead. FLAME provides a scalable foundation for secure, privacy-preserving AI in medicine, bridging legal, technical, and organizational requirements for multi-institutional collaboration.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3233/shti260485","URL":"https://doi.org/10.3233/shti260485","source":"pubmed"},{"id":"doi:10.3389/fdgth.2026.1812254","type":"article-journal","title":"A privacy-preserving federated learning framework for generalizable CBCT to synthetic CT translation in head and neck.","abstract":"Cone-beam computed tomography (CBCT) has become a widely adopted modality for image-guided radiotherapy (IGRT). However, CBCT is characterized by increased noise, limited soft-tissue contrast, and artifacts. These issues result in unreliable Hounsfield unit (HU) values, which limits electron density estimation for direct dose calculation. These issues have been addressed by deriving synthetic CT (sCT) from CBCT, particularly by adopting deep learning (DL) methods. However, existing DL approaches are hindered by institutional heterogeneity, scanner-dependent variations, and data privacy regulations that prevented multi-center data sharing.","author":[{"family":"Cb","given":"Raggio"},{"family":"Mf","given":"Spadea"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/fdgth.2026.1812254","URL":"https://doi.org/10.3389/fdgth.2026.1812254","source":"pubmed"},{"id":"doi:10.3791/72164","type":"article-journal","title":"Improving Cross-Center Generalization for Multi-modal MRI Meningioma Segmentation via Glioma-Pretrained Federated Learning.","abstract":"Automated meningioma segmentation on multi-modal MRI remains challenging when models are transferred across institutions, because scanner protocols, image characteristics, and annotation styles may differ between centers. Federated learning (FL) provides a privacy-preserving strategy for multi-center model development, but standard aggregation may not fully overcome cross-center domain shift. This study aimed to quantify the external generalization gap in MRI meningioma segmentation and evaluate whether Glioma-pretrained FL could improve robustness without centralized data pooling. A UMamba 2D architecture was used for binary meningioma segmentation using T1, T1c, and T2 MRI as model inputs. The protocol included 450 BraTS2023-Men cases as the source-domain meningioma dataset and 174 independent clinical cases from our institute as the external validation cohort. Three final meningioma segmentation strategies were quantitatively evaluated under the same external validation setting: centralized training on BraTS2023-Men, meningioma-pretrained FL across three simulated clients, and Glioma-pretrained FL initialized from a BraTS2023-Gli source model before federated optimization. Centralized training on Glioma was used only to generate the glioma-pretrained initialization and was not reported as an independently evaluated final meningioma segmentation strategy. Model performance was assessed using the Dice similarity coefficient (DSC) and Intersection over Union (IoU). Centralized training on BraTS2023-Men showed a clear external generalization drop, with DSC decreasing from 0.8958 on the internal BraTS2023-Men test set to 0.7452 on the SPHS cohort. Meningioma-pretrained FL yielded lower external performance (DSC = 0.7122), whereas Glioma-pretrained FL achieved comparable performance to centralized training on BraTS2023-Men and improved over meningioma-pretrained FL (DSC = 0.7503; IoU = 0.6301; Holm-adjusted p &lt; 0.001). These results suggest that glioma-pretrained initialization provides a more robust starting point for federated meningioma segmentation and may improve external generalization while preserving institutional data privacy.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3791/72164","URL":"https://doi.org/10.3791/72164","source":"pubmed"},{"id":"doi:10.1371/journal.pone.0348359","type":"article-journal","title":"A cloud-edge-end collaborative intelligent caching method based on incremental federated learning algorithms.","abstract":"In a cloud-edge-end collaborative system, data generated by terminal devices often contains users' sensitive information and is constantly generated and changing, leading to potential data privacy leaks in caches. Additionally, due to the inability to promptly capture these dynamic changes and the failure to consider the actual capabilities of nodes, caching strategies become outdated, resulting in reduced cache hit rates and cache imbalance issues. Therefore, this study proposes a cloud-edge-end collaborative intelligent caching method based on an incremental federated learning algorithm. First, the federated learning algorithm is used to aggregate data from terminal devices to the cloud, enabling collaborative data processing while protecting data privacy. Second, incremental learning methods are employed to continuously update terminal data, with the updated data aggregated to the cloud, thereby enabling real-time tracking of data trends and allowing cache strategies to rapidly adapt to dynamic changes in terminal data. Finally, considering the actual capabilities of nodes, the popularity of aggregated data and the weights of edge and terminal nodes are calculated. Data is cached in edge and terminal nodes in descending order of popularity and weight. When cache space is insufficient, data replacement is performed based on the importance of data within nodes, thereby completing intelligent data caching. Experimental results demonstrate that this method achieves good performance in data update aggregation, with high data caching balance and cache hit rates.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1371/journal.pone.0348359","URL":"https://doi.org/10.1371/journal.pone.0348359","source":"pubmed"},{"id":"doi:10.3390/e28040423","type":"article-journal","title":"Bias-Corrected Federated Learning for Video Recommendation over Stochastic Communication Links.","abstract":"With the increasing demand for privacy-preserving and real-time personalized services in large-scale video platforms, designing robust federated recommendation frameworks over practical communication networks has become increasingly important. To this end, this paper proposes a bias-corrected federated learning framework tailored for video recommendation over stochastic communication links. At the local training stage, a bias-corrected mechanism is introduced to explicitly account for video duration and user activity, mitigating feature-level bias and enabling the learned representations to more accurately reflect users' intrinsic preferences. To meet the timeliness requirements of real-time federated learning, the successful upload probability of local model transmission is analytically characterized under time-varying channel conditions. Building upon this probabilistic model, a statistically corrected global aggregation strategy is designed to preserve the unbiasedness of the global update with respect to the ideal fully reliable FedAvg scheme, even when a subset of local nodes fails to upload their models within the specified delay constraint. Comprehensive experimental evaluations validate that the proposed framework significantly improves recommendation accuracy and maintains robustness against communication unreliability in practical distributed environments.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/e28040423","URL":"https://doi.org/10.3390/e28040423","source":"pubmed"},{"id":"doi:10.1038/s41598-026-59224-0","type":"article-journal","title":"BC-AWFedAvg: blockchain-assisted adaptive federated learning for secure RAN Slicing in beyond-5G networks.","abstract":"The evolution of beyond-5G networks introduces new challenges for radio resource management, particularly for heterogeneous service requirements across multiple virtual network operators. This work presents BC-AWFedAvg, a layered framework for federated deep reinforcement learning in O-RAN network slicing that integrates adaptive aggregation, blockchain-based governance, secure aggregation, and differential privacy. The proposed design separates learning, governance, and storage functions to support coordinated training while preserving privacy and limiting exposure of individual updates. Simulation results in the considered setting indicate that the proposed framework improved robustness in the considered setting under several adversarial scenarios while maintaining acceptable quality-of-service performance. These findings suggest that combining complementary mechanisms may be a promising direction for secure federated learning in next-generation wireless networks.","author":[{"family":"Je","given":"Hajlaoui"},{"family":"As","given":"Aldalbahi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-59224-0","URL":"https://doi.org/10.1038/s41598-026-59224-0","source":"pubmed"},{"id":"doi:10.1038/s41598-026-51138-1","type":"article-journal","title":"BFLAFD: blockchain-enabled federated learning framework for adaptive fire detection in IIoT networks.","abstract":"The Industrial Internet of Things (IIoT) is a network of interconnected sensors, devices, and control systems in the oil and gas sectors that has been developed to make the industries automated and continuously monitored. However, there are challenges in fire detection in such environments, including the unreliable nature of the sensor data, privacy issues, communications delays, and the lack of a generalized model across locations in a distributed solution. To overcome the above problems, BFLAFD, Blockchain-assisted Federated Learning framework for Adaptive Fire Detection is introduced. Unlike centralized methods, BFLAFD makes use of Federated Learning (FL), where local edge servers train models directly on-device in which data confidentiality is upheld and less data is transmitted. Hierarchical aggregation process is effective in maximizing global performance, in addition to being sensitive to sensor drift and device heterogeneity. To make sure of trust and resilience, BFLAFD combines the permissioned blockchain with smart contracts providing access control, transparency of logs and no tampering of models or insider manipulation. Furthermore, Personalized Federated Learning (PFL) makes it possible to create a customized fire detection model, effectively enhancing the accuracy in varying conditions. Experimental evaluations have shown that BFLAFD has 98.2% detection accuracy, false alarm rate of 2.7%, and a 100-150 ms inference latency, and blockchain validation time of 1-2&#xa0;s. In addition, the cost of communication was reduced by 82.3% compared to centralized training. Overall, BFLAFD offers critical IIoT environments fast, accurate, and secure fire detection solutions.","author":[{"family":"Sk","given":"Singh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-51138-1","URL":"https://doi.org/10.1038/s41598-026-51138-1","source":"pubmed"},{"id":"doi:10.1177/00368504261456965","type":"article-journal","title":"Asynchronous federated learning with partial weights aggregation for energy consumption forecasting.","abstract":"Accurate energy forecasting is essential for grid stability, demand-side management, and efficient renewable integration. However, energy consumption data collected from smart meters may expose sensitive user information, thus raising privacy concerns. Federated Learning (FL) offers a privacy-preserving mechanism for collaborative model training without sharing raw data. However, conventional synchronous FL suffers from training delays caused by heterogeneous client availability and computational capabilities, while frequent exchange of model parameters can lead to communication overheads. To address these challenges, this paper proposes an asynchronous federated learning framework for energy forecasting that enables continuous global model updating without waiting for all clients to complete local training. We introduce a federated asynchronous adaptive aggregation mechanism, where client-specific learning rates are dynamically adjusted based on both update staleness and model performance contribution. A partial aggregation strategy is defined for a Long Short-Term Memory (LSTM) forecasting model that splits the local models' layers, allowing clients to exchange only a subset of the weights with the server. The proposed solution is evaluated using real-world energy consumption data from multiple consumers. Experimental results demonstrate that the proposed asynchronous adaptive strategy outperforms the classic FedAvg approach and maintains prediction accuracy relative to personalised FedAvg, while reducing communication costs. Additionally, the proposed method outperforms the classic FedAsync algorithm across all client groups, with statistically significant improvements in most cases.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1177/00368504261456965","URL":"https://doi.org/10.1177/00368504261456965","source":"pubmed"},{"id":"doi:10.64898/2026.07.30.26359337","type":"article-journal","title":"Use of Federated Learning for validating and updating privacy-preserving decentralized multi-study prognostic models in Traumatic Brain Injury","abstract":"Developing modern clinical prediction models (CPMs) and advanced analytics requires large datasets, often necessitating data from different studies. Privacy regulations may hinder data sharing, especially across countries. Decentralized federated data infrastructures, where data remain in their original location and analyses are run only in a shared, secure environment, may address these challenges. We implemented a privacy-preserving federated learning (FL) infrastructure and evaluated and updated the IMPACT prognostic models for traumatic brain injury (TBI) using 2 studies. A multi-continental federated infrastructure was established between 2 large-scale studies (TRACK-TBI from the United States and CENTER-TBI from Europe and Israel). Three IMPACT prognostic models for post-TBI 6-month mortality and unfavorable outcomes were evaluated, followed by model updates through 2 FL approaches trained across the TRACK-TBI and CENTER-TBI studies. Internal validation, external cross-validation, and sub-study validations were performed. CPMs were evaluated for discrimination and calibration. The federated cohort included 1616 participants (TRACK-TBI: n=441, CENTER-TBI: n=1175). Both FL performed well, with comparable coefficient estimates, AUCs (area under the receiver operating characteristics curve) between 0.77-0.88, and calibrated probabilities. Compared to the original IMPACT and single-study models, both federated models presented similar discrimination (AUC), were well-calibrated, were more efficient (higher precision), and reduced the impact of missing data in model estimation. FL is feasible for privacy-preserving development and evaluation of CPMs, and can enable validation and updating across large, virtually analyzed datasets while overcoming regulatory constraints on data combination. Federated infrastructures can facilitate global collaboration to advance data-hungry analytical methods, such as artificial intelligence.","author":[{"family":"Jc","given":"Wong"},{"family":"He","given":"Hinson"},{"family":"Tb","given":"Kuipers"},{"family":"Bpt","given":"Hoekstra"},{"family":"Jk","given":"Yue"},{"family":"Hf","given":"Lingsma"},{"family":"Aj","given":"Markowitz"}],"issued":{"date-parts":[[2026]]},"DOI":"10.64898/2026.07.30.26359337","URL":"https://doi.org/10.64898/2026.07.30.26359337","source":"pubmed"},{"id":"doi:10.3390/s26134016","type":"article-journal","title":"Context-Aware Online Model Splitting and Device Association for Semi-Decentralized Federated Learning in Internet of Things.","abstract":"As a distributed approach to Artificial Intelligence (AI) model construction over wireless networks, federated learning (FL) based on multi-device collaborative training can protect data privacy, as well as increase the computing load of local model updates. In contrast, split learning (SL) with proper model splitting can adapt to the computation and transmission capabilities among devices. In this paper, while taking advantage of FL and SL, we concentrate on a semi-decentralized hybrid federated split learning (SD-HFSL) framework, in which we surpass the limitations of a single central server and allow the shared split models to be aggregated among multiple edge servers. To verify the importance of latency optimization for training efficiency, we analyze the convergence performance of SD-HFSL while jointly considering the limited computation and communication resources. Then, aiming at maximizing the long-term training efficiency, we propose an online optimization problem that includes local model splitting and device association. Considering that the training latency is unknown to the system a priori, a context-aware online training algorithm with sublinear regret is proposed based on the framework of contextual multi-armed bandit (CMAB), where the edge servers can observe the context information of device sites for latency estimation, followed by the iterative optimization based on the evaluated information in different contexts. Experiments on several neural network models show that the proposed algorithm reduces training latency and improves test accuracy compared with the selected benchmarks.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26134016","URL":"https://doi.org/10.3390/s26134016","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-9138351/v1","type":"article-journal","title":"Governance-Aware Federated Learning for Trustworthy and Compliant Decentralized Infrastructure Monitoring","abstract":"Abstract The paper introduces a federated learning (GFL) architecture that is governed to monitor decentralized infrastructure in a legally heterogeneous setting. The framework incorporates legal compliance limits, auditability controls, and policy alignment as governed by LLM into the federated optimization process, thus creating the possibility of trustful and policy-oriented deployment of AI to the nodes of the public sector.This is to develop a mathematical model that embodies multi-agent lawful infestations, metadata audit rating, and also dynamic trust stabilization within the limits of regulations. Empirical analysis with a synthesized dataset of 20 European jurisdiction nodes reveals that the GFL model performs better than governance-free baselines in the accuracy of their predictions +1 3.5 %) and concurrently enhances transparency, explainability, and auditability. Clients that follow a legal and semantic protocol of governance always provide quality and more consistent updates. The findings show that rather than deterring model convergence and trust in the decentralized systems, the enforcement of governance improves such aspects. The paper is another addition to the intersection of federated AI in digital governance, providing a scalable strategy to institutional compliance in critical areas of infrastructure, such as smart grids, urban mobility, and environmental sensing","author":[{"family":"Qasemabadi","given":"Seyed"},{"family":"Sangchouli","given":"Mohammadmahdi"},{"family":"Shadman","given":"Fatemeh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9138351/v1","URL":"https://doi.org/10.21203/rs.3.rs-9138351/v1","source":"europepmc"},{"id":"doi:10.1038/s41598-026-46244-z","type":"article-journal","title":"FedLiverNet: a federated learning framework for privacy-preserving and efficient liver cancer detection.","abstract":"Liver cancer continues to be a significant health issue on the global front, and proper segmentation of the liver and the tumor formed from the computed tomography is essential in the early diagnosis and subsequent treatment strategies. Although deep learning models can be trained to perform well in segmentation, the optimal way to train a strong model is with large, diverse datasets that may be distributed across institutions and cannot be centralized due to privacy and other regulatory restrictions. Federated learning enables joint training without exchanging patient data; however, performance may be poor on non-independent and identically distributed (non-IID) data, and privacy is a concern in optimization. This paper presents FedLiverNet, a communication-efficient and privacy-guaranteed federated liver and tumor segmentation system. FedLiverNet is a variant of the U-Net segmentation architecture that incorporates a modified backbone, differential privacy aggregation, and clustered federated learning with local adaptation to promote personalized support among heterogeneous clients. Simulation-based experiments indicate that FedLiverNet achieves a 0.89&#x2009;&#xb1;&#x2009;0.03 tumor Dice score and a 23% reduction in communication cost and is more effective than either federated averaging or local-only training under heterogeneous data distributions. These findings make FedLiverNet a viable solution to privacy-constrained, multi-center liver cancer detection and segmentation.","author":[{"family":"Za","given":"Shaikh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-46244-z","URL":"https://doi.org/10.1038/s41598-026-46244-z","source":"pubmed"},{"id":"doi:10.1016/j.dib.2026.113111","type":"article-journal","title":"Dataset of round-level federated learning and layer-2 blockchain overhead from ten raspberry Pi edge clients.","abstract":"This data article describes round-level records from 40 federated learning sessions run on a physical testbed of ten Raspberry Pi 4 Model B devices acting as edge clients, with an Ubuntu workstation as the aggregation server. Each session trains a one-dimensional Squeeze-and-Excitation ResNet classifier over ten aggregation rounds using the Flower framework and commits every round to a smart contract on Base mainnet, a public Ethereum-compatible Layer-2 chain, while pinning the model and metric artifacts to IPFS through Pinata. The sessions cover two human activity recognition benchmarks, MHEALTH and UCI-HAR, and two middleware configurations: a baseline that performs blockchain and IPFS operations synchronously with per-round gas estimation, and an optimised variant that overlaps IPFS uploads with local training and reuses cached gas parameters, giving 20 sessions per benchmark. For every round the dataset records server-side global metrics (accuracy, F1, AUC, and the confusion matrix), per-device training telemetry (loss, training time, CPU, memory, and temperature), per-client evaluation records, and the IPFS content identifiers and transaction hashes that link each round to Base mainnet. For every session, millisecond-resolution logs record the latency of each blockchain and storage operation. The repository contains 12 CSV files, two archives of raw session directories, and a data dictionary. Researchers can reuse the convergence traces to benchmark edge federated learning, the device telemetry to characterise resource use on constrained hardware, and the latency distributions to parameterise simulations of blockchain-enabled federated learning without physical deployments or transaction fees. All 1240 blockchain transactions remain publicly verifiable on the Base mainnet explorer.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.dib.2026.113111","URL":"https://doi.org/10.1016/j.dib.2026.113111","source":"pubmed"},{"id":"doi:10.5281/zenodo.21553574","type":"article-journal","title":"Privacy-Preserving Techniques for Secure Cloud Computing : A Survey of Recent Advances","abstract":"Cloud computing has gained immense popularity in recent years due to its on-demand and scalable computing resources. However, with the growth of cloud computing, privacy and security concerns have also increased. The primary concern is how to ensure the confidentiality and integrity of data in the cloud, as the data is stored on third-party servers. To address these concerns, various privacy-preserving techniques have been proposed, which allow users to store and process their data in the cloud without compromising privacy and security. We provide a thorough overview of current developments in privacy-preserving methods for safe cloud computing in this study. We start by giving a general review of cloud computing and the security issues it presents. Then, we go over a variety of privacy-preserving methods, such as differential privacy, homomorphic encryption, secure outsourcing, and secure multi-party computation. We also highlight their advantages and limitations. Finally, we conclude with some future research directions in privacy-preserving cloud computing.","author":[{"family":"Savitha","given":"N"},{"family":"Kiran","given":"Dr"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.21553574","URL":"https://doi.org/10.5281/zenodo.21553574","source":"datacite"},{"id":"doi:10.5281/zenodo.21553575","type":"article-journal","title":"Privacy-Preserving Techniques for Secure Cloud Computing : A Survey of Recent Advances","abstract":"Cloud computing has gained immense popularity in recent years due to its on-demand and scalable computing resources. However, with the growth of cloud computing, privacy and security concerns have also increased. The primary concern is how to ensure the confidentiality and integrity of data in the cloud, as the data is stored on third-party servers. To address these concerns, various privacy-preserving techniques have been proposed, which allow users to store and process their data in the cloud without compromising privacy and security. We provide a thorough overview of current developments in privacy-preserving methods for safe cloud computing in this study. We start by giving a general review of cloud computing and the security issues it presents. Then, we go over a variety of privacy-preserving methods, such as differential privacy, homomorphic encryption, secure outsourcing, and secure multi-party computation. We also highlight their advantages and limitations. Finally, we conclude with some future research directions in privacy-preserving cloud computing.","author":[{"family":"Savitha","given":"N"},{"family":"Kiran","given":"Dr"}],"issued":{"date-parts":[[2023]]},"DOI":"10.5281/zenodo.21553575","URL":"https://doi.org/10.5281/zenodo.21553575","source":"datacite"},{"id":"doi:10.5281/zenodo.21307218","type":"article-journal","title":"Companion artifact: Spectrum Sensing from Classical Detection to Federated Learning — A Taxonomy, Benchmarking-Gap Analysis, and Deployment Guide","abstract":"Companion artifact for the survey \"Spectrum Sensing from Classical Detection to Federated Learning: A Taxonomy, Benchmarking-Gap Analysis, and Deployment Guide\". Contains three components: (1) a reference implementation of four classical spectrum-sensing detectors (energy detection, matched filtering, single-cycle cyclostationary feature detection, and the maximum-minimum-eigenvalue detector) under one fully specified reference setup, evaluated under both ideally known noise and a realistic ±1 dB noise-uncertainty condition; (2) a meta-analysis script that recomputes the cross-paradigm variance decomposition, reproducing the group means, the eta-squared variance partition, and the modern-band standard deviation, together with a sensitivity analysis over grouping choices; and (3) the machine-readable 41-study corpus table with each study's taxonomy cell, reported accuracy, SNR operating point, dataset, and hardware-validation flag. The survey argues that the central weakness of the spectrum-sensing literature is unverifiable evaluation. This artifact is released so that every survey-internal number is auditable rather than asserted. The classical tier requires only NumPy and needs no dataset download.","author":[{"family":"Mohd Ali","given":"Yazan"},{"family":"Kaymih","given":"Nour"},{"family":"Bany Salameh","given":"Haythem"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21307218","URL":"https://doi.org/10.5281/zenodo.21307218","source":"datacite"},{"id":"doi:10.5281/zenodo.21307219","type":"article-journal","title":"Companion artifact: Spectrum Sensing from Classical Detection to Federated Learning — A Taxonomy, Benchmarking-Gap Analysis, and Deployment Guide","abstract":"Companion artifact for the survey \"Spectrum Sensing from Classical Detection to Federated Learning: A Taxonomy, Benchmarking-Gap Analysis, and Deployment Guide\". Contains three components: (1) a reference implementation of four classical spectrum-sensing detectors (energy detection, matched filtering, single-cycle cyclostationary feature detection, and the maximum-minimum-eigenvalue detector) under one fully specified reference setup, evaluated under both ideally known noise and a realistic ±1 dB noise-uncertainty condition; (2) a meta-analysis script that recomputes the cross-paradigm variance decomposition, reproducing the group means, the eta-squared variance partition, and the modern-band standard deviation, together with a sensitivity analysis over grouping choices; and (3) the machine-readable 41-study corpus table with each study's taxonomy cell, reported accuracy, SNR operating point, dataset, and hardware-validation flag. The survey argues that the central weakness of the spectrum-sensing literature is unverifiable evaluation. This artifact is released so that every survey-internal number is auditable rather than asserted. The classical tier requires only NumPy and needs no dataset download.","author":[{"family":"Mohd Ali","given":"Yazan"},{"family":"Kaymih","given":"Nour"},{"family":"Bany Salameh","given":"Haythem"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21307219","URL":"https://doi.org/10.5281/zenodo.21307219","source":"datacite"},{"id":"doi:10.26187/deakin.33390286","type":"article-journal","title":"Defense Against Information Integrity Attacks in Federated IoT Systems Using Inertial Momentum-Aware IALM-RPCA","abstract":"While federated learning offers a decentralized approach to model training, ensuring the integrity of the information from each IoT client remains a challenge. This work delves into the dynamics of multi-stage federated learning, its susceptibility to information integrity attacks, and how to defend against such threats. A comprehensive understanding of data uncertainty and the challenges of poisoning attacks is discussed, laying a solid groundwork for the proposed defense mechanisms. At its core, this paper introduces a novel multi-stage federated learning model that segments the federated learning process into distinct phases with a novel approach of inertial momentum-aware Inexact Augmented Lagrange Multiplier Robust PCA with constant momentum factor and unaltered norm of the traditional one, each tailored to optimize for both efficiency and security. This robust framework is then tested against data injection-based poisoning attacks, using sparse noise, and demonstrates the effectiveness of the proposed recovery techniques like Robust PCA. Performance results highlight the resilience and efficiency of the introduced model with novel reconstruction algorithm, emphasizing the importance of this approach in real-world IoT settings. Data analysis, model summaries, and impacts of adversarial attacks further reinforce the findings, which are evaluated using rigorous statistical metrics and machine learning algorithms. The paper concludes by acknowledging its efficiency in detection and recovery from data poisoning attacks, improving robustness and data reconstruction in IoT environments while highlighting opportunities for further security enhancements.","author":[{"family":"Tanmoy","given":"Oudarja"},{"family":"Hasan","given":"Sakib"},{"family":"Anwar","given":"Adnan"},{"family":"Mamun","given":"Md"},{"family":"Hasan","given":"Abm"},{"family":"Rahman","given":"Akhlaqur"}],"issued":{"date-parts":[[2026]]},"DOI":"10.26187/deakin.33390286","URL":"https://doi.org/10.26187/deakin.33390286","source":"datacite"},{"id":"doi:10.5281/zenodo.22175626","type":"article-journal","title":"Application of Artificial Intelligence Frameworks in Development of Medical Imaging Diagnosis Systems: A Comprehensive Review of Novel Methodologies, Clinical Validation, Performance Optimization, and Future Perspectives","abstract":"Artificial intelligence (AI) has revolutionized medical imaging with automated disease detection, image segmentation, diagnosis, prognosis prediction, and clinical decision support across various imaging modalities, such as X-ray, computed tomography (CT), magnetic resonance imaging (MRI), ultrasound, positron emission tomography (PET), retinal imaging, and digital pathology. Recent advances in deep learning, transformer architectures, multimodal learning and foundation models have significantly improved the diagnostic accuracy and reduced the reliance on handcrafted feature engineering. However, challenges like data heterogeneity, model interpretability, external validation, privacy preservation, computational efficiency, and regulatory compliance still hinder the widespread clinical implementation.In this paper, this review presents a comprehensive study on the evolution of modern AI frameworks in medical imaging by integrating recent methodological advances with perspectives on clinical translation. The review covers the latest deep learning architectures, such as convolutional neural networks, Vision Transformers, hybrid CNN–Transformer models, multimodal learning frameworks, generative artificial intelligence, diffusion models, federated learning, privacy-preserving learning, and medical foundation models. In addition, the review covers the cutting-edge explainable AI techniques, including Grad-CAM, SHAP, LIME, and attention visualization, for boosting transparency and clinician confidence. The review also discusses the state-of-the-art performance optimization strategies, including transfer learning, active learning, domain adaptation, neural architecture search, hyperparameter optimization, model compression, and computational resource optimization. Equally important, the latest developments in clinical validation, external evaluation, robustness assessment, fairness, uncertainty estimation, regulatory considerations, and deployment frameworks are critically analyzed to underscore their role in facilitating safe clinical implementation.The review analysis concludes with the identification of key research challenges and directions for the future including multimodal foundation models, vision-language systems, retrieval-augmented generation,","author":[{"family":"Sur","given":"Susreeti"},{"family":"Rakesh Kumar","given":"Mandal"},{"family":"Debanil","given":"Chanda"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.22175626","URL":"https://doi.org/10.5281/zenodo.22175626","source":"datacite"},{"id":"doi:10.21203/rs.3.rs-9754545/v1","type":"article-journal","title":"Federated Learning for Intrusion Detection in Internet of Medical Things (IoMT): A PRISMA-Based Systematic Literature Review","abstract":"Abstract The rapid expansion of Internet of Medical Things (IoMT) devices has introduced new opportunities for real-time healthcare monitoring but also increased cybersecurity risks. Traditional centralized intrusion detection systems (IDS) struggle to scale across distributed IoMT networks while preserving sensitive patient data. Federated Learning (FL) offers a privacy-preserving alternative, enabling collaborative model training without sharing raw data. This study presents a PRISMA-based systematic literature review of FL applications for IoMT intrusion detection, analyzing 86 peer-reviewed studies published between 2015 and 2025 from Scopus, IEEE Xplore, Web of Science, and PubMed. Findings indicate that FL-based IDS frameworks achieve competitive accuracy (85–93%) and F1-scores (~ 0.87), yet face challenges in handling non-IID heterogeneous data, optimizing energy and communication-efficiency, and ensuring robust privacy guarantees. Based on the synthesis, we propose a conceptual framework integrating privacy-preserving aggregation, adaptive federated optimization, and edge-intelligence mechanisms, tailored to IoMT constraints. This work provides a comprehensive foundation for developing scalable, privacy-aware, and resource-efficient IoMT intrusion detection systems, while identifying key research gaps and future directions for enhancing robustness, personalization, and real-world deployment.","author":[{"family":"Nwokoro","given":"Ifeanyi"},{"family":"Osaghae","given":"Edgar"},{"family":"Kayode","given":"Saheed"},{"family":"Sibe","given":"Tombari"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9754545/v1","URL":"https://doi.org/10.21203/rs.3.rs-9754545/v1","source":"europepmc"},{"id":"doi:10.21203/rs.3.rs-8883453/v1","type":"article-journal","title":"Decoupled Text-Guided Distillation for Efficient Federated Learning on Edge Devices","abstract":"Abstract Federated Learning (FL) enables the collaborative training of models across heterogeneous edge devices while preserving data privacy; however, its performance degrades significantly under domain shift. While integrating Vision-Language Models (VLMs) can mitigate this, existing prompt-tuning methods typically remain coupled to the massive VLM backbone during inference, rendering them impractical for resource-constrained edge devices. To address this challenge, we propose CLIP-assisted Domain-Invariant Federated Learning (CDIFed), which decouples the VLM from the deployment model to enhance robustness without incurring high inference latency. This framework integrates a Text-Guided Domain Adapter, implemented as a parameter-efficient bottleneck module, which aligns visual features with invariant text-based anchors to filter domain-specific noise while maintaining class-discriminative semantics. CDIFed operates through a communication-efficient two-phase framework: clients first adapt a frozen CLIP teacher, and then the adapted teacher supervises the training of a lightweight student network via feature knowledge distillation. Unlike previous approaches, the heavy VLM is discarded after adaptation, and only the student model parameters are transmitted to the server for aggregation. Experiments on the Digits and Office-Caltech benchmarks demonstrate that CDIFed significantly outperforms state-of-the-art methods in federated domain generalisation while maintaining the inference efficiency required for heterogeneous edge devices.","author":[{"family":"Kim","given":"Younghan"},{"family":"Park","given":"Yongjae"},{"family":"Cho","given":"Jae"},{"family":"Cho","given":"Jungchan"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-8883453/v1","URL":"https://doi.org/10.21203/rs.3.rs-8883453/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-8681063/v1","type":"article-journal","title":"Standardized API Call Protocols for implementing Federated Learning in FAIRDatabase","abstract":"Abstract The rapid expansion of machine learning methodologies in biomedical research has intensified the tension between the demand for large scale data analysis and the stringent privacy regulations governing sensitive health data. The integration of federated learning with FAIR-compliant databases necessitates a carefully engineered application programming interface (API) that reconciles multiple, partially competing requirements: preservation of the data governance, provenance, and access control mechanisms mandated by the FAIR principles; support for the iterative and stateful communication patterns inherent to federation learning protocols; maintenance of modularity to enable independent evolution and replacement of both database and machine learning components; and adherence to standards that promote long term interoperability and facilitate future extensions and ecosystem development. In this context, we propose a systematic methodology for designing and implementing standardised APIs that enable FAIR data repositories to support collaborative machine learning while respecting the governance, access control, and compliance requirements of the underlying database systems. This work contributes to a replicable framework that can be applied to other databases that seek to enable collaborative science at scale while maintaining the privacy protections essential for sensitive health information.","author":[{"family":"Regt","given":"Sem"},{"family":"Bumbuc","given":"Roland"},{"family":"Sheraton","given":"Vivek"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-8681063/v1","URL":"https://doi.org/10.21203/rs.3.rs-8681063/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-8389432/v1","type":"article-journal","title":"PSO-Driven Client Selection for Federated Learning in IDS applications","abstract":"Abstract Federated Learning has emerged as a highly promising distributed and collaborative learning paradigm, enabling local clients to train models without sharing their raw data. However, the global model’s performance is often impacted by heterogeneity and variability at the client level, making the selection of clients in each training round a critical factor for overall model effectiveness.In this paper, we propose a novel client selection strategy within the federated learning framework that leverages Particle Swarm Optimization (PSO) to identify the most suitable participants for training a high-performing model in fewer communication rounds. The PSO-based algorithm dynamically adjusts the contribution weights of client models based on their performance and stability metrics, aiming to improve global model accuracy, accelerate convergence, and enable smarter, adaptive model aggregation.We apply this approach specifically to develop a collaborative intrusion detection system for cybersecurity in Edge IIoT environments. Clients are selected based on key criteria such as data quality and individual model performance. Experimental evaluation on a real-world Edge IIoT dataset demonstrates that our solution achieves over 91% accuracy in the global model while reducing the number of communication rounds by approximately 40% compared to the traditional Federated Averaging (FedAvg) method.These results highlight that the proposed approach not only significantly boosts the accuracy of the global model but also accelerates its convergence, making it a robust and efficient solution for Intrusion Detection Solutions (IDS) applications.","author":[{"family":"Wali","given":"Aymen"},{"family":"Boughdiri","given":"Maher"},{"family":"Mrabet","given":"Hichem"},{"family":"Jemai","given":"Abderrazek"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-8389432/v1","URL":"https://doi.org/10.21203/rs.3.rs-8389432/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-9316077/v1","type":"article-journal","title":"Hybrid Transformer-Based Recommender System with LLM-Assisted Semantic Modeling for Sequential and Federated Learning","abstract":"Abstract Sequential recommender systems based on transformers have shown good performance to model the user interaction dynamics, although they usually depend mainly on the interaction data and fail to use rich semantic data that exists in item metadata. This is a major constraint especially in cases of sparsity and cold-start when the interaction histories are not enough to model preference accurately. This paper presents the suggestion of LLMTransRec, a hybrid Transformer-based recommendation model combining collaborative interaction cues, semantic representations generated by pretrained language models, and content-based features into a single model. In order to successfully integrate these heterogeneous modalities, we propose a context-sensitive gating system, which dynamically weighs the contributions of these modalities on the context of the sequence of interactions between a user. We test the suggested framework on a variety of benchmark datasets having different sparsity and richness of content. The experimental findings prove that the model attains stable improvements when compared to strong sequential and graph-based baselines on an integrated sampled evaluation protocol. Other studies, such as ablation analysis and cold-start analysis indicate that semantic features and adaptive fusion can help enhance the robustness of recommendations. On the whole, the findings indicate that the integration of semantic representations into sequential recommendation models is a viable way of improving the performance in the sparse-data context.","author":[{"family":"Maddala","given":"Lakshmi"},{"family":"Pamula","given":"Rajendra"},{"family":"Subbarao","given":"Katteda"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9316077/v1","URL":"https://doi.org/10.21203/rs.3.rs-9316077/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-8724094/v1","type":"article-journal","title":"Standardized API Design for Privacy-Preserving Federated Learning in FAIR-Compliant Biomedical Databases","abstract":"Abstract The rapid expansion of machine learning methodologies in biomedical research has intensified the tension between the demand for large scale data analysis and the stringent privacy regulations governing sensitive health data. The integration of federated learning with FAIR-compliant databases necessitates a carefully engineered application programming interface (API) that reconciles multiple, partially competing requirements: preservation of the data governance, provenance, and access control mechanisms mandated by the FAIR principles; support for the iterative and stateful communication patterns inherent to federation learning protocols; maintenance of modularity to enable independent evolution and replacement of both database and machine learning components; and adherence to standards that promote long term interoperability and facilitate future extensions and ecosystem development. In this context, we propose a systematic methodology for designing and implementing standardised APIs that enable FAIR data repositories to support collaborative machine learning while respecting the governance, access control, and compliance requirements of the underlying database systems. This work contributes to a replicable framework that can be applied to other databases that seek to enable collaborative science at scale while maintaining the privacy protections essential for sensitive health information.","author":[{"family":"Regt","given":"Sem"},{"family":"Bumbuc","given":"Roland"},{"family":"Sheraton","given":"Vivek"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-8724094/v1","URL":"https://doi.org/10.21203/rs.3.rs-8724094/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-8240248/v1","type":"article-journal","title":"Fair Client Selection Method for Federated Learning Based on Discretized Firefly Algorithm","abstract":"Abstract Federated Learning (FL) enables collaborative model training without exchanging sensitive local data, ensuring privacy and advancing distributed machine learning. However, in edge scenarios, FL faces challenges of data heterogeneity, device resource constraints, and fairness imbalance among small-data clients, making it difficult to balance performance, efficiency, and fairness. To tackle this, we propose the Discrete Firefly Algorithm (DFA) for fair client selection in FL, mapping clients to fireflies, retaining the brightness attraction\"core while adapting to discrete selection. DFA quantifies brightness through data volume and historical contributions, optimizes efficiency with selective sampling, and guarantees fairness for small-data clients. Experiments on MNIST, Fashion-MNIST, and CIFAR-10 demonstrate DFA outperforms baselines: achieving 73.12\\((%)\\) accuracy on CIFAR-10 (2.56\\((%)\\) and 10.82\\((%)\\) higher than random selection and Power-of-Choice), with lower overhead, 2.4\\((%)\\) performance improvement for small-data clients, and compliance with the principle of contribution-matching benefit.","author":[{"family":"Li","given":"Xiaoye"},{"family":"Zhang","given":"Yangyang"},{"family":"Sun","given":"Zhenlong"},{"family":"Zhao","given":"Wei"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-8240248/v1","URL":"https://doi.org/10.21203/rs.3.rs-8240248/v1","source":"europepmc"},{"id":"doi:10.21203/rs.3.rs-8161081/v1","type":"article-journal","title":"Federated Learning Enhanced YOLOv8 for Privacy-Preserving Student Classroom Behavior Recognition","abstract":"Abstract In the field of smart education, real-time object detection for analyzing student classroom behaviors provides valuable, objective data to help teachers optimize teaching methods and improve student learning experiences. This supports a positive and engaging classroom environment. However, models like YOLOv8 face challenges in real-world settings, such as varying object scales (\"far small near large\"), frequent occlusions, class imbalances, and privacy concerns when sharing data across institutions. Current datasets are often simulated, small in scale, and lack diversity, which limits their ability to reflect actual classroom conditions. To address these issues, this paper introduces FedYOLO-Behavior, a federated learning (FL) enhanced YOLOv8 framework that ensures privacy while recognizing behaviors effectively. We build a large, real-world database covering educational stages from kindergarten to university, combined with open-source data for augmentation. Local models are improved with multi-head self-attention (MHSA) for better context understanding, Ghost Convolution for efficiency, and Focal-EIoU loss to handle imbalances and small objects. FL with differential privacy allows safe collaboration between schools without sharing raw data. Experiments show significant improvements: mAP from 81.8% to 85.6%, precision from 77.9% to 82.3%, recall from 75.4% to 79.1%, inference speed increased by 18.2% (reaching 112 FPS), and parameters reduced by 25.1%, with a privacy budget of ε=0.9. This work promotes innovative, secure AI applications in education, contributing to national goals in technology and harmonious learning environments.","author":[{"family":"Ma","given":"Shuai"},{"family":"Chang","given":"Heyou"},{"family":"Han","given":"Jian"},{"family":"Zheng","given":"Hao"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-8161081/v1","URL":"https://doi.org/10.21203/rs.3.rs-8161081/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-8376153/v1","type":"article-journal","title":"ProtoMFL: A Robust Multimodal Federated Learning Framework via Cross-Modal Prototype Integration","abstract":"Abstract Multimodal federated learning (MFL) has made substantial progress in aggregating multimodal knowledge across distributed environments. However, it still encounters persistent challenges caused by modality-missing data at the client level. Traditional knowledge distillation–based approaches provide limited performance in handling these modality-missing scenarios. To mitigate the performance degradation caused by modality dropout, this paper proposes a prototype-based multimodal federated learning framework, termed Prototype-based Multimodal Federated Learning (ProtoMFL). By replacing sample-level representations with category-level prototypes as knowledge carriers, ProtoMFL enables more efficient cross-modal knowledge aggregation. The ProtoMFL framework consists of three core components. Cross-Modal Prototype Regularisation reduces distributional discrepancies between client and global models. Cross-Modal Prototype Contrast enhances the aggregation of similar prototypes and separation of dissimilar ones through contrastive learning. Cross-Modal Alignment enforces semantic alignment between modalities at the feature level, thereby mitigating the adverse effects of modality dropout. Experimental results show that ProtoMFL significantly outperforms existing methods in both accuracy and robustness across multiple benchmark datasets. Even under severe modality dropout, ProtoMFL maintains stable performance, achieving an average improvement of approximately 2.8% over the baseline CreamFL model without prototype mechanisms. This improvement effectively mitigates model drift issues caused by heterogeneous modalities.","author":[{"family":"Zhang","given":"Junsun"},{"family":"Sun","given":"Chaochao"},{"family":"Peng","given":"Yuan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-8376153/v1","URL":"https://doi.org/10.21203/rs.3.rs-8376153/v1","source":"europepmc"},{"id":"doi:10.22541/au.176479439.99762263/v1","type":"article-journal","title":"Federated Learning Approach Using Transfer Learning Architectures for Lung Cancer Detection","abstract":"There have been many advancements in the field of medical imaging but even then, accurate cancer detection remains a challenge because of limited labelled data. In the research field, a lot of work is already done for this task, using pre-trained features from prominent architectures like VGG16, ResNet50, MobileNetV2, InceptionV3, and DensNet121. These approaches face the issue of privacy of patients' sensitive information and unnecessary latency of exchange of data from nodes to sever, so in this paper, we use Federated Learning that enables collaborative learning across geographically distributed medical institutions. Additionally, we implement differential privacy techniques to obscure the patients' identities which would further enhance privacy protection. This paper also presents the evaluation of effectiveness of different transfer learning architectures within the FL setting, comparing their performance with centralized learning and standalone transfer learning approaches. This work adds a new direction to cancer detection with improved privacy protection, leading to earlier intervention for cancer detection.","author":[{"family":"Choure","given":"Purvi"},{"family":"Prajapat","given":"Shaligram"},{"family":"Berwal","given":"Krishan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.22541/au.176479439.99762263/v1","URL":"https://doi.org/10.22541/au.176479439.99762263/v1","source":"europepmc"},{"id":"doi:10.21203/rs.3.rs-7978012/v1","type":"article-journal","title":"A Robust Federated Learning Method for Data Heterogeneity with Enhanced Momentum-guided Aggregation","abstract":"Abstract Federated learning (FL) is an emerging distributed machine learning paradigm that enables multiple edge devices to collaboratively train a model for a specific task while preserving privacy. Yet, due to the non-independent and identically distributed (Non-IID) data dispersed on edge devices, FL suffers from slow convergence and low accuracy. This paper focuses on this problem, and proposes a robust method through enhanced data sharing. Specifically, this paper adopts feature distillation to obtain the performance sensitive features, which are used to generating proxy data for initializing public data before FL training. Meanwhile, We use a momentum-guided strategy for parameter aggregation. In order to evaluate the performance of the method, this paper also conducts many experiments. As demonstrated by the results, the method could outperform the state-of-the-art methods by 5.31% in terms of performance.","author":[{"family":"Nie","given":"Linhai"},{"family":"Wang","given":"Jin"},{"family":"Hu","given":"Naixuan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7978012/v1","URL":"https://doi.org/10.21203/rs.3.rs-7978012/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7109247/v1","type":"article-journal","title":"Model Poisoning Attacks to Federated Learning based on Fake Clients","abstract":"Abstract The increasing use of decentralized and anonymous networks creates vast amounts of darknet traffic, offering opportunities to enhance network security by detecting threats, filtering malicious activity, and identifying anomalies through improved traffic classification. Federated Learning (FL) presents a promising approach for decentralized data processing, allowing models to be trained across distributed devices while preserving data privacy. However, FL is vulnerable to poisoning attacks, where adversarial clients can degrade the performance of the global model. In this paper, we utilize a rich dataset that captures encrypted darknet traffic to develop new methods for defending against model poisoning attacks.We propose novel attack strategies based on fake clients and gradient inversion: Model Poisoning Attack based on Fake Clients (MPAF), Gradient Descent Inversion Attack (GDIA), and Selective Aggregation Poisoning Attack (SAPA). Alongside these attacks, we introduce two defense strategies: Adaptive Weighting in Aggregation (AWA) and Statistical Outlier Filtering (SOF). Experimental results show that attacks like GDIA can drastically reduce accuracy to 0% and the MPAF attack reduces accuracy to approximately 32.38%. The AWA defense notably restores accuracy under GDIA to around 80.95% and under MPAF to about 93.33%, clearly outperforming SOF. After refining the attack implementations by strengthening the base model for MPAF and reducing the intensity of GDIA, MPAF became significantly stronger, bringing accuracy down to 0%. However, GDIA exhibited more controlled degradation, with AWA defense still effectively stabilizing accuracy at approximately 72.06%.","author":[{"family":"Ghahremani","given":"Mani"},{"family":"Metwally","given":"Alan"},{"family":"Taheri","given":"Rahim"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7109247/v1","URL":"https://doi.org/10.21203/rs.3.rs-7109247/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-6920944/v1","type":"article-journal","title":"MMVO-SHFL: A Fair and Efficient Hierarchical Federated Learning","abstract":"Abstract Federated learning (FL) enables collaborative model training without centralizing data. However, the traditional FL framework is cloud-based and suffers from high communication latency. On the other hand, the edge-based FL framework, although reducing communication latency by leveraging edge servers, suffers from degraded model accuracy due to the limited data access of these servers. To overcome these limitations, this work introduces a novel hierarchical federated learning framework named MMVO - SHFL. It incorporates a bandwidth prediction based on LSTM, a unique MAB - Driven dynamic client selection strategy and an MVO - Guided model parameter optimization mechanism. Extensive experiments show that MMVO-SHFL significantly improves model convergence speed while also enhancing model accuracy. Compared to traditional methods, MMVO-SHFL not only reduces energy consumption but also significantly improves the fairness of client participation. MMVO-SHFL outperforms existing methods across various configurations, highlighting its great potential for large-scale heterogeneous federated learning scenarios. Moreover, through grid search optimization of hyperparameters G and β , the optimal combination ( G = 0.8, β = 0.1) is determined to maximize its performance. MMVO-SHFL provides a more efficient, energy-saving, and fair solution for large-scale heterogeneous FL scenarios.","author":[{"family":"Liu","given":"Xia"},{"family":"Wang","given":"Jianping"},{"family":"Chen","given":"Danyang"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-6920944/v1","URL":"https://doi.org/10.21203/rs.3.rs-6920944/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7339691/v1","type":"article-journal","title":"FedGlu: A personalized federated learning-based glucose forecasting algorithm for improved performance in glycemic excursion regions FedGlu: Personalized federated-learning based glucose forecasting algorithm","abstract":"Abstract Background: Continuous glucose monitoring (CGM) devices allow real-time glucose readings leading to improved glycemic control. However, glucose predictions in the lower (hypoglycemia) and higher (hyperglycemia) extremes, referred as glycemic excursions, remain challenging due to their rarity. Moreover, limited access to sensitive patient data hampers the development of robust machine learning models even with advanced deep learning algorithms available. Methods: We propose to simultaneously provide accurate glucose predictions in the excursion regions while addressing data privacy concerns. To tackle excursion prediction, we propose a novel Hypo-Hyper (HH) loss function that penalizes errors based on the underlying glycemic range with a higher penalty at the extremes over the normal glucose range. On the other hand, to address privacy concerns, we propose FedGlu, a machine learning model trained in a federated learning (FL) framework. FL allows collaborative learning without sharing sensitive data by training models locally and sharing only model parameters across other patients. The HH loss combined within FedGlu addresses both the challenges at the same time. Results: The HH loss function demonstrates a 46% improvement over mean-squared error (MSE) loss across 125 patients. Compared to local models, FedGlu improved glycemic excursion detection by 35% compared to local models. This improvement translates to enhanced performance in predicting both, hypoglycemia and hyperglycemia, for 105 out of 125 patients. Conclusions: These results underscore the effectiveness of the proposed HH loss function in augmenting the predictive capabilities of glucose predictions. Moreover, implementing models within a federated learning framework not only ensures better predictive capabilities but also safeguards sensitive data concurrently.","author":[{"family":"Darpit","given":"Dave"},{"family":"Vyas","given":"Kathan"},{"family":"Jayagopal","given":"Jagadish"},{"family":"Garcia","given":"Alfredo"},{"family":"Erraguntla","given":"Madhav"},{"family":"Lawley","given":"Mark"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7339691/v1","URL":"https://doi.org/10.21203/rs.3.rs-7339691/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7644110/v1","type":"article-journal","title":"Efficient Federated Learning Based On Domain Adaptation and Knowledge Distillation Losses","abstract":"Abstract Numerous devices nowadays generate vast amounts of data for learning. Traditional centralized learning necessitates transmitting all data to a central site, which conducts the model training. However, much of these data may be sensitive, leading customers to refuse to share it. Federated Learning (FL) addresses this dilemma by employing a distributed learning framework where multiple local users collaborate to train a shared model via the central server's coordination. Nevertheless, reducing communication costs with respect to computational costs and efficiently handling non-independent and identically distributed (non-IID) problems still present significant struggles. Therefore, we propose an efficient FL method using domain adaptation and knowledge distillation losses to solve the abovementioned issues. Experimental results implemented on MNIST, CIFAR-10, and CIFAR-100 datasets demonstrate that our method can achieve almost the same accuracy as the other well-known FL methods using fewer communication rounds, particularly for non-IID situations.","author":[{"family":"Liu","given":"Jui"},{"family":"Ku","given":"Cooper"},{"family":"Wang","given":"Shao"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7644110/v1","URL":"https://doi.org/10.21203/rs.3.rs-7644110/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7866368/v1","type":"article-journal","title":"Privacy-Preserving and Communication-Efficient Federated Learning for Cloud-Scale Distributed Intelligence","abstract":"Abstract This study focuses on privacy protection and multi-party collaborative optimization in cloud computing environments. A federated learning framework is proposed, integrating differential privacy mechanisms and communication compression strategies. The framework adopts a layered architecture consisting of local computing nodes, a compression module, and a privacy-enhancing module. It enables global model training without exposing raw data, ensuring both model performance and data security. During the training process, the framework uses the federated averaging algorithm as the basis for global aggregation. A Gaussian noise perturbation mechanism is introduced to enhance the model's resistance to inference attacks. To address bandwidth limitations in practical cloud computing scenarios, a lightweight communication compression strategy is designed. This helps reduce the overhead and synchronization pressure caused by parameter exchange. The experimental design includes sensitivity analysis from multiple dimensions, such as network bandwidth constraints, client count variation, and data distribution heterogeneity. These experiments validate the adaptability and robustness of the proposed method under various complex scenarios. The results show that the method outperforms existing approaches in several key metrics, including accuracy, communication rounds, and model size. The proposed approach demonstrates strong engineering deployability and system-level security. It provides a novel technical path for building efficient and trustworthy distributed intelligent systems.","author":[{"family":"Liu","given":"Heyao"},{"family":"Kang","given":"Yue"},{"family":"Liu","given":"Yuchen"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7866368/v1","URL":"https://doi.org/10.21203/rs.3.rs-7866368/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7633900/v1","type":"article-journal","title":"Lightweight Federated Learning with Genetic Optimization for PM 2.5 Forecasting in IoT Networks","abstract":"Abstract This study presents a lightweight and privacy-preserving federated learning framework designed for resource-constrained IoT sensor networks, emphasizing efficient distributed computation across heterogeneous devices. The framework enables decentralized training of LSTM models directly on IoT nodes, eliminating the need for centralized data aggregation while ensuring full data privacy. To enhance computational efficiency and reduce communication overhead, a Genetic Algorithm-based model compression method is applied, pruning redundant weights and achieving approximately a 37% reduction in model size. The framework is evaluated on a real-world time-series forecasting task, achieving 66.3% classification accuracy across multiple categories, while also demonstrating low latency and high scalability in large-scale heterogeneous deployments. Furthermore, this approach supports real-time operation and effective management of hardware resource constraints, enabling practical deployment in distributed networks. These results highlight the potential of combining federated deep learning with evolutionary optimization to build efficient, secure, and scalable distributed IoT systems, providing a robust blueprint for future grid and edge computing applications.","author":[{"family":"Nazari","given":"Hadi"},{"family":"Farjami","given":"Yaghoub"},{"family":"Taeizadeh","given":"Ali"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7633900/v1","URL":"https://doi.org/10.21203/rs.3.rs-7633900/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7250878/v1","type":"article-journal","title":"FedCKD: Cluster-Aware Knowledge Distillation for Heterogeneous Medical Federated Learning","abstract":"Abstract In medical knowledge systems, federated learning provides a promising paradigm for collaborative knowledge extraction while preserving data privacy. However, inherent heterogeneity in medical information—stemming from variations in disease distribution, imaging protocols, and patient demographics—severely degrades the performance of traditional federated frameworks. To address this challenge, we propose FedCKD , a knowledge-driven federated framework tailored for heterogeneous medical information systems. FedCKD introduces three key innovations: (1) a label-driven knowledge clustering mechanism that partitions medical nodes based on disease-specific knowledge representations, ensuring intra-cluster semantic consistency; (2) a two-stage adaptive aggregation strategy for knowledge-oriented model fusion within each cluster, balancing local specialization and cluster-level consistency; (3) a cross-cluster knowledge distillation protocol that enables privacy-preserving transfer of complementary knowledge across specialized medical domains via weighted teacher ensembles. By simulating interoperability in distributed medical systems, FedCKD achieves cross-domain knowledge integration while respecting statistical heterogeneity. Comprehensive experiments on multiple datasets demonstrate that FedCKD significantly outperforms state-of-the-art methods, establishing it as an effective solution for knowledge extraction and integration in privacy-sensitive, heterogeneous medical ecosystems.","author":[{"family":"Xu","given":"Yi"},{"family":"Chen","given":"Kun"},{"family":"Luo","given":"Haoyu"},{"family":"Liu","given":"Xiao"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7250878/v1","URL":"https://doi.org/10.21203/rs.3.rs-7250878/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7008997/v1","type":"article-journal","title":"HyBloFED: A Hybrid Blockchain Integrated Federated Learning Approach for Brain Tumor Classification","abstract":"Abstract Brain tumors, complex and potentially devastating, demand precise classification for effective patient prognosis and treatment planning. This paper introduces a novel approach to automate brain tumor classification using deep learning techniques, particularly convolutional neural networks (CNNs). However, conventional centralized methods compromise patient privacy and data security. To address this issue, federated learning (FL), a collaborative paradigm enabling model training across multiple institutions and aggregating models at a central server while preserving the confidentiality of sensitive medical data, is proposed. Moreover, an aggregation function at the central server is modified to identify the effect of aggregation on global model training. In addition to that, Blockchain technology is also integrated with FL architecture to enhance privacy preservation and trust, to ensure the integrity and immutability of patient data. By synergistically integrating modified FL, CNNs, and Blockchain technology, the proposed approach achieves accuracy (98%) and security in brain tumor classification. Through this, it aims to advance the field of medical imaging while prioritizing patient privacy and data security (through Blockchain technology) in brain tumor diagnosis and treatment.","author":[{"family":"Shrimali","given":"Bela"},{"family":"Joshi","given":"Sarthak"},{"family":"Patel","given":"Hiren"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7008997/v1","URL":"https://doi.org/10.21203/rs.3.rs-7008997/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7627809/v1","type":"article-journal","title":"Hierarchical Personalized Continual Federated Learning for Real Time Risk Prediction of Chronic Diseases","abstract":"Abstract The most prevalent global morbidity and mortality is chronic illnesses like cardiovascular diseases, diabetes, and respiratory diseases which require the proper prediction of risks and in a timely manner so as to have preventative measures. Nevertheless, predictive systems in real-time have been found to be severely limited by fragmented healthcare data, patient population heterogeneity, non-independent and identically distributed (non-IID) data distributions, and strict privacy policies that cannot allow direct data sharing. Current centralized systems tend to perform poorly when it comes to generalizing across dissimilar healthcare locations, and traditional federated learning algorithms have a scalability bottleneck, suboptimal communication, inadequate personalization, and susceptibility to data drift with time, rendering them unsuitable to real-world application. To overcome these obstacles, we suggest a Hierarchical Personalized Continual Federated Learning (HiPerC-FL) model of real-time risk prediction of chronic diseases that incorporates multi-modal input of wearable, electronic health records, and imaging data without having to reveal raw patient data. The system uses a hierarchical aggregation topology between edge devices, hospital servers, and global coordinators to reduce the latency and communication and a personalized meta-learning module coupled with client clustering helps to address the impact of data heterogeneity. Moreover, on-device adaptation that happens continuously allows local models to be immune to concept drift, and a causal feature regularizer makes predictions more interpretable and reliable. Secure aggregation, differential privacy, and verifiable audit trails are the means of implementing privacy and governance, and both adhere to clinical standards. Benchmark healthcare simulation Experimental results on benchmark healthcare data show that, compared to baseline federated methods, HiPerC-FL always yields progress of 7–10 percent in predictive accuracy, is 50 percent more cost-effective in communication, and GUI remains stable under extended distribution shifts. This evidence confirms that the given framework proves to be not only effective but also practically deployable to the real-time chronic disease monitoring process, which can serve as a scalable and ethically-acceptable roadmap to the precision of the healthcare provision.","author":[{"family":"Ghoshal","given":"Abhigyan"},{"family":"Ali","given":"Mohammad"},{"family":"Sambath","given":"M"},{"family":"Balraj","given":"E"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7627809/v1","URL":"https://doi.org/10.21203/rs.3.rs-7627809/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7406272/v1","type":"article-journal","title":"Optimal Uncertainty Budget Allocation for Robust Federated Learning under Byzantine Attacks","abstract":"Abstract Federated learning (FL) has revolutionized the development of machine learning models by enabling decentralized training while safeguarding user privacy. However, the presence of Byzantine adversaries introduces significant vulnerabilities, as malicious clients can disrupt the learning process by providing misleading updates. This paper addresses the critical challenge of allocating an uncertainty budget across heterogeneous clients to enhance the robustness of federated learning systems against such adversarial attacks. We introduce the Uncertainty Budget Allocation Problem (UBAP), formulating it as a mixed-integer nonlinear program (MINLP) aimed at optimizing resource distribution for improved model convergence and stability. Our framework not only rethinks traditional assumptions about client contributions but also presents a novel mathematical analysis underlying the relationship between uncertainty allocation and adversarial strength. Extensive empirical evaluations on standard benchmarks demonstrate substantial improvements in model performance and resistance to attacks, showcasing the practical efficacy of our approach. Through this work, we underscore the importance of optimal uncertainty budget allocation to foster resilience in federated learning systems, paving the way for further innovations in this domain and enhancing the security of decentralized AI applications.","author":[{"family":"Lian","given":"Weiwei"},{"family":"Tao","given":"Jun"},{"family":"Mei","given":"Xinjun"},{"family":"Fang","given":"Yu"},{"family":"Shen","given":"Zhou"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7406272/v1","URL":"https://doi.org/10.21203/rs.3.rs-7406272/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-6827986/v1","type":"article-journal","title":"Enhanced Security Verifiable Secure Aggregation Scheme in Federated Learning","abstract":"Abstract Federated Learning(FL) enables multiple participants to build a loosely coupled distributed machine learning system under the coordination of a central server. Existing FL models typically assume that the server aggregating data is semi-honest, but this assumption does not align with the complexities of real-world application environments, where the server may carry out collusion attacks or replay attacks. VerifyNet is a representative federated learning protocol for verifiable secure aggregation. In this paper, we analyze the security of VerifyNet, identify two shortcomings: low tolerance to collusion attacks and inability to resist combinatorial replay attacks. Furthermore, we have experimentally confirmed the existence of these two security vulnerabilities. To address the issue of low tolerance for collusion attacks, we have constructed a secure homomorphic hash function key generator using a randomized approach to prevent malicious servers from obtaining shared keys and forging data. To address the issue of being unable to resist replay attacks, we have constructed a secure additional verification information generation algorithm using AES-CTR encryption mode, which prevents malicious servers from obtaining increments from historical data and constructing combinatorial replay attacks. Security analysis shows that our scheme effectively achieves privacy protection and aggregation verification. We tested the performance of the scheme in a local area network environment. Experimental data indicates that when the number of clients is 500 and the number of gradients per client is 5000, our scheme only requires an additional 5.76‰ computational overhead and 3.46% communication overhead compared to the VerifyNet protocol, and eliminates the security vulnerabilities of collusion attacks and combinatorial replay attacks.","author":[{"family":"Yao","given":"Wujun"},{"family":"Han","given":"Yiliang"},{"family":"Zhou","given":"Tanping"},{"family":"Wang","given":"Xiaolin"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-6827986/v1","URL":"https://doi.org/10.21203/rs.3.rs-6827986/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7052572/v1","type":"article-journal","title":"Medical support platform for melanoma analysis and detection based on Federated Learning","abstract":"Abstract Advances in computer science and medicine have led to the emergence of artificial intelligence as a key tool in the medical and scientific fields. Its application in the diagnosis and treatment of diseases, such as cancer, has proven to be fundamental in improving early detection and saving lives. This article presents a proposal based on Deep Learning to develop a model capable of detecting melanomas in the skin from clinical images. The aim is to provide doctors with a tool to support early identification of this type of cancer, considering additional factors such as sun exposure and the patient's skin tone. To optimize diagnostic accuracy and avoid information dispersion, a collaborative learning technique called Federated Learning is implemented. This technique allows models trained locally by doctors to be synchronized with a global model that will be updated periodically, ensuring continuous improvement of the system without compromising the privacy of patient data. In addition, a web application is presented to manage and process the information efficiently, making it easier for doctors to consult and analyze the results.","author":[{"family":"Laso","given":"Sergio"},{"family":"Herrera","given":"Juan"},{"family":"Flores-Martin","given":"Daniel"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7052572/v1","URL":"https://doi.org/10.21203/rs.3.rs-7052572/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-6979939/v1","type":"article-journal","title":"Federated Learning for Secure and Privacy- Preserving Edge AI in Smart Cities","abstract":"Abstract The rapid expansion of smart cities has led to the integration of Artificial Intelligence (AI) at the edge, enabling real-time decision-making for intelligent urban infrastructure. However, conventional centralized AI models pose critical challenges, including data privacy risks, security vulnerabilities, and high computational overhead. This paper investigates Federated Learning (FL) as a transformative paradigm to enhance security, privacy, and efficiency in edge AI systems for smart cities. Unlike traditional AI training methods, to cyber threats while ensuring compliance with data protection regulations. To address key challenges in heterogeneous smart city environments, we propose a hybrid optimization framework integrating differential privacy, secure multi-party computation (SMPC), and blockchain-based authentication. This approach strengthens resilience against adversarial attacks while ensuring secure model updates. Additionally, we introduce an adaptive aggregation mechanism, which dynamically adjusts model updates based on device reliability, data distribution, and network conditions, optimizing both learning efficiency and energy consumption in edge AI networks. Extensive experimentation on real-world smart city datasets demonstrates that the proposed framework enhances model accuracy, robustness, and privacy preservation compared to conventional AI approaches. Our findings establish Federated Learning as a cornerstone for secure, scalable, and privacy-aware AI in smart cities, facilitating trustworthy deployment of intelligent urban infrastructure. This research provides valuable insights for policymakers, researchers, and industry professionals, paving the way for next-generation AI-driven smart cities with enhanced security, privacy, and efficiency.","author":[{"family":"Joshi"},{"family":"Fatima","given":"Shahin"},{"family":"Hanirvesh","given":"Kesani"},{"family":"Siddiqui","given":"Shadab"},{"family":"Hazra","given":"Sumit"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-6979939/v1","URL":"https://doi.org/10.21203/rs.3.rs-6979939/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-5858510/v1","type":"article-journal","title":"Quantization-Based Chained Privacy-Preserving Federated Learning","abstract":"Abstract Federated Learning (FL) is an advanced distributed machine learning framework crucial in protecting data privacy and security. By enabling multiple participants to train models while keeping their data local collaboratively, FL effectively mitigates the risks associated with centralized storage and sharing of raw data. However, traditional FL schemes face significant challenges regarding communication efficiency, computational costs, and privacy preservation. For instance, its communication and computational overhead in edge computing scenarios is often excessively high, hindering real-time applications. This paper proposes an innovative federated learning framework, Q-Chain FL, integrating quantization compression techniques into a chained FL architecture. This Q-Chain FL scheme adopts efficient compression and transmission of model parameter differences at the user node and executes seamless decompression and aggregation at the server node. Experiments on several publicly available datasets, including MNIST, CIFAR-10, and CelebA, demonstrate low communication and computational overhead, fast convergence speed, and high security of Q-Chain FL. Compared to traditional FedAvg and Chain-PPFL, Q-Chain FL reduces communication overhead by approximately 62.5\\% and 44.7\\%, respectively. These results underscore the robustness and adaptability of Q-Chain FL in various datasets and real-world learning scenarios.","author":[{"family":"Liu","given":"Ya"},{"family":"Wu","given":"Shumin"},{"family":"Li","given":"Yibo"},{"family":"Zhao","given":"Fengyu"},{"family":"Ren","given":"Yanli"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-5858510/v1","URL":"https://doi.org/10.21203/rs.3.rs-5858510/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-6658077/v1","type":"article-journal","title":"Resisting Against  Targeted Poisoning Attacks in Lightweight Privacy-Preserving Federated Learning","abstract":"Abstract Federated learning is a distributed computing paradigm designed to protect client privacy. However, its distributed nature makes it vulnerable to targeted poisoning attacks.Although existing solutions can effectively mitigate such attacks, they often struggle to handle statistical heterogeneity.Moreover, privacy attacks often coexist with targeted poisoning attacks in federated learning, further increasing the difficulty of defense.To address the above challenges, this paper proposes a lightweight privacy-preserving federated learning framework, named FedSP, to defend against targeted poisoning attacks. The key idea is to design a protocol between two servers to detect and aggregate model updates submitted by clients in a perturbed form. Specifically, we design an adaptive clustering strategy during aggregation to mitigate inconsistencies of model updates caused by statistical heterogeneity.Additionally, we employ a dimensionality reduction to identify a plausible model update, eliminating assumptions regarding the proportion of malicious clients and the root dataset.Theoretical analysis demonstrates the privacy preservation and convergence of FedSP.Extensive experiments show that FedSP effectively defends against targeted poisoning attacks without compromising privacy.","author":[{"family":"Zhang","given":"Hongliang"},{"family":"Xie","given":"Haojie"},{"family":"Lv","given":"Jiandong"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-6658077/v1","URL":"https://doi.org/10.21203/rs.3.rs-6658077/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-6090375/v1","type":"article-journal","title":"Verifiable Secure Aggregation Scheme for Privacy Protection in Federated Learning","abstract":"Abstract Federated learning enables multiple participants to construct a distributed machine learning system coordinated by server. Most existing solutions assume a semi-honest system, considering each participant to be honest but curious, which does not align with the complex real-world environment. In reality, servers might be malicious, potentially tampering with or forging aggregation results. To verify the integrity of server aggregation computations while protecting the privacy of clients, this paper introduces a privacy-preserving verifiable secure aggregation scheme for federated learning networks. Initially, we construct a functional reuse private key ring generation algorithm, enabling clients to encrypt and protect their private gradients using the private key ring. Subsequently, leveraging the discrete logarithm difficulty problem, we devise a commitment protocol where clients commit to their encrypted private gradients. Upon receiving the aggregation result from the server, they collaboratively unlock the commitment, thereby verifying the aggregation result. Security analysis demonstrates that our solution effectively ensures privacy protection. We simulated consumer electronic products on the Raspberry Pi and tested the performance of the solution. Experimental data reveals that, with 100 clients, our scheme demonstrates that the overhead for proof generation and verification computations are 39.9% and 34.1% of the existing scheme, respectively, highlighting its lightweight nature.","author":[{"family":"Yao","given":"Wujun"},{"family":"Zhou","given":"Tanping"},{"family":"Han","given":"Yiliang"},{"family":"Wang","given":"Xiaolin"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-6090375/v1","URL":"https://doi.org/10.21203/rs.3.rs-6090375/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7907693/v1","type":"article-journal","title":"DEFEND: Intelligent Temporal Backdoor Detection and Mitigation in Federated Learning via Reinforcement Learning-Coordinated Multi-Layer Defense","abstract":"Abstract Collaborative machine learning in financial systems faces an escalating security threat: temporal backdoor attacks that exploit multi-round dependencies to systematically compromise fraud detection and risk assessment models—a challenge that existing static defense mechanisms cannot adequately counter. This paper presents DEFEND (DEep Federated Ensemble Network Defense), a comprehensive framework that integrates multi-layer defense with reinforcement learning-based adaptive coordination to counter sophisticated temporal backdoor strategies in federated learning environments. The framework introduces four key innovations: (1) Temporal Behavioral Analysis Layer employing multi-scale statistical profiling with dynamic time warping for attack pattern recognition across communication rounds, (2) Byzantine-Robust Statistical Aggregation using geometric median estimation with adaptive outlier detection, (3) Multi-Scale Validation Protocol with automated model rollback mechanisms, and (4) MDP-based Defense Coordination formulating security decisions as a Markov Decision Process optimized via Proximal Policy Optimization to dynamically balance robustness and utility. Extensive experiments on the FinMultiTime dataset across three distinct market periods (2009-2025) demonstrate superior performance over state-of-the-art baselines, achieving defense success rates of 95.6\\%$\\pm$1.0\\% for ResNet-18 and 94.0\\%$\\pm$1.2\\% for MobileNet-V2 while maintaining clean accuracy above 85\\%. Ablation studies reveal that the MDP-based coordination provides the largest individual contribution (8.2\\% defense success rate improvement), while the complete multi-layer architecture achieves up to 18.7\\% improvement over single-layer baselines. Cross-period generalization analysis demonstrates robust transferability with less than 6\\% performance degradation across different market regimes, validating practical deployment viability in dynamic financial environments.","author":[{"family":"Liu","given":"Wenan"},{"family":"Yang","given":"Qixuan"},{"family":"Gong","given":"Weihang"},{"family":"Yin","given":"Rongji"},{"family":"Li","given":"Zheng"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7907693/v1","URL":"https://doi.org/10.21203/rs.3.rs-7907693/v1","source":"europepmc"},{"id":"doi:10.21203/rs.3.rs-6848848/v1","type":"article-journal","title":"BlockFed: Blockchain-based Privacy Preserving Federated Learning for 5G-assisted Healthcare Ecosystems","abstract":"Abstract With the rapid adoption of 5G networks and the growing reliance on digital healthcare, the need for secure, efficient, and privacy-aware data processing has become increasingly critical. This paper presents a novel approach BlockFed , that integrates Federated Learning (FL) with Blockchain (BC) technology to ensure data privacy and model integrity in 5G-assisted healthcare ecosystems. In the proposed system, Patient Health Record (PHR) remains at local Healthcare Entities (HE) such as hospitals and research centers, and only encrypted model updates are shared, effectively preserving user privacy. BC is employed to record and verify model weight transactions, providing tamper-proof integrity and transparency among participating HE. To mitigate the high storage demands of BC, the InterPlanetary File System (IPFS) is utilized for off-chain storage of model weights. Additionally, a lightweight homomorphic encryption scheme is incorporated to protect model parameters during aggregation and transmission. This integrated approach offers a scalable and trustworthy solution for collaborative healthcare intelligence while safeguarding sensitive PHR. Experimental insights and theoretical validation demonstrate the system’s potential for practical deployment in next-generation healthcare infrastructures.","author":[{"family":"Verma","given":"Ashwin"},{"family":"Pathak","given":"Sunil"},{"family":"Bhattacharya","given":"Pronaya"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-6848848/v1","URL":"https://doi.org/10.21203/rs.3.rs-6848848/v1","source":"europepmc"},{"id":"doi:10.3389/frai.2026.1825067","type":"article-journal","title":"Structural impact of non-IID heterogeneity on federated behavioral anomaly detection in IoT and IoMT systems.","abstract":"The expansion of Internet of Things (IoT) and Internet of Medical Things (IoMT) infrastructures has increased the generation of multivariate sensor streams that reflect complex operational behaviors in industrial and clinical environments. Centralized anomaly detection approaches face limitations in IoMT due to privacy constraints, latency, and device heterogeneity. Federated learning (FL) enables distributed model training without data centralization; however, its behavior under highly non-Independent and Identically Distributed (non-IID) conditions remains insufficiently understood. This study proposes a trace-level behavioral modeling approach combined with federated training via FedAvg to analyze the impact of non-IID heterogeneity on anomaly detection. An Integrated Hybrid Dataset (IHD) comprising 71,980 behavioral traces, with 22,698 used for evaluation, was constructed from Edge-IIoTset, TON_IoT, and IoMT data. The centralized model achieved F 1 = 0.981 and Recall = 0.993, while the federated model preserved discriminative capacity (AUC-ROC = 0.995) but reduced Recall to 0.530. Degradation is concentrated in IoMT (Recall = 0.290), with increased Brier Score and Expected Calibration Error, showing that preserved discrimination does not ensure operational effectiveness.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/frai.2026.1825067","URL":"https://doi.org/10.3389/frai.2026.1825067","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-9642506/v1","type":"article-journal","title":"Federated Meta-Learning in the Time-Series Domain: a Scoping Review","abstract":"Abstract Wearables, industrial sensors, and other connected systems can create time-series data that is used in making decisions, predictions, and monitoring. Nevertheless, such information can be highly sensitive, and it involves the behaviour of the users, their habits, and health status, which pose a serious privacy risk. Federated learning (FL) has become a powerful tool to improve this issue, enabling models to be trained cooperatively and retaining raw data in local devices. Simultaneously, meta-learning methods have also been developed to provide quick adaptation of models to new users or tasks, and this is especially useful in low-data settings. Despite the extensive research on both FL and meta-learning, their combination as federated meta-learning (FedMeta) to time-series applications has received comparatively limited literature coverage and is scattered across the research. This paper provides a review of the literature in which FedMeta is used in terms of time series, where the interest is specifically on privacy preservation. We do not just look at the performance of the algorithm, but we look at how these methods cope with the major real-world problems, such as client-level personalisation, data distribution heterogeneity, and performance limitations imposed by limited computation or communication resources. We further examine the privacy mechanisms that have been used in previous studies and identify instances where privacy assumptions or threat models are not well defined. The analysis is based on the PRISMA-ScR approach. Peer-reviewed publications in the past five years (2019-2024) were identified and searched in seven major academic databases, including Scopus, IEEE Xplore, SpringerLink, ACM, ScienceDirect, Web of Science, and PubMed. Out of a total of 1,551 records, 21 studies met the required inclusion criteria after screening. The review shows that majority of the available literature applies FedMeta methods within a simulated or small scale experimental system and this limits the applicability of the results. Also, the selection of datasets, metrics of evaluation, or assumed adversarial models is not standard and hence, it is not easy to compare studies. In order to deal with client heterogeneity, numerous works resort to adaptive aggregation techniques or client attention techniques, whereas efficiency is often enhanced with the help of asynchronous updates or lightweight training. Differential privacy or secure aggregation is commonly used to provide privacy protection, often with a significant decrease in model accuracy. Altogether, the studies reviewed suggest that FedMeta is a good prospect for personalized and privacy-sensitive time-series modeling. Nevertheless, the area remains in its infancy, and additional advancements will necessitate a set of common standards, a better description of privacy assumptions, and a confirmation of the results based on large-scale, longitudinal studies that would take place under real-world conditions of deployment.","author":[{"family":"Qamar","given":"Suleman"},{"family":"Zhou","given":"Ian"},{"family":"Tofigh","given":"Farzad"},{"family":"Lipman","given":"Justin"},{"family":"Abolhasan","given":"Mehran"},{"family":"Piccardi","given":"Massimo"}],"issued":{"date-parts":[[2026]]},"DOI":"10.21203/rs.3.rs-9642506/v1","URL":"https://doi.org/10.21203/rs.3.rs-9642506/v1","source":"europepmc"},{"id":"doi:10.3389/frai.2026.1807960","type":"article-journal","title":"U-SplitDoRA: an improved privacy-preserved U-shaped split parameter-efficient fine-tuning framework through weight decomposition for large language models.","abstract":"As large language models (LLMs) are getting bigger with respect to the parameter count, ranging from a few million to billions, methods like parameter-efficient fine-tuning (PEFT) have emerged as a crucial approach for adapting these LLMs, such as GPT, Llama, and DeepSeek, to resource-constrained and privacy-sensitive environments. The robustness of large language models (LLMs) while operating on complex tasks and with large datasets makes them feasible for various application domains. This also demands the availability of more public datasets to train LLMs in the future. The federated learning (FL) technique, where several entities collaboratively train a machine learning model without sharing their data, is a widely adopted decentralized training framework. This is followed by a central server, which aggregates the models to create a global model. FL LLM fine-tuning has gained attention recently to overcome the aforementioned training data scarcity issue. LLMs are collaboratively fine-tuned by several data owners without disclosing their private data. The large number of trainable parameters has a direct effect on training such complex models on the client side. The split learning technique, through model partitioning, solves the training overhead by offloading certain training tasks to the server side. Previous research based on the split learning approach for FL LLM fine-tuning, namely SplitLoRA and HSpliLoRA, sets the foundation for further research in this direction. Frameworks like SplitLoRA have already enabled collaborative fine-tuning through model partitioning, but privacy preservation and adaptation quality remain open research challenges. U-SplitDoRA-an improved privacy preserved U-shaped split parameter-efficient fine-tuning framework through weight decomposition for large language models-is proposed. U-SplitDoRA harnesses the parallelization power of FL through the split learning approach, using weight-decomposed low-rank adaptation (DoRA) as the PEFT technique. To further address privacy concerns, the U-shaped paradigm is adapted while splitting the model. By partitioning the model into three parts (head, body, and tail), with the head and tail remaining on the client side while the body is on the server side, it ensures that neither raw data nor labels are exposed to the server, thus providing strong privacy. Additionally, replacing low-rank adaptation (LoRA) with DoRA as the PEFT method further enhances adaptation, as it updates both the magnitude and direction of weights, resulting in superior expressiveness and reducing the gap between PEFT fine-tuning and full parameter fine-tuning to a minimal margin. Experiments are conducted using GPT-2-S and GPT-2-M trained on the E2E benchmark dataset. The simulation results confirm that U-SplitDoRA attains better accuracy scores and convergence speed than other SOTA LLM fine-tuning frameworks. Thus, the proposed method addresses key gaps in privacy and adaptation quality, paving the way for efficient, robust, and privacy-preserving fine-tuning of LLM models in a distributed setting.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/frai.2026.1807960","URL":"https://doi.org/10.3389/frai.2026.1807960","source":"pubmed"},{"id":"doi:10.3389/fnins.2026.1827009","type":"article-journal","title":"Federated training of spiking neural networks on edge hardware for audio processing.","abstract":"Spiking Neural Networks have caught significant attention recently for their potential for energy-efficient computation on neuromorphic hardware and their event-driven processing. Spiking Neural networks employ spike-based learning paradigms, which require specialized training procedures such as Surrogate Gradient Descent. At the same time, Federated Learning allows collaborative model training on decentralized devices with preservation of data privacy protection. However, to date, few research has examined the suitability of Federated learning with ARM-based hardware. This work primarily investigates whether Federated Spiking Neural Networks training on ARM-based hardware is feasible with the Raspberry Pi 5 as a widely available and low-cost edge computing device for audio signal processing tasks. We perform a comparative analysis of federated Spiking Neural Network and federated convolutional neural networks on ARM processors and evaluate their performance on different data partitioning strategies using Dirichlet-based splits and various federated averaging algorithms. Using Federated learning, this work investigates the impact of data heterogeneity and aggregation strategies on model convergence, communication overhead, and latency in distributed training paradigms. The results provided showcases the important insights into the trade-offs of FL-SNN implementations on Von Neumann architectures and their applications in decentralized neuromorphic computing for audio processing.","author":[{"family":"Ss","given":"Kaimal"},{"family":"Ss","given":"Reka"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/fnins.2026.1827009","URL":"https://doi.org/10.3389/fnins.2026.1827009","source":"pubmed"},{"id":"doi:10.3390/healthcare14121612","type":"article-journal","title":"From Integrated Care to Learning Systems.","abstract":"Integrated care is increasingly shaped by digital infrastructures, data governance, and AI-enabled analytics, yet the relevant literature remains fragmented across health-services research, digital health, and machine learning. This article reports a scoping review, conducted in line with PRISMA-ScR guidance, that maps how integrated care models have evolved conceptually, what digital and AI-enabled infrastructures support them, how their clinical, economic, and equity impacts can be evaluated, and what current implementations imply for sustainable scaling. We searched PubMed, Scopus, Semantic Scholar, and Crossref (retrieval date 31 October 2025; forward screening to 31 March 2026) and added grey literature from named policy bodies. The searches identified 15,189 records, reducing to 11,789 after intra- and cross-source deduplication and grey-literature integration; 620 full texts were assessed and 192 were included in the synthesis. Four domains were synthesised: conceptual foundations of integrated care, AI and multimodal analytics, implementation barriers, and digital-governance foundations. We chart the field using a Type I-V maturity scheme (disease, cohort, whole-system, digital-integrated, learning), benchmarked against the Rainbow, MacColl, EMRAM/AMAM, and NHS ICS models. Most deployments cluster at digitally integrated but only weakly adaptive Type IV; recurrent failure modes-temporal blind spots, maintenance debt, semantic drift, and governance gaps-block progression to Type V, and high-profile clinical-AI failures illustrate the cost of attempting Type V analytics on Type IV-or-worse infrastructure. A walk through nine world regions maps each to its current Type I-V position and shows that organisational and payment integration-not digital sophistication alone-is currently the dominant driver of progress. The COMFORTage Integrated Care Model Library is positioned as a workflow of AI agents orchestrating predictive, preventive, and personalised care across the integrated-care lifecycle rather than as a single federated-learning programme. The review positions AI-enabled integrated care less as a finished model than as an emerging design space requiring longitudinal data assets, stewarded model lifecycles, accountable governance, and outcome-based contracting for clinically useful, equitable, and trustworthy learning systems.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/healthcare14121612","URL":"https://doi.org/10.3390/healthcare14121612","source":"pubmed"},{"id":"doi:10.3389/fnbot.2026.1649168","type":"article-journal","title":"Robust federated learning for UAV object detection: a joint self-distillation and drift compensation approach.","abstract":"The rapid advancement of unmanned aerial vehicles (UAVs) in disaster response and environmental monitoring has underscored the growing importance of real-time object detection within UAV swarm networks. However, the non-independent and identically distributed (non-IID) characteristics of data in UAV networks present significant challenges to model convergence and adaptability. To tackle these challenges, this study introduces a robust federated UAV object detection framework tailored for non-IID data distributions. The framework aims to enhance adaptability across clients, thereby improving both detection performance and convergence speed. Our approach includes a self-distillation mechanism that leverages personalized knowledge from local model historical states to guide current local training, striking a balance between specialization and adaptability. Additionally, we propose a drift compensation mechanism to synchronize local and global model updates, mitigating model drift. We conducted extensive experiments on the VisDrone2019-DET dataset, comparing our method to baseline models. Results demonstrate that our approach accelerates convergence speed by approximately 2.2 times and enhances detection performance by around 3%, offering an efficient and robust solution for UAV-based object detection under non-IID conditions.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/fnbot.2026.1649168","URL":"https://doi.org/10.3389/fnbot.2026.1649168","source":"pubmed"},{"id":"doi:10.1038/s41598-026-46141-5","type":"article-journal","title":"FeXAI: Federated and Explainable AI for cyber threat detection in IoT-enabled smart transportation systems.","abstract":"The rapid development of smart cities, fueled by the growth of the Internet of Things (IoT) and interconnected systems, has greatly enhanced urban infrastructure, especially in transportation and energy management. However, this increased connectivity also raises the risk of cyberattacks, threatening service availability, financial stability, and public safety. This study introduces a resilient cybersecurity framework designed to detect and classify various cyber threats, including DoS, DDoS, Reconnaissance, Sybil, Replay, and Spoofing attacks, targeting critical transportation systems such as the Internet of Vehicles (IoV), electric vehicle (EV) charging networks, and Vehicular Ad hoc Networks (VANETs). By combining machine learning with Federated Learning (FL), the framework effectively tackles key challenges like high computational costs, dependence on centralized data, and scalability across different IoT systems. FL improves data privacy by keeping sensitive information on edge devices, reducing concerns over centralized data storage. Moreover, TreeSHAP, an interpretability technique, is utilized to provide transparency and deeper insights into attack detection. The proposed system achieves high F1 scores of 0.980, 0.982, and 0.99 on the CICIoV2024, CICEVSE2024, and VeReMi Extension datasets, respectively, demonstrating its effectiveness on multiple IoT security datasets relevant to smart city transportation and energy systems. while safeguarding user privacy.","author":[{"family":"Aj","given":"Rufus"},{"family":"Cc","given":"Columbus"},{"family":"Ck","given":"Aravind"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-46141-5","URL":"https://doi.org/10.1038/s41598-026-46141-5","source":"pubmed"},{"id":"doi:10.1038/s41598-026-49490-3","type":"article-journal","title":"TwinGuard-Sec: a federated blockchain-enabled AI framework for standardized security and privacy in cross-domain digital twin ecosystems over 6G.","abstract":"The fast rate of cross-domain Digital Twins (DT) ecosystem growth in 6G-based scenarios poses unresolved security and privacy issues into the scope of existing frameworks. This study examines the inherent constraints of existing methods and presents TwinGuard-Sec, a new federated blockchain-based AI system expressly aimed at providing a set of standardized security and privacy of data in a heterogeneous realm of DT. The methodology comprises a dual-layered systematic architectural framework, comprising an AI-governed threat intelligence unit and zero-knowledge identity verifications and a distributed ledger technology layer that is domain-interoperable with lightweight consensus algorithms ensuring synchronous operation in real time. The framework fills three essential gaps in research including: absence of standardized cross-domain security protocols, inadequate privacy preserving mechanisms applied to sensitive inter-organizational data sharing and lack of scalable consensus algorithms to be used in DT-specific needs. We show on the rigorous test of a comprehensive 6G virtual twin testbed that includes 50 distributed nodes and five application domains (smart mobility, e-health, industrial IoT, smart cities, and autonomous systems) that the performance is significantly improved: 27.4% increase in threat detection accuracy (reaching 95.0% vs. 76.4% base) can be improved, better privacy preservation with a differential privacy parameter&#x2009;=&#x2009;0.94 (62% improvement), 21.2% reduction in latency down to 147&#xa0;ms. The system achieves Precision&#x2009;=&#x2009;0.968, Recall&#x2009;=&#x2009;0.959, and F1-score&#x2009;=&#x2009;0.963 (macro-average), with AUC-ROC&#x2009;=&#x2009;0.989 across eight attack categories. These results confirm that TwinGuard-Sec is an innovative means of ensuring the safety of cross-domain DT coordination, equipping it with both theoretical frameworks and implementation channels of the next generation intelligent infrastructure systems.","author":[{"family":"Mm","given":"Alnfiai"},{"family":"Rm","given":"Alotaibi"},{"family":"Fa","given":"Alotaibi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-49490-3","URL":"https://doi.org/10.1038/s41598-026-49490-3","source":"pubmed"},{"id":"doi:10.3389/frai.2026.1752124","type":"article-journal","title":"Systematic review of trends in deep learning for UAV cybersecurity.","abstract":"Unmanned Aerial Vehicles (UAVs) operate in navigation, sensing, and communication environments that are frequently degraded or adversarial. Their attack surface spans flight-control and payload software, radio links, and swarm coordination. This PRISMA-aligned systematic review synthesizes peer-reviewed studies published between 2015 and 2025 and organizes the evidence using an OSI-inspired threat taxonomy that maps spoofing, jamming, intrusion, and malware to system touchpoints and observable anomalies. We compare deep learning architectures, training targets, feature representations, evaluation practice, and deployment constraints relevant to single UAVs and swarms. Across the literature, convolutional and recurrent models dominate intrusion and anomaly detection pipelines, while attention-based, graph, and generative models appear in newer work targeting multi-agent settings and limited labels. Evidence most often relies on protocol traffic and onboard telemetry, whereas RF inputs are used less frequently and are typically represented as raw samples or spectrograms when datasets allow. Studies increasingly report efficiency-oriented deployment using pruning, quantization, distillation, or split inference to meet onboard compute and energy limits. Federated and multi-agent approaches are evaluated for scalability and robustness under poisoned updates, and blockchain-integrated designs are discussed under bandwidth and power constraints. Key gaps persist in shared datasets, repeatable adversarial stress testing, uncertainty and explainability reporting, privacy preservation, and certification-ready assurance cases for aviation regulation.","author":[{"family":"Ta","given":"Ahanger"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/frai.2026.1752124","URL":"https://doi.org/10.3389/frai.2026.1752124","source":"pubmed"},{"id":"doi:10.3390/jimaging12050195","type":"article-journal","title":"On Vision Transformer Explainability for Personal Protective Equipment Detection: A Qualitative and Quantitative Analysis.","abstract":"The safety of workers in industrial settings is ensured through the correct use of Personal Protective Equipment (PPE). The use of such equipment can be monitored using Deep Learning (DL). Federated Machine Learning (FML) is a technique that can be used in this context to preserve the privacy of sensitive information and provide explainability for the models adopted. Explainability techniques are an essential resource for interpreting the classification performed by the model. In this regard, this study aims to evaluate, through the adoption of specific similarity indices, the robustness and consistency of the explainability algorithms adopted to identify the areas of the images that are decisive for PPE classification. The dataset consists of 1600 real images representing work environments, in which staff are portrayed both with and without Personal Protective Equipment; specifically, there are workers wearing helmets, workers wearing reflective vests, workers wearing both devices and, finally, workers without any PPE. SSIM, VIF and SCC are the most relevant indices involved in the study. In the experimental phase, their mean values stand at 0.99, 0.96 and 0.96 for the intra-client study, and 0.96, 0.91 and 0.71 in the inter-client analysis.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/jimaging12050195","URL":"https://doi.org/10.3390/jimaging12050195","source":"pubmed"},{"id":"doi:10.1038/s41598-026-49344-y","type":"article-journal","title":"GNN-ML-FRL: a graph-enhanced meta-adaptive federated learning framework for scalable pest identification and ernvironmental modeling.","abstract":"Heterogeneous agro-ecological factors, insect breeding, and climate change are serious challenges to sustainable agricultural management. The study proposes a graph-enhanced meta-adaptive federated learning framework (GNN-ML-FRL) to address the challenges in precision agriculture. The proposed framework integrates Federated Learning (FL) for collaborative training of models in a decentralized manner across geographically distributed farms, Meta-Learning (ML) for rapid adaptation to changing environmental factors, and Graph Neural Networks (GNNs) for capturing spatial dependencies among agricultural entities. A comprehensive multivariate IoT environmental dataset with 52.56&#xa0;million time-series observations gathered from 500 dispersed sensors over a 12-month period, the IP102 insect pest recognition benchmark (75,222 images across 102 species), and curated genomic datasets from MaizeGDB and the Rice Annotation Project Database for genotype-informed modeling are the three standardized datasets used to assess the framework. Experimental results show statistically significant improvements (p&#x2009;&lt;&#x2009;0.01) over CNN and graph-based baselines, achieving 89.3% Top-1 accuracy, 7.8% higher generalization performance, and 12.4% reduction in prediction loss across geographically unseen farms. SHAP-based explainability further indicate that environmental accuracy-related features contributed nearly 63% positive influence, while loss-related factors contributed 37% negative influence, validating model robustness. Geographic generality is confirmed by site-out validation using IoT data, and resilience is improved under varied crop conditions by genotype-informed graph modeling. The findings show that a scalable and statistically sound framework for data-driven pest identification and environmental modeling in precision agriculture may be achieved by combining spatial graph reasoning, meta-adaptive learning, and decentralized training.","author":[{"family":"Mahalakshmi"},{"family":"Sk","given":"Mathivanan"},{"family":"Rb","given":"Joseph"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-49344-y","URL":"https://doi.org/10.1038/s41598-026-49344-y","source":"pubmed"},{"id":"doi:10.1371/journal.pone.0342953","type":"article-journal","title":"FedEmoNet: Privacy-preserving federated learning with TCN-Transformer fusion for cross-corpus speech emotion recognition.","abstract":"Federated learning offers a promising path toward privacy-preserving speech emotion recognition, yet existing approaches remain confined to single-corpus evaluation, lack formal differential privacy guarantees, and provide no mechanism for model interpretability. Meanwhile, cross-corpus generalization continues to challenge even centralized systems, with typical accuracy drops of 20-40% on unseen datasets due to domain shift in recording conditions, speaker demographics, and cultural expression norms. This paper introduces FedEmoNet, a unified framework that jointly addresses these open problems by combining FedProx-based distributed optimization, a hybrid Temporal Convolutional Network-Transformer (TCN-Transformer) architecture, Particle Swarm Optimization (PSO) feature selection, and calibrated ([Formula: see text])-differential privacy. Five heterogeneous clients-two German-speech (EmoDB), two English-speech (RAVDESS), and one mixed-collaborate under non-IID conditions (Dirichlet [Formula: see text]) without exchanging raw audio. Each client extracts multi-scale phase space reconstructions at micro (25&#x2009;ms), meso (250&#x2009;ms), and macro (2.5&#x2009;s) temporal resolutions alongside spectral and handcrafted features, which are fused through multi-head attention across the TCN-Transformer branches. On held-out, speaker-independent test sets the framework achieves 99.07%&#x2009;&#xb1;&#x2009;0.35% accuracy on EmoDB (107 samples) and 98.96%&#x2009;&#xb1;&#x2009;0.42% on RAVDESS (288 samples). Zero-shot cross-corpus evaluation on CREMA-D (1,488 samples) yields 68.15%&#x2009;&#xb1;&#x2009;1.23% overall, with a clear arousal-dependent pattern: high-arousal emotions (angry, happy, sad) transfer at 71.9% versus 62.1% for low-arousal categories (neutral, disgust, fear). Ablation experiments confirm that PSO selection (+2.80%), Transformer blocks (+2.10%), and the FedProx protocol (+2.62%) each contribute significantly, and a monotonic reduced-data curve rules out memorization. Membership inference attack resistance drops to near-chance levels (AUC&#x2009;&#x2009;=&#x2009;&#x2009;0.52) under differential privacy while retaining 98.5% accuracy. A dual SHAP-LIME explainability analysis reveals high inter-method agreement (r&#x2009;=&#x2009;0.997) and confirms that prosodic features-particularly fundamental frequency statistics-serve as language-invariant emotion indicators across all three corpora (r&#x2009;=&#x2009;0.94 cross-corpus consistency).","author":[{"family":"Ra","given":"Obeidat"},{"family":"Na","given":"Aljarrah"},{"family":"Hh","given":"Shehadeh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1371/journal.pone.0342953","URL":"https://doi.org/10.1371/journal.pone.0342953","source":"pubmed"},{"id":"doi:10.3389/fmed.2026.1759016","type":"article-journal","title":"Federated learning and Data Lakehouse for healthcare analytics: a knowledge transfer initiative between Germany and Tunisia.","abstract":"Healthcare institutions worldwide generate growing volumes of heterogeneous clinical data, yet legal, ethical, and infrastructural constraints often prevent these data from being centralized for analysis. Federated learning approaches offer a promising solution by enabling multi-site computation without transferring sensitive patient information, but require well-designed cross-site data harmonization. Modern Data Lakehouse architectures address this requirement by providing a scalable, governed foundation for multimodal clinical dataset integration through unified storage, metadata-rich governance, and FAIR-aligned data access. Despite increasing interest in such technologies across the Middle East and North Africa (MENA) region, operational deployments remain limited due to fragmented infrastructures, insufficient data governance, and gaps in practical expertise. This perspective article reports on a German-Tunisian knowledge and technology transfer initiative conducted within the DAAD Ta'ziz Partnership programme. As mentioned in the, 'the Arabic word 'Ta'ziz' means 'strengthening/consolidation' and has been chosen to clearly express the intended outcome of the programme' [https://www.daad.de/en/information-services-for-higher-education-institutions/further-information-on-daad-programmes/taziz-partnership/ (visited on November 28th, 2025)]. The collaboration between the University Hospital of Cologne and the University of Sfax introduced and implemented federated learning concepts via the Personal Health Train paradigm, and explored the design of a Data Lakehouse tailored to emerging healthcare ecosystems in Tunisia. Through an internship programme, hands-on MLOps training, and a large-scale workshop, the project built technical capacity in containerized analytics workflows, data governance, FAIR data management, and lakehouse engineering. We synthesize lessons learned regarding infrastructural limitations, data governance maturity, interoperability challenges, and institutional readiness, and outline considerations for sustainable adoption of distributed analytics in the MENA region. The findings highlight the critical importance of capacity building, bidirectional knowledge exchange, proof-of-concept validation, and administrative engagement for deploying trustworthy AI and modern data infrastructures in sensitive healthcare environments. We by emphasizing the need for further developments regarding federated learning and Data Lakehouse adoption in Tunisia, and how cross-regional partnerships can accelerate responsible, privacy-preserving digital health innovation.","author":[{"family":"Mah","given":"Taieb"},{"family":"Ma","given":"Bouri"},{"family":"Mb","given":"Abdallah"},{"family":"Mh","given":"Kammoun"},{"family":"Fk","given":"Tang"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/fmed.2026.1759016","URL":"https://doi.org/10.3389/fmed.2026.1759016","source":"pubmed"},{"id":"doi:10.1038/s41598-026-49826-z","type":"article-journal","title":"Federated learning-driven intelligent framework for multi-center radiotherapy dose distribution prediction oriented toward linear accelerators.","abstract":"High-quality radiotherapy dose distribution prediction for linear accelerators remains a labor-intensive process constrained by inter-planner variability and institutional data silos. While deep learning has shown promise in automating dose distribution prediction, most existing methods are trained on single-center datasets, limiting their generalizability. Direct multi-center data pooling is hindered by stringent privacy regulations and heterogeneous clinical protocols. This paper proposes a federated learning-driven framework that enables collaborative model training across geographically distributed institutions without exchanging raw patient data. The framework comprises a multi-scale attention U-Net for three-dimensional dose prediction and an adaptive weighted federated aggregation strategy that dynamically balances data volume and local model quality to address non-independent and non-identically distributed data challenges. A layered privacy protection mechanism integrating gradient-clipped differential privacy with secure aggregation provides privacy-enhancing protections with quantifiable bounds during parameter exchange. Experiments conducted across four clinical centers on head-and-neck and abdominal IMRT cases demonstrate that the proposed approach achieves a mean Gamma pass rate of 96.8%, closely approaching the centralized training upper bound of 97.5% while significantly outperforming single-center models and standard federated averaging. Ablation studies confirm the individual contributions of adaptive weighting and dual attention modules, and robustness analyses validate fault tolerance under client dropout and adversarial conditions. The proposed framework offers a practical and privacy-preserving pathway for breaking data silos in AI-driven radiotherapy research.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-49826-z","URL":"https://doi.org/10.1038/s41598-026-49826-z","source":"pubmed"},{"id":"doi:10.3390/s26092864","type":"article-journal","title":"A Privacy-Preserving Artificial Intelligence-Driven Sensing System for Distributed Multimodal Risk Detection.","abstract":"Withthe widespread deployment of intelligent terminals, mobile payment platforms, and Internet of Things devices, security systems are being progressively transformed from traditional transaction outcome analysis toward an intelligent perception paradigm centered on user behavior, device states, and environmental context. To address the challenges of multimodal data heterogeneity, non-independent and identically distributed data across nodes, and the difficulty of centralized modeling under privacy constraints in distributed scenarios, an artificial intelligence-driven federated multimodal security perception framework, namely FMS-LLM, is proposed. At its core, the framework introduces a Non-IID adaptive federated fusion mechanism that achieves dual-level alignment-structural alignment via parameter-level masks and semantic alignment via feature consistency constraints-to effectively mitigate cross-node distribution discrepancies. Additionally, an LLM-driven semantic enhancement module is developed, utilizing trend-guided token selection and inertia-suppression to map low-level sensing features into high-level risk semantic representations, thereby supporting logical reasoning and explainable decision-making. This framework takes user behavioral sensing data, device state information, environmental context data, and transaction behavior data as inputs, and constructs an integrated security analysis pipeline of \"perception-collaboration-reasoning\". Experimental results on the distributed multimodal security perception task demonstrate that the proposed method achieves an Accuracy of 91.62%, a Precision of 91.04%, a Recall of 90.37%, an F1-score of 90.70%, and a ROC-AUC of 94.73%, consistently outperforming baseline methods including Logistic Regression, Random Forest, LSTM, the centralized multimodal deep model, FedAvg, FedProx, and MOON. Under strongly Non-IID conditions, when &#x3b1;=0.1, the model still maintains an Accuracy of 88.47% and an F1-score of 87.11%, demonstrating stronger cross-node robustness. The ablation study further indicates that the complete model attains the best classification performance while reducing communication cost to 18.92 MB/Round. These results demonstrate that the proposed method can effectively fuse multi-source sensing information under privacy-preserving conditions and support intelligent security perception tasks with higher accuracy, stronger robustness, and improved interpretability.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26092864","URL":"https://doi.org/10.3390/s26092864","source":"pubmed"},{"id":"doi:10.2147/rmhp.s606165","type":"article-journal","title":"Privacy, Security &amp; Governance Frameworks for AI-Powered Wearable Internet of Health Things in Elderly Care: A Comprehensive Review.","abstract":"The global aging population is expanding at an unprecedented rate, with projections indicating that 1.4 billion people will be aged 60 years or older by 2030 and 2.1 billion by 2050, placing immense pressure on healthcare systems worldwide. Artificial intelligence (AI)-powered wearable Internet of Health Things (IoHT) devices - including smartwatches, biosensors, and continuous health monitors - have emerged as transformative tools for real-time elderly health monitoring, fall detection, and predictive analytics. However, the massive collection of sensitive biometric data by these devices raises critical concerns regarding privacy, security, and governance that remain insufficiently addressed, particularly for elderly populations. This comprehensive review synthesizes evidence from 333 peer-reviewed articles published between 2018 and 2025 cross PubMed, Scopus, Web of Science, IEEE Xplore, and Google Scholar to identify, analyze, and compare governance frameworks for AI-powered wearable IoHT in elderly care. The analysis reveals significant regulatory fragmentation across jurisdictions: while the European Union's General Data Protection Regulation (GDPR) and AI Act provide the most comprehensive rights-based framework, the United States relies on a patchwork of sector-specific regulations with notable gaps for consumer wearables, and Asia-Pacific nations exhibit highly variable approaches ranging from mature (Singapore, Japan) to nascent (Indonesia, Malaysia). Elderly-specific provisions remain conspicuously absent across all regulatory regimes examined. This review proposes a novel five-layer integrative governance framework - the first to unify technical security, privacy protection, ethical AI governance, regulatory compliance, and person-centered governance specifically designed for elderly care contexts. The framework addresses unique vulnerabilities associated with cognitive decline, reduced digital literacy, and caregiver dependency. Findings underscore the urgent need for harmonized, age-sensitive regulatory approaches and privacy-preserving technologies such as federated learning and differential privacy to ensure that AI-powered wearable IoHT fulfills its promise of enhancing elderly healthcare without compromising dignity, autonomy, or data security.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.2147/rmhp.s606165","URL":"https://doi.org/10.2147/rmhp.s606165","source":"pubmed"},{"id":"doi:10.1016/j.isci.2026.115729","type":"article-journal","title":"FPECGNET: A deep learning framework based on federated learning and prototype learning for interpretable ECG classification with privacy.","abstract":"Under everyday operational conditions, data privacy regulations prohibit sharing information between hospitals or institutions, contributing to the emergence of data silos. This scarcity of data hinders improvements in model performance. Another major obstacle to AI applications in medicine is the lack of interpretability, where the decision-making process of models cannot be reflected. We propose a deep learning framework based on federated learning and prototype learning that simulates the reality of hospital data silos while endowing the model with interpretability. Using the publicly available PTB-XL dataset, we divided it into three subsets by category and trained the model using the federated learning framework, achieving superior performance on these subsets. When tested on the aggregated model using the PTB-XL dataset, it demonstrated performance comparable to the current state-of-the-art model.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.isci.2026.115729","URL":"https://doi.org/10.1016/j.isci.2026.115729","source":"pubmed"},{"id":"doi:10.3389/fpubh.2026.1769078","type":"article-journal","title":"FedMal-XAI: an explainable federated vision transformer leveraging knowledge distillation for privacy-preserving malaria detection.","abstract":"Plasmodium parasites are the cause of malaria, a deadly illness that continues to pose a serious danger to world health, especially in areas with low resources where subjectivity, complexity, along with privacy issues make it difficult to employ traditional diagnostic techniques like microscopy and quick diagnostic testing. To overcome these specific challenges of diagnostic subjectivity, logistical complexity, and data privacy, this paper suggests a privacy-preserving federated learning system that uses sophisticated Vision Transformers (ViTs) for automated malaria identification from blood smear images. This paper suggests a privacy-preserving federated learning system that uses sophisticated Vision Transformers (ViTs) for automated malaria identification from segmented red blood cell (RBC) images in order to get around these issues. This architecture successfully addresses important privacy and logistical restrictions by enabling cooperative training among decentralized institutions without exchanging sensitive data. Prominent centralized convolutional neural networks (CNNs) are matched in diagnostic accuracy by the federated ViT models, which include ViT-B/16, DeiT-Tiny, Swin-T, and DINOv2. Interestingly, the federated transformer variations outperform the CNN ensemble (ResNet50&#x202f;+&#x202f;VGG16) with an accuracy of 98.15%, FedDistill-DeiT achieving 97.79%, FedAvg-Swin-T reaching 97.75%, and FedDistill-Swin-T achieving a high ROC-AUC of 0.9977. These findings show that, even in the presence of diverse data distributions, federated Vision Transformers provide a reliable, scalable, and interpretable malaria screening solution that combines high accuracy with solid privacy guarantees.","author":[{"family":"Ta","given":"Bhuiyan"},{"family":"Fi","given":"Khan"},{"family":"Fm","given":"Noori"},{"family":"Akm","given":"Masum"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/fpubh.2026.1769078","URL":"https://doi.org/10.3389/fpubh.2026.1769078","source":"pubmed"},{"id":"doi:10.3390/s26082418","type":"article-journal","title":"Artificial Intelligence-Driven Multimodal Sensor Fusion for Complex Market Systems via Federated Transformer-Based Learning.","abstract":"In highly digitalized and networked modern trading systems, large volumes of heterogeneous data are continuously generated from multiple sources during market operations. However, due to the complexity of data structures, significant differences in temporal scales, and constraints imposed by data privacy protection, traditional single-source modeling approaches are unable to fully exploit multisource information. To address this issue, a federated multimodal prediction framework for complex market systems, termed Federated Market-Sensor Transformer (FMST), is proposed. In this framework, data originating from different information sources are uniformly modeled as multimodal time series. A multimodal market-sensor representation module is constructed to perform unified feature encoding, and a cross-modal Transformer fusion architecture is employed to characterize dynamic interaction relationships among different information sources. Meanwhile, a federated collaborative learning mechanism is introduced during the training phase, enabling multiple data nodes to perform collaborative model optimization without sharing raw data. In this manner, data privacy can be preserved while improving the cross-region generalization capability of the model. Systematic experimental evaluation is conducted on the constructed multimodal market-sensor dataset. The experimental results demonstrate that the proposed method consistently outperforms traditional statistical models and deep learning approaches across multiple evaluation metrics. In the main prediction experiment, FMST achieves a root mean square error (RMSE) of 0.1136, a mean absolute error (MAE) of 0.0832, and a coefficient of determination R2 of 0.8517, while the direction prediction accuracy reaches 74.56%, clearly outperforming baseline models including ARIMA, LSTM, Temporal CNN, Transformer, and FedAvg-LSTM. In the cross-region generalization experiment, FMST maintains strong performance, achieving an RMSE of 0.1242, an MAE of 0.0908, an R2 value of 0.8261, and a direction prediction accuracy of 72.48%. The ablation study further indicates that the three core components-multimodal market-sensor representation, cross-modal Transformer fusion, and federated collaborative learning-each make important contributions to the overall model performance. These experimental findings demonstrate that the proposed method can effectively integrate multisource market information and significantly enhance the prediction capability for complex market dynamics, providing a new technical pathway for the application of artificial intelligence-driven multimodal sensing systems in economic data analysis.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26082418","URL":"https://doi.org/10.3390/s26082418","source":"pubmed"},{"id":"doi:10.1038/s41598-026-50690-0","type":"article-journal","title":"Lightweight and Energy-Aware Intrusion Detection for Industrial IoT Using TinyML and Edge AI.","abstract":"Industrial Internet of Things (IIoT) ecosystems are expanding rapidly. Scalable and reliable intrusion detection systems (IDS) are needed to protect critical infrastructures from evolving cyber threats. This study proposes a hybrid IDS framework that combines Graph Attention Networks (GAT) and Bidirectional Gated Recurrent Units (BiGRU) for privacy&#x2011;preserving distributed detection. The model is optimized with the Grey Wolf Optimizer (GWO) and enhanced through Federated Learning (FL). In IIoT traffic, GAT captures complex structural links, while BiGRU analyzes bidirectional temporal patterns, enabling accurate anomaly detection. GWO automates hyperparameter tuning and offers faster convergence than traditional methods such as Ant Colony Optimization. FL trains models locally on distributed IIoT devices, preserving data privacy and supporting decentralized deployment. The framework demonstrates improved scalability potential through decentralized training and reduced communication overhead (20% lower in a 10-node simulation), achieving detection accuracies of up to 95% across diverse attack scenarios, including Distributed Denial of Service (DDoS), Advanced Persistent Threats (APTs), and Zero&#x2011;Day exploits. It has been evaluated on the Edge&#x2011;Industrial Internet of Things dataset (Edge&#x2011;IIoTset), Canadian Institute for Cybersecurity - Internet of Things 2023 Dataset (CICIoT2023), and Real&#x2011;Time Internet of Things 2022 (RT&#x2011;IoT2022) datasets. An Explainable AI (XAI) module further improves interpretability by leveraging GAT's attention mechanism. Overall, this technology demonstrates competitive offline performance on the EDGE-IIoTset, CICIoT2023, and RT-IoT2022 benchmark datasets, achieving F1-scores of up to 0.94, and shows promising scalability potential through decentralized Federated Learning with 20% lower communication overhead in a 10-node simulation. However, inference latency on resource-constrained edge hardware remains a challenge (e.g., 120-180&#xa0;ms per sample on Raspberry Pi 4), which limits its strict real-time feasibility in mission-critical environments. Therefore, further model compression, adversarial robustness testing, and real-world deployment validation are required before practical edge-level applicability can be confirmed.","author":[{"family":"Mi","given":"Alghamdi"},{"family":"Sb","given":"Chaabane"},{"family":"Wm","given":"Alawad"},{"family":"Oh","given":"Albalawi"},{"family":"Oi","given":"Alqaisi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-50690-0","URL":"https://doi.org/10.1038/s41598-026-50690-0","source":"pubmed"},{"id":"doi:10.3389/fpubh.2026.1859276","type":"article-journal","title":"Localized AI for stroke care in LMICs: a framework to overcome structural and diagnostic barriers.","abstract":"Low- and middle-income countries (LMICs) bear a disproportionate share of the global stroke burden, driven not only by resource limitations but also by systemic inefficiencies in workforce distribution, diagnostic access, and prehospital care coordination. While advances in artificial intelligence (AI) have demonstrated significant potential in stroke diagnosis and management, many existing solutions remain poorly aligned with the infrastructural and policy realities of LMIC health systems, limiting their scalability and long-term impact. This study presents a comprehensive narrative review of literature published between January 2015 and March 2026, synthesizing evidence across digital health, stroke systems of care, and AI deployment models. We identify three persistent structural barriers-workforce shortages, diagnostic centralization, and fragmented care pathways-that collectively constrain timely intervention in acute stroke. In response, we propose a \"Localized AI + Policy\" framework that integrates lightweight AI models, edge computing, and federated learning within context-specific health system and governance structures. This approach emphasizes decentralized computation, data sovereignty, and alignment with national health policies, enabling more resilient and scalable deployment of AI in resource-constrained environments. By shifting the focus from technology-centric innovation to system-integrated implementation, this framework highlights a pathway for translating AI advances into sustainable public health impact. The findings underscore the importance of embedding digital health solutions within broader strategies for health system strengthening, universal health coverage, and global health equity.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/fpubh.2026.1859276","URL":"https://doi.org/10.3389/fpubh.2026.1859276","source":"pubmed"},{"id":"doi:10.7759/cureus.107501","type":"article-journal","title":"Machine Learning-Based Data Extraction Tools in Healthcare: A Systematic Review.","abstract":"The healthcare industry's digital transformation has led to an unprecedented volume of multimodal data. Machine learning (ML)-based extraction tools offer promising solutions for managing this data explosion, particularly when integrated with federated database systems. If a large language model (LLM) is trained to extract data from this multimodal information and ensure high accuracy while remaining affordable, the potential to improve the data extraction process within the medical field would be limitless, reducing costs and manpower across the board. A systematic review was conducted following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, searching major databases for studies published between 2018 and 2024, supplemented by grey literature sources. Analysis focused on the performance and implementation costs of ML-based extraction tools in healthcare settings. From 1,247 initial records, 21 studies met the inclusion criteria. ML-based extraction demonstrated superior accuracy, ranging from 61% to 98%, compared to traditional methods. Implementation costs averaged between $500,000 and $2.5 million. Two primary categories of tools emerged: image-based and text-oriented. ML-based extraction tools show significant promise in healthcare data management, though successful implementation requires careful consideration of costs, security protocols, and regulatory compliance. The development of a dedicated LLM capable of efficiently extracting data from various medical sources could revolutionize healthcare by streamlining data management and reallocating resources toward patient care and research advancements.","author":[{"family":"Zi","given":"Khalpey"},{"family":"Fh","given":"Khaliel"}],"issued":{"date-parts":[[2026]]},"DOI":"10.7759/cureus.107501","URL":"https://doi.org/10.7759/cureus.107501","source":"pubmed"},{"id":"doi:10.3389/fpubh.2026.1785170","type":"article-journal","title":"Cognitive sovereignty and decolonial public health: reclaiming epistemic authority in the global AI era.","abstract":"As artificial intelligence (AI) becomes critical infrastructure for global health, it reproduces colonial patterns of extraction, mining data from the Global South to train models owned by the Global North. While international bodies like the WHO emphasize \"ethical AI,\" they often overlook the structural violence of this digital colonialism. This perspective argues that true health equity requires more than bias mitigation; it demands cognitive sovereignty: the right of communities to govern not just their data but also the epistemic logic, interpretive frameworks, and algorithmic reasoning of the systems that analyze it. Drawing from Indigenous data governance principles (OCAP/CARE) and concrete implementation cases from Kenya, Nigeria, Rwanda, and Latin America, we demonstrate how cognitive sovereignty extends beyond data sovereignty to encompass control over knowledge production itself. By anchoring this political vision in specific technical architectures, federated learning, and community-led surveillance, we can move from extractive \"AI for good\" to a decolonial future of autonomous health intelligence. Recent cases from Kenya's AI health deployments and pathogen genomics illustrate both the urgency and feasibility of this transformation.","author":[{"family":"Ef","given":"Agyemang"},{"family":"Sk","given":"Srivastav"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/fpubh.2026.1785170","URL":"https://doi.org/10.3389/fpubh.2026.1785170","source":"pubmed"},{"id":"doi:10.3390/jimaging12060228","type":"article-journal","title":"A Comprehensive Review of Artificial Intelligence for Brain Tumor Analysis: Taxonomy, Robustness, and Open Challenges in Neuro-Oncology.","abstract":"Detecting brain tumors can be challenging as a clinical problem because of tumor heterogeneity and reliance on manual neuroimaging interpretation, which can be prone to human error. Artificial intelligence (AI) has shown strong potential as a clinical decision-support tool, assisting radiologists in improving diagnostic accuracy and supporting the interpretation of neuroimaging data. AI using machine learning (ML) and deep learning (DL) algorithms has performed credibly in tumor detection, segmentation, and classification tasks. Challenges such as dataset bias, limited generalization, lack of explainability, and high computational costs must be addressed before clinical application. This article provides a comprehensive review of AI methods applied to brain tumor imaging, with a primary focus on adult diffuse gliomas and secondary coverage of brain metastases, meningiomas, and pediatric tumors where relevant. The major contribution of this review is a new three-factor (diagnostic tasks, learning strategies, and data modalities) taxonomy. Beyond accuracy-based metrics, we provide a qualitative assessment of robustness, generalization, and the principal barriers to clinical adoption identified in the published literature, while acknowledging that comprehensive clinical utility evidence remains an open research direction.","author":[{"family":"Mh","given":"Qasem"},{"family":"Tma","given":"Sariera"},{"family":"Sm","given":"Alshraah"},{"family":"Ass","given":"Mufleh"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/jimaging12060228","URL":"https://doi.org/10.3390/jimaging12060228","source":"pubmed"},{"id":"doi:10.1038/s41598-026-56630-2","type":"article-journal","title":"Hybrid electricity management system for residential power block applications.","abstract":"Developing countries have been facing electricity crisis for decades, mainly due to insufficient electricity generating capacity as compared to its demand and extensive growth in population. This research outlines Electric_Fed, a decentralized electricity management model that incorporates blockchain and federated learning (FL). This model allows multiple clients, defined as houses or blocks, to privately and efficiently collaborate on electricity consumption and production forecasting without needing to exchange raw, shared data. The model achieves collaboration on a global model. Automated electricity trading is made possible through the use of Blockchain and smart contracts, which allows clients to securely and transparently participate in real-time peer-to-peer energy exchange. In this article, electricity production model using cross device FL technique and blockchain has been presented. In the proposed model, FL has been used to provide a secured client server model to facilitate clients having surplus electricity or facing electricity shortfall within the network. Central Server issues alert to the clients having surplus electricity to start Smart contract with the client facing electricity shortfall implementing blockchain. Blockchain ensures safe transaction and reliable transfer of payments to the client selling electricity. Electricity production and consumption dataset from Jan 2024 to Dec 2024 has been processed. Experimental validation proved that the proposed model is capable enough to fulfill 89.2% electricity need of the selected region. The experimental setup results (from June 2024) validate the model's effectiveness. Client A reached a surplus of +&#x2009;909 kWh, Client B&#x2009;+&#x2009;12 kWh, Client C&#x2009;+&#x2009;230 kWh and Client D fell short by -&#x2009;693 kWh. The system automatically recognized surpluses and deficits, trading energy to balance the needs of Clients A and C, which allowed them to reach full self-sufficiency. This model lowered national grid dependency and reduced overall electricity costs. As compared to other models presented for electricity supply, our proposed model is not only more effective in fulfilling the regional electricity demands but it also suggested a secured mechanism to compensate the electricity selling client participating in electricity trade. This demonstrates, in and of itself, the effectiveness of the model proposed while providing intelligent decentralized energy management in a system that is, as a whole, efficient, scalable, and privacy preserving.","author":[{"family":"Ar","given":"Hamraz"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-56630-2","URL":"https://doi.org/10.1038/s41598-026-56630-2","source":"pubmed"},{"id":"doi:10.1038/s41598-026-40881-0","type":"article-journal","title":"F-Transformer: a federated transformer for efficient and privacy-preserving sequence generation.","abstract":"Transformer models have demonstrated remarkable success in natural language processing (NLP) tasks, but their deployment in distributed environments faces critical challenges including high computational demands, large memory footprints, and privacy concerns when handling sensitive data. Existing federated learning (FL) implementations with transformers often sacrifice either model performance for privacy or resource efficiency for accuracy. We propose the F-Transformer, a lightweight federated transformer framework that addresses these limitations through integrated privacy-preserving mechanisms and architectural optimizations. Our framework employs a compact architecture with 4 attention heads, 4 layers, and 64 embedding dimensions, achieving only 0.87 million trainable parameters. We implement an incremental FL strategy where local clients continuously train on newly arriving data while the global model aggregates updates through the FedAvg algorithm. The framework integrates privacy objectives directly into the optimization process through a novel regularization formulation. We evaluate the F-Transformer on the WikiText-2 dataset using validation perplexity as the primary metric, along with central processing unit (CPU) utilization, memory consumption, and training loss convergence. Our results demonstrate a validation perplexity of 5.9894, surpassing state-of-the-art (SOTA) models including BERT-Large, GPT-2, and SparseGPT while using significantly fewer parameters. The framework achieves 40% reduction in CPU utilization and 34% reduction in memory consumption compared to centralized training. These results establish the F-Transformer as an effective solution for privacy-preserving sequence generation in resource-constrained federated environments.","author":[{"family":"Nk","given":"Jadav"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-40881-0","URL":"https://doi.org/10.1038/s41598-026-40881-0","source":"pubmed"},{"id":"doi:10.1016/j.artmed.2026.103481","type":"article-journal","title":"State-of-the-art TinyML approaches for colorectal cancer detection: Current advances, challenges, and future directions.","abstract":"Colorectal cancer (CRC) remains a leading cause of cancer-related mortality worldwide, with diagnostic disparities, particularly pronounced in resource-constrained and decentralized healthcare settings. Recent advances in TinyML machine learning models optimized for ultra-low-power, memory-constrained embedded devices have created new opportunities for scalable on-device CRC screening and diagnostics. This review presents a systematic and CRC-centric analysis of TinyML technologies across the diagnostic continuum, including capsule endoscopy, histopathology, breath analysis, and biosignal-based screening. Unlike existing surveys that address TinyML from a general healthcare perspective, this study focuses specifically on the technical, clinical, and deployment challenges unique to CRC diagnostics. We propose a structured taxonomy encompassing model compression techniques, hardware-software co-design strategies, and clinical deployment paradigms, and critically analyze the accuracy-latency-energy trade-offs across representative platforms. This review further synthesizes recent (2024-2025) advances in TinyML compilers, hardware accelerators, and edge-cloud integration, highlighting their implications for real-world clinical translation. By consolidating current evidence, identifying benchmarking and regulatory gaps, and outlining a forward-looking research roadmap, this survey clarifies the role of TinyML as a viable enabler of real-time, privacy-preserving, resource-efficient CRC diagnostics. These findings provide actionable insights for researchers, clinicians, and system designers seeking to deploy TinyML solutions in equitable and clinically meaningful cancer care.","author":[{"family":"Sa","given":"Bhat"},{"family":"Mc","given":"Chen"},{"family":"Nf","given":"Huang"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.artmed.2026.103481","URL":"https://doi.org/10.1016/j.artmed.2026.103481","source":"pubmed"},{"id":"doi:10.1088/2057-1976/ae74d6","type":"article-journal","title":"A comprehensive review of deep learning applications in the segmentation and classification of skin cancer.","abstract":"Skin cancer (SC) is one of the most prevalent forms of cancer worldwide. Both melanoma and non-melanoma types pose major challenges for early detection, accurate diagnosis, and proper treatment. Conventional diagnostic approaches, such as biopsy and visual examination, are often time-consuming, subjective, and prone to human error. Recent advances in artificial intelligence (AI) and deep learning (DL) have greatly improved the accuracy of SC diagnosis. This systematic review explores the applications of DL techniques in the segmentation and classification of skin lesions between 2014 and 2024. Following the Preferred Reporting Items for Systematic reviews and Meta-Analyses guidelines and applying predefined inclusion and exclusion criteria, a total of 77 experimental studies out of 540 were analyzed from major databases, including Scopus, IEEE, PubMed, and MDPI. Convolutional neural networks (CNNs) were identified as the most widely used for classification, while U-Net and its variants dominated segmentation tasks. Hybrid and ensemble frameworks demonstrated superior performance on benchmark datasets such as the International Skin Imaging Collaboration (ISIC) archive and HAM10000. Moreover, this work incorporates a formal risk-of-bias analysis, revealing critical concerns about class imbalance and data leakage. Almost all reviewed studies for the classification task achieved an average accuracy of 96% for the ISIC dataset, while the HAM10000 dataset attained an average accuracy of 93%. Despite these advances, challenges such as class imbalance, limited dataset diversity, and insufficient clinical validation persist. Addressing these issues through data augmentation, explainable AI, and federated learning could further enhance the generalizability and clinical applicability of AI-driven diagnosis systems. Additionally, this study identifies a clear paradigm shift from standalone CNNs to hybrid frameworks and multi-source feature fusion strategies, aiming to improve SC diagnosis.","author":[{"family":"Ar","given":"Shaaban"},{"family":"Am","given":"Salaheldin"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1088/2057-1976/ae74d6","URL":"https://doi.org/10.1088/2057-1976/ae74d6","source":"pubmed"},{"id":"doi:10.2174/0115734056400866250923175325","type":"article-journal","title":"Federated Deep Learning Approaches for Detecting Ocular Diseases in Medical Imaging: A Systematic Review.","abstract":"Artificial intelligence has significantly enhanced disease diagnosis in healthcare, particularly through Deep Learning (DL) and Federated Learning (FL) approaches. These technologies have shown promise in detecting ocular diseases using medical imaging while addressing challenges related to data privacy and security. FL enables collaborative learning without sharing sensitive medical data, making it an attractive solution for healthcare applications. This systematic review aims to analyze the advancements in AI-driven ocular disease detection, with a particular focus on FL-based approaches. The article evaluates the evolution, methodologies, challenges, and effectiveness of FL in enhancing diagnostic accuracy while ensuring data confidentiality.","author":[],"issued":{"date-parts":[[2025]]},"DOI":"10.2174/0115734056400866250923175325","URL":"https://doi.org/10.2174/0115734056400866250923175325","source":"pubmed"},{"id":"doi:10.21037/jtd-2026-0804","type":"article-journal","title":"Evolution, hotspots, and future directions of artificial intelligence in asthma research: a Web of Science-based bibliometric analysis [2016-2026].","abstract":"While artificial intelligence (AI) offers unprecedented capabilities for predictive modeling and precision asthma management, there is an urgent clinical necessity to successfully translate these rapid algorithmic innovations into real-world respiratory care. The exponential growth of cross-disciplinary AI literature has paradoxically created information overload for clinicians, obscuring underlying translational friction and hindering evidence-based implementation. Consequently, bibliometric analysis serves as the optimal quantitative vehicle to decode this vast scientific architecture. This study aims to objectively map the evolutionary trajectory, global research landscape, and emerging hotspots of AI in asthma, providing actionable roadmaps to reconcile computational development with clinical practice.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.21037/jtd-2026-0804","URL":"https://doi.org/10.21037/jtd-2026-0804","source":"pubmed"},{"id":"doi:10.3390/cancers18091322","type":"article-journal","title":"Artificial Intelligence-Enhanced Multiparametric MRI and VI-RADS in Bladder Cancer: Current Evidence, Clinical Opportunities and Barriers to Translation.","abstract":"Accurate distinction between non-muscle-invasive bladder cancer (NMIBC) and muscle-invasive bladder cancer (MIBC) remains the key local staging problem in bladder cancer because treatment intensity, timing of radical therapy, and suitability for bladder-preserving strategies all depend on it. Multiparametric magnetic resonance imaging (mpMRI) and the Vesical Imaging-Reporting and Data System (VI-RADS) now provide a standardized imaging framework for local staging and increasingly support MRI-first clinical pathways. Artificial intelligence (AI) has emerged as an additional decision-support layer, but the evidence base remains methodologically uneven. In this structured narrative review, we synthesized peer-reviewed literature from January 2020 to March 2026, while retaining foundational VI-RADS studies from 2018 to 2019, and prioritized guideline documents, meta-analyses, prospective cohorts, multicenter and externally validated AI studies, response-assessment studies, and papers addressing implementation and reporting quality. Current evidence shows that radiomics and deep learning models can achieve high discrimination for MIBC detection on MRI, and that the most plausible incremental value of AI lies in equivocal VI-RADS lesions, reader support outside high-volume expert settings, and multimodal risk stratification. However, most studies remain retrospective, highly selected, segmentation-dependent, and vulnerable to reference-standard bias, domain shift, and poor calibration. This review therefore emphasizes several translational issues that are often underreported: lesion-level versus patient-level inference, the distortive effect of TURBT-based labels, the need to evaluate false-negative consequences in VI-RADS 3 tumors, and the distinction between diagnostic support and broader pathway redesign. We also discuss response assessment, nacVI-RADS, segmentation automation, multicenter and federated infrastructure, workflow ownership, and the limits of imaging-only models in a biologically heterogeneous disease. The most credible near-term role of AI is not autonomous diagnosis, but augmentation of standardized mpMRI and VI-RADS within multidisciplinary care. Future progress will depend on prospective utility studies, site-held-out validation, transparent reporting, and the integration of imaging with molecular and cellular heterogeneity through radiogenomic and multi-omics approaches.","author":[{"family":"Cg","given":"Popescu"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/cancers18091322","URL":"https://doi.org/10.3390/cancers18091322","source":"pubmed"},{"id":"doi:10.1038/s41598-025-13519-w","type":"article-journal","title":"Decentralized federated deep Q-learning for IoMT security: leveraging MK-VQFHE and blockchain with IPFS.","abstract":"The explosion of patient data, the demand for real-time insights, and the critical importance of data security drive healthcare innovation. Medical Internet of Things (IoMT) offers a promising solution, connecting medical devices, sensors, and healthcare systems to improve patient care. However, managing and securing vast amounts of complex data produced by IoMT devices remains a significant challenge. Existing approaches often fall short of providing comprehensive solutions. To address this, this research proposes a novel approach combining Multi-Key Verifiable Quaternion Fully Homomorphic Encryption (MK-VQFHE) and blockchain for decentralized Federated Learning (FL) to enhance data management capabilities. Initially, the proposed study collects IoT healthcare data from publicly available datasets after that, Multi-Key Verifiable Quaternion Fully Homomorphic Encryption for encryption is used to safeguard data security. Encrypted data is then used in a collaborative learning model enabled by blockchain technology. The proposed Federated Deep Q-learning (FDQL) model enhances privacy protection by training inputs and evaluating threats. The data is securely stored using Interplanetary File System (IPFS) technology within the blockchain network. This study introduces a Practical Byzantine Fault Tolerant (PBFT) consensus technique to verify proposed structure&#x2019;s integrity. Performance metrics demonstrate proposed approach produce superior accuracy ranging from 99.2% to 99.4% across multiple datasets compared to existing approaches. Meanwhile the existing models such as BiLSTM, CNN, DNN, and ANN are attained accuracy of below 99%. The proposed research aims to enhance IoMT data security, introduce advanced encryption strategies, ensure efficient data storage, analyze privacy protection effectiveness, validate blockchain framework efficiency, and evaluate performance metrics for comprehensive insights.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-025-13519-w","URL":"https://doi.org/10.1038/s41598-025-13519-w","source":"pubmed"},{"id":"doi:10.1038/s41598-026-44162-8","type":"article-journal","title":"A multi-layer AI decision support system for startup success prediction and risk assessment using knowledge graphs and federated learning.","abstract":"Since start-ups have grown so quickly in recent decades, it is more important than ever to determine what elements contribute to their success or failure. Numerous factors, such as market conditions, product differentiation, finance availability, and managerial methods, influence these results. However, precise forecasting is a constant issue due to the intricacy of business ecosystems and the interaction of non-financial and financial aspects. A multi-layer AI-driven prediction model that incorporates early-stage start-ups&#x2019; financial and non-financial characteristics is presented in this paper. A Graph Convolutional Network (GCN) creates feature embeddings at Layer 1, whereas a knowledge graph documents the connections between affecting factors. Federated learning is used to safely combine dispersed knowledge while maintaining privacy. Layer 2 uses a deep neural network (DNN) to forecast success or failure based on the fused features. Layer 3 offers risk assessment and interpretability, detecting survival variables including team dynamics and product differentiation. An experimental sample of 20 start-ups was thoroughly examined as part of the model&#x2019;s evaluation on a dataset collected from Crunchbase that included approximately 623,000 companies, 799,000 founders, and 227,000 funding events. The findings show increased forecast accuracy and emphasize the value of non-financial elements in addition to conventional financial measurements. Through the integration of sophisticated AI approaches with organized domain knowledge, our work connects theoretical frameworks with empirical data. The suggested model contributes to start-up research and the real-world implementation of AI in business analytics by offering entrepreneurs, investors, and legislators a strong decision-support tool.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-44162-8","URL":"https://doi.org/10.1038/s41598-026-44162-8","source":"pubmed"},{"id":"doi:10.3390/s26082558","type":"article-journal","title":"Efficient Medical Image Segmentation in Multisensor Imaging: A Survey in the Era of Mamba and Foundation Models.","abstract":"Deep learning has revolutionized medical image segmentation; however, the clinical deployment of state-of-the-art models is severely impeded by their quadratic computational complexity and substantial resource demands, particularly in multisensor and multimodal imaging scenarios. In response, the field is undergoing a paradigm shift towards efficiency, characterized by the rise of linear-complexity architectures and the optimization of foundation models. This paper presents a comprehensive survey of efficient medical image segmentation methodologies, systematically reviewing the evolution from heavy, accuracy-driven models to lightweight, deployment-ready paradigms. In particular, we highlight the growing importance of efficient segmentation in multisensor medical imaging, where heterogeneous data sources such as CT, MRI, ultrasound, and infrared imaging introduce additional challenges in scalability and computational cost. We propose a novel taxonomy that categorizes these advancements into four distinct streams: (1) Mamba and State Space Models, which leverage selective scanning mechanisms to achieve global receptive fields with linear complexity; (2) Efficient Adaptation of Foundation Models, focusing on parameter-efficient fine-tuning and knowledge distillation to tailor the Segment Anything Model (SAM) for medical domains; (3) Advanced Lightweight Architectures, covering the resurgence of large-kernel CNNs and the emergence of Kolmogorov-Arnold Networks (KANs); and (4) Data-Efficient Strategies, including semi-supervised and federated learning to address annotation scarcity. Furthermore, we conduct a rigorous comparative analysis of representative algorithms on mainstream benchmarks, providing a granular evaluation of the trade-offs between segmentation accuracy and computational overhead. The survey also discusses key challenges in multisensor and multimodal settings, including modality heterogeneity, data fusion complexity, and resource constraints. Finally, we identify critical challenges and outline future research directions, serving as a roadmap for the development of next-generation efficient and scalable medical image analysis systems.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26082558","URL":"https://doi.org/10.3390/s26082558","source":"pubmed"},{"id":"doi:10.3389/fmicb.2026.1705116","type":"article-journal","title":"STROBE-causal machine learning for the human microbiome: systematic review on methodological innovations and validation frameworks.","abstract":"The reproducibility crisis in causal microbiome research necessitates robust validation frameworks. Current studies often face inconsistent validation methods, limited interpretability, and a lack of standardized reporting, creating a gap in reliable causal inference. This systematic review evaluates over 60 peer-reviewed studies published between 2015 and 2024 to: (1) establish benchmarking standards leveraging synthetic data and biological plausibility assessments; (2) compare advanced causal machine learning (ML) methodologies, including Double/Debiased ML, Deep Instrumental Variables (Deep IV), and Directed Acyclic Graphs (DAGs), in their application to microbiome-host systems; and (3) propose the STROBE-CML (Strengthening the Reporting of Observational Studies in Epidemiology-Causal Machine Learning) guidelines to standardize reporting practices. We emphasize critical innovations such as federated validation pipelines and time-series causal discovery frameworks that address these gaps by facilitating scalable, privacy-preserving, and reproducible inference across heterogeneous cohorts. A decision support tool is introduced to guide researchers in selecting appropriate causal ML approaches based on data structure, research question, and computational constraints. By synthesizing methodological advances with rigorous validation paradigms, this review provides a roadmap for generating reliable, biologically interpretable, and clinically translatable causal claims in microbiome science.","author":[{"family":"Ai","given":"Shehata"},{"family":"Mf","given":"El"},{"family":"Ama","given":"Mohamed"},{"family":"Mf","given":"Abouelenein"},{"family":"He","given":"Degha"},{"family":"Ii","given":"Teiba"},{"family":"Ss","given":"Mahmoud"}],"issued":{"date-parts":[[2026]]},"DOI":"10.3389/fmicb.2026.1705116","URL":"https://doi.org/10.3389/fmicb.2026.1705116","source":"pubmed"},{"id":"doi:10.4132/jptm.2026.04.27","type":"article-journal","title":"What's new in digital and computational pathology 2026: advances in adoption, standards, AI technologies, and clinical integration.","abstract":"Digital and computational pathology are expanding rapidly worldwide, driven by advances in whole-slide imaging, AI algorithms, multimodal data integration, and improved digital infrastructure. Adoption continues to accelerate in the United States and internationally, supported by professional guidelines, emerging reimbursement pathways, and the growing need for remote workflows and collaborative diagnostics. Progress in interoperability standards, regulatory frameworks, and FDA approvals has strengthened the foundation for clinical deployment, while large-scale data repositories and federated learning approaches enable more robust and privacy-preserving model development. Foundation models, multimodal AI systems, and LLM-based copilots are reshaping diagnostic support, prognostication, workflow efficiency, clinical trials and drug discovery.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.4132/jptm.2026.04.27","URL":"https://doi.org/10.4132/jptm.2026.04.27","source":"pubmed"},{"id":"doi:10.5445/ir/1000194123","type":"article-journal","title":"Towards robust neurocomputing model in efficient federated brain tumour segmentation with sparsification and weights clustering","abstract":"Brain tumour segmentation is a key application of AI in neuroimaging. Recently, federated learning (FL) has emerged as a strategic and increasingly relevant paradigm in neural computing due to its ability to address key challenges in large-scale neural network training, such as data access, privacy, collaborative learning, and model robustness. However, its adoption is currently hindered by high communication costs and the heterogeneity of client data. In this study, we investigated an efficient FL framework for brain tumour segmentation based on communication-aware optimization. We evaluated FedWSOComp, which integrates sparsification, quantiza tion, and entropy-based encoding, in combination with a 3D U-Net architecture under both homogeneous and heterogeneous data distributions. The multi-institutional FeTS 2024 dataset was employed and partitioned into independent and identically distributed (IID) and non-IID settings, with an independent test set of 67 patients. An overall of 18 configurations combined sparsification rates and quantization levels. Performance was measured using Dice Similarity Coefficient (DSC) and 95th percentile Hausdorff Distance (HD95). Experimental results demonstrated that aggressive compression caused severe degradation in segmentation quality, with HD95 ex ceeding 60 mm. In contrast, higher retention with finer quantization achieved the best balance between efficiency and accuracy, reaching a DSC of 0.98±0.09 and HD95 of 10.40±15.54 mm on the test set under non-IID conditions. The findings demonstrated that, when configured with moderate-to-fine quantization and high sparsification re tention, FedWSOComp enabled accurate and communication-efficient federated brain tumour segmentation. This study provides quantitative evidence and practical guidance for the deployment of FL-based segmentation models in privacy-sensitive and bandwidth-constrained clinical settings.","author":[{"family":"Raza","given":"Asaf"},{"family":"Raggio","given":"Ciro"},{"family":"Guzzo","given":"Antonella"},{"family":"Spadea","given":"Maria"},{"family":"Fortino","given":"Giancarlo"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5445/ir/1000194123","URL":"https://doi.org/10.5445/ir/1000194123","source":"datacite"},{"id":"doi:10.53348/jsst1s","type":"article-journal","title":"DNA BASED FEDERATED LEARNING","abstract":"The Deoxyribonucleic Acid (DNA) can be considered as one of the most effective biometrics. It can be used in different fields such as personal identification, parental verification and relativity. In this study, federated learning is suggested. It is a type of machine learning which recently attracts researchers' attentions. Here it has been suggested to be employed for controlling the DNA samples over large areas. That is, the DNA centers have DNA samples, each DNA center can deal with its samples by using a single machine learning method. Then, according to the federated learning machine learnings information are combined by a main machine learning. In this case, all DNA centers can effectively share their information. So, there is no need to waste long correspondence time between the DNA centers to share required data. Also, there may no need to re- train any machine learning again for new DNA samples Keywords: DNA, Machine Learning, Pattern Recognition","author":[{"family":"Al-Nima","given":"Raid"},{"family":"Yahya","given":"Maan"},{"family":"Qaba","given":"Azzah"}],"issued":{"date-parts":[[2026]]},"DOI":"10.53348/jsst1s","URL":"https://doi.org/10.53348/jsst1s","source":"crossref"},{"id":"doi:10.2174/9789815322224125030005","type":"article-journal","title":"Federated Learning-Based Data Dissemination Systems for IoVs","abstract":"Federated learning-based data dissemination solutions for Internet of Vehicles (IoVs) are gaining interest owing to their capacity to increase data dissemination performance and privacy. This chapter examines the current state of the art in federated learning-based data dissemination systems for IoVs, as well as the obstacles and possibilities associated with their deployment. A literature study, analysis of data dissemination needs in IoVs, and assessment of performance and privacy implications of alternative federated learning techniques are all part of the process for creating and assessing federated learning-based data dissemination systems in IoVs. The findings of a literature analysis and tests evaluating the performance and privacy of federated learning-based data dissemination systems in IoVs reveal that these systems have the potential to increase data dissemination performance and privacy, but various problems must be addressed. This chapter adds to the current literature by offering a thorough examination of the state-of-the-art federated learning-based data distribution systems for IoVs. The chapter discusses important obstacles and possibilities, as well as insights into the approach used to create and evaluate these systems. The chapter explores the consequences for IoVs of federated learning-based data dissemination systems, such as better data dissemination performance and privacy. The chapter focuses on possible applications in smart transportation, urban planning, and public safety. The chapter investigates the implications of federated learning-based data dissemination systems for IoVs, such as improved data dissemination performance and privacy. The chapter focuses on smart transportation, urban planning, and public safety applications.","author":[{"family":"Negi","given":"Gaurav"},{"family":"Krishna","given":"Gopal"},{"family":"Gupta","given":"Jitendra"}],"issued":{"date-parts":[[2025]]},"DOI":"10.2174/9789815322224125030005","URL":"https://doi.org/10.2174/9789815322224125030005","source":"crossref"},{"id":"doi:10.2174/9789815322224125030007","type":"article-journal","title":"Federated Learning-Based Vehicle Number Plate Recogntion in IoVs","abstract":"Artificial intelligence is widely used in a variety of industries. AI technology drives much of what we do. In a similar vein, as AI-based technologies advance, smart automobiles and the Smart Transport system will likewise experience revolutionary transformation. Different techniques are applied to create a system that is used to manage traffic and increase security inside the transportation network, different techniques are used. The automatic number recognition system (ANPR) described in this research can extract an image of a vehicle license plate by employing image processing methods. To make things easier, the proposed system may be operated without the installation of any extra GPS-like devices. The suggested system consists of image processing techniques, such as filters to eliminate blur and noise when distantly acquired photographs of moving vehicles are taken. To obtain the region of interest, its edges are detected, and an image is cropped. The procedure for better outcomes includes normalization, localization, image enhancement, restoration, and character retention approaches. Its effectiveness may be negatively impacted by the state of the license plate, unconventional formats, complex vision, camera quality, camera position, tolerance for distortion, motion blur, contrast-related issues, reflections, limitations in a processing unit, environmental factors, indoor/outdoor or time-independent shots, software tools, or other hardware-based restrictions.Even with the greatest algorithms, a successful ANPR system implementation might need extra computer hardware to boost the proposed System’s accuracy.","author":[{"family":"Pathak","given":"Disha"},{"family":"Srivastava","given":"Somya"},{"family":"Gupta","given":"Shelly"}],"issued":{"date-parts":[[2025]]},"DOI":"10.2174/9789815322224125030007","URL":"https://doi.org/10.2174/9789815322224125030007","source":"crossref"},{"id":"doi:10.4018/979-8-3373-3306-9.ch002","type":"article-journal","title":"Advancing AI Integration in Healthcare Using Federated Learning","abstract":"With the growing demand for secure, intelligent, and collaborative healthcare systems, Federated Learning (FL) has emerged as a transformative approach for developing AI models without exposing sensitive patient data. This chapter provides a detailed overview of FL's core architecture and its specialized variants—such as hierarchical, asynchronous, and personalized FL—tailored to healthcare's distributed and privacy-sensitive landscape. It explores privacy-enhancing mechanisms including local differential privacy, secure aggregation, and homomorphic encryption, all critical for training in heterogeneous environments. The integration of FL with blockchain, edge computing, and explainable AI (XAI) demonstrates how these technologies strengthen traceability, transparency, and real-time intelligence. Key challenges such as regulatory inconsistencies, infrastructural constraints, and model drift are critically examined. The chapter envisions a federated AI ecosystem promoting equitable, privacy-aware innovation through institutional collaboration and scalable deployment strategies.","author":[{"family":"Singh","given":"Mandeep"},{"family":"Kaur","given":"Gaganjot"},{"family":"Saxena","given":"Pranshu"},{"family":"Sharma","given":"Megha"}],"issued":{"date-parts":[[2025]]},"DOI":"10.4018/979-8-3373-3306-9.ch002","URL":"https://doi.org/10.4018/979-8-3373-3306-9.ch002","source":"crossref"},{"id":"doi:10.1515/zwf-2024-0131","type":"article-journal","title":"Federated Learning in der Arbeitsplanung","abstract":"Abstract Die Aufbereitung von praxisnahen Trainingsdatensätzen für Deep Learning in der Arbeitsplanung ist eine große Herausforderung. Die Datengrundlage aktueller Ansätze basiert auf synthetisch erstellten 3D-Modellen. Eine solche synthetisierte Generierung von Trainingsdaten bildet jedoch nur sehr begrenzt die industrielle Praxis ab. Vor diesem Hintergrund haben Ansätze ein hohes Potenzial, bei denen aus den Daten mehrerer Unternehmen eine ausreichend große Datengrundlage gebildet werden kann, ohne dass diese an eine zentrale Stelle übertragen werden müssen. Eine im beschriebenen Kontext vielversprechende Methode ist das Federated Learning (FL), für dessen Anwendung in der Arbeitsplanung in diesem Beitrag ein Ansatz beschrieben wird.","author":[{"family":"Hussong","given":"Marco"},{"family":"Klar","given":"Matthias"},{"family":"Aurich","given":"Jan"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1515/zwf-2024-0131","URL":"https://doi.org/10.1515/zwf-2024-0131","source":"crossref"},{"id":"doi:10.2139/ssrn.5194150","type":"manuscript","title":"DISTILLED ONE-SHOT FEDERATED LEARNING: A HIGHLY EFFICIENT AND SECURE APPROACH FOR REDUCING COMMUNICATION COSTS IN FEDERATED EDGE LEARNING","abstract":"Federated Edge Learning enables dispersed edge nodes to train a global model in the Artificial Internet of Things (AIoT), advancing cloud computing. Nevertheless, existing federated learning techniques suffer from communication inefficiencies, requiring many rounds to transmit large model weights, particularly under uneven data distribution. To address this, we propose Distilled One-Shot Federated Learning (DOSFL), which notably reduces communication costs while maintaining high performance. DOSFL allows clients to distill their local datasets into smaller synthetic datasets in just one communication round, sending this data to the server for global model training. The distilled data, resembling noise, becomes useless after the model updates, eliminating the need to transmit large model weights and gradients. As a result, DOSFL reduces communication costs by up to three orders of magnitude compared to traditional methods, while retaining up to 99% of the execution of centralized training across vision and language tasks using models like CNNs, LSTMs, and Transformers. Furthermore, DOSFL enhances security by preventing attackers from building effective models using leaked distilled data. In summary, DOSFL offers an efficient and secure solution for federated learning, achieving 0.1 percent or less of the traditional methods' communication costs while maintaining high accuracy.","author":[{"family":"Shudapreyaa","given":"RS"}],"issued":{"date-parts":[[2026]]},"DOI":"10.2139/ssrn.5194150","URL":"https://doi.org/10.2139/ssrn.5194150","source":"crossref"},{"id":"doi:10.22541/essoar.175807020.09297050/v1","type":"article-journal","title":"Federated Reinforcement Learning Framework for Privacy Preserving Few Shot Learning Authors","abstract":"This study introduces a federated reinforcement learning framework for few-shot learning (FRL-FSL), aiming to address the dual challenges of data scarcity and privacy preservation in distributed environments. The proposed framework integrates policy gradient optimization with secure aggregation and introduces validator nodes to ensure the authenticity of both data and model updates. Experiments were conducted on the Omniglot and FC100 datasets under 1-shot and 5-shot conditions, with comparisons against FedAvg, FedFSL and traditional supervised baselines. Results demonstrate that FRL-FSL achieved an average accuracy of 87.3% on Omniglot (5-shot), improving by 25.9% over FedAvg and 13.8% over FedFSL, while maintaining 72.6% accuracy in 1-shot tasks. On the FC100 dataset, FRL-FSL reached 59.8% accuracy in 5-shot learning, outperforming FedAvg by 18.6% and FedFSL by 7.1%, and achieved 46.3% in 1-shot learning. The framework also reduced the privacy risk index by 37% relative to FedAvg, with convergence accelerated by nearly 30% compared to baselines. These findings confirm that FRL-FSL achieves a practical balance between accuracy, convergence, and privacy, offering a promising solution for real-world, privacy-sensitive applications.","author":[{"family":"Kang","given":"Minsoo"},{"family":"Park","given":"Jihye"},{"family":"Choi","given":"Donghyun"},{"family":"Kim","given":"Seoyeon"}],"issued":{"date-parts":[[2025]]},"DOI":"10.22541/essoar.175807020.09297050/v1","URL":"https://doi.org/10.22541/essoar.175807020.09297050/v1","source":"crossref"},{"id":"doi:10.1371/journal.pone.0355601","type":"article-journal","title":"FedMamba-IoMT: Federated state space models with differential privacy and byzantine resilience for privacy-preserving intrusion detection in Internet of Medical Things.","abstract":"The proliferation of Internet of Medical Things (IoMT) devices has created critical cybersecurity challenges demanding intrusion detection systems that achieve high accuracy across diverse attack taxonomies while preserving patient privacy across institutional boundaries. Existing federated learning (FL) approaches face an inherent tension: Transformer-based architectures achieve strong detection performance but incur quadratic computational complexity and substantial communication overhead, while lightweight classifiers sacrifice representational capacity. Moreover, most FL-based intrusion detection systems lack formal privacy guarantees and robustness against adversarial participants. This paper introduces FedMamba-IoMT, the first federated State Space Model framework for privacy-preserving intrusion detection in IoMT networks, incorporating differential privacy (DP-SGD), Byzantine-resilient aggregation, and multi-level explainability. The proposed architecture reformulates tabular network traffic features as pseudo-sequential tokens processed through stacked selective State Space Model (Mamba) blocks with gated residual connections, achieving linear computational complexity &#x1d4aa;(n) with 78% fewer parameters than Transformer alternatives. We design a novel FedMamba aggregation strategy that weights client contributions by a convex combination of dataset proportion and inverse validation loss, augmented with a cosine similarity-based Byzantine filter that detects and excludes malicious model updates. Integration of DP-SGD with R&#xe9;nyi differential privacy accounting provides formal privacy guarantees (&#x3b5;&#x2208;{1.0,2.0,3.0,5.0,8.0}, &#x3b4;=10-5) while maintaining competitive accuracy. Comprehensive evaluation across three benchmark datasets-Edge-IIoTset (2,219,201 samples, 15 classes), CICIoMT2024 (3,204,537 samples, 19 classes), and Gotham Dataset 2025 (496,191 samples, 8 high-level traffic categories)-demonstrates that FedMamba-IoMT achieves 99.47&#xb1;0.04%, 99.52&#xb1;0.04%, and 98.90&#xb1;0.04% multiclass accuracy without DP, and 98.52%, 98.18%, and 97.16% at &#x3b5;=3.0, surpassing all prior federated IDS approaches. Byzantine resilience experiments demonstrate that the proposed defense maintains &gt;95% accuracy under 30% malicious clients across label-flipping, model poisoning, and free-rider attacks. Gradient inversion analysis confirms that FedMamba's compact parameterization (135K parameters, 0.52 MB) provides 2&#xd7; higher reconstruction error compared to Transformer-based FL, and the integrated SHAP and LIME explainability framework supports regulatory compliance with the FDA's 2023 cybersecurity guidance for medical devices.","author":[{"family":"Ym","given":"Al"},{"family":"Am","given":"Almadani"},{"family":"Ah","given":"Abdelhaliem"},{"family":"Is","given":"Fathi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1371/journal.pone.0355601","URL":"https://doi.org/10.1371/journal.pone.0355601","source":"pubmed"},{"id":"doi:10.1038/s41598-026-55768-3","type":"article-journal","title":"Federated ConvNeXt-swin temporal fusion network for malware and botnet detection in IoT systems.","abstract":"The rapid expansion of Internet of Things (IoT) infrastructures has significantly increased the exposure of edge devices to malware and botnet attacks. Conventional intrusion detection systems are largely centralized and struggle to operate effectively in decentralized, heterogeneous, and privacy-sensitive IoT environments, thereby limiting scalability and robustness. To address these challenges, this study proposes the Federated ConvNeXt-Swin Temporal Fusion Network (F-CSTFNet), a federated deep learning framework designed for distributed IoT malware and botnet detection. The proposed architecture integrates ConvNeXt-based convolutional feature extraction with Swin Transformer temporal attention to capture both local traffic patterns and long-range behavioral dependencies within network flows. This hybrid convolution-attention design enables the detection of short-term anomalies as well as evolving attack dynamics directly from network telemetry. In addition, a channel-adaptive feature recalibration mechanism enhances robustness when learning from heterogeneous and noisy client data. The model is trained using a federated learning paradigm that enables multiple IoT clients to collaboratively learn a global model without sharing raw data, thereby preserving data privacy and locality. Extensive experiments conducted on the IoT-23 and N-BaIoT datasets demonstrate that F-CSTFNet outperforms several state-of-the-art centralized and federated baselines in terms of detection accuracy, convergence stability, and client-level fairness. The framework also achieves low performance variance across clients, a high Jain's Fairness Index (JFI), and reduced inequality during distributed training. These results demonstrate the effectiveness of the proposed architecture as a scalable, privacy-preserving, and resilient intrusion detection framework for next-generation IoT security systems.","author":[{"family":"Fs","given":"Alsubaei"},{"family":"Aa","given":"Almazroi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-55768-3","URL":"https://doi.org/10.1038/s41598-026-55768-3","source":"pubmed"},{"id":"doi:10.1016/j.artmed.2026.103422","type":"article-journal","title":"Centralized pooling and federated learning for Canadian patient-level data sharing in multicenter medical AI: A scoping review.","abstract":"Algorithms that support screening, triage, and treatment decisions depend on training data drawn from patient populations. Limited access to patient-level records across institutions and jurisdictions can reduce representation and contribute to uneven model performance across populations. Canada's federated health system, where provinces and territories manage separate datasets and privacy regimes, limits multicenter medical AI research. We conducted a scoping review to map how Canadian researchers share patient-level data in multicenter medical AI collaborations. We searched PubMed, IEEE Xplore, ACM Digital Library, Scopus, and Web of Science from 2018 to February 2025 and implemented a human-in-the-loop large language model process to support screening and extraction, with reviewer validation. Among 3100 included studies, 160 reported multicenter patient-level data collection. Centralized pooling dominated this subset, with 95% of studies using centralized storage and 5% (n = 8) reporting decentralized approaches, including federated learning, sequential model transfer, and distributed feature sharing. Governance requirements were frequently described as multi-site and sequential, and 81.8% of multicenter collaborations reported parallel ethics approvals from three or more institutional review boards. Only one decentralized collaboration operated entirely within Canada. International partnerships comprised 80% of multicenter studies, and many cohorts included non-Canadian sites or non-Canadian data. Our findings support adoption of distributed model development protocols and interoperable governance that limit central pooling while enabling consistent training, validation, and reporting across sites, as only 1 of 160 multicenter studies reported a decentralized approach with Canadian patient data only.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.artmed.2026.103422","URL":"https://doi.org/10.1016/j.artmed.2026.103422","source":"pubmed"},{"id":"doi:10.1007/s10278-026-02069-w","type":"article-journal","title":"Advanced Deep Learning Architectures in MRI-Based Brain Tumor Classification: A Systematic Review Focused on Meningiomas.","abstract":"Deep learning (DL) is increasingly applied to automate brain tumor classification from magnetic resonance imaging (MRI), yet meaningful clinical deployment remains limited by tumor heterogeneity, dataset bias, and incomplete tumor-specific validation. This systematic review synthesizes developments from 2016 to 2025 in advanced DL-based MRI brain tumor classification, with a specific emphasis on meningioma-focused classification and subtyping, given their persistent underrepresentation in AI research. Fifty-six eligible studies were analyzed and organized into five methodological categories: transformer-based models; transformer-based feature extraction pipelines; attention-enhanced convolutional neural networks (CNNs); federated learning approaches; and emerging or unconventional DL strategies. Among studies reporting class-wise metrics (n&#x2009;=&#x2009;19), attention-enhanced CNNs and hybrid CNN-transformer architectures showed high overall accuracy with fewer extreme drops in meningioma performance (F1-score, 89.0% to 99.0%) than end-to-end transformers (F1-score, 79.0% to 97.0%), despite the latter achieving high peak accuracies. Our analysis also showed a marked gap between general tumor categorization and clinically actionable subtyping or grade classification, particularly for meningiomas, reflecting both limited meningioma-targeted tasks and reduced transparency in per-class reporting. Study design and reporting practices limited cross-study comparability, with heavy reliance on a small number of publicly available datasets, frequent class imbalance disadvantaging meningioma representation, and approximately 60% of studies not specifying MRI sequence details. Although predictive performance improved and some studies incorporated interpretability or&#xa0;clinical decision-support&#xa0;components, reporting remained sparse, reinforcing the translational gap between methodological progress and deployment readiness. Finally, publication activity accelerated sharply after 2023 but remained geographically concentrated, raising concerns about representativeness. Our findings call for standardized per-class reporting, greater&#xa0;diverse datasets, interpretability components,&#xa0;and clinically aligned meningioma evaluations to support effective translation into practice.","author":[{"family":"Sj","given":"Holdsworth"},{"family":"Ja","given":"Correia"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1007/s10278-026-02069-w","URL":"https://doi.org/10.1007/s10278-026-02069-w","source":"pubmed"},{"id":"doi:10.1177/00369330261449407","type":"article-journal","title":"ICDMS: Integrating multimodal data for intelligent clinical decision-making in healthcare: Current trends and future directions.","abstract":"Intelligent Clinical Decision-Making Systems have become a cornerstone of modern healthcare by enabling accurate diagnosis, prognosis and treatment planning through data-driven insights. With the growing availability of heterogeneous healthcare data such as medical images, clinical records, physiological signals and textual reports, multimodal learning has emerged as a powerful paradigm for integrating diverse data sources. This review presents a comprehensive and systematic analysis of multimodal approaches for intelligent clinical decision-making by leveraging machine learning, deep learning, transfer learning and natural language processing techniques. A structured literature search was conducted using IEEE, Elsevier, Wiley Online Library and Springer databases, focusing on peer-reviewed studies published between 2020 and 2025. The selected articles were analysed based on data modalities, learning strategies, healthcare applications, datasets and performance evaluation metrics. This review highlights the effectiveness of multimodal frameworks in addressing key challenges such as class imbalance, disease prediction, patient monitoring and treatment planning. Additionally, it discusses the benefits, open challenges and limitations of existing intelligent clinical decision frameworks, including scalability, interpretability and real-world deployment issues. Finally, the review outlines future research directions emphasizing the integration of Internet of Things-enabled healthcare data, federated learning for privacy preservation and blockchain-based secure data sharing to enhance the reliability and clinical adoption of intelligent decision-making systems.","author":[{"family":"Shehnaz"},{"family":"As","given":"Albaqami"},{"family":"Ku","given":"Nisa"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1177/00369330261449407","URL":"https://doi.org/10.1177/00369330261449407","source":"pubmed"},{"id":"doi:10.1177/09287329261466487","type":"article-journal","title":"Mapping the knowledge landscape and research trends of artificial intelligence in breast cancer diagnosis and treatment: A bibliometric analysis.","abstract":"BackgroundArtificial intelligence (AI) has generated rapidly growing research in breast cancer diagnosis and treatment, yet its intellectual structure, collaboration patterns, and thematic evolution remain unmapped.ObjectiveTo analyze the global research landscape of AI in breast cancer from 2001 to 2025, identifying publication trends, contributors, collaboration networks, thematic clusters, and translational gaps.MethodsA bibliometric analysis of 7673 Web of Science articles. VOSviewer was used for co-authorship networks and keyword co-occurrence with overlay visualization; CiteSpace for burst detection.ResultsThe field grew exponentially (29.20% annual growth), with 88.41% of publications from 2020-2025. China (31.63%) and the United States (21.10%) dominated output, but co-authorship networks revealed limited international collaboration for China (22.4% non-Chinese co-authors) and structural exclusion of low- and middle-income countries. Seven keyword clusters showed persistent separation between technology-centric and clinically oriented terms. Temporal overlay revealed a shift from computer-aided detection (pre-2015) to deep learning (2015-2022) and explainable AI, federated learning, and vision transformers (2022-2025). Burst detection confirmed \"feature selection\" (strength=19.53) and \"computer aided detection\" (strength=20.24) as historical hotspots; limited recent bursts indicate emerging frontiers are still accumulating citation impact.ConclusionsAI research in breast cancer has expanded rapidly with evolving themes, yet a translational gap persists between innovation and clinical integration. Future efforts should prioritize prospective validation, international collaboration with underrepresented regions, and standardized frameworks for integrating explainable and privacy-preserving AI into workflows.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1177/09287329261466487","URL":"https://doi.org/10.1177/09287329261466487","source":"pubmed"},{"id":"doi:10.3390/s26103148","type":"article-journal","title":"AdaFed-LDR: Adaptive Federated Learning with Layerwise Dynamics Regularization for Robust Wi-Fi Localization.","abstract":"Wi-Fi Channel State Information (CSI)-based indoor localization enables high-precision positioning, but its deployment across multiple environments faces two major challenges: privacy concerns from centralizing CSI data, and severe statistical heterogeneity (non-IID) arising from the strong environment-dependency of CSI. This heterogeneity creates a stability-plasticity trade-off in federated learning-maintaining precision in known environments (stability) while adapting to unseen domains (plasticity). To address this trade-off, we propose AdaFed-LDR, which combines server-side Confidence-Weighted Adaptive Aggregation with client-side Layerwise Dynamics Regularization (LDR). The aggregation recalibrates client contributions based on feature covariance changes, while LDR imposes depth-dependent constraints-stronger constraints on shallow layers to preserve environment-agnostic features and weaker constraints on deeper layers to allow environment-specific adaptation. Evaluated across 8 indoor environments using Leave-One-Out Cross-Validation and 5 random seeds, AdaFed-LDR achieved a mean localization error (MLE) of 0.41 cm in known environments, corresponding to an 88.2% reduction compared with FedAvg. In domain generalization to unseen environments, AdaFed-LDR achieved an MLE of 218.2&#xb1;2.8 cm, demonstrating an improvement over FedPos (257.6&#xb1;14.04 cm). With one adaptation sample per reference point, MLE improved to 21 cm. Ablation experiments confirmed that combining the two proposed components achieved the highest improvement (83.9%) compared with applying them individually, supporting AdaFed-LDR as a reproducible approach to the stability-plasticity trade-off in federated CSI-based localization.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.3390/s26103148","URL":"https://doi.org/10.3390/s26103148","source":"pubmed"},{"id":"doi:10.1038/s41598-026-42051-8","type":"article-journal","title":"LaED: a novel lightweight, edge-aware and explainable deep learning model for privacy-preserving facial attendance tracking in resource-constrained educational environments.","abstract":"Facial recognition is increasingly adopted for automated classroom attendance; however, real-world deployment in schools remains constrained by privacy risks, ethical obligations, demographic bias, spoofing threats, and limited computational resources. Recent incidents involving Microsoft Teams in New South Wales in 2025 and Chelmer Valley High School in the United Kingdom show how poorly governed systems violate student rights and regulatory compliance. Despite growing adoption, many existing attendance systems focus narrowly on recognition accuracy or efficiency, while overlooking spoof resistance, open-set identity handling, fairness mitigation, auditability, and privacy protection. This paper presents LaED, a lightweight, edge-aware, and explainable deep learning framework for privacy-preserving classroom attendance in resource-constrained educational environments. The framework combines multimodal spoof detection, open-set facial recognition, and fairness-aware representation learning within a unified edge-based design. Spoofing attacks, including replay and deepfake attempts, are mitigated through the fusion of physiological and temporal facial cues, while unknown identities are explicitly rejected to reduce proxy attendance. To support responsible deployment, LaED incorporates federated learning with differential privacy, ensuring that biometric data remain local to schools while enabling accountable model updates. Experimental evaluation on CASIA-FASD, CelebA-Spoof, DFDC, FairFace, and a consent-driven classroom dataset shows that LaED achieves over 97.8% recognition accuracy, APCER and BPCER values below 2%, demographic fairness gaps under 2%, and inference latency below 150 milliseconds on edge hardware. Additional tests confirm reliable operation under realistic classroom conditions. These results demonstrate that regulation-aligned and trustworthy facial attendance is feasible on low-cost devices, offering a practical pathway for responsible biometric AI in education.","author":[{"family":"Eo","given":"Abiodun"},{"family":"Oi","given":"Abiodun"},{"family":"Ba","given":"Shawar"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-42051-8","URL":"https://doi.org/10.1038/s41598-026-42051-8","source":"pubmed"},{"id":"doi:10.1038/s41598-026-48845-0","type":"article-journal","title":"dsLassoCov: a federated Lasso approach incorporating covariate control.","abstract":"Machine learning has been widely adopted in biomedical research, fueled by the increasing availability of data. However, integrating datasets across institutions is challenging due to legal restrictions and data governance complexities. Federated learning allows the direct, privacy preserving training of machine learning models using geographically distributed datasets, but faces the challenge of how to appropriately control for covariate effects. The naive implementation of conventional covariate control methods in federated learning scenarios is often impractical due to the substantial communication costs, particularly with high-dimensional data. To address this issue, we introduce dsLassoCov, a machine learning approach designed to control for covariate effects and allow an efficient training in federated setting. In biomedical analysis, this may support the identification of biomarker candidates while accounting for the effect of confounding. Using simulated data, we demonstrate that dsLassoCov can efficiently and effectively manage confounding effects during model training. In our real-world data analysis, we replicated a large-scale exposome analysis using data from six geographically distinct databases, achieving results consistent with previous studies. By addressing the challenge of covariate control, our proposed approach can accelerate the application of federated learning in large-scale biomedical studies.","author":[{"family":"Jr","given":"Gonzalez"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-48845-0","URL":"https://doi.org/10.1038/s41598-026-48845-0","source":"pubmed"},{"id":"doi:10.1038/s41598-026-47175-5","type":"article-journal","title":"CRAFT: cold-start recommender with attention and federated training.","abstract":"One of the main challenges in recommender systems is the cold-start problem, in which recommendation systems struggle to recommend new or rarely visited items. The traditional methods usually comprise centralized data merging or collaborative filtering techniques, which are not easily applicable in the decentralized settings. The current federated recommendation techniques like FedMF and FedGN have limited support for cold-start personalization, particularly in situations where the metadata of the items is sparse or non-existent. To overcome these drawbacks, we propose a new federated learning-based model, CRAFT (Cold-start Recommender with Attention and Federated Training), that improves cold-start recommendations without compromising the privacy of the user. CRAFT proposes an attention mechanism to highlight salient user-item interaction patterns to enhance the inference of user preferences. Every client then trains a personalized model locally, where the updates are collectively aggregated through Federated Averaging (FedAvg) so that the collective intelligence is obtained without losing the sensitive information. CRAFT provides very personalized suggestions by adding time-varying dynamics and rich interaction histories. CRAFT can also be scaled to be deployed across distributed environments with the use of NVFlare platform. As indicated by experimental results on three real world datasets, including MovieLens 1M, Amazon Movies &amp; TV and CiteULike, CRAFT is able to achieve nDCG 20 in cold-start scenarios up to 16.8 better than state of the art baselines, with strong privacy guarantees.","author":[{"family":"Rs","given":"John"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-47175-5","URL":"https://doi.org/10.1038/s41598-026-47175-5","source":"pubmed"},{"id":"doi:10.1038/s41598-026-45454-9","type":"article-journal","title":"Privacy-preserving federated learning with optimized ensemble weighting and knowledge distillation for COVID-19 detection from non-IID medical imaging data.","abstract":"Medical imaging enables rapid and accurate diagnosis of COVID-19, with CT scans proving especially effective. However, data privacy concerns limit collaborative model development across hospitals. To address this issue, we introduce a novel federated learning framework. It is referred to as Independent Knowledge Distillation with post-Ensemble Federated Learning (IKDEFL). Differential Privacy (DP) is integrated into the framework to improve privacy guarantees. Three DP mechanisms are evaluated. These include Fixed Gaussian, Gaussian Adaptive, and Tree Adaptive. The evaluation has been conducted on heterogeneous and Non-Independent and Identically Distributed (Non-IID) datasets. These datasets reflect real-world hospital scenarios. Results show that IKDEFL significantly outperforms existing federated knowledge distillation methods. It achieves a generalization performance with accuracy up to 82.79% and an F1-score of 82.78% on Non-IID data. Among the DP methods, the Tree Adaptive mechanism has consistently provided the best trade-off between privacy and prediction quality. Peak accuracy reaches 76.62% under strict privacy constraints, where [Formula: see text] and [Formula: see text]. This result is close to the performance of models without privacy protections. These findings demonstrate that adaptive DP techniques can be effectively applied in federated healthcare models. They support the development of privacy-preserving AI systems for clinical diagnostics.","author":[],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-45454-9","URL":"https://doi.org/10.1038/s41598-026-45454-9","source":"pubmed"},{"id":"doi:10.1016/j.mex.2026.103898","type":"article-journal","title":"A fully homomorphic encryption federated learning architecture for privacy preserving in industrial internet of things.","abstract":"We are currently entering the fifth revolution of industry - Industry 5.0. IIoT is the domain where massive quantities of data are flourished by the associated devices in an industry on a daily basis. To realize industry 4.0, the Industrial internet of things is considered a prominent one. Federated Learning, also known as collaborative learning, employs a decentralized approach in its applicability while maintaining data privacy, but many existing frameworks struggle with handling privacy of gradients which are transferred to Federated servers. Unlike conventional approaches, proposed fully homomorphic encryption based Federated Learning-FHEEFL ensures that raw gradients never leave local IoT nodes; instead, only updates or changes in model secured with encryption techniques are transmitted. The framework achieves a very strong privacy as well as data security. FHEEFL is intended for lightweight to moderate-capacity models characteristically labouring in Edge-IIoTset analytics, where privacy guarantees must be well-adjusted with computational feasibility. While CKKS-based encrypted aggregation incurs additional overhead, the framework establishes practical applicability for privacy-critical industrial tasks under realistic resource constraints. It is proven as a privacy-centric federated learning solution, setting a new benchmark in tackling key challenges in data security and privacy. The proposed method is implemented using Edge-IIoTset dataset. The proposed fully homomorphic encryption based Federated Learning-FHEEFL method is tested in IIOT scenarios.&#x2022;FHEEFL method provides better privacy with better memory usage, CPU usage and throughput parameters.&#x2022;Performance is analysed with 4 variations of models- Tiny, small, medium and large.","author":[{"family":"Subhedar","given":"Shraddha"},{"family":"Parasar","given":"Deepa"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1016/j.mex.2026.103898","URL":"https://doi.org/10.1016/j.mex.2026.103898","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-8300312/v1","type":"article-journal","title":"Sparsity-Aware Edge Caching in IoVs with Asynchronous Federated and Deep Reinforcement Learning","abstract":"Abstract The edge content caching technology of the Internet of Vehicles (IoVs) is a key technology to reduce the latency of content access. However, within the hotspot area, with a large number of content access requests generated by vehicle users, the rapid changes in user interests and the explosive dissemination of high-value content have led to limited transmission delay. Therefore, to reduce the delay, accurately predicting and timely updating popular content as well as exploring high-value content have become critical yet challenging. To solve this problem, a sparsity-aware edge caching (SAEC) scheme is proposed. Firstly, aiming at the sparsity problem of VU data, a sparse self-encoder based on self-attentive (SAE-ELA) model is proposed. By extracting the potential features of sparse data of vehicle users and capturing the historical preference associations of users, the accuracy of content prediction was improved. Secondly, this paper adopts the asynchronous federated learning (AFL) framework to solve the problem of low cache update efficiency, thereby shortening the model training time to improve the real-time performance of content update. Finally, in order to solve the problem of insufficient exploration of potential high-value content, a Dueling Deep Q-network based on Intrinsic Curiosity Module (ICM-DDQN) algorithm is proposed. By organically combining traditional value function learning with curiosity driven active exploration, the exploration efficiency of cached content has been improved, thereby reducing the Content Transmission Delay (CTD). Simulation results show that the proposed SAEC is significantly superior to the existing methods in terms of Cache Hit Ratio(CHR) and CTD.","author":[{"family":"Gao","given":"Jing"},{"family":"Chen","given":"Jiahui"},{"family":"Huan","given":"Yanqi"},{"family":"Wu","given":"Liuyang"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-8300312/v1","URL":"https://doi.org/10.21203/rs.3.rs-8300312/v1","source":"europepmc"},{"id":"doi:10.1038/s41598-026-47631-2","type":"article-journal","title":"Federated CT foundation models for multi-center detection of lymph node metastasis in pancreatic cancer.","abstract":"Pancreatic ductal adenocarcinoma (PDAC) remains one of the most lethal malignancies, with prognosis strongly influenced by the presence of lymph node metastasis (LNM). However, preoperative LNM assessment from computed tomography (CT) is limited by low sensitivity, high inter-observer variability, and substantial heterogeneity across imaging protocols. This retrospective multi-center study (546 patients from three institutions) introduces a privacy-preserving deep learning framework that integrates large-scale CT foundation model pre-training with heterogeneity-aware federated optimization to improve LNM detection in PDAC. A CT Vision Foundation Model, pre-trained on 148,000 volumetric CT scans using contrastive self-supervised learning, is fine-tuned to generate transferable 3D representations for patient-level LNM classification. To enable decentralized model training while mitigating inter-institutional variability, we extend federated aggregation to jointly account for label-distribution discrepancies and representation-level divergence across clients. The centralized model achieved a balanced accuracy of 0.601 and a diagnostic odds ratio (DOR) of 3.45, outperforming classical machine learning baselines and prior PDAC LNM approaches. Under federated settings, the proposed heterogeneity-aware strategy consistently outperformed standard FedAvg, recovering a substantial proportion of the centralized model's performance while preserving strict data privacy. In particular, it improved balanced accuracy by 12.6% over FedAvg and demonstrated superior discriminative ability across all participating cohorts. These findings indicate that combining foundation model pre-training with discrepancy-aware federated learning enhances generalization, robustness, and clinical relevance for multi-center PDAC LNM detection. The proposed framework offers a scalable and privacy-preserving pathway for deploying deep learning models across distributed healthcare systems.","author":[{"family":"Dd","given":"Gaviria"},{"family":"Asa","given":"Hosseini"}],"issued":{"date-parts":[[2026]]},"DOI":"10.1038/s41598-026-47631-2","URL":"https://doi.org/10.1038/s41598-026-47631-2","source":"pubmed"},{"id":"doi:10.21203/rs.3.rs-7612387/v1","type":"article-journal","title":"Multi agent federated reinforcement learning for Distributed MPTCP Agents","abstract":"Abstract The integration of Multi access Edge Computing (MEC) with low Earth orbit (LEO) satellite constellations is a promising paradigm for global, low latency connectivity. However, the dynamic topology and heterogeneous link qualities of satellite networks pose significant challenges for efficient multipath transport protocol (MPTCP) scheduling. Traditional schedulers, often based on heuristics or designed for fixed size inputs, struggle to adapt to the variable number of available paths. We propose a novel Multi Agent Federated Reinforcement Learning (MAFRL) framework that leverages Set Transformers for permutation invariant encoding of variable path sets. Each agent learns a local policy using Proximal Policy Optimization (PPO), augmented with a soft fairness constraint to ensure equitable performance. A federated learning scheme, using FedProx aggregation, enables collaborative training across distributed agents without sharing raw data, preserving privacy and improving robustness to non IID data. Extensive emulation experiments show our approach outperforms heuristic and learning based baselines in aggregate throughput, latency, and fairness, particularly under path variability. This work demonstrates the viability of set based learning and federated optimization for intelligent resource management in next generation satellite terrestrial networks.","author":[{"family":"Suarez","given":"Jorge"},{"family":"Jia","given":"Min"},{"family":"Anyembe","given":"CS"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7612387/v1","URL":"https://doi.org/10.21203/rs.3.rs-7612387/v1","source":"europepmc"},{"id":"doi:10.5281/zenodo.21260159","type":"article-journal","title":"Artificial Intelligence-Enhanced Pharmacovigilance for Diethylene Glycol Contamination in Paediatric Medicines: A Systematic Narrative Review and Proposed Computational Framework for Real-Time Outbreak Detection and Supply Chain Surveillance","abstract":"Diethylene glycol (DEG) contamination of paediatric oral liquid medicines constitutes a recurrent, preventable global public health crisis. Since 2022 alone, WHO-confirmed outbreaks across The Gambia, Uzbekistan, Indonesia, and India have claimed the lives of more than 300 children. Despite over eight decades of documented incidents— beginning with the 1937 Elixir Sulfanilamide disaster—conventional pharmacovigilance systems have repeatedly failed to detect, contain, or prevent these outbreaks with sufficient rapidity to avert mass mortality. Artificial intelligence (AI) and machine learning (ML) offer transformative potential for real-time signal detection, supply chain anomaly identification, and predictive risk stratification in pharmaceutical quality surveillance.Objectives: To: (i) synthesise the global epidemiological, clinical, and toxicological evidence on DEG poisoning across twelve major outbreak events spanning 1937–2025; (ii) characterise the systemic pharmacovigilance and regulatory failures that have enabled repeated tragedies; and (iii) propose a validated, five-component AI-driven computational framework specifically architected to address identified failure domains.Methods: We conducted a pre-specified systematic narrative review of peer-reviewed literature, WHO Medical Product Alerts, regulatory communications, epidemiological field reports, and grey literature pertaining to DEG and ethylene glycol (EG) pharmaceutical contamination from 1937 to June 2025. Search strategies were applied across PubMed/MEDLINE, Scopus, WHO IRIS, Cochrane Library, and Google Scholar using MeSH-mapped terms. Supplementary AI/ML pharmacovigilance literature was systematically searched to inform framework design. Data were extracted across standardised domains: outbreak epidemiology, pathophysiological mechanisms, clinical presentation, management outcomes, regulatory responses, and identified failure points.Results: Twelve major DEG outbreak events were identified across ten countries, accounting for over 800 documented deaths, with children under five years disproportionately represented. Consistent root causes across all outbreaks encompassed six recurring failure domains: excipient adulteration, over-reliance on unverified Certificates of Analysis (CoAs), absence of mandatory in-house testing, inadequate regulatory inspections, opaque supply chains, and delayed alert and recall systems. Clinically, DEG produces a characteristic biphasic illness progressing to high-anion-gap metabolic acidosis, proximal tubular necrosis, and multiorgan failure; mortality is strongly time-dependent. Early fomepizole and haemodialysis significantly reduce case fatality rates. No outbreak reviewed employed AI-assisted surveillance or supply chain monitoring. The proposed AI framework comprises: (1) NLP-based real-time pharmacovigilance signal mining; (2) Graph Neural Network supply chain traceability; (3) CNN-assisted portable spectroscopic quality screening; (4) Federated Learning for cross-jurisdictional surveillance; and (5) LLM-powered regulatory alert synthesis.Conclusions: DEG poisoning remains a race against death in which current pharmacovigilance systems consistently fail to intervene in time. The proposed AI-driven framework, if deployed within an international regulatory architecture, offers a credible, scalable, and privacy preserving pathway to preventing future paediatric fatalities from preventable pharmaceutical contamination. Immediate investment in digital pharmacovigilance infrastructure—particularly in low- and middle-income countries—is a global health imperative.","author":[{"family":"Chennamchetty","given":"Vijay"},{"family":"Nallagonda","given":"Ravindra"},{"family":"Mv","given":"Raghavendra"},{"family":"Sankuru","given":"Daniel"},{"family":"Dasari","given":"Sumedha"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21260159","URL":"https://doi.org/10.5281/zenodo.21260159","source":"datacite"},{"id":"doi:10.5281/zenodo.21260160","type":"article-journal","title":"Artificial Intelligence-Enhanced Pharmacovigilance for Diethylene Glycol Contamination in Paediatric Medicines: A Systematic Narrative Review and Proposed Computational Framework for Real-Time Outbreak Detection and Supply Chain Surveillance","abstract":"Diethylene glycol (DEG) contamination of paediatric oral liquid medicines constitutes a recurrent, preventable global public health crisis. Since 2022 alone, WHO-confirmed outbreaks across The Gambia, Uzbekistan, Indonesia, and India have claimed the lives of more than 300 children. Despite over eight decades of documented incidents— beginning with the 1937 Elixir Sulfanilamide disaster—conventional pharmacovigilance systems have repeatedly failed to detect, contain, or prevent these outbreaks with sufficient rapidity to avert mass mortality. Artificial intelligence (AI) and machine learning (ML) offer transformative potential for real-time signal detection, supply chain anomaly identification, and predictive risk stratification in pharmaceutical quality surveillance.Objectives: To: (i) synthesise the global epidemiological, clinical, and toxicological evidence on DEG poisoning across twelve major outbreak events spanning 1937–2025; (ii) characterise the systemic pharmacovigilance and regulatory failures that have enabled repeated tragedies; and (iii) propose a validated, five-component AI-driven computational framework specifically architected to address identified failure domains.Methods: We conducted a pre-specified systematic narrative review of peer-reviewed literature, WHO Medical Product Alerts, regulatory communications, epidemiological field reports, and grey literature pertaining to DEG and ethylene glycol (EG) pharmaceutical contamination from 1937 to June 2025. Search strategies were applied across PubMed/MEDLINE, Scopus, WHO IRIS, Cochrane Library, and Google Scholar using MeSH-mapped terms. Supplementary AI/ML pharmacovigilance literature was systematically searched to inform framework design. Data were extracted across standardised domains: outbreak epidemiology, pathophysiological mechanisms, clinical presentation, management outcomes, regulatory responses, and identified failure points.Results: Twelve major DEG outbreak events were identified across ten countries, accounting for over 800 documented deaths, with children under five years disproportionately represented. Consistent root causes across all outbreaks encompassed six recurring failure domains: excipient adulteration, over-reliance on unverified Certificates of Analysis (CoAs), absence of mandatory in-house testing, inadequate regulatory inspections, opaque supply chains, and delayed alert and recall systems. Clinically, DEG produces a characteristic biphasic illness progressing to high-anion-gap metabolic acidosis, proximal tubular necrosis, and multiorgan failure; mortality is strongly time-dependent. Early fomepizole and haemodialysis significantly reduce case fatality rates. No outbreak reviewed employed AI-assisted surveillance or supply chain monitoring. The proposed AI framework comprises: (1) NLP-based real-time pharmacovigilance signal mining; (2) Graph Neural Network supply chain traceability; (3) CNN-assisted portable spectroscopic quality screening; (4) Federated Learning for cross-jurisdictional surveillance; and (5) LLM-powered regulatory alert synthesis.Conclusions: DEG poisoning remains a race against death in which current pharmacovigilance systems consistently fail to intervene in time. The proposed AI-driven framework, if deployed within an international regulatory architecture, offers a credible, scalable, and privacy preserving pathway to preventing future paediatric fatalities from preventable pharmaceutical contamination. Immediate investment in digital pharmacovigilance infrastructure—particularly in low- and middle-income countries—is a global health imperative.","author":[{"family":"Chennamchetty","given":"Vijay"},{"family":"Nallagonda","given":"Ravindra"},{"family":"Mv","given":"Raghavendra"},{"family":"Sankuru","given":"Daniel"},{"family":"Dasari","given":"Sumedha"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21260160","URL":"https://doi.org/10.5281/zenodo.21260160","source":"datacite"},{"id":"doi:10.34734/fzj-2026-03473","type":"article-journal","title":"When to Harmonize? Evaluating Stage-Specific Harmonization in Federated Brain Age Estimation","abstract":"Federated learning (FL) is a promising solution for healthcare Artificial Intelligence (AI), striking a balance between patient privacy and the need for diverse datasets. FL enables collaborative model training across institutions, preserving confidentiality and advancing clinical tasks such as diagnosis and treatment planning. However, a key challenge in this setting is the inherent heterogeneity of medical datasets acquired in different institutions, which can undermine the generalizability and performance of the model. This issue is particularly pronounced in neuroimaging applications, such as Magnetic resonance imaging (MRI), where site-specific biases arise from variations in scanner hardware, acquisition protocols, and preprocessing pipelines. These differences introduce non-biological variability that can jeopardize downstream analyses and model training. To remove the effect of sites, harmonization techniques are essential tools to improve robustness and reliability. Harmonization techniques are usually applied at the feature level; however, given the limited access to the data possessed by FL schemes, feature-level harmonization may not be enough to remove site effects. In this work, we propose two complementary harmonization strategies within the FL framework: (1) the traditional feature harmonization, by applying ComBat to directly correct the MRI-derived features; and (2) gradient harmonization, which aligns local model updates, particularly the gradients of fully connected layers, across sites to mitigate inter-site distributional shifts before global aggregation. Together, these approaches aim to improve cross-site consistency and improve the model's overall performance in federated medical imaging tasks.","author":[{"family":"Halder","given":"Tanurima"},{"family":"Deo","given":"Kunal"},{"family":"Nieto","given":"Nicolás"},{"family":"Patil","given":"Kaustubh"},{"family":"Jadhav","given":"Kshitij"}],"issued":{"date-parts":[[2025]]},"DOI":"10.34734/fzj-2026-03473","URL":"https://doi.org/10.34734/fzj-2026-03473","source":"datacite"},{"id":"doi:10.5281/zenodo.21702373","type":"article-journal","title":"AI-Guided Precision Nutrition and Pharmacogenomics: A New Frontier in Personalized Cancer Prevention and Therapy","abstract":"Precision nutrition and pharmacogenomics, driven by artificial intelligence (AI) technologies, are becoming an integral part of the science of precision oncology and are providing a pathway to highly personalized cancer prevention, diagnosis, and treatment. Current cancer therapeutics make use of standard treatment regime that do not consider the interindividual variation in genetic, metabolic, microbiome, immune regulatory, and drug sensitivity. The shift towards a personalized oncology model that combines genomic, transcriptomic, proteomic, and metabolomic and clinical information to guide therapeutic decision has been fast-tracked with the advent of recent technologies and advances in machine learning, multi-omics, and computational biology. Precision nutrition is an approach to dietary interventions based on molecular profile, and pharmacogenomics is the study of the influence of genetic variation on drug metabolism, efficacy, toxicity, and therapeutic resistance. In the last six years (2020 to 2025), artificial intelligence (AI) has been used increasingly over the past five years to: forecast cancer risk, discover predictive biomarkers, optimize chemotherapy and immunotherapy treatments, minimize side effects of drugs, and improve nutritional care in cancer Moreover, the increasing body of evidence underscores the importance of the gut microbiome, inflammation control, immune response, and therapeutic response, thereby enhancing the precision medicine capabilities of AI. New technologies like digital twin, explainable AI, federated learning, and real-time wearable biosensors will be expected to take cancer adoptive management to the next level, as well as in the field of cancer prevention. This chapter critically reviews recent advances and prospects of AI-driven precision nutrition and pharmacogenomics, highlighting their translational applications, clinical relevance, technological advances for cancer prevention and treatment.","author":[{"family":"Sadia","given":"Khan"},{"family":"Hafsa","given":"Asif"},{"family":"Hina","given":"Nawab"},{"family":"Amna","given":"Arif"},{"family":"Natasha","given":"Iqbal"},{"family":"Aiman","given":"Zahid"},{"family":"Rubia","given":"Anwer"},{"family":"Abida","given":"Shamim"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21702373","URL":"https://doi.org/10.5281/zenodo.21702373","source":"datacite"},{"id":"doi:10.5281/zenodo.21702374","type":"article-journal","title":"AI-Guided Precision Nutrition and Pharmacogenomics: A New Frontier in Personalized Cancer Prevention and Therapy","abstract":"Precision nutrition and pharmacogenomics, driven by artificial intelligence (AI) technologies, are becoming an integral part of the science of precision oncology and are providing a pathway to highly personalized cancer prevention, diagnosis, and treatment. Current cancer therapeutics make use of standard treatment regime that do not consider the interindividual variation in genetic, metabolic, microbiome, immune regulatory, and drug sensitivity. The shift towards a personalized oncology model that combines genomic, transcriptomic, proteomic, and metabolomic and clinical information to guide therapeutic decision has been fast-tracked with the advent of recent technologies and advances in machine learning, multi-omics, and computational biology. Precision nutrition is an approach to dietary interventions based on molecular profile, and pharmacogenomics is the study of the influence of genetic variation on drug metabolism, efficacy, toxicity, and therapeutic resistance. In the last six years (2020 to 2025), artificial intelligence (AI) has been used increasingly over the past five years to: forecast cancer risk, discover predictive biomarkers, optimize chemotherapy and immunotherapy treatments, minimize side effects of drugs, and improve nutritional care in cancer Moreover, the increasing body of evidence underscores the importance of the gut microbiome, inflammation control, immune response, and therapeutic response, thereby enhancing the precision medicine capabilities of AI. New technologies like digital twin, explainable AI, federated learning, and real-time wearable biosensors will be expected to take cancer adoptive management to the next level, as well as in the field of cancer prevention. This chapter critically reviews recent advances and prospects of AI-driven precision nutrition and pharmacogenomics, highlighting their translational applications, clinical relevance, technological advances for cancer prevention and treatment.","author":[{"family":"Sadia","given":"Khan"},{"family":"Hafsa","given":"Asif"},{"family":"Hina","given":"Nawab"},{"family":"Amna","given":"Arif"},{"family":"Natasha","given":"Iqbal"},{"family":"Aiman","given":"Zahid"},{"family":"Rubia","given":"Anwer"},{"family":"Abida","given":"Shamim"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21702374","URL":"https://doi.org/10.5281/zenodo.21702374","source":"datacite"},{"id":"doi:10.5281/zenodo.21701718","type":"article-journal","title":"SEAScale: Predictive Serverless Autoscaling Architecture For High Demand Ride Sharing Platforms","abstract":"Modern networks produced enormous volumes of data which made them increasingly vulnerable to advanced cyber threats. Traditional scanning tools and standalone anomaly detection systems fell short in identifying evolving or zero day attacks. This study proposed a hybrid framework that integrated active network scanning with machine learning based anomaly detection to deliver an adaptive and automated solution. The framework combined Nmap based host scanning tcpdump based traffic capture and Zeek based feature extraction with a Random Forest classifier and an autoencoder trained on normal traffic to flag deviations. Recent advancements in supervised and unsupervised anomaly detection together with integrated intrusion detection systems published between 2020 and 2025 were reviewed and synthesized to position the proposed approach within the broader field. Experimental comparison across five learning models showed that the Convolutional Neural Network achieved the highest accuracy of 93 percent followed by Long Short Term Memory at 91 percent while Random Forest balanced accuracy and interpretability at 89 percent. The hybrid correlation mechanism that combined scan derived signals with model predictions reduced false alarms and strengthened real time threat visibility compared with single method systems. The study concluded that combining active scanning with machine learning significantly improved detection accuracy reduced false positives and enabled actionable reporting for security analysts. Future directions identified included adaptive learning methods lightweight models suited for edge devices and secure distributed detection through federated learning.","author":[{"family":"Pradeep"},{"family":"Zaidi","given":"Husain"},{"family":"Singh","given":"Sakshi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21701718","URL":"https://doi.org/10.5281/zenodo.21701718","source":"datacite"},{"id":"doi:10.5281/zenodo.21701717","type":"article-journal","title":"SEAScale: Predictive Serverless Autoscaling Architecture For High Demand Ride Sharing Platforms","abstract":"Modern networks produced enormous volumes of data which made them increasingly vulnerable to advanced cyber threats. Traditional scanning tools and standalone anomaly detection systems fell short in identifying evolving or zero day attacks. This study proposed a hybrid framework that integrated active network scanning with machine learning based anomaly detection to deliver an adaptive and automated solution. The framework combined Nmap based host scanning tcpdump based traffic capture and Zeek based feature extraction with a Random Forest classifier and an autoencoder trained on normal traffic to flag deviations. Recent advancements in supervised and unsupervised anomaly detection together with integrated intrusion detection systems published between 2020 and 2025 were reviewed and synthesized to position the proposed approach within the broader field. Experimental comparison across five learning models showed that the Convolutional Neural Network achieved the highest accuracy of 93 percent followed by Long Short Term Memory at 91 percent while Random Forest balanced accuracy and interpretability at 89 percent. The hybrid correlation mechanism that combined scan derived signals with model predictions reduced false alarms and strengthened real time threat visibility compared with single method systems. The study concluded that combining active scanning with machine learning significantly improved detection accuracy reduced false positives and enabled actionable reporting for security analysts. Future directions identified included adaptive learning methods lightweight models suited for edge devices and secure distributed detection through federated learning.","author":[{"family":"Pradeep"},{"family":"Zaidi","given":"Husain"},{"family":"Singh","given":"Sakshi"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21701717","URL":"https://doi.org/10.5281/zenodo.21701717","source":"datacite"},{"id":"doi:10.21203/rs.3.rs-7585019/v1","type":"article-journal","title":"Adaptive Spectrum Sensing and Management in Cognitive Radio Networks Using Federated Deep Reinforcement Learning","abstract":"Abstract The dynamic and unexpected character of settings in wireless communication calls for sophisticated spectrum sensing techniques for cognitive radio networks. Building on the work of earlier ensemble machine learning approaches, this study presents a state-of-the-art framework for real-time spectrum management using federated deep reinforcement learning (FDRL). The combination of reinforcement learning's strategic decision-making process with Deep Belief Networks' (DBN) and Long Short-Term Memory's (LSTM) architectures is at the heart of this methodology. This approach, which operates inside a federated learning paradigm, gives user privacy and data locality, guaranteeing a reliable and private solution. Through processing signal vectors under different noise situations, the FDRL model repeatedly learns the best spectrum allocation strategies, improving its comprehension over time. This novel approach offers effective adaptability to the ever-changing wireless environment, improving network speed and spectrum utilization while protecting user privacy. Effectively separating idle from active channels, it continuously adjusts to variations in signal-to-noise ratios and user demands. This sophisticated technology is shown through thorough simulations to provide a significant improvement in both spectrum efficiency and user throughput. Because of its scalability and decentralization, it presents a viable answer to the changing wireless network environment, which is marked by an increasing need for autonomy and data-driven operations. This approach's proven ability to reduce interference and improve service quality indicates a major step forward for intelligent and autonomous spectrum sensing methods, which are critical in the age of ubiquitous wireless communication. The suggested FDRL-DBN-LSTM approach was implemented in Python and achieves an accuracy of 98.4%.","author":[{"family":"Msaraswathi"},{"family":"Dlakshminarayana"},{"family":"Pvaishnavidevi"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7585019/v1","URL":"https://doi.org/10.21203/rs.3.rs-7585019/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-7616375/v1","type":"article-journal","title":"Federated Reinforcement Learning Framework for Privacy Preserving Few Shot Learning","abstract":"Abstract This study introduces a federated reinforcement learning framework for few-shot learning (FRL-FSL), aiming to address the dual challenges of data scarcity and privacy preservation in distributed environments. The proposed framework integrates policy gradient optimization with secure aggregation and introduces validator nodes to ensure the authenticity of both data and model updates. Experiments were conducted on the Omniglot and FC100 datasets under 1-shot and 5-shot conditions, with comparisons against FedAvg, FedFSL and traditional supervised baselines. Results demonstrate that FRL-FSL achieved an average accuracy of 87.3% on Omniglot (5-shot), improving by 25.9% over FedAvg and 13.8% over FedFSL, while maintaining 72.6% accuracy in 1-shot tasks. On the FC100 dataset, FRL-FSL reached 59.8% accuracy in 5-shot learning, outperforming FedAvg by 18.6% and FedFSL by 7.1%, and achieved 46.3% in 1-shot learning. The framework also reduced the privacy risk index by 37% relative to FedAvg, with convergence accelerated by nearly 30% compared to baselines. These findings confirm that FRL-FSL achieves a practical balance between accuracy, convergence, and privacy, offering a promising solution for real-world, privacy-sensitive applications.","author":[{"family":"Kang","given":"Minsoo"},{"family":"Park","given":"Jihye"},{"family":"Choi","given":"Donghyun"},{"family":"Kim","given":"Seoyeon"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-7616375/v1","URL":"https://doi.org/10.21203/rs.3.rs-7616375/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-6674711/v1","type":"article-journal","title":"Federated Defense: A Privacy-Preserving Deep Learning Model for IoT Malware Detection","abstract":"Abstract As the Internet of Things (IoT) continues to expand, securing the vast network of IoT devices, particularly in Machine-to-Machine (M2M) communication, has become a critical concern. Traditional security approaches often fall short, particularly in protecting privacy and ensuring scalability across IoT systems' diverse and vast landscapes. This paper introduces the Federated Defense Model, a novel approach that harnesses federated learning (FL) to enhance IoT malware detection while preserving user privacy. Unlike centralized models, the proposed FL framework processes data locally on IoT devices, avoiding transmitting sensitive information and reducing bandwidth demands. We developed and evaluated a lightweight, one-dimensional convolutional neural network (CNN) optimized for the typical environment of IoT devices. Using the IoT-23 dataset, a collection of labeled network traffic representing various malware and benign scenarios, the experiments demonstrate that the proposed Federated Defense Model achieves superior accuracy and precision compared to the signature, heuristic, and traditional machine learning-based security models. Moreover, the Federated Defense Model is compared with the existing state-of-the-art FL models. The findings suggest that integrating FL with deep learning techniques bolsters IoT security and mitigates privacy risks and scalability challenges inherent in centralized approaches. This work contributes to the ongoing evolution of privacy protection strategies in the IoT domain, emphasizing the role of privacy-preserving methodologies in developing resilient digital ecosystems.","author":[{"family":"Abbas","given":"Sohail"},{"family":"Abrar","given":"Mohammad"},{"family":"Jan","given":"Mian"},{"family":"Abul","given":"Osman"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-6674711/v1","URL":"https://doi.org/10.21203/rs.3.rs-6674711/v1","source":"preprints"},{"id":"doi:10.21203/rs.3.rs-6819153/v1","type":"article-journal","title":"A Federated Meta-Learning Aided Intelligent Edge Framework by Using the Parameter Optimization Approach","abstract":"Abstract Edge intelligence can enable fast intelligent services by integrating edge computing with machine learning, thereby facilitating intelligent information processing for Internet of Things (IoT) devices on the edge. However, intelligent data processing at the edge may expose IoT devices to the risk of private information leakage. To mitigate this issue, we propose a federal meta-learning-aided data processing framework to cope with complex tasks in edge IoT networks. Unfortunately, communications between edge IoT devices and edge servers in federated frameworks incur significant overhead. To address this challenge, we propose a parameter optimization algorithm that alleviates communication costs between edge IoT devices and edge servers, thereby reducing classification errors induced by parameter optimization. Moreover, the convergence of the federated meta-learning method is derived, which theoretically confirms the feasibility of the proposed approach. Simulation results demonstrate that the error minimization-based quantization compression optimization algorithm can substantially enhance communication efficiency while incurring only negligible precision losses.","author":[{"family":"Zhu","given":"Xiaofeng"},{"family":"Fan","given":"Qiaosong"},{"family":"Peng","given":"Jiaqiang"},{"family":"Qian","given":"Yuwen"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21203/rs.3.rs-6819153/v1","URL":"https://doi.org/10.21203/rs.3.rs-6819153/v1","source":"preprints"},{"id":"doi:10.5281/zenodo.18983099","type":"article-journal","title":"Die ZuMult-Plattform als Instrument für sprachvergleichende Analysen auf mündlichen Daten","abstract":"“[Tools widely used by corpus linguists] all offer a different user-experience, because each tool is created in isolation and thus offers a different user interface, control flow, and functionality.” (Anthony 2009) Nur wenige größere Korpora sind aus sich heraus auf sprachvergleichende Analysen angelegt. Zu den von Vorneherein als „Comparable Corpus“ konzipierten Ausnahmen gehören das International Comparable Corpus (ICC, Čermáková et al. 2021) oder das mehrsprachige GeWiss-Korpus (Fandrych et al. 2017). Ein alternativer oder ergänzender Ansatz sind virtuelle vergleichbare Korpora wie EuReCo (Trawínski & Kupietz 2021), die über Föderation verteilter, auf vergleichbarer technischer Basis stehender Korpora und Korpusplattformen ermöglichen, das Deutsche kontrastiv mit anderen europäischen Sprachen in Beziehung zu setzen; EuReCo ist allerdings auf schriftsprachliche Daten beschränkt. Unser Beitrag stellt die ZuMult-Plattform (Fandrych et al. 2023) und deren Potential, Vergleichbares für mündliche Korpora zu leisten, vor. ZuMult ist eine offene, flexible, auf etablierten Standards und Technologien basierende Architektur für den Zugang zu audiovisuellen Korpora. Neben Korpusrecherchen mit CQP unterstützt ZuMult die für eine Analyse gesprochener Sprache notwendige erweiterte Kontextualisierung von Suchergebnissen (Frick & Schmidt 2025) sowie interaktive Transkriptanalysen (Schmidt et al. 2023), die insbesondere für qualitativ orientierte Analysen, z.B. in der Gesprächsforschung, und für didaktische Anwendungen, z.B. in der DaF/DaZ-Lehre, genutzt werden. An der Universität Duisburg-Essen erfolgt eine Erweiterung, die darauf zielt, Sprache auch in ihrem multimodalen Zusammenspiel mit weiteren körperlichen Ressourcen – wie Blick, Gestik etc. – für korpuslinguistische Herangehensweisen zu erschließen. Über eine ZuMult-Instanz am Archiv für Gesprochenes Deutsch sind bereits seit 2021 neben dem GeWiss-Korpus zwei der wichtigsten Referenzkorpora des gesprochenen Deutsch – FOLK (Deppermann & Hartung 2011, Schmidt 2016, Reineke et al. 2023) und Deutsch Heute (Kleiner 2015) – zugänglich. Mit der Veröffentlichung einer ZuMult-Instanz für das französische ESLO-Korpus (Abouda & Baude 2006, Baude & Dugua 2011, Eshkol-Taravella 2012, Schmidt 2025) ergeben sich nun erste Möglichkeiten für sprachvergleichende Analysen zwischen dem Deutschen und dem Französischen. Wie unser Beitrag anhand von Proof-Of-Concept-Implementierungen zeigen wird, sind beispielsweise auch das Griffith Corpus of Spoken Australian English (Haugh & Chang 2013), das TIGR Corpus des gesprochenen Italienisch (Miecznikowski-Fuenfschilling et al. i.V.), das Training Corpus of Spoken Slovenian (Verdonik 2024) und die Kollektionen aus Oral History Digital (Pagenstecher 2024) in ZuMult integrierbar. Gleiches gilt für das zwölfsprachige EXMARaLDA-Demokorpus, das derzeit im Rahmen eines Text+-Kooperationsprojekt für eine Publikation in einer ZuMult-Instanz an der Universität Hamburg aufbereitet wird. Auf ähnliche Weise eröffnet ZuMult neue Möglichkeiten der „vergleichenden Sprachinselforschung“ (Boas 2016), also der (korpusgestützten) Analyse von Sprachkontaktphänomenen und Sprachentwicklung von Deutsch als Minderheitensprache im Kontakt mit dominanten anderen (i.d.R. europäischen) Sprachen. Seit Dezember 2024 macht eine ZuMult-Instanz an der University of Texas in Austin (Boas et al. 2025) die Daten des Texas German Dialect Projects auf ähnliche Weise verfügbar, wie die IDS-Instanz Zugriff etwa auf Daten zum Australiendeutsch (Clyne 1981), Deutsch in Namibia (Zimmer et al. 2020) oder Mennonitendeutsch in Amerika (Kaufmann et al. 2023) bietet. Unser Beitrag wird diese verschiedenen Anwendungen vorstellen und illustrieren, sowie einige methodische Herausforderungen der Mehrsprachigkeit (z.B. sprachübergreifendes POS-Tagging) und technische Ansätze zur Aggregation von Suchanfragen an mehrere ZuMult-Instanzen (CLARIN Federated Content Search) thematisieren. Keywords: gesprochene Sprache; Ko","author":[{"family":"Schmidt","given":"Thomas"},{"family":"Abouda","given":"Lotfi"},{"family":"Badin","given":"Flora"},{"family":"Boas","given":"Hans"},{"family":"Blevins","given":"Margaret"},{"family":"Bührig","given":"Kristin"},{"family":"Dugua","given":"Céline"},{"family":"Fandrych","given":"Christian"},{"family":"Ferger","given":"Anne"},{"family":"Frick","given":"Elena"},{"family":"Kompiel","given":"Peter"},{"family":"Miecznikowski-Fuenfschilling","given":"Johanna"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18983099","URL":"https://doi.org/10.5281/zenodo.18983099","source":"datacite"},{"id":"doi:10.5281/zenodo.18983100","type":"article-journal","title":"Die ZuMult-Plattform als Instrument für sprachvergleichende Analysen auf mündlichen Daten","abstract":"“[Tools widely used by corpus linguists] all offer a different user-experience, because each tool is created in isolation and thus offers a different user interface, control flow, and functionality.” (Anthony 2009) Nur wenige größere Korpora sind aus sich heraus auf sprachvergleichende Analysen angelegt. Zu den von Vorneherein als „Comparable Corpus“ konzipierten Ausnahmen gehören das International Comparable Corpus (ICC, Čermáková et al. 2021) oder das mehrsprachige GeWiss-Korpus (Fandrych et al. 2017). Ein alternativer oder ergänzender Ansatz sind virtuelle vergleichbare Korpora wie EuReCo (Trawínski & Kupietz 2021), die über Föderation verteilter, auf vergleichbarer technischer Basis stehender Korpora und Korpusplattformen ermöglichen, das Deutsche kontrastiv mit anderen europäischen Sprachen in Beziehung zu setzen; EuReCo ist allerdings auf schriftsprachliche Daten beschränkt. Unser Beitrag stellt die ZuMult-Plattform (Fandrych et al. 2023) und deren Potential, Vergleichbares für mündliche Korpora zu leisten, vor. ZuMult ist eine offene, flexible, auf etablierten Standards und Technologien basierende Architektur für den Zugang zu audiovisuellen Korpora. Neben Korpusrecherchen mit CQP unterstützt ZuMult die für eine Analyse gesprochener Sprache notwendige erweiterte Kontextualisierung von Suchergebnissen (Frick & Schmidt 2025) sowie interaktive Transkriptanalysen (Schmidt et al. 2023), die insbesondere für qualitativ orientierte Analysen, z.B. in der Gesprächsforschung, und für didaktische Anwendungen, z.B. in der DaF/DaZ-Lehre, genutzt werden. An der Universität Duisburg-Essen erfolgt eine Erweiterung, die darauf zielt, Sprache auch in ihrem multimodalen Zusammenspiel mit weiteren körperlichen Ressourcen – wie Blick, Gestik etc. – für korpuslinguistische Herangehensweisen zu erschließen. Über eine ZuMult-Instanz am Archiv für Gesprochenes Deutsch sind bereits seit 2021 neben dem GeWiss-Korpus zwei der wichtigsten Referenzkorpora des gesprochenen Deutsch – FOLK (Deppermann & Hartung 2011, Schmidt 2016, Reineke et al. 2023) und Deutsch Heute (Kleiner 2015) – zugänglich. Mit der Veröffentlichung einer ZuMult-Instanz für das französische ESLO-Korpus (Abouda & Baude 2006, Baude & Dugua 2011, Eshkol-Taravella 2012, Schmidt 2025) ergeben sich nun erste Möglichkeiten für sprachvergleichende Analysen zwischen dem Deutschen und dem Französischen. Wie unser Beitrag anhand von Proof-Of-Concept-Implementierungen zeigen wird, sind beispielsweise auch das Griffith Corpus of Spoken Australian English (Haugh & Chang 2013), das TIGR Corpus des gesprochenen Italienisch (Miecznikowski-Fuenfschilling et al. i.V.), das Training Corpus of Spoken Slovenian (Verdonik 2024) und die Kollektionen aus Oral History Digital (Pagenstecher 2024) in ZuMult integrierbar. Gleiches gilt für das zwölfsprachige EXMARaLDA-Demokorpus, das derzeit im Rahmen eines Text+-Kooperationsprojekt für eine Publikation in einer ZuMult-Instanz an der Universität Hamburg aufbereitet wird. Auf ähnliche Weise eröffnet ZuMult neue Möglichkeiten der „vergleichenden Sprachinselforschung“ (Boas 2016), also der (korpusgestützten) Analyse von Sprachkontaktphänomenen und Sprachentwicklung von Deutsch als Minderheitensprache im Kontakt mit dominanten anderen (i.d.R. europäischen) Sprachen. Seit Dezember 2024 macht eine ZuMult-Instanz an der University of Texas in Austin (Boas et al. 2025) die Daten des Texas German Dialect Projects auf ähnliche Weise verfügbar, wie die IDS-Instanz Zugriff etwa auf Daten zum Australiendeutsch (Clyne 1981), Deutsch in Namibia (Zimmer et al. 2020) oder Mennonitendeutsch in Amerika (Kaufmann et al. 2023) bietet. Unser Beitrag wird diese verschiedenen Anwendungen vorstellen und illustrieren, sowie einige methodische Herausforderungen der Mehrsprachigkeit (z.B. sprachübergreifendes POS-Tagging) und technische Ansätze zur Aggregation von Suchanfragen an mehrere ZuMult-Instanzen (CLARIN Federated Content Search) thematisieren. Keywords: gesprochene Sprache; Ko","author":[{"family":"Schmidt","given":"Thomas"},{"family":"Abouda","given":"Lotfi"},{"family":"Badin","given":"Flora"},{"family":"Boas","given":"Hans"},{"family":"Blevins","given":"Margaret"},{"family":"Bührig","given":"Kristin"},{"family":"Dugua","given":"Céline"},{"family":"Fandrych","given":"Christian"},{"family":"Ferger","given":"Anne"},{"family":"Frick","given":"Elena"},{"family":"Kompiel","given":"Peter"},{"family":"Miecznikowski-Fuenfschilling","given":"Johanna"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18983100","URL":"https://doi.org/10.5281/zenodo.18983100","source":"datacite"},{"id":"doi:10.5281/zenodo.18507990","type":"article-journal","title":"Efficient Waste Classification in Recycling Systems Using Contrast and Attention Enhanced Deep Learning","abstract":"This repository contains the source code, trained models, and implementation details corresponding to the study “Efficient Waste Classification in Recycling Systems Using Contrast and Attention Enhanced Deep Learning.” It includes the full implementation of the proposed Class-Adaptive Color-Sensitive (CASC) filtering algorithm, the Dense Class-Adaptive Color-Sensitive DenseNet-121 (DenseCSENet-121) architecture, preprocessing scripts, and model evaluation routines. The materials allow reproduction of the reported experimental results on the nine-category waste classification dataset described in the paper.","author":[{"family":"Shyamala Devi","given":"M"},{"family":"Natarajan","given":"Yuvaraj"},{"family":"K R","given":"Sri"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18507990","URL":"https://doi.org/10.5281/zenodo.18507990","source":"datacite"},{"id":"doi:10.5281/zenodo.18507991","type":"article-journal","title":"Efficient Waste Classification in Recycling Systems Using Contrast and Attention Enhanced Deep Learning","abstract":"This repository contains the source code, trained models, and implementation details corresponding to the study “Efficient Waste Classification in Recycling Systems Using Contrast and Attention Enhanced Deep Learning.” It includes the full implementation of the proposed Class-Adaptive Color-Sensitive (CASC) filtering algorithm, the Dense Class-Adaptive Color-Sensitive DenseNet-121 (DenseCSENet-121) architecture, preprocessing scripts, and model evaluation routines. The materials allow reproduction of the reported experimental results on the nine-category waste classification dataset described in the paper.","author":[{"family":"Shyamala Devi","given":"M"},{"family":"Natarajan","given":"Yuvaraj"},{"family":"K R","given":"Sri"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.18507991","URL":"https://doi.org/10.5281/zenodo.18507991","source":"datacite"},{"id":"doi:10.5281/zenodo.17926787","type":"article-journal","title":"Privacy-Preserving Machine Learning: Techniques, Frameworks, and Future Directions","abstract":"Machine Learning that Preserves Privacy (PPML) facilitates model training and analysis while safeguarding sensitive data, model parameters, and user privacy. This survey reviews advances from 2019–2024 with a focus on four major techniques: Homomorphic Encryption (HE), Differential Privacy (DP), Secure Multi-Party Computation (MPC), and Federated Analytics (FA) with secure aggregation. Additionally, it offers a unified taxonomy connecting these methods to adversary models, deployment patterns, and practical applications. Drawing on benchmark outcomes and system evaluations from 2019–2024, this survey assesses PPML systems regarding efficiency, accuracy, deployment costs, and ROI. It emphasizes where each approach excels—such as HE for encrypted inference, DP for secure model release, MPC for collaborative training across silos, and FA for extensive client-side analytics—and details critical engineering trade-offs in industries like healthcare, finance, telecommunications, and IoT.The paper also proposes a practical research roadmap emphasizing hybrid pipelines that combine cryptographic methods with DP, hardware–software co-design to accelerate HE/MPC, and standardized benchmarks for privacy–utility–cost evaluation. Additionally, it stresses the need for operational auditing, explainability. The survey subsequently dives into the latest PPML trends like the merger of hardware-bound TEEs with cryptographic protocols, which would give a dual advantage of higher performance and security. It mentions the use of federated learning in edge and IoT devices, which is increasing but poses unique challenges due to limited computing power and unstable connectivity. Through the analysis of the practical installations, the survey brings out the major causes of communication overload, non-scalable systems, and the risk of losing privacy, which, being articulated in the form of guidelines, help to conquer the mentioned issues and the like in the design of PPML pipelines in the different environments of heterogeneous resources. The paper, lastly, insists upon the role of ethical and regulatory considerations in the acceptance of PPML. Organizations are required to align their technical solutions with the legal requirements as data privacy regulations like GDPR, HIPAA, and CCPA are the main determinants of the handling of sensitive information. Privacy impact assessments are suggested by the survey to be detailed, behavior modeling to be constantly checked, and privacy-improving activities to be disclosed in a public way so that the stakeholders' trust can be earned. It is through the collaboration of the technical thoroughness and ethical supervision that the PPML will pave the way for the secure and responsible use of AI in different sectors.","author":[{"family":"Fatima","given":"Ruksar"},{"family":"Siddiqua","given":"Ayesha"},{"family":"Mahvash","given":"Aliza"},{"family":"Sheeba","given":"Syeda"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17926787","URL":"https://doi.org/10.5281/zenodo.17926787","source":"datacite"},{"id":"doi:10.5281/zenodo.17926788","type":"article-journal","title":"Privacy-Preserving Machine Learning: Techniques, Frameworks, and Future Directions","abstract":"Machine Learning that Preserves Privacy (PPML) facilitates model training and analysis while safeguarding sensitive data, model parameters, and user privacy. This survey reviews advances from 2019–2024 with a focus on four major techniques: Homomorphic Encryption (HE), Differential Privacy (DP), Secure Multi-Party Computation (MPC), and Federated Analytics (FA) with secure aggregation. Additionally, it offers a unified taxonomy connecting these methods to adversary models, deployment patterns, and practical applications. Drawing on benchmark outcomes and system evaluations from 2019–2024, this survey assesses PPML systems regarding efficiency, accuracy, deployment costs, and ROI. It emphasizes where each approach excels—such as HE for encrypted inference, DP for secure model release, MPC for collaborative training across silos, and FA for extensive client-side analytics—and details critical engineering trade-offs in industries like healthcare, finance, telecommunications, and IoT.The paper also proposes a practical research roadmap emphasizing hybrid pipelines that combine cryptographic methods with DP, hardware–software co-design to accelerate HE/MPC, and standardized benchmarks for privacy–utility–cost evaluation. Additionally, it stresses the need for operational auditing, explainability. The survey subsequently dives into the latest PPML trends like the merger of hardware-bound TEEs with cryptographic protocols, which would give a dual advantage of higher performance and security. It mentions the use of federated learning in edge and IoT devices, which is increasing but poses unique challenges due to limited computing power and unstable connectivity. Through the analysis of the practical installations, the survey brings out the major causes of communication overload, non-scalable systems, and the risk of losing privacy, which, being articulated in the form of guidelines, help to conquer the mentioned issues and the like in the design of PPML pipelines in the different environments of heterogeneous resources. The paper, lastly, insists upon the role of ethical and regulatory considerations in the acceptance of PPML. Organizations are required to align their technical solutions with the legal requirements as data privacy regulations like GDPR, HIPAA, and CCPA are the main determinants of the handling of sensitive information. Privacy impact assessments are suggested by the survey to be detailed, behavior modeling to be constantly checked, and privacy-improving activities to be disclosed in a public way so that the stakeholders' trust can be earned. It is through the collaboration of the technical thoroughness and ethical supervision that the PPML will pave the way for the secure and responsible use of AI in different sectors.","author":[{"family":"Fatima","given":"Ruksar"},{"family":"Siddiqua","given":"Ayesha"},{"family":"Mahvash","given":"Aliza"},{"family":"Sheeba","given":"Syeda"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17926788","URL":"https://doi.org/10.5281/zenodo.17926788","source":"datacite"},{"id":"doi:10.5281/zenodo.17926849","type":"article-journal","title":"Resource Management in the Edge–Cloud Continuum: Trends, Algorithms, and Open Challenges","abstract":"The continuum of edge and cloud computing has emerged as a vital computing model for enabling latency-sensitive, data-heavy, and geographically scattered applications. As billions of devices generate massive volumes of data, efficient resource management across heterogeneous, distributed infrastructures has become essential. This study presents a systematic review of 68 research articles published between 2019 and 2024 that address resource distribution, task delegation, scheduling, orchestration, and optimization within the edge–cloud continuum. The paper highlights emerging themes such as AI-driven orchestration, multi-agent reinforcement learning, federated optimization, and serverless edge computing. We evaluate the performance, precision, scalability, and flexibility of traditional heuristics, mathematical models, and RL techniques under varying workloads. Although these significant advancements have been made, several open problems still exist—mobility-aware scheduling, cross-layer security integration with over-the-air encrypted computation results, and the absence of general ML models and benchmarks along with large-scale real-world deployment. The paper also ends by emphasizing the future research directions required to create and implement intelligent, autonomous, and scalable resource management frameworks that are designed for 6G/enhanced mobile broadband (eMBB), IoT/operating on devices over Bluetooth, autonomous systems/enabled by local cloudlets, and immersive applications/such as immersive gaming.","author":[{"family":"Fatima","given":"Ruksar"},{"family":"Anjum","given":"Suhana"},{"family":"Junaidi","given":"Shaista"},{"family":"Rafa","given":"Ruqayya"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17926849","URL":"https://doi.org/10.5281/zenodo.17926849","source":"datacite"},{"id":"doi:10.5281/zenodo.17926850","type":"article-journal","title":"Resource Management in the Edge–Cloud Continuum: Trends, Algorithms, and Open Challenges","abstract":"The continuum of edge and cloud computing has emerged as a vital computing model for enabling latency-sensitive, data-heavy, and geographically scattered applications. As billions of devices generate massive volumes of data, efficient resource management across heterogeneous, distributed infrastructures has become essential. This study presents a systematic review of 68 research articles published between 2019 and 2024 that address resource distribution, task delegation, scheduling, orchestration, and optimization within the edge–cloud continuum. The paper highlights emerging themes such as AI-driven orchestration, multi-agent reinforcement learning, federated optimization, and serverless edge computing. We evaluate the performance, precision, scalability, and flexibility of traditional heuristics, mathematical models, and RL techniques under varying workloads. Although these significant advancements have been made, several open problems still exist—mobility-aware scheduling, cross-layer security integration with over-the-air encrypted computation results, and the absence of general ML models and benchmarks along with large-scale real-world deployment. The paper also ends by emphasizing the future research directions required to create and implement intelligent, autonomous, and scalable resource management frameworks that are designed for 6G/enhanced mobile broadband (eMBB), IoT/operating on devices over Bluetooth, autonomous systems/enabled by local cloudlets, and immersive applications/such as immersive gaming.","author":[{"family":"Fatima","given":"Ruksar"},{"family":"Anjum","given":"Suhana"},{"family":"Junaidi","given":"Shaista"},{"family":"Rafa","given":"Ruqayya"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17926850","URL":"https://doi.org/10.5281/zenodo.17926850","source":"datacite"},{"id":"doi:10.5281/zenodo.17926595","type":"article-journal","title":"Federated Learning: Advances, Privacy Mechanisms, and Real-World Deployments","abstract":"Federated Learning (FL) has emerged as a powerful paradigm that enables collaborative model training across decentralized clients while preserving data privacy. Instead of aggregating sensitive data in a central server, FL coordinates local training on distributed devices—ranging from smartphones to IoT sensors and institutional servers—and collects only model updates. This design addresses major privacy, ethical, and security concerns associated with centralized data storage. Between 2019 and 2024, extensive research has focused on core FL challenges such as non-IID data distributions, device and system heterogeneity, resource limitations, privacy risks arising from gradient leakage, and practical deployment barriers in fields like healthcare and edge IoT. In this paper, we review recent advances across four themes: (1) distinctions and best practices for cross-device vs. cross-silo FL; (2) privacy-preserving mechanisms, including differential privacy and secure aggregation; (3) communication- and model-compression techniques for reducing bandwidth usage; and (4) real-world deployments in healthcare and edge-IoT environments. We analyze these works based on efficiency, accuracy, privacy trade-offs, and deployment-level considerations such as resource savings and regulatory alignment. Our synthesis shows that modern compression techniques—such as quantization, sparsification, and knowledge distillation—can significantly reduce communication costs with minimal accuracy loss, making FL feasible for resource-constrained devices. Privacy mechanisms remain essential for sensitive domains, though they commonly introduce accuracy and utility trade-offs. Cross-silo deployments demonstrate performance close to centralized baselines while maintaining data locality, yet full-scale adoption still depends on standardization, infrastructure readiness, and clearer ROI evidence. We highlight open gaps such as limited convergence theory for private and compressed FL under non-IID data, lack of unified benchmarks, insufficient empirical ROI studies, and the challenges of scaling FL to large modern models. To support clarity, we provide comparative tables, research-gap matrices, and conceptual diagrams illustrating accuracy-efficiency-privacy tradeoffs, followed by prioritized directions for future research.","author":[{"family":"Fatima","given":"Ruksar"},{"family":"Mahvash","given":"Aliza"},{"family":"Siddiqua","given":"Ayesha"},{"family":"Sheeba","given":"Syeda"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17926595","URL":"https://doi.org/10.5281/zenodo.17926595","source":"datacite"},{"id":"doi:10.5281/zenodo.17926594","type":"article-journal","title":"Federated Learning: Advances, Privacy Mechanisms, and Real-World Deployments","abstract":"Federated Learning (FL) has emerged as a powerful paradigm that enables collaborative model training across decentralized clients while preserving data privacy. Instead of aggregating sensitive data in a central server, FL coordinates local training on distributed devices—ranging from smartphones to IoT sensors and institutional servers—and collects only model updates. This design addresses major privacy, ethical, and security concerns associated with centralized data storage. Between 2019 and 2024, extensive research has focused on core FL challenges such as non-IID data distributions, device and system heterogeneity, resource limitations, privacy risks arising from gradient leakage, and practical deployment barriers in fields like healthcare and edge IoT. In this paper, we review recent advances across four themes: (1) distinctions and best practices for cross-device vs. cross-silo FL; (2) privacy-preserving mechanisms, including differential privacy and secure aggregation; (3) communication- and model-compression techniques for reducing bandwidth usage; and (4) real-world deployments in healthcare and edge-IoT environments. We analyze these works based on efficiency, accuracy, privacy trade-offs, and deployment-level considerations such as resource savings and regulatory alignment. Our synthesis shows that modern compression techniques—such as quantization, sparsification, and knowledge distillation—can significantly reduce communication costs with minimal accuracy loss, making FL feasible for resource-constrained devices. Privacy mechanisms remain essential for sensitive domains, though they commonly introduce accuracy and utility trade-offs. Cross-silo deployments demonstrate performance close to centralized baselines while maintaining data locality, yet full-scale adoption still depends on standardization, infrastructure readiness, and clearer ROI evidence. We highlight open gaps such as limited convergence theory for private and compressed FL under non-IID data, lack of unified benchmarks, insufficient empirical ROI studies, and the challenges of scaling FL to large modern models. To support clarity, we provide comparative tables, research-gap matrices, and conceptual diagrams illustrating accuracy-efficiency-privacy tradeoffs, followed by prioritized directions for future research.","author":[{"family":"Fatima","given":"Ruksar"},{"family":"Mahvash","given":"Aliza"},{"family":"Siddiqua","given":"Ayesha"},{"family":"Sheeba","given":"Syeda"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17926594","URL":"https://doi.org/10.5281/zenodo.17926594","source":"datacite"},{"id":"doi:10.5281/zenodo.17549886","type":"article-journal","title":"D3.2 Competence Centres concepts and activities (pre)existing in EOSC","abstract":"This deliverable presents an overview of EOSC-related activities and projects that could be taken into account for the design and implementation of Competence Centres (CCs), positioned as key instruments to support data-intensive, FAIR-compliant, and interdisciplinary research within the European Open Science Cloud (EOSC). It synthesises existing practices, conceptual frameworks, and policy recommendations drawn from ongoing and past projects. CCs are understood by most of the research communities as decentralised, composable structures that may consolidate community expertise, support training and guidance for data sharing and reuse or provide embedded services across diverse research contexts. The deliverable outlines the different types of contributions of the domain-specific clusters. Each science cluster intends to align its CC strategies on either thematic priorities, governance approaches, training assets etc, and reflect on how they then could align within the OSCARS CC design and definition proposed in the framework of OSCARS WP1 (Bodera Sempere et al., 2024). The result of this landscaping highlights existing or in development principles, acknowledges heterogeneous implementations, foster cross-community learning and will lay the groundwork for a future inter-OSCARs project and inter-community paper on all kind of Competence Centres that can act in the framework of EOSC (discipline specific or thematic, local, regional, national…). The document identifies key interdisciplinary challenges such as multimodal data integration and large-scale metadata analysis emphasising the need for cultural change, capacity building, and embedded support mechanisms close to research practice, Challenges identified demand robust infrastructures, sustained collaboration, and the realisation of the “FAIR web of data,” a central EOSC ambition. The OSCARS CC model builds upon these insights to propose a federated and scalable ecosystem of competence. The models offer a practical roadmap to foster uptake, interoperability, and sustainability of Open Science across European research communities.","author":[{"family":"David","given":"Romain"},{"family":"Hienola","given":"Anca"},{"family":"Schmidt-Tremmel","given":"Friederike"},{"family":"Van Der Lek","given":"Iulianna"},{"family":"Nentwich","given":"Melanie"},{"family":"Bodera Sempere","given":"Jordi"},{"family":"Kalaitzi","given":"Vasso"},{"family":"Guerrieri","given":"Giovanni"},{"family":"Wolff-Boenisch","given":"Bonnie"},{"family":"Guezennec","given":"Cécile"},{"family":"Draščić","given":"Martina"},{"family":"Vipavc Brvar","given":"Irena"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17549886","URL":"https://doi.org/10.5281/zenodo.17549886","source":"datacite"},{"id":"doi:10.5281/zenodo.17549887","type":"article-journal","title":"D3.2 Competence Centres concepts and activities (pre)existing in EOSC","abstract":"This deliverable presents an overview of EOSC-related activities and projects that could be taken into account for the design and implementation of Competence Centres (CCs), positioned as key instruments to support data-intensive, FAIR-compliant, and interdisciplinary research within the European Open Science Cloud (EOSC). It synthesises existing practices, conceptual frameworks, and policy recommendations drawn from ongoing and past projects. CCs are understood by most of the research communities as decentralised, composable structures that may consolidate community expertise, support training and guidance for data sharing and reuse or provide embedded services across diverse research contexts. The deliverable outlines the different types of contributions of the domain-specific clusters. Each science cluster intends to align its CC strategies on either thematic priorities, governance approaches, training assets etc, and reflect on how they then could align within the OSCARS CC design and definition proposed in the framework of OSCARS WP1 (Bodera Sempere et al., 2024). The result of this landscaping highlights existing or in development principles, acknowledges heterogeneous implementations, foster cross-community learning and will lay the groundwork for a future inter-OSCARs project and inter-community paper on all kind of Competence Centres that can act in the framework of EOSC (discipline specific or thematic, local, regional, national…). The document identifies key interdisciplinary challenges such as multimodal data integration and large-scale metadata analysis emphasising the need for cultural change, capacity building, and embedded support mechanisms close to research practice, Challenges identified demand robust infrastructures, sustained collaboration, and the realisation of the “FAIR web of data,” a central EOSC ambition. The OSCARS CC model builds upon these insights to propose a federated and scalable ecosystem of competence. The models offer a practical roadmap to foster uptake, interoperability, and sustainability of Open Science across European research communities.","author":[{"family":"David","given":"Romain"},{"family":"Hienola","given":"Anca"},{"family":"Schmidt-Tremmel","given":"Friederike"},{"family":"Van Der Lek","given":"Iulianna"},{"family":"Nentwich","given":"Melanie"},{"family":"Bodera Sempere","given":"Jordi"},{"family":"Kalaitzi","given":"Vasso"},{"family":"Guerrieri","given":"Giovanni"},{"family":"Wolff-Boenisch","given":"Bonnie"},{"family":"Guezennec","given":"Cécile"},{"family":"Draščić","given":"Martina"},{"family":"Vipavc Brvar","given":"Irena"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17549887","URL":"https://doi.org/10.5281/zenodo.17549887","source":"datacite"},{"id":"doi:10.5281/zenodo.17670445","type":"article-journal","title":"Trusted Research Environments for Healthcare AI: State of the Art Global Landscape Report","abstract":"This global landscape report provides a comprehensive overview of the current state of Trusted Research Environments (TREs) worldwide, with a particular focus on the Nordic countries and the United Kingdom. TREs—also known as data safe havens—are secure platforms that enable researchers to analyse sensitive data while complying with legal, ethical, and contractual obligations. The report is part of the TRE4HealthAI project, which aims to develop next-generation TRE capabilities tailored for healthcare AI applications. A key finding of the report is that the DARE UK Five Safes framework and the Standard Architecture for Trusted Research Environments (SATRE) have emerged as the de facto standard for TRE implementations in both the UK and across Europe. Originally developed within the UK legal and research context, the Five Safes framework—comprising Safe People, Safe Projects, Safe Settings, Safe Data, and Safe Outputs—has been widely adopted and adapted by TREs internationally. The DARE UK Blueprint, which operationalises this framework, is technology-agnostic and aligns with FAIR principles (Findable, Accessible, Interoperable, Reusable), making it suitable for diverse implementations. In the UK, various research and commercial projects such as the Turing Data Safe Haven, Treehouse, and the Aridhia digital research environment have all been evaluated for conformity with the Five Safes-based SATRE specification. These implementations demonstrate a high degree of standardisation, transparency, and interoperability, setting a benchmark for TRE development. The report also highlights how European initiatives, particularly the EOSC-ENTRUST project, have mapped their TRE architectures and requirements to the DARE UK Blueprint. TREs in Norway (NORTRE) and Finland (CSC SD) have been shown to align closely with the DARE UK model, despite differences in governance structures—such as the role of data controllers and output approval processes. The EOSC-ENTRUST project has adopted the DARE UK infrastructure layer and extended it to accommodate European legal and organisational contexts. Furthermore, the report outlines the growing demand for next-generation TRE capabilities, including support for AI model development, federated learning, and scalable cloud infrastructure. These capabilities are being mapped onto the Five Safes framework to ensure continued compliance and security and enable new capabilities for supervisory authorities to learn from the evidence generated in a TRE. Projects like GRAIMATTER and various empirical studies (eg, Kavianpour et al. 2022) provide guidance on integrating AI safely within TREs. In conclusion, the DARE UK Five Safes framework has become the foundational model for TREs, influencing both national and international implementations. Its adaptability, combined with a strong emphasis on governance and interoperability, positions it as the cornerstone for future TRE development in support of secure, ethical, and innovative research that also enhances evidence-based regulatory learnings. This work is funded as part of the Advanced Digitalisation programme by Vinnova, the Swedish Innovation Agency, project reference 2024-01412.","author":[{"family":"Emanuilov","given":"Ivo"},{"family":"Larsson","given":"Björn"},{"family":"Dubber","given":"Andrew"},{"family":"Magas","given":"Michela"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17670445","URL":"https://doi.org/10.5281/zenodo.17670445","source":"datacite"},{"id":"doi:10.5281/zenodo.17670446","type":"article-journal","title":"Trusted Research Environments for Healthcare AI: State of the Art Global Landscape Report","abstract":"This global landscape report provides a comprehensive overview of the current state of Trusted Research Environments (TREs) worldwide, with a particular focus on the Nordic countries and the United Kingdom. TREs—also known as data safe havens—are secure platforms that enable researchers to analyse sensitive data while complying with legal, ethical, and contractual obligations. The report is part of the TRE4HealthAI project, which aims to develop next-generation TRE capabilities tailored for healthcare AI applications. A key finding of the report is that the DARE UK Five Safes framework and the Standard Architecture for Trusted Research Environments (SATRE) have emerged as the de facto standard for TRE implementations in both the UK and across Europe. Originally developed within the UK legal and research context, the Five Safes framework—comprising Safe People, Safe Projects, Safe Settings, Safe Data, and Safe Outputs—has been widely adopted and adapted by TREs internationally. The DARE UK Blueprint, which operationalises this framework, is technology-agnostic and aligns with FAIR principles (Findable, Accessible, Interoperable, Reusable), making it suitable for diverse implementations. In the UK, various research and commercial projects such as the Turing Data Safe Haven, Treehouse, and the Aridhia digital research environment have all been evaluated for conformity with the Five Safes-based SATRE specification. These implementations demonstrate a high degree of standardisation, transparency, and interoperability, setting a benchmark for TRE development. The report also highlights how European initiatives, particularly the EOSC-ENTRUST project, have mapped their TRE architectures and requirements to the DARE UK Blueprint. TREs in Norway (NORTRE) and Finland (CSC SD) have been shown to align closely with the DARE UK model, despite differences in governance structures—such as the role of data controllers and output approval processes. The EOSC-ENTRUST project has adopted the DARE UK infrastructure layer and extended it to accommodate European legal and organisational contexts. Furthermore, the report outlines the growing demand for next-generation TRE capabilities, including support for AI model development, federated learning, and scalable cloud infrastructure. These capabilities are being mapped onto the Five Safes framework to ensure continued compliance and security and enable new capabilities for supervisory authorities to learn from the evidence generated in a TRE. Projects like GRAIMATTER and various empirical studies (eg, Kavianpour et al. 2022) provide guidance on integrating AI safely within TREs. In conclusion, the DARE UK Five Safes framework has become the foundational model for TREs, influencing both national and international implementations. Its adaptability, combined with a strong emphasis on governance and interoperability, positions it as the cornerstone for future TRE development in support of secure, ethical, and innovative research that also enhances evidence-based regulatory learnings. This work is funded as part of the Advanced Digitalisation programme by Vinnova, the Swedish Innovation Agency, project reference 2024-01412.","author":[{"family":"Emanuilov","given":"Ivo"},{"family":"Larsson","given":"Björn"},{"family":"Dubber","given":"Andrew"},{"family":"Magas","given":"Michela"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5281/zenodo.17670446","URL":"https://doi.org/10.5281/zenodo.17670446","source":"datacite"},{"id":"doi:10.3217/n659p-tr557","type":"article-journal","title":"Metacampus as a Base for a Unite! Blended Intensive Programme (BIP): Competencies for Collaborative Teaching in Joint Programme (Report 04/2025)","abstract":"The report “Metacampus as a Base for a Unite! Blended Intensive Programme (BIP): Competencies for Collaborative Teaching in Joint Programmes” describes the Erasmus + Blended Intensive Program (BIP) held from June to October 2024 under the Unite! University Network. The six‑week English‑language course aimed to build skills for collaborative teaching in joint programmes among academic staff, lecturers and researchers. The programme combined an online introductory module (June 17 2024), a five‑day on‑site week at KTH Royal Institute of Technology (August 26‑30 2024), and a final online presentation (October 9 2024). Unite!'s federated learning management system (LMS) Metacampus acted as the central LMS, offering secure, institution‑based access to materials, while public channels (KTH website, Unite! portal) provided open information such as schedules. Workshops covered case studies of joint programmes, digital tools for collaboration, administrative challenges (double‑degrees, European Degree), large‑scale teaching strategies, and multicultural learning. Experts from the Unite! community led the sessions. Feedback highlighted the value of the intensive format, the need for hands‑on practice, the complexity of administrative procedures, and the importance of informal networking. Key lessons stress a hybrid digital environment that blends secure Metacampus collaboration with open public communication, and the necessity of a dedicated digital liaison to coordinate local and central IT tools.Recommendations for future BIPs include adopting this hybrid model, clarifying integration of local tools with Metacampus, and assigning a central facilitator to streamline technical and administrative support. The series “Unite! Digital Teaching and Learning Success Story Report” was initiated by Cm.2 Digital Campus, led by TU Graz (Martin Ebner) as part of the idea to collect, spread and enhance the possibilities, usages and lessons learned of teaching and learning offers using the federated learning management system Metacampus in 11/2024 till the end of the current Erasmus+ funding period 10/2026. All reports are available under open license and originally published at the TU Graz repository. Copyright holders are the authors.","author":[{"family":"Acosta-Garcia","given":"Marcela"},{"family":"Galante","given":"Lorenzo"},{"family":"Kauppinen","given":"Tomi"},{"family":"Keller","given":"Elizabeth"},{"family":"Knutsson","given":"Karin"},{"family":"Pears","given":"Arnold"},{"family":"Tucker Smith","given":"Madeleine"}],"issued":{"date-parts":[[2025]]},"DOI":"10.3217/n659p-tr557","URL":"https://doi.org/10.3217/n659p-tr557","source":"datacite"},{"id":"doi:10.36227/techrxiv.174495330.08787592/v1","type":"article-journal","title":"Federated Learning: Recent Advances and Future Directions","abstract":"This review examines federated learning as an innovative approach enabling distributed machine learning while preserving data privacy. We explore recent developments addressing four core challenges: data heterogeneity across participants, optimizing communication between devices, strengthening privacy protections, and accommodating diverse system capabilities. Our analysis identifies substantial algorithmic progress while highlighting persistent gaps in practical implementations across varied devices, balancing privacy with model performance, and expanding applications in vertical federated learning contexts. The survey concludes with strategic recommendations for overcoming these limitations and accelerating real-world adoption of federated learning techniques, particularly focusing on scalability across heterogeneous environments, enhanced privacy-utility balance, and improved cross-organizational implementations that could substantially expand federated learning's practical impact.","author":[{"family":"Indrani","given":"Lakshmi"},{"family":"Gadiraju","given":"Deepika"},{"family":"Baligodugula","given":"Vishnu"}],"issued":{"date-parts":[[2025]]},"DOI":"10.36227/techrxiv.174495330.08787592/v1","URL":"https://doi.org/10.36227/techrxiv.174495330.08787592/v1","source":"crossref"},{"id":"doi:10.2174/9789815322224125030006","type":"article-journal","title":"Breaking the Centralization Barrier: Exploring Decentralized Federated Learning for Vehicle Number Plate Recognition in IoV","abstract":"The development of effective and safe machine learning systems for vehicle number plate recognition (VNPR) is now necessary due to the emergence of the Internet of Vehicles (IoV). However, traditional centralised techniques run into issues with data privacy, communication overhead, and centralised data access restrictions. This study explores the possibilities of decentralized federated learning for VNPR in the IoV to solve these constraints. Decentralised federated learning, which overcomes the centralization barrier, allows local model training on the edge devices of participating cars, protecting data privacy and cutting down on communication overhead. The ramifications of this paradigm change are examined in this research, including improved data privacy and security, shared intelligence, and resilience against errors and assaults. It also looks at the trade-off between performance and decentralisation while emphasising the balance attained via improved model aggregation and resource use. Additionally covered is the difficulty of consensus algorithms and blockchain-based networks, highlighting the need for further investigation and development. Decentralised federated learning has been identified as a possible strategy for overcoming the centralization barrier in VNPR systems, opening the door to the implementation of efficient, secure, and private machine learning in the IoV.","author":[{"family":"Panwar","given":"Arvind"},{"family":"Gaba","given":"Priyanka"},{"family":"Sugandh","given":"Urvashi"},{"family":"Bohra","given":"Navdeep"}],"issued":{"date-parts":[[2025]]},"DOI":"10.2174/9789815322224125030006","URL":"https://doi.org/10.2174/9789815322224125030006","source":"crossref"},{"id":"doi:10.1093/database/baaf016","type":"article-journal","title":"A comprehensive experimental comparison between federated and centralized learning","abstract":"Abstract Federated learning is an upcoming machine learning paradigm which allows data from multiple sources to be used for training of classifiers without the data leaving the source it originally resides. This can be highly valuable for use cases such as medical research, where gathering data at a central location can be quite complicated due to privacy and legal concerns of the data. In such cases, federated learning has the potential to vastly speed up the research cycle. Although federated and central learning have been compared from a theoretical perspective, an extensive experimental comparison of performances and learning behavior still lacks. We have performed a comprehensive experimental comparison between federated and centralized learning. We evaluated various classifiers on various datasets exploring influences of different sample distributions as well as different class distributions across the clients. The results show similar performances under a wide variety of settings between the federated and central learning strategies. Federated learning is able to deal with various imbalances in the data distributions. It is sensitive to batch effects between different datasets when they coincide with location, similar to central learning, but this setting might go unobserved more easily. Federated learning seems to be robust to various challenges such as skewed data distributions, high data dimensionality, multiclass problems, and complex models. Taken together, the insights from our comparison gives much promise for applying federated learning as an alternative to sharing data. Code for reproducing the results in this work can be found at: https://github.com/swiergarst/FLComparison","author":[{"family":"Garst","given":"Swier"},{"family":"Dekker","given":"Julian"},{"family":"Reinders","given":"Marcel"}],"issued":{"date-parts":[[2025]]},"DOI":"10.1093/database/baaf016","URL":"https://doi.org/10.1093/database/baaf016","source":"crossref"},{"id":"doi:10.21032/jhis.2025.50.1.31","type":"article-journal","title":"Comparison of Federated Learning and Fair Federated Learning for Pneumonia Patient Classification","abstract":"Objectives: This study aims to compare the performance of federated learning (FL) and fair federated learning (FFL) in classifying pneumonia patients based on chest X-ray data. The primary focus is on assessing the accuracy and fairness of these models in handling imbalanced and distributed data in real-world healthcare settings.Methods: We used a large chest X-ray dataset to evaluate the performance of FL and FFL models. The models were built using the ResNet50 architecture, and experiments were conducted under both independent and identically distributed (IID) and non-IID data conditions. The FFL approach applied optimized loss functions to address data imbalance and ensure fair contribution from each client, regardless of the local data distribution.Results: Our findings indicate that FFL consistently outperforms traditional FL models, particularly in non-IID environments. The FFL model demonstrated higher accuracy in pneumonia classification, achieving a significant improvement in model fairness and performance across different client datasets. The use of the ResNet50 architecture further enhanced the model’s ability to handle complex X-ray image patterns.Conclusions: FFL offers a superior solution for handling imbalanced medical data compared to conventional FL models. Its ability to maintain fairness while improving classification accuracy makes it an ideal approach for decentralized healthcare systems, ensuring better patient outcomes while preserving data privacy.","author":[{"family":"Na","given":"Kyungmin"},{"family":"Kim","given":"Dohyoung"},{"family":"Lee","given":"Youngho"}],"issued":{"date-parts":[[2025]]},"DOI":"10.21032/jhis.2025.50.1.31","URL":"https://doi.org/10.21032/jhis.2025.50.1.31","source":"crossref"},{"id":"doi:10.2174/9789815322224125030003","type":"article-journal","title":"Federated Learning on Wheels: A Decentralized Approach to Privacy-Enhanced Data Collection in Internet of Vehicles","abstract":"Due to privacy issues and the scattered nature of data produced by vehicles, the Internet of Vehicles (IOV) poses considerable hurdles for data collecting. In this chapter, we examine the idea of “Federated Learning on Wheels” (FLoW), which provides a decentralised method for IOV data collection with a focus on privacy. FLOW makes use of the onboard computer resources of cars to carry out model training locally, making sure that private information stays on the cars and is not shared with a centralised server. This strategy overcomes the shortcomings of conventional centralised data collecting approaches while simultaneously protecting user privacy. We examine the fundamentals of federated learning and how they relate to IOV, highlighting the advantages of maintaining privacy. We also look at secure aggregation procedures and confidentiality safeguards as additional methods for privacy-enhanced data acquisition in FLOW. Additionally, we emphasise the significance of accuracy and performance issues in decentralised contexts and use examples that illustrate FLOW's usefulness. We also explore security and trust issues, talking about possible weaknesses and methods to secure the reliability of participants and model updates. We also consider how blockchain technology may be incorporated for improved security and openness. We conclude by discussing FLOW future directions, difficulties, and ethical issues in order to shed light on its possible significance and legal ramifications. Overall, this chapter clarifies the relevance of Federated Learning to Wheels as a ground-breaking approach to data collecting with increased privacy in the Internet of Vehicles.","author":[{"family":"Sharma","given":"Neha"},{"family":"Sugandh","given":"Urvashi"},{"family":"Agarwal","given":"Jyoti"},{"family":"Panwar","given":"Arvind"},{"family":"Gaba","given":"Priyanka"}],"issued":{"date-parts":[[2025]]},"DOI":"10.2174/9789815322224125030003","URL":"https://doi.org/10.2174/9789815322224125030003","source":"crossref"},{"id":"doi:10.36227/techrxiv.175393442.27040019/v1","type":"article-journal","title":"Splitting Smarter: Differential Privacy for Secure Healthcare Federated Learning","abstract":"Split Federated Learning (SplitFed) has emerged as a decentralized method of training ML models that enables multiple healthcare parties to collaboratively models without sharing their raw data. This method is, however, vulnerable against label inference attacks which can compromise patient privacy. In this paper, we have investigated the vulnerability of SplitFed models to label inference attacks in biomedical imaging. In addition, we have proposed a solution that incorporates differential privacy (DP) into SplitFed to protect against label inference attacks. Results indicate the efficacy of the SplitFed model under multiple conditions and found that the label inference accuracy changes from 100% (No-DP) to 0% (with DP). This indicates that integration of DP offers a robust mechanism for protecting patient privacy. Additionally, the usage of Cauchy noise in DP provides the best protection out of all noise categories.","author":[{"family":"Yetunde","given":"Munirat"},{"family":"Shukla","given":"Raj"},{"family":"Das","given":"Tapadhir"}],"issued":{"date-parts":[[2025]]},"DOI":"10.36227/techrxiv.175393442.27040019/v1","URL":"https://doi.org/10.36227/techrxiv.175393442.27040019/v1","source":"crossref"},{"id":"doi:10.5957/tos-2025-019","type":"article-journal","title":"Federated Learning for the Maritime and Offshore Industries: Training Machine Learning Models Without Data Sharing","abstract":"In the fast-evolving field of machine learning, data privacy and security are critical concerns, especially in sensitive and heavily regulated sectors such as offshore energy and maritime operations. Traditional centralized machine learning approaches require consolidating all data in a single location for model training, which can introduce substantial privacy risks and often conflict with regulatory requirements. Federated Learning (FL) allows data to remain decentralized while enabling collaborative and individual model development in industry. FL is particularly valuable for two key use cases. First, multiple companies with a shared interest in developing a generalized model can collaborate through FL while reducing data privacy concerns. Second, a single company with distributed data—whether due to regulatory restrictions or logistical challenges—can leverage FL to train a unified model across its datasets without centralizing data on one server. By keeping data localized, FL enhances privacy, regulatory compliance, and operational flexibility, making it well-suited for the needs of the maritime and offshore industries. Popularized by Google, FL not only lowers privacy concerns but also strengthens model robustness and accuracy by drawing on diverse, distributed data sources. This capability is particularly valuable in offshore and maritime applications, where varied data from different installations, vessels, and sea conditions can inform more reliable predictive insights. In this paper, we discuss the challenges and opportunities of applying Federated Learning to the offshore and maritime industries. We present an overview of different applications, including examples from recent studies and toy cases that demonstrate FL's potential for these industries. We illustrate the value of FL further through the case of a container ship navigating various seas, where we aim to predict the vessel's performance under different conditions. Additionally, we describe a working procedure using an open-source package to implement FL, tested in a real-world setting with participants retaining data locally. A server hosted by a cloud provider was used to coordinate the Federated Learning process. This example demonstrates the strength of Federated Learning in developing generalized and robust models—a crucial need in the offshore and maritime sectors, where diverse and adaptable models are essential to address complex and variable real-world conditions. By maintaining data locally for each participant, FL emerges as a promising approach for developing regulatory-compliant, data-driven solutions for critical maritime applications.","author":[{"family":"Ducamp","given":"Gaspard"},{"family":"Kim","given":"Peter"},{"family":"Ruth","given":"Eivind"}],"issued":{"date-parts":[[2025]]},"DOI":"10.5957/tos-2025-019","URL":"https://doi.org/10.5957/tos-2025-019","source":"crossref"},{"id":"doi:10.22541/au.173695875.55142963/v1","type":"article-journal","title":"PrivacyShepherd: Federated Learning for SNP-Based Sheep Breed Identification","abstract":"This study presents an innovative federated learning framework that addresses the challenge of identifying the breeds of Iranian sheep using an SNP-based genotype dataset which contains the SNP values of four breeds of Iranian sheep. In the first phase of the research, an SNP selection phase is performed using the Particle Swarm Optimization algorithm to find the best subset of SNPs which will result in the best possible classification accuracy. In this phase, PSO detected 5565 SNPs among 46000 which has resulted in 98% classification accuracy. The second phase then uses a federated learning framework with the aggregation algorithm kfedAvg to train different local learning models with different private local datasets. The result achieved from this phase indicates, on average, the accuracy of 85% by the clients on the local test data.","author":[{"family":"Nourmohammdi","given":"Reza"},{"family":"Moradi","given":"Mohammad"},{"family":"Behravan","given":"Iman"}],"issued":{"date-parts":[[2025]]},"DOI":"10.22541/au.173695875.55142963/v1","URL":"https://doi.org/10.22541/au.173695875.55142963/v1","source":"crossref"},{"id":"doi:10.36227/techrxiv.176159757.71660556/v1","type":"article-journal","title":"Energy-Aware Adaptive Federated Learning for IoT Security in 6G","abstract":"Artificial intelligence (AI) and machine learning (ML) are widely adopted in sixth generation (6G) mobile networks. However, the deployment of AI in communication networks will require huge amounts of resources, such as computing, memory, bandwidth, and, as a result, energy. Certain use cases that are associated with resource-constrained devices, for instance, the internet of things (IoT), necessitate designing resource-aware and adaptable AI/ML techniques. In this article, a decentralized energy-aware federated learning (FL) model is proposed for IoT devices that allows the deployment of AI-based cybersecurity operations in 6G. We employ an ordered dropout (OD) mechanism to construct nested submodels from a larger neural network (NN), enabling dynamic adaptation to the energy availability of the system and reducing the overall energy footprint. The experimental evaluations show that the proposed energy-aware model extends the operational lifetime of the deployment framework from 82% to 135% for different datasets, while reducing inference time per sample by up to 50% for the smallest submodel.","author":[{"family":"Rumesh","given":"Yasintha"},{"family":"Porambage","given":"Pawani"},{"family":"Ahmad","given":"Ijaz"}],"issued":{"date-parts":[[2025]]},"DOI":"10.36227/techrxiv.176159757.71660556/v1","URL":"https://doi.org/10.36227/techrxiv.176159757.71660556/v1","source":"crossref"},{"id":"doi:10.5281/zenodo.21701383","type":"article-journal","title":"Hybrid Machine Learning Based Integrated Network  Scanning And Anomaly Framework","abstract":"Modern networks produced enormous volumes of data which made them increasingly vulnerable to advanced cyber threats. Traditional scanning tools and standalone anomaly detection systems fell short in identifying evolving or zero day attacks. This study proposed a hybrid framework that integrated active network scanning with machine learning based anomaly detection to deliver an adaptive and automated solution. The framework combined Nmap based host scanning tcpdump based traffic capture and Zeek based feature extraction with a Random Forest classifier and an autoencoder trained on normal traffic to flag deviations. Recent advancements in supervised and unsupervised anomaly detection together with integrated intrusion detection systems published between 2020 and 2025 were reviewed and synthesized to position the proposed approach within the broader field. Experimental comparison across five learning models showed that the Convolutional Neural Network achieved the highest accuracy of 93 percent followed by Long Short Term Memory at 91 percent while Random Forest balanced accuracy and interpretability at 89 percent. The hybrid correlation mechanism that combined scan derived signals with model predictions reduced false alarms and strengthened real time threat visibility compared with single method systems. The study concluded that combining active scanning with machine learning significantly improved detection accuracy reduced false positives and enabled actionable reporting for security analysts. Future directions identified included adaptive learning methods lightweight models suited for edge devices and secure distributed detection through federated learning.","author":[{"family":"Jain","given":"Sarthak"},{"family":"Suyash"},{"family":"Kumar","given":"Upendra"},{"family":"Chauhan","given":"Utsav"},{"family":"Kumar","given":"Ashish"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21701383","URL":"https://doi.org/10.5281/zenodo.21701383","source":"datacite"},{"id":"doi:10.5281/zenodo.21701384","type":"article-journal","title":"Hybrid Machine Learning Based Integrated Network  Scanning And Anomaly Framework","abstract":"Modern networks produced enormous volumes of data which made them increasingly vulnerable to advanced cyber threats. Traditional scanning tools and standalone anomaly detection systems fell short in identifying evolving or zero day attacks. This study proposed a hybrid framework that integrated active network scanning with machine learning based anomaly detection to deliver an adaptive and automated solution. The framework combined Nmap based host scanning tcpdump based traffic capture and Zeek based feature extraction with a Random Forest classifier and an autoencoder trained on normal traffic to flag deviations. Recent advancements in supervised and unsupervised anomaly detection together with integrated intrusion detection systems published between 2020 and 2025 were reviewed and synthesized to position the proposed approach within the broader field. Experimental comparison across five learning models showed that the Convolutional Neural Network achieved the highest accuracy of 93 percent followed by Long Short Term Memory at 91 percent while Random Forest balanced accuracy and interpretability at 89 percent. The hybrid correlation mechanism that combined scan derived signals with model predictions reduced false alarms and strengthened real time threat visibility compared with single method systems. The study concluded that combining active scanning with machine learning significantly improved detection accuracy reduced false positives and enabled actionable reporting for security analysts. Future directions identified included adaptive learning methods lightweight models suited for edge devices and secure distributed detection through federated learning.","author":[{"family":"Jain","given":"Sarthak"},{"family":"Suyash"},{"family":"Kumar","given":"Upendra"},{"family":"Chauhan","given":"Utsav"},{"family":"Kumar","given":"Ashish"}],"issued":{"date-parts":[[2026]]},"DOI":"10.5281/zenodo.21701384","URL":"https://doi.org/10.5281/zenodo.21701384","source":"datacite"}]