Publication: Safe deep reinforcement learning for resource allocation with peak age of information violation guarantees
| dc.contributor.coauthor | Gunes Reyhan, B. | |
| dc.contributor.coauthor | Coleri, S. | |
| dc.date.accessioned | 2026-08-14T11:26:39Z | |
| dc.date.issued | 2025 | |
| dc.description.abstract | In Wireless Networked Control Systems (WNCSs), control and communication systems must be co-designed due to their strong interdependence. This paper presents a novel optimization theory-based safe deep reinforcement learning (DRL) framework for ultra-reliable WNCSs, ensuring constraint satisfaction while optimizing performance, for the first time in the literature. The approach minimizes power consumption under key constraints, including Peak Age of Information (PAoI) violation probability, transmit power, and schedulability in the finite blocklength regime. PAoI violation probability is uniquely derived by combining stochastic maximum allowable transfer interval (MATI) and maximum allowable packet delay (MAD) constraints in a multi-sensor network. The framework consists of two stages: optimization theory and safe DRL. The first stage derives optimality conditions to establish mathematical relationships among variables, simplifying and decomposing the problem. The second stage employs a safe DRL model where a teacher-student framework guides the DRL agent (student). The control mechanism (teacher) evaluates compliance with system constraints and suggests the nearest feasible action when needed. Extensive simulations show that the proposed framework outperforms rule-based and other optimization theory based DRL benchmarks, achieving faster convergence, higher rewards, and greater stability. | |
| dc.description.harvestedfrom | Manual | |
| dc.description.indexedby | WOS | |
| dc.description.indexedby | Scopus | |
| dc.description.publisherscope | International | |
| dc.description.readpublish | N/A | |
| dc.description.sponsoredbyTubitakEu | TÜBİTAK | |
| dc.description.sponsorship | Sinem Coleri acknowledges the support of the Scientific and Technological Research Council of Turkey 2247-A National Leaders Research Grant 121C314. | |
| dc.description.version | Published Version | |
| dc.identifier.ScopusPercentile | 94 | |
| dc.identifier.ScopusQuartile | Q1 | |
| dc.identifier.WoSPercentile | 90,4 | |
| dc.identifier.WoSQuartile | Q1 | |
| dc.identifier.doi | 10.1109/tcomm.2025.3591167 | |
| dc.identifier.eissn | 1558-0857 | |
| dc.identifier.embargo | N/A | |
| dc.identifier.endpage | 14211 | |
| dc.identifier.grantno | 121C314 | |
| dc.identifier.issn | 0090-6778 | |
| dc.identifier.issue | 12 | |
| dc.identifier.scopus | 2-s2.0-105011486053 | |
| dc.identifier.startpage | 14197 | |
| dc.identifier.uri | http://doi.org/10.1109/tcomm.2025.3591167 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14288/34621 | |
| dc.identifier.volume | 73 | |
| dc.identifier.wos | 001649720000035 | |
| dc.keywords | Reinforcement learning | |
| dc.keywords | Computer science | |
| dc.keywords | Resource allocation | |
| dc.keywords | Computer network | |
| dc.keywords | Distributed computing | |
| dc.keywords | Artificial intelligence | |
| dc.keywords | Optimization | |
| dc.keywords | Sensors | |
| dc.keywords | Ultra reliable low latency communication | |
| dc.keywords | Resource management | |
| dc.keywords | Delays | |
| dc.keywords | Mathematical models | |
| dc.keywords | Wireless sensor networks | |
| dc.keywords | Networked control systems | |
| dc.keywords | Information age | |
| dc.keywords | Data models | |
| dc.keywords | Wireless networked control systems (WNCS) | |
| dc.keywords | Ultra-reliable low latency communication (URLLC) | |
| dc.keywords | Finite blocklength (FBL) | |
| dc.keywords | Optimization theory | |
| dc.keywords | Safe deep reinforcement learning (DRL) | |
| dc.keywords | Teacher student framework | |
| dc.keywords | Peak age of information (PAoI) violation probability | |
| dc.language | eng | |
| dc.publisher | IEEE | |
| dc.relation.affiliation | Koç University | |
| dc.relation.collection | Koç University Institutional Repository | |
| dc.relation.ispartof | IEEE Transactions on Communications | |
| dc.relation.openaccess | N/A | |
| dc.rights | N/A | |
| dc.rights.uri | N/A | |
| dc.subject | Physical sciences | |
| dc.subject | Computer science | |
| dc.subject | Computer networks and communications | |
| dc.subject | Telecommunications | |
| dc.title | Safe deep reinforcement learning for resource allocation with peak age of information violation guarantees | |
| dc.type | Journal Article | |
| dspace.entity.type | Publication |
