Publication:
Safe deep reinforcement learning for resource allocation with peak age of information violation guarantees

dc.contributor.coauthorGunes Reyhan, B.
dc.contributor.coauthorColeri, S.
dc.date.accessioned2026-08-14T11:26:39Z
dc.date.issued2025
dc.description.abstractIn Wireless Networked Control Systems (WNCSs), control and communication systems must be co-designed due to their strong interdependence. This paper presents a novel optimization theory-based safe deep reinforcement learning (DRL) framework for ultra-reliable WNCSs, ensuring constraint satisfaction while optimizing performance, for the first time in the literature. The approach minimizes power consumption under key constraints, including Peak Age of Information (PAoI) violation probability, transmit power, and schedulability in the finite blocklength regime. PAoI violation probability is uniquely derived by combining stochastic maximum allowable transfer interval (MATI) and maximum allowable packet delay (MAD) constraints in a multi-sensor network. The framework consists of two stages: optimization theory and safe DRL. The first stage derives optimality conditions to establish mathematical relationships among variables, simplifying and decomposing the problem. The second stage employs a safe DRL model where a teacher-student framework guides the DRL agent (student). The control mechanism (teacher) evaluates compliance with system constraints and suggests the nearest feasible action when needed. Extensive simulations show that the proposed framework outperforms rule-based and other optimization theory based DRL benchmarks, achieving faster convergence, higher rewards, and greater stability.
dc.description.harvestedfromManual
dc.description.indexedbyWOS
dc.description.indexedbyScopus
dc.description.publisherscopeInternational
dc.description.readpublishN/A
dc.description.sponsoredbyTubitakEuTÜBİTAK
dc.description.sponsorshipSinem Coleri acknowledges the support of the Scientific and Technological Research Council of Turkey 2247-A National Leaders Research Grant 121C314.
dc.description.versionPublished Version
dc.identifier.ScopusPercentile94
dc.identifier.ScopusQuartileQ1
dc.identifier.WoSPercentile90,4
dc.identifier.WoSQuartileQ1
dc.identifier.doi10.1109/tcomm.2025.3591167
dc.identifier.eissn1558-0857
dc.identifier.embargoN/A
dc.identifier.endpage14211
dc.identifier.grantno121C314
dc.identifier.issn0090-6778
dc.identifier.issue12
dc.identifier.scopus2-s2.0-105011486053
dc.identifier.startpage14197
dc.identifier.urihttp://doi.org/10.1109/tcomm.2025.3591167
dc.identifier.urihttps://hdl.handle.net/20.500.14288/34621
dc.identifier.volume73
dc.identifier.wos001649720000035
dc.keywordsReinforcement learning
dc.keywordsComputer science
dc.keywordsResource allocation
dc.keywordsComputer network
dc.keywordsDistributed computing
dc.keywordsArtificial intelligence
dc.keywordsOptimization
dc.keywordsSensors
dc.keywordsUltra reliable low latency communication
dc.keywordsResource management
dc.keywordsDelays
dc.keywordsMathematical models
dc.keywordsWireless sensor networks
dc.keywordsNetworked control systems
dc.keywordsInformation age
dc.keywordsData models
dc.keywordsWireless networked control systems (WNCS)
dc.keywordsUltra-reliable low latency communication (URLLC)
dc.keywordsFinite blocklength (FBL)
dc.keywordsOptimization theory
dc.keywordsSafe deep reinforcement learning (DRL)
dc.keywordsTeacher student framework
dc.keywordsPeak age of information (PAoI) violation probability
dc.languageeng
dc.publisherIEEE
dc.relation.affiliationKoç University
dc.relation.collectionKoç University Institutional Repository
dc.relation.ispartofIEEE Transactions on Communications
dc.relation.openaccessN/A
dc.rightsN/A
dc.rights.uriN/A
dc.subjectPhysical sciences
dc.subjectComputer science
dc.subjectComputer networks and communications
dc.subjectTelecommunications
dc.titleSafe deep reinforcement learning for resource allocation with peak age of information violation guarantees
dc.typeJournal Article
dspace.entity.typePublication

Files