Publication:
Safe DRL for resource allocation with PAoI violation guarantees

dc.contributor.advisorErgen, Sinem Çöleri
dc.contributor.kuauthorReyhan, Berire Güneş
dc.contributor.programElectrical and Electronics Engineering
dc.contributor.refereeErdoğan, Alper Tunga
dc.contributor.refereeÖzdemir, Mehmet Kemal
dc.contributor.schoolcollegeinstituteGRADUATE SCHOOL OF SCIENCES AND ENGINEERING
dc.coverage.spatialİstanbul
dc.date.accessioned2026-09-10T09:04:37Z
dc.date.issued2025
dc.description.abstractAbstract: In Wireless Networked Control Systems (WNCSs), control and communication systems must be co-designed due to their strong interdependence. This thesis presents a novel optimization theory-based safe deep reinforcement learning (DRL) framework for ultra-reliable WNCSs, ensuring constraint satisfaction while optimizing performance, for the first time in the literature. The approach minimizes power consumption under key constraints, including Peak Age of Information (PAoI) violation probability, transmit power, and schedulability in the finite blocklength (FBL) regime. PAoI violation probability is uniquely derived by combining stochastic maximum allowable transfer interval (MATI) and maximum allowable packet delay (MAD) constraints in a multi-sensor network. The framework consists of two stages: optimization theory and safe DRL. The first stage derives optimality conditions to establish mathematical relationships among variables, simplifying and decomposing the problem. The second stage employs a safe DRL model where a teacher-student framework guides the DRL agent (student). The control mechanism (teacher) evaluates compliance with system constraints and suggests the nearest feasible action when needed. Extensive simulations show that the proposed framework outperforms rule-based and other optimization theory based DRL benchmarks, achieving faster convergence, higher rewards, and greater stability.
dc.description.abstractÖzet: Kablosuz Ağ Tabanlı Kontrol Sistemlerinde (WNCS), kontrol ve iletişim sistemleri arasındaki güçlü karşılıklı bağımlılık, bu iki yapının eşzamanlı olarak tasarlanmasını zorunlu kılmaktadır. Bu tez, literatürde ilk kez, performansı optimize ederken sistem kısıtlamalarını sağlamayı garantileyen ultra güvenilir WNCS'ler için yeni bir optimizasyon teorisi tabanlı güvenli Derin Pekiştirmeli Öğrenme (DRL) metodu sunmaktadır. Önerilen algoritma, sonlu blok uzunluğu (FBL) rejiminde, Bilginin Tepe Yaşı (PAoI) ihlal olasılığı, iletim gücü (transmit power) ve planlanabilirlik (schedulability) dahil olmak üzere temel kısıtlamalar altında güç tüketimini en aza indirir. PAoI ihlal olasılığı, bir çoklu sensör ağında, stokastik maksimum izin verilen aktarım aralığı (MATI) ve maksimum izin verilen paket gecikmesi (MAD) kısıtlamalarının birleştirilmesiyle benzersiz bir şekilde türetilir. Önerilen metot iki aşamadan oluşur: optimizasyon teorisi ve güvenli DRL. İlk aşama, değişkenler arasında matematiksel ilişkiler kurmak, sorunu basitleştirmek ve ayrıştırmak için optimallik koşulları (optimality conditions) türetilir. İkinci aşama, bir öğretmen-öğrenci çerçevesinin (teacher-student framework) DRL etmenini (öğrenci) yönlendirdiği güvenli bir DRL modeli kullanır. Kontrol mekanizması (öğretmen), seçilen eylemin sistem kısıtlamalarına uyumunu değerlendirir ve gerektiğinde sistem gereksinimlerini sağlayan en yakın eylemi önerir. Kapsamlı simülasyonlar, önerilen metodun kural tabanlı ve diğer optimizasyon teorisi tabanlı DRL kıyaslama ölçütlerinden daha iyi performans gösterdiğini, daha hızlı yakınsama, daha yüksek ödüller ve daha büyük istikrar elde ettiğini göstermektedir.
dc.description.fulltextYes
dc.format.extentxii, 45 pages ;; 30 cm.
dc.identifier.embargoNo
dc.identifier.endpage57
dc.identifier.filenameinventorynoT_2025_073_GSSE
dc.identifier.urihttps://hdl.handle.net/20.500.14288/35171
dc.identifier.yoktezid974516
dc.identifier.yoktezlinkhttps://tez.yok.gov.tr/UlusalTezMerkezi/TezGoster?key=V-oEQd0LkkqRGCXNzJWCTToVkSKaI5S4gMb2nCIKp-3BT2QqK2l51H7rJIvbvPI7
dc.keywordsSafe Deep Reinforcement Learning (Safe RL)
dc.keywordsResource allocation
dc.keywordsPeak Age of Information (PAoI)
dc.keywordsAge of Information (AoI)
dc.keywordsConstrained Markov Decision Process (CMDP)
dc.keywordsWireless communication networks
dc.keywordsMachine learning for communications
dc.keywordsGüvenli Derin Pekiştirmeli Öğrenme (Safe DRL)
dc.keywordsKaynak tahsisi
dc.keywordsEn yüksek bilgi yaşı (PAoI)
dc.keywordsBilgi yaşı (AoI)
dc.keywordsKısıtlı Markov Karar Süreci (CMDP)
dc.keywordsKablosuz haberleşme ağları
dc.keywordsHaberleşmede makine öğrenmesi
dc.languageeng
dc.publisherKoç University
dc.relation.collectionKoç University Theses & Dissertations Collection
dc.rightsrestrictedAccess
dc.rights.copyrightsnote© All Rights Reserved. Accessible to Koç University Affiliated Users Only!
dc.subjectWireless communication systems
dc.subjectTelecommunication systems
dc.subjectDeep learning (Machine learning)
dc.titleSafe DRL for resource allocation with PAoI violation guarantees
dc.typeThesis
dspace.entity.typePublication
relation.isAdvisorOfThesisa8eafa7b-40c0-4cb5-befc-9276c16e46dd
relation.isAdvisorOfThesis.latestForDiscoverya8eafa7b-40c0-4cb5-befc-9276c16e46dd
relation.isParentOrgUnitOfPublication434c9663-2b11-4e66-9399-c863e2ebae43
relation.isParentOrgUnitOfPublication.latestForDiscovery434c9663-2b11-4e66-9399-c863e2ebae43

Files