
DRLD2D: Deep Reinforcement Learning-Based Optimized Routing for 5G Device-to-Device Net-works
In context of 5G network, having an efficient communication between the devices in Device to Device (D2D) commu- nication network is a challenging task. The optimised routing in this dynamic and decentralized environment of D2D is crucial for maintaining efficient communication. To overcome these limitations this paper presents Deep Reinforcement Learning (DRL) based routing for D2D 5G network. DRLD2D uses DRL to dynamically optimize the route choice in the network. The problem of routing in DRLD2D is modelled as Markov Decision Process (MDP) and uses the approach of DRL to enable intelligent, adaptive and efficient routing decisions. Th e optimal paths are learnt by Reinforcement Learning (RL) by analysing network state, node load, Signal -to-Noise Ratio (SNR) and not relying on the static routing tables which is done in the conventional routing protocols. Extensive simulations are conduct ed using Network Simulator 3 Artificial Intelligence (NS3 AI) model for the proposed DRLD2D protocol which out performs the prevailing routing protocols such as Ad hoc On -Demand Distance Vector (AODV) and Multipath Energy and Quality of Service -Aware Optim ized Link State Routing Protocol Version 2 (MEQSA -OLSRv2) when the performance is measured in terms of End -to-End (E2E) delay, Packet Delivery Ratio (PDR), Throughput and Energy Efficiency. The obtained results validate the effectiveness of RL in addressin g the routing complexities of D2D communication network which makes it a promising solution for future 5G networks and beyond.
[1] Abou -Rjeily C and El -Zahr S., “Deep Reinforce- ment Learning Based Relaying for Buffer -Aided Co -Operative Communications, ” Physical Com- munication , vol. 59, pp. 1 -33., 2023. https://doi.org/10.1016/j.phycom.2023.102086
[2] Adu -Manu K., Amoako E. , and Engmann F., “Ad- vancements in Machine Learning ‐Enhanced Green Wireless Sensor Networks: A Comprehen- sive Survey on Energy Efficiency, Network Per- formance, and Future Directions, ” Journal of Sen- sors , vol. 2025, no. 1, pp. 1 -25, 2025. https://doi.org/10. 1155/2025/5242517
[3] Agyekum K, Boakye A., Opoku J., Agyemang J., and et al., ” Resource Allocation in D2D ‐Enabled 5G Networks Using Multiagent Reinforcement Learning, ” Journal of Computer Networks and Communications , vol. 2024, no. 1, pp. 1 -14, 2024. https:// doi.org/10.1155/2024/2780845
[4] Arya G., Bagwari A., and Chauhan D., “Perfor- mance Analysis of Deep Learning -Based Routing Protocol for an Efficient Data Transmission in 5G WSN Communication, ” IEEE Access , vol. 10, pp. 9340 -9356, 2022. DOI: 10.1109/AC- CESS.2022 .3142082
[5] Bunu S., Saraee M., and Alani O., “Machine Learning -Based Optimized Link State Routing Protocol for D2D Communication in 5G/B5G, ” in Proceedings of the International Conference on Electrical Engineering and Informatics , Banda Aceh, pp. 7 -12, 2022. https://doi.org/10.1109/ICELTICs56128.2022.99 32126
[6] Chen P., Liu S., Wang X. and Kamwa I., “Physics - Guided Multi -Agent Deep Reinforcement Learn- ing for Robust Active Voltage Control in Electrical Distribution Systems, ” IEEE Transactions on Cir- cuits and Syst ems I: Regular Papers , vol. 71, no. 2, pp. 922 -933, 2023. DOI: 10.1109/TCSI.2023.3340691
[7] Cui Y., Zhang Q., Feng Z., Wei Z., and et al., “To- pology -Aware Resilient Routing Protocol for FANETs: An Adaptive Q -Learning Approach, ” IEEE Internet of Things Journal , vol. 9, no. 19, pp. 18632 -18649, 2022. DOI: 10.1109/JIOT.2022.3162849
[8] Duong T and Binh L., “An improved method of AODV routing protocol using reinforcement learning for en -suring QoS in 5G -based mobile ad -hoc networks, ”. ICT Express , Vol. 10, no. 1, pp. 97 -103, 2024. https://doi.org/10.1016/j.icte.2023.07.002
[9] Elleuch W., Sondi P., Meddahi A., and Lecomte S., “Evaluation of 5G Relay -Empowered and Device - to-Device Communications for Rescue Mission, ” IEEE Access , vol. 13, pp. 104614 -104629, 2025. https://doi .org/10.1109/ACCESS.2025.3579872
[10] Goswami M., Panda N., Mohanty S., and Pattnaik P., “Machine Learning Tech -niques and Routing Protocols in 5G and 6G Mobile Network Commu- nication System -An Overview, ” in Proceedings of the 7 th International Conference on Tre nds in Elec- tronics and Informatics , Tirunelveli, pp. 1094 - 1101, 2023. DOI: 10.1109/ICOEI56765.2023.10125697
[11] Gottam S and Kar U., “Graph Attention Trans- former -Based Meta -Reinforcement Learning for Secure and Low -Latency D2D Routing in 6G Net - works, ” TechRxi v, 2025. https://doi.org/10.36227/techrxiv.174803075.576 48343/v1
[12] Guo H., Zhou X., Liu J., and Zhang Y., “Vehicular Intelligence in 6G: Networking, Communications, and Computing, ” Vehicular Communications , vol. 33, no. c, pp. 100399, 2022. https://doi.org/1 0.1016/j.vehcom.2021.100399
[13] Huang J., Yang C., Zhang S., Yang F., and et al., “Reinforcement Learning Based Resource Man- agement for 6G -Enabled mIoT with Hypergraph Interference Model, ” IEEE Transactions on Com- munications , vol. 72, no. 7, pp. 4179 -4192, 202 4. DOI: 10.1109/TCOMM.2024.3372892
[14] Huang Z., Li T, Song C, Li Z, and et al., “Joint Spectrum and Power Al -Location Scheme Based on Value Decomposition Networks in D2D Com- munication Networks, ” EURASIP Journal on Wireless Communications and Networking , vol. 2024, no. 1, pp. 1 -16, 2024. https://doi.org/10.1186/s13638 -024 -02393 -1
[15] Huang, J., Yang Y., Gao Z., He D., and Ng D., “Dynamic Spectrum Access for D2D -Enabled In- ternet of Things: A Deep Reinforcement Learning Approach, ” IEEE Internet of Things Journal , vol . 9 no. 18, pp. 17793 -17807, 2022. DOI: 10.1109/JIOT.2022.3160197
[16] Jayakumar S and Nandakumar S., “Distributed Resource Optimisation Using the Q -Learning Al- gorithm, in Device -to-Device Communication: A Reinforcement Learning Paradigm ” Results in En- gineering , vol. 23, pp. 1 -14, 2024. https://doi.org/10.1016/j.rineng.2024.102462
[17] Kumari P and Sahana S., “Swarm Based Hybrid ACO -PSO Meta -Heuristic (HAPM) for QoS Mul- ticast Routing Optimization in MANETs, ” Wire- less Personal Communications , vol. 123, pp. 1145 -1167, 2022. https://doi.org/10.1007/s11277 - 021 -09174 -9
[18] Lauric D , Luxey -Bitri A, Raes R, Rouvoy R, and et al., “A Critical Review of Mobile Device -To - Device Communication, ” in proceedings of the 25 th IFIP WG 6.1 International Conference, held as Part of the 20 th International Federated Con- ference on Distributed Computing Techniques , Lille, pp. 1 -24, 2025. https://doi.org/10.1007/978 - 3-031 -95728 -4_1
[19] Long D., Wu Q, Fan Q, Fan P, and et al., “A Power Allocation Scheme for MIMO -NOMA and D2D Vehicular Edge Computing Ba sed on Decentral- ized DRL, ” Sensors , vol. 23, no. 7, pp. 1 -19, 2023. https://doi.org/10.3390/s23073449
[20] Lu Y., Wang X., Li F., and Huang M., “RLbR: A Reinforcement Learning Based V2V Routing Framework for Offloading 5G Cellular IoT, ” IET Communications , vol. 16, no. 4, pp. 303 -313, 2022. https://doi.org/10.1049/cmu2.12346
[21] Malathy S., Jayarajan P., Hindia M., Tilwari V., and et al., “Routing Constraints in The Device -To - Device Communication for Beyond Iot 5G Net- works: A Review, ” Wireless Networks , vol. 27, no. 5, pp. 3207 -3231, 2021. https://doi.org/10.1007/s11276 -021 -02641 -y
[22] Malik T., Malik K., Afzal A., Wang L., and et al., “RL -IoT: Reinforcement Learning -Based Routing Approach for Cognitive Radio -Enabled IoT Com- munications, ” IEEE Internet of Things Journal , vol. 10, no. 2, pp. 1836 -1847, 2023. DOI: 10.1109/JIOT.2022.3210703
[23] Malta S., Pinto P., and Fernandez -Veiga M., “Us- ing Reinforcement Learning to Reduce Energy Consumption of Ultra -Dense Networks with 5G Use Cases Requirements, ” IEEE Access , vol. 11, pp. 54 17 -5428, 2023. DOI: 10.1109/AC- CESS.2023.3236980
[24] Mangal A., Rizvi M., and Khan S., Advanced Wireless Sensing Techniques for 5G Networks, Chapman and Hall/CRC. 2018. DOI:10.1201/9781351021746 -1
[25] Moghaddasi K., Rajabi S., Gharehchopogh F., and Ghaffari A., “An Advanced Deep Rein -Forcement Learning Algorithm for Three -Layer D2D -Edge - Cloud Computing Architecture for Efficient Task Offloading in The Internet of Things, ” Sustainable Computing: Informatics and Systems , vol. 43, pp. 100992. 2024. https://doi.org/10.1 016/j.sus- com.2024.100992
[26] Nauman A., Jamshed M., Qadri Y., Ali R., and DRLD2D: Deep Reinforcemen t Learning -Based Optimized Routing for 5G Device -to-Device Networks 1057 Kim S., “Reliability Optimization in Narrowband Device -to-Device Communication for 5G and Be- yond -5G Networks, ”. IEEE Access , vol. 9, pp. 157584 -157596, 2021. DOI: 10.1109/AC- CESS.2021.312 9896
[27] Pandey S., Vyas M., and Mangal A., “Exploring 5G: Technologies, Challenges, and 5G for En- hancement of IoT, ” Grenze International Journal of Engineering and Technology , vol. 6, no. 2., pp. 92 -100, 2020. https://research.eb- sco.com/c/z3bf3p/search/details/xfim7gowtr?re- quest -context=plink&db=aci
[28] Quy V., Chehri A., Quy N., Nguyen V., and Ban N., “An Efficient Routing Algorithm for Self -Or- ganizing Networks in 5G -Based Intell igent Trans- portation Systems, ” IEEE Transactions on Con - sumer Electronics , vol. 70, no. 1, pp. 1757 -1765, 2024. DOI: 10.1109/TCE.2023.3329390
[29] Ron D and Lee J., “Learning -Based Joint Op - Timization of Mode Selection and Transmit Power Control for D2D Communi cation Underlaid Cel- lular Networks, ” Expert Systems with Applications , vol. 198, pp. 116725, 2022. https://doi.org/10.1016/j.eswa.2022.116725
[30] Sahu R., Rizvi M., and Mishra M., “Routing Over- head Performance Study and Evaluation of Zone Based Energy Efficien t Multi -Path Routing Proto- cols in Manets by Using NS2. 35” in proceedings of the International Conference on Advances in Technology, Management and Education , Bhopal, pp. 88 -93, 2021. DOI: 10.1109/ICATME50232.2021.9732762.
[31] Sahu R., Sharma S., and Rizvi M., “Energy Con- sumption Evaluation of ZBLE, AOMDV and AODV Routing Protocols in Mobile Ad -hoc Net- works, ” Advances in Wireless Communications and Net -works , vol. 5, no. 2 pp. 41 -51, 2019. https://doi.org/10.11648/j.awcn.20190502.11
[32] Sahu R., Sharma S., and Rizv i M., “ZBLE: Energy Efficient Zone -Based Leader Election Multipath Routing Protocol for MANETs, ” International Journal Innovative Technological Exploration En- gineering , vol. 8, no. 9, pp. 2231 -2237, 2019. https://doi.org/10.35940/ijitee.H6916.078919
[33] Sahu R ., Sharma S., and Rizvi M., “ZBLE: Zone Based Efficient Energy Multipath Protocol for Routing in Mobile Ad Hoc Networks, ” Wireless Personal Com -munications , vol. 113, pp. 2641 - 2659., 2020. https://doi.org/10.1007/s11277 -020 - 07345 -8
[34] Sahu R., Sharma S., and Rizvi M., “ZBLE: Zone based Leader Selection Energy Constrained AOMDV Routing Protocol, ” International Jour- nal of Computer Networks and Applications , vol. 6, no. 3, pp. 56 -68, 2019. https://doi.org/10.5815/ijwmt.2019.05.05
[35] Sahu R., Sharma S., and Rizvi M., “ZBLE: Zone Based Leader Selection Protocol, ” International Journal of Wireless and Microwave Technologies , vol. 9, no. 4, pp.11 -25, 2019. https://doi.org/10.5815/ijwmt.2019.04.02
[36] Sahu V and Sahu R. , “ Energy Efficient Multipath Routing in Zone -based Mobile Ad -hoc Networks: Mathematical Formulation, ” International Jour- nal of Mathematical Sciences and Computing , vol. 8, no. 3, pp. 37 -48, 2022. https://doi.org/10.5815/ijmsc.2022.03.04
[37] Sharma V., Mohapat ra S., Shitharth S., Yonbawi S., and et al., “An Optimization -Based Machine Learning Technique for Smart Home Security Us- ing 5G, ” Computers and Electrical Engineering , vol. 104, pp. 10843, 2022, https://doi.org/10.1016/j.compele- ceng.2022.108434
[38] Ssengonzi C ., Kogeda O., and Olwal T., “A Sur- vey of Deep Reinforcement Learning Application in 5G and Beyond Network Slicing and Virtualiza- tion, ” Array , vol. 14, pp. 1 -27, 2022. https://doi.org/10.1016/j.array.2022.100142
[39] Tilwari V., Dimyati K., Hindia M., Amiri I., and Mohmed Noor Izam T., “EMBLR: A High -Perfor- mance Optimal Routing Approach for D2D Com- munications in Large -Scale Iot 5G Network, ” Symmetry , vol. 12, no. 3, pp.1 -23, 2020. https://doi.org/10.3390/sym12030438
[40] Wang C., Chai X., Peng S., Yuna Y, and Li G., “Deep Reinforcement Learning with Entropy and Attention Mechanism for D2D -Assisted Task Of- floading in Edge Computing, ” IEEE Transactions on Services Computing , vol. 17, no. 6, pp. 3317 - 3329, 2024. DOI: 10.1109/TSC.2024.3495503
[41] Wang X., Hu J., Lin H., Garg S ., and et al., “QoS and Privacy -Aware Routing for 5G -Enabled In- dustrial Internet of Things: A Federated Reinforce- ment Learning Approach, ” IEEE Transactions on Industrial Informatics , vol. 18, no. 6, pp. 4189 - 4197, 2022. DOI: 10.1109/TII.2021.3124848
[42] Yang X ., Yan J., Wang D., Xu Y., and Hua G., “WOAD3QN -RP: An Intelligent Routing Proto- col in Wireless Sensor Networks -A Swarm Intelli- gence and Deep Reinforcement Learning Based Ap -Proach, ” Expert Systems with Applications , vol. 246, pp.123089, 2024. https://doi. org/10.1016/j.eswa.2023.123089
[43] Yu S and Lee J., “Deep Reinforcement Learning Based Resource Allocation for D2D Communica- tions Underlay Cellular Networks, ” Sensors , vol. 22, no. 23, pp. 1 -19, 2022. https://doi.org/10.3390/s22239459
[44] Yu Y and Tang X., “Hybrid Centralized -Distrib- uted Resource Allocation Based on Deep Rein- forcement Learning for Cooperative D2D Com- munications, ”. IEEE Access , vol. 12, pp. 196609 - 196623, 2024. https://doi.org/10.1109/AC- CESS.2024.3521590
[45] Zhang C., Wu C., Lin M., Lin Y., and Liu W., “Proximal Policy Optimization for Efficient D2D - Assisted Computation Offloading and Resource Allocation in Multi -Access Edge Computing, ” Future Internet , vol. 16, no. 1, pp. 1 -17, 2024. https://doi.org/10.3390/fi16010019
[46] Zhi Y., Deng X., Tian J., Qiao J., and Lu D., “Deep Reinforcement Learning -Based Resource Alloca- tion for D2D Communications in Heterogeneous Cellular Networks, ” Digital Communications and Networks , vol. 8, no. 5, pp. 834 -842. 2022. https://doi.org/10.1016/j.dcan.2021.09.013