IRIS Catalogo Istituzionale della Ricerca dell'Università degli Studi del Molise

The use of Artificial Intelligence principles represents the next research challenge to support future network applications in the upcoming 6G era. In this work, we propose a novel approach: exploiting the principles of Reinforcement Learning (RL) and the availability of programmable switches to implement a new forwarding mechanism in the data plane of the 6G core network. More in detail, we define a Q-learning-based forwarding mechanism that acts at packet level and is able to select the minimum latency path at line rate. Our solution, referred to as Q-Learning-based Queue Length Routing in DAta Plane ((QL)2-RODAP), is fully decentralized and exploits in-band network telemetry to distribute network states among network nodes. We show that, either in random and real network topologies, our (QL)2-RODAP algorithm promptly reacts to sudden traffic bursts, and allows reducing the peak of queuing delays of about 65 − 85% with respect to other RL based approaches, thus cutting off the long tail of end-to-end latency that is critical for delay sensitive applications.

In-Network Q-Learning-Based Packet Forwarding for Delay Sensitive Applications

Polverini M.;Cianfrani A.;Listanti M.;Caiazzi T.;Scazzariello M.

2025-01-01

Abstract

The use of Artificial Intelligence principles represents the next research challenge to support future network applications in the upcoming 6G era. In this work, we propose a novel approach: exploiting the principles of Reinforcement Learning (RL) and the availability of programmable switches to implement a new forwarding mechanism in the data plane of the 6G core network. More in detail, we define a Q-learning-based forwarding mechanism that acts at packet level and is able to select the minimum latency path at line rate. Our solution, referred to as Q-Learning-based Queue Length Routing in DAta Plane ((QL)2-RODAP), is fully decentralized and exploits in-band network telemetry to distribute network states among network nodes. We show that, either in random and real network topologies, our (QL)2-RODAP algorithm promptly reacts to sudden traffic bursts, and allows reducing the peak of queuing delays of about 65 − 85% with respect to other RL based approaches, thus cutting off the long tail of end-to-end latency that is critical for delay sensitive applications.

Scheda breve

Scheda completa

Scheda completa (DC)

	Codice DOI
	
				https://dx.doi.org/10.1109/MNET.2025.3552929
			
	Codice Scopus
	
				2-s2.0-105000542067
			
	Appare nelle tipologie:
	
				1.1 Articolo in rivista

File in questo prodotto:

Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11695/147670

Citazioni

ND

1

ND

social impact