🤖 AI 资讯

· ·
← 返回列表

Differentiable Policy Transport over Multi-Layer Network Feasibility Geometry

arXiv cs.LG2026-09-23 04:00:00扩散模型,强化学习,微调蒸馏,招聘HR,论文原文 ↗

arXiv:2609.26068v1 Announce Type: cross

Abstract: Learning-based control is increasingly central to automating network operations. A learned policy, however, must satisfy cross-layer constraints on interference, power-rate coupling, flow conservation, service chains, capacity, latency, and reliability. Existing methods typically account for only a subset of this geometry and only indirectly, e.g., through reward penalties, Lagrange multipliers, or post-hoc repairs. This paper proposes \emph{Network Feasibility Geometry Reinforcement Learning} (NFG-RL), which models coupled constraints via transport theory and the residual inclusion $\bphi_{\mathfrak{N}}(x,a)\in\cK_{\mathfrak{N}}$, defining the executed policy as the pushforward of a proto-policy through a feasibility-transport map. NFG-RL compiles heterogeneous constraints into typed residual blocks and transports proto-actions through a differentiable variational operator, letting active constraints shape execution, exploration, and actor gradients. Our analysis shows that exact transport yields almost-sure feasible execution, while active constraints contract exploration onto the feasible tangent space. It further establishes a nonnegative first-order gain from critic-tilted transport over plain projection and recovers backpressure scheduling as the gradient of a lifted drift residual. In two public-trace-conditioned wireless-edge surrogate environments, NFG-RL improves feasible utility by \textbf{37.5--41.5\%} over the strongest non-NFG method in each environment, reduces raw-action violation by \textbf{48.5--60.8\%}, and lowers P99 delay by \textbf{57.0--75.5\%}, outperforming a range of optimization and learning baselines.