🤖 AI 资讯

· · ↗
← 返回列表

Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression

arXiv cs.LG2026-09-23 04:00:00大模型,Transformer,论文原文 ↗
arXiv:2609.25710v1 Announce Type: cross Abstract: The statistical accuracy of neural networks depends on both their approximation power and the complexity of the class fitted from data. While increasing network size is a natural way to improve approximation, parameter magnitude provides another resource whose role must be quantified in both respects. We establish a sharp width--magnitude tradeoff at fixed depth using one elementary bounded $1$-Lipschitz Dyadic--Triangular Activation. For the unit $\beta$-H\"older ball on $[0,1]^d$ with $0<\beta\leq1$, the optimal $L^p$ approximation error for $0

<\infty$ is of order $[N^2\log(eNT)]^{-\beta/d}$ when the network width satisfies $N\geq2d+3$ and the parameter magnitudes are bounded by $T\geq1$. Matching lower bounds hold for every fixed globally H\"older activation; its H\"older exponent affects the constants but not the rate. Under bounded design densities and independent centered sub-Gaussian noise, approximate least squares over the full clipped class at depth $23$ attains the classical H\"older minimax risk $\mathcal{O}(M^{-\frac{2\beta}{2\beta+d}})$ without logarithmic loss whenever $N^2\log(eNT)\asymp M^{\frac{d}{2\beta+d}}$, where $M$ is the sample size. This yields a continuum of statistically optimal choices, ranging from unit parameter radius to fixed network size. At fixed size, four hidden layers with at most $8d+7$ nonzero parameters give a near-optimal radius, while six layers with at most $8d+27$ attain the optimal order $\log T=\mathcal{O}(\eta^{-d/\beta})$ at approximation error $\eta$. The same decoding method also yields fixed-size Transformer approximation.