Understanding and Enhancing Kimi Delta Attention [R]
| TLDR: We demonstrate and explain the difference in expressivity of Gated Deltanet (GDN) and Kimi Delta Attention (KDA). We show how the full diagonal gate in KDA can act as a reflection allowing 2D rotations to be carried out in a single step, but only if the range of the gates is extended to [-1,1] and the delta rule learning rate is extended to [0, 2] which we call Complex KDA (CKDA). Our theory demonstrates that this form allows us to express any orthogonal diagonal-plus-rank-one matrix and track the S3, S4, and A5 groups, but not S5. Our experiments show that CKDA can learn S3 and S4, shows promising results on Audio continuation and it can train stably and be competitive with standard KDA on language modelling. [link] [comments] |
Context-Adaptive Thresholding for Conditionally Representative Monitoring and Classification
Financially Guided Deep Portfolio Optimization
Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction
CAffNet: Hard Constraint-Affine Neural Networks
Practical Scaling Laws: Converting Compute into Performance in a Data-Constrained World
Diffusion-Induced Spatial Attention Overlapping Community Detection
Unlocking Cross-Scenario Physical Layer Security: A Mixture-of-Experts Framework with Generative Diffusion Models
GTR: Gated Token Recurrence for Efficient Dense Prediction
PatchKV: Efficient KV Cache Recovery for Dynamically Edited LLM Contexts
Faithful Faithfulness Evaluations: Challenges & Pitfalls Learned from a Breast MRI Case Study
Statistical Gains from Looped Estimation under Parameter Budgets
Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression
<\infty$ is of order $[N^2\log(eNT)]^{-\beta/d}$ when the network width satisfies $N\geq2d+3$ and the parameter magnitudes are bounded by $T\geq1$. Matching lower bounds hold for every fixed globally H\"older activation; its H\"older exponent affects the constants but not the rate. Under bounded design densities and independent centered sub-Gaussian noise, approximate least squares over the full clipped class at depth $23$ attains the classical H\"older minimax risk $\mathcal{O}(M^{-\frac{2\beta}{2\beta+d}})$ without logarithmic loss whenever $N^2\log(eNT)\asymp M^{\frac{d}{2\beta+d}}$, where $M$ is the sample size. This yields a continuum of statistically optimal choices, ranging from unit parameter radius to fixed network size. At fixed size, four hidden layers with at most $8d+7$ nonzero parameters give a near-optimal radius, while six layers with at most $8d+27$ attain the optimal order $\log T=\mathcal{O}(\eta^{-d/\beta})$ at approximation error $\eta$. The same decoding method also yields fixed-size Transformer approximation.
Learning from Humans for Proactive Assistance in Human-Robot Collaborative Transport
PreGS: A Parameter-Transfer-Based Multi-Expert Graph Neural Network for Node Classification
Information-Theoretic Decoupled Prompt Tuning for Continual Learning
Can You Delete a Year of Market Data? Machine Unlearning Against Exact Retraining Oracles
Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity
Spectral Tail Interventions in Decoder-Only Language Models: Reasoning-Sensitive Weight Structure from Controlled Surgery