同日推出"廉价模型",OpenAI和Anthropic开打"价格战"
中国让开源大模型工作更长时更“聪明” - banyuetan.org
Kimi K3登陆AWS:中国大模型开始"收租",月之暗面把开源玩成了收费站 - 36kr.com
Arguing about arguments
QontoFAQ: A better Information Retrieval Benchmark [R]
Retrieval benchmarks sometimes feel benchmaxxed by models, so we wanted to find a way to tie it as close as possible to my objective: finding the article that answers a product question right.
We worked on a new metric which seems more proportional to document relevance, and built up a benchmarking dataset to measure embedding models.
Here is an article on the approach: https://medium.com/qonto-way/qontofaq-benchmarking-information-retrieval-acd89600ebe1
and the associated code: https://github.com/qonto/qonto-faq-benchmark
[link] [comments]
LinearSolveBench: new benchmark for linear solvers [P]
LinearSolverBench measures the ability of a model or harness to write fast, accurate, and general numerical solvers for large sparse linear systems in C.
The goal is to encourage algorithmic advances in numerical methods for solving linear systems of equations.
[link] [comments]
Unreal Agent
CAffNet: Hard Constraint-Affine Neural Networks
Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
CID: Measuring Feature Importance Through Counterfactual Distributions
GTR: Gated Token Recurrence for Efficient Dense Prediction
SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation
SPARC: SuperPixel-Aware Region Contrastive Learning for Self-Supervised Dense Prediction