QontoFAQ: A better Information Retrieval Benchmark [R]
Retrieval benchmarks sometimes feel benchmaxxed by models, so we wanted to find a way to tie it as close as possible to my objective: finding the article that answers a product question right.
We worked on a new metric which seems more proportional to document relevance, and built up a benchmarking dataset to measure embedding models.
Here is an article on the approach: https://medium.com/qonto-way/qontofaq-benchmarking-information-retrieval-acd89600ebe1
and the associated code: https://github.com/qonto/qonto-faq-benchmark
[link] [comments]
LinearSolveBench: new benchmark for linear solvers [P]
LinearSolverBench measures the ability of a model or harness to write fast, accurate, and general numerical solvers for large sparse linear systems in C.
The goal is to encourage algorithmic advances in numerical methods for solving linear systems of equations.
[link] [comments]
Unreal Agent
Three Routes to One Answer: Reconciling AIPW, TMLE, and Double Machine Learning for Applied Researchers
Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging
Rethinking Post-Hoc Calibration in Semantic Segmentation
The Virtue of Sparsity in Complexity
Unified Multimodal Uncertain Inference
Converge to Surprise: Evolutionary Self-supervised Image Clustering
Communication-Efficient Byzantine-Robust Federated Conformal Prediction via Partial Sharing
Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
Hierarchical Sparse Bayesian Multitask Learning for Disease Prediction in Pooled Microbiome Studies
Diffusion-Induced Spatial Attention Overlapping Community Detection
On Basis Function Selection for Sparse Gaussian Process Regression