🤖 AI 资讯

每日 05:00 更新 · 09-23 · 主站 liuch.name ↗
全部标签 →
筛选标签:扩散模型 · 返回个性化推荐 · 清空筛选
AI 资讯

Serious editors are a commitment (at least for me)

Lobsters

Comments

· 大模型,AI应用,Google,Agent智能体,搜索RAG,扩散模型,强化学习,招聘HR,榜单评测

Arguing about arguments

Lobsters

Comments

2026-09-21T00:00:00.000Z · 大模型,算力芯片,AI应用,开源,Anthropic,Agent智能体,推理思考,扩散模型,强化学习,招聘HR,开发者生态
AI 资讯

Plain-text files are at risk

Lobsters

Comments

· 大模型,算力芯片,AI应用,Google,Microsoft,Agent智能体,搜索RAG,扩散模型,强化学习,端侧AI,招聘HR,网络安全,榜单评测
AI 资讯

QontoFAQ: A better Information Retrieval Benchmark [R]

Reddit r/MachineLearning

Retrieval benchmarks sometimes feel benchmaxxed by models, so we wanted to find a way to tie it as close as possible to my objective: finding the article that answers a product question right.

We worked on a new metric which seems more proportional to document relevance, and built up a benchmarking dataset to measure embedding models.

Here is an article on the approach: https://medium.com/qonto-way/qontofaq-benchmarking-information-retrieval-acd89600ebe1

and the associated code: https://github.com/qonto/qonto-faq-benchmark

submitted by /u/espadrine
[link] [comments]
2026-09-22 13:45:18 · 开源,扩散模型,模型评测,向量数据库,招聘HR,网络安全
AI 资讯

Simulating fault tolerance with stage skipping in pipeline-parallel training [R]

Reddit r/MachineLearning

Our most recent work at Templar explores fault tolerance in Crucible, our distributed pre-training platform. The goal is to keep healthy workers training when another pipeline stage goes offline.

Crucible combines data-parallel replicas with pipeline parallelism. Each replica holds a copy of the model, split into stages on separate workers. SparseLoCo exchanges compressed updates between replicas, while pipeline compression reduces the communication across stage boundaries.

We combine those methods with stage skipping. When an inner stage goes offline, activations and gradients bypass it for multiple steps. Healthy stages keep processing tokens instead of waiting for recovery. The bypass omits the unavailable stage’s computation.

The simulations use a 178M model, eight replicas and four stages per replica. At a 1% per-replica failure probability per global step, validation loss stayed close to the no-failure baseline, even though each simulated outage removed a stage for six global steps. Each configuration is compared with its own no-failure run.

Fixed projections shared across layers improve robustness further when using pipeline compression. This suggests that shared projectors align representations across stage boundaries, making bypasses less disruptive. The alignment explanation remains a hypothesis.

These results point toward training on a broader pool of compute, including unreliable workers and spot instances. This is a simulation of the learning effects of stage failures, rather than a measurement of physical worker replacement or production cost savings.

The article includes the setup, comparisons and figures:

https://www.tplr.ai/publications/blog/skipping-stages-with-fixed-projections

submitted by /u/covenant_ai
[link] [comments]
2026-09-22 15:47:47 · 具身智能,扩散模型,模型安全对齐,招聘HR
AI 资讯

LinearSolveBench: new benchmark for linear solvers [P]

Reddit r/MachineLearning

LinearSolverBench measures the ability of a model or harness to write fast, accurate, and general numerical solvers for large sparse linear systems in C.

The goal is to encourage algorithmic advances in numerical methods for solving linear systems of equations.

https://github.com/hgarud/LinearSolveBench

submitted by /u/hgarud
[link] [comments]
2026-09-22 15:34:58 · AI应用,开源,搜索RAG,扩散模型,模型评测,招聘HR

Understanding and Enhancing Kimi Delta Attention [R]

Reddit r/MachineLearning
Understanding and Enhancing Kimi Delta Attention [R]

TLDR: We demonstrate and explain the difference in expressivity of Gated Deltanet (GDN) and Kimi Delta Attention (KDA). We show how the full diagonal gate in KDA can act as a reflection allowing 2D rotations to be carried out in a single step, but only if the range of the gates is extended to [-1,1] and the delta rule learning rate is extended to [0, 2] which we call Complex KDA (CKDA). Our theory demonstrates that this form allows us to express any orthogonal diagonal-plus-rank-one matrix and track the S3, S4, and A5 groups, but not S5. Our experiments show that CKDA can learn S3 and S4, shows promising results on Audio continuation and it can train stably and be competitive with standard KDA on language modelling.
Paper title: Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention

https://i.redd.it/v7oqopy3v1rh1.gif

submitted by /u/Yossarian_1234
[link] [comments]
2026-09-22 10:34:43 · 大模型,月之暗面,Transformer,扩散模型,招聘HR

Play social multiplayer games against frontier AI models and see if you can beat them! [D]

Reddit r/MachineLearning
Play social multiplayer games against frontier AI models and see if you can beat them! [D]

play here

Play games like poker, risk, diplomacy with friends or alone against AI models, guess what, you can talk to them and change strategies and outcomes. itss fun!

submitted by /u/Expert_Cobbler8984
[link] [comments]
2026-09-22 16:49:55 · 扩散模型,招聘HR,网络安全
AI 资讯

Fulvid

Product Hunt

A standalone desktop editor for Markdown and MDX

Discussion | Link

2026-09-20 04:12:19 · 扩散模型,招聘HR,榜单评测
AI 资讯

Clueso MCP

Product Hunt

Create and edit videos by chatting

Discussion | Link

2026-09-16 09:35:16 · Agent智能体,扩散模型,招聘HR

Obscura: VPN that can't log your activity

Hacker NewsComments
· AI应用,开源,搜索RAG,扩散模型,强化学习,招聘HR,榜单评测

The Softness of Metal

Hacker NewsComments
2026-09-21T10:00:00.000Z · AI应用,具身智能,Meta,快手,文生视频,搜索RAG,扩散模型,强化学习,招聘HR,榜单评测

MUNI Heritage Weekend in San Francisco

Hacker NewsComments
· 具身智能,自动驾驶,Google,Meta,语音音频,扩散模型,招聘HR,榜单评测,开发者生态
AI 资讯

Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived

Hacker NewsComments
· AI应用,开源,Microsoft,代码生成,扩散模型,招聘HR,榜单评测,开发者生态
AI 资讯

All the ways a cell can die — and why the variety matters

Nature

Nature, Published online: 22 September 2026; doi:10.1038/d41586-026-02929-z

Scientists are poking holes in established ideas about how cells live and die, and are harnessing these discoveries to fight diseases from cancer to autoimmune conditions.
2026-09-22 00:00:00 · 扩散模型,招聘HR
继续滚动加载更多…