🤖 AI 资讯

每日 05:00 更新 · 09-23 · 主站 liuch.name ↗
全部标签 →
筛选标签:榜单评测 · 返回个性化推荐 · 清空筛选

AI“减速”讨论未歇,OpenAI、Anthropic同日上新,竞逐更低成本

澎湃新闻
· 大模型,AI应用,开源,融资,政策监管,OpenAI,Anthropic,xAI,代码生成,Agent智能体,推理思考,模型评测,提示工程,模型安全对齐,招聘HR,模型发布,合作,榜单评测,开发者生态

还未见顶:2-4个月后超强厄尔尼诺才迎峰值,关键海区已逼近历史极值

澎湃新闻
· 政策监管,模型评测,招聘HR,榜单评测
AI 资讯

Inside AI Prompt Security: Why Stopping Every LLM Exploit Is Impossible - PCMag

2026-09-22 19:14:17 · 大模型,Google,提示工程,招聘HR,榜单评测
AI 资讯

Inside AI Prompt Security: Why Stopping Every LLM Exploit Is Impossible - PCMag UK

2026-09-22 19:00:16 · 大模型,OpenAI,Google,推理思考,提示工程,招聘HR,榜单评测
AI 资讯

Serious editors are a commitment (at least for me)

Lobsters

Comments

· 大模型,AI应用,Google,Agent智能体,搜索RAG,扩散模型,强化学习,招聘HR,榜单评测
AI 资讯

Plain-text files are at risk

Lobsters

Comments

· 大模型,算力芯片,AI应用,Google,Microsoft,Agent智能体,搜索RAG,扩散模型,强化学习,端侧AI,招聘HR,网络安全,榜单评测
AI 资讯

Fulvid

Product Hunt

A standalone desktop editor for Markdown and MDX

Discussion | Link

2026-09-20 04:12:19 · 扩散模型,招聘HR,榜单评测
AI 资讯

FeedsBar

Product Hunt

A quiet, continuous news/media ticker for your Mac desktop

Discussion | Link

2026-08-30 11:29:45 · 招聘HR,榜单评测
AI 资讯

PewCB

Product Hunt

Vibe-routing is coming. Desktop PCB Factory is here.

Discussion | Link

2026-09-21 14:35:40 · 招聘HR,榜单评测
AI 资讯

2BA.AI

Product Hunt

Stop waiting for tokens, and start shipping

Discussion | Link

2026-09-16 13:13:32 · 招聘HR,榜单评测
AI 资讯

Googlebook

Product Hunt

The laptop your Android phone has been waiting for

Discussion | Link

2026-09-21 16:42:06 · Google,招聘HR,榜单评测

Obscura: VPN that can't log your activity

Hacker NewsComments
· AI应用,开源,搜索RAG,扩散模型,强化学习,招聘HR,榜单评测

The Softness of Metal

Hacker NewsComments
2026-09-21T10:00:00.000Z · AI应用,具身智能,Meta,快手,文生视频,搜索RAG,扩散模型,强化学习,招聘HR,榜单评测

MUNI Heritage Weekend in San Francisco

Hacker NewsComments
· 具身智能,自动驾驶,Google,Meta,语音音频,扩散模型,招聘HR,榜单评测,开发者生态
AI 资讯

Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived

Hacker NewsComments
· AI应用,开源,Microsoft,代码生成,扩散模型,招聘HR,榜单评测,开发者生态
AI 资讯

Tight Sample Complexity Bounds for Entropic Best Policy Identification

arXiv cs.LGarXiv:2605.13717v2 Announce Type: replace Abstract: We study best-policy identification for finite-horizon risk-sensitive reinforcement learning under the entropic risk measure. Recent work established a constant gap in the exponential horizon dependence between lower and upper bounds on the number of samples required to identify an approximately optimal policy. Precisely, known lower bounds scale in $\Omega(e^{|\beta| H})$ where $H$ is the horizon of the MDP, while the state-of-the-art upper bound achieves at best $O(e^{2|\beta| H})$ (arXiv:2506.00286v2) using a generative model. We show that this extra exponential factor can be traced to overly loose concentration control for exponential utilities. To close this open gap, we revisit the analysis of this problem through a forward-model based algorithm building on KL-based exploration bonuses that we adapt to the entropic criterion. The improvement we get is due to two main novel technical innovations. We leverage the smoothness properties of the exponential utility to derive sharper concentration bounds, and we propose a new stopping rule that exploits further this tightness to obtain a sample complexity that matches the lower bound.
2026-09-23 04:00:00 · AI应用,搜索RAG,强化学习,微调蒸馏,招聘HR,榜单评测,论文
继续滚动加载更多…