🤖 AI 资讯

每日 05:00 更新 · 09-23 · 主站 liuch.name ↗
全部标签 →
筛选标签:办公效率 · 返回个性化推荐 · 清空筛选

英美首次从无人潜航器试射鱼雷,奥库斯推进水下无人战力发展

澎湃新闻
· AI应用,具身智能,政策监管,Agent智能体,办公效率,政务,工业制造,招聘HR,基础设施,合作

白宫推出流媒体频道“特朗普电视”,此前五大电视网暂停报道特朗普

澎湃新闻
· 办公效率,政务,招聘HR,模型发布,产品更新,版权诉讼

言短意长|机关食堂开放:定向是智慧,克制是边界

澎湃新闻
· 政策监管,办公效率,医疗健康,教育学习,政务,招聘HR,财报,版权诉讼,论文

OpenAI新举措:允许第三方机构在模型训练、开发阶段进行安全评估

华尔街见闻OpenAI周二表示,将允许外部机构在模型训练、评估及部署的全流程中开展技术安全评估,而非仅限于发布前的例行审查。OpenAI负责外部安全评审工作的Lama Ahmad表示,对于涉及最敏感内容的评估工作,可能邀请外部评估人员进入其办公室开展工作。
· 政策监管,OpenAI,办公效率,模型发布

办公Agent不再比谁会做PPT,开始争夺“谁最懂这家公司”

华尔街见闻当AI从"桌面工具"变成"有工牌的同事",真正的竞争才刚开始。办公Agent核心战场是企业上下文——谁能让AI长期驻留组织、读懂业务规则、记住历史决策,谁才握有下一轮护城河。模型可以替换,公司记忆无法一夜重建。
· AI应用,Agent智能体,办公效率
AI 资讯

发布“企业上下文”,千问办公终于定义清楚了Ag 10:35

网易科技
2026-09-23T10:35:14+08:00 · 阿里巴巴,办公效率,模型发布
AI 资讯

对话千问办公副总裁束骏亮:传统办公软件会退守 12:28

网易科技
2026-09-23T12:27:43+08:00 · 阿里巴巴,对话助手,办公效率
AI 资讯

千问办公发布新AI硬件QwenNote A2,CEO陈宇森:可能卖出千万台

新浪科技
2026-09-22T14:30:58+08:00 · 大模型,阿里巴巴,办公效率,模型发布

千问办公把“企业上下文”做成了产品

品玩
· 阿里巴巴,办公效率

Video Artificial intelligence and sports - ABC News - Breaking News, Latest News and Videos

Google News AI (英文)Video Artificial intelligence and sports  ABC News - Breaking News, Latest News and Videos
· Google,办公效率,招聘HR
AI 资讯

Gap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy

arXiv cs.LGarXiv:2609.26508v1 Announce Type: new Abstract: Streaming principal component analysis (PCA) seeks to recover a leading spectral subspace in a single pass over a data stream. We give a new analysis of the ubiquitous Oja's algorithm [Oja82] for the most general, gap-free variant of this problem, where no eigengap assumptions are made on the underlying mean matrix, complemented by a nearly-matching lower bound. Prior works achieving near-optimal rates for streaming PCA either required gap assumptions [JJK+16, HNWW21], or were limited to rank-one updates [AZL17, Lia23]. Our proof only uses a second moment bound on the individual stochastic updates, bypassing the almost sure bounds needed by prior near-optimal analyses, and the analogous offline matrix Bernstein bound. We also extend our result to a Rayleigh quotient notion of approximate PCA, addressing an open question of [JJK+16]. As our main application, we give gap-free differentially private PCA guarantees for sub-Gaussian data, settling Conjecture 1.1 of [Bro26] up to logarithmic factors.
2026-09-23 04:00:00 · 办公效率,论文
AI 资讯

Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs

arXiv cs.CLarXiv:2509.13813v3 Announce Type: replace Abstract: Large language models are known to hallucinate, generating linguistically plausible but incorrect answers to questions. Uncertainty quantification has been proposed as a strategy to detect such behaviour, but existing methods lack a unified framework to assess reliability at both the prompt and answer level. We introduce a geometric framework which quantifies language model uncertainty at both levels by explicitly modelling a prompt-conditioned semantic distribution in answer embedding space. Our approach is black-box and sampling-based; we generate multiple answers per prompt, and use archetypal analysis to estimate a geometric support for the answer distribution. At the prompt level, we approximate the distribution entropy to quantify uncertainty; for each individual answer, we then use notions of atypicality to assess its reliability relative to the batch. We employ our framework to not only detect hallucinations but correct them, by selecting the batch example deemed most reliable. Experiments show that our framework performs comparably to or better than prior methods on short form question-answering datasets, and achieves superior results on medical datasets where hallucinations carry particularly critical risks. Beyond pure performance, we suggest the theoretical grounding of our work provides support for semantic distributions as useful objects of study for language model uncertainty.
2026-09-23 04:00:00 · 大模型,办公效率,扩散模型,强化学习,向量数据库,提示工程,模型安全对齐,论文
AI 资讯

Semantic Abstraction for Natural Language Inference: a Methodological Framework for Discovering and Compensating Semantic Knowledge and Reasoning Gaps in Large Language Models

arXiv cs.CLarXiv:2609.26610v1 Announce Type: new Abstract: Despite their outstanding performance on many NLP tasks, LLMs face serious challenges related to semantic abstraction. In this study, we are interested in understanding how LLMs leverage abstract semantic knowledge in natural language inference (NLI), which requires sophisticated linguistic capabilities to interpret implicit meanings, contextual conceptual relationships, and semantic connections between words and phrases. To this end, we propose a methodological framework for constructing new semantic knowledge at a higher level of abstraction, which we define under the notions of semantic compatibility and incompatibility for NLI. In this framework, the meaning of the lexical-semantic relations between the premise and the hypothesis is reconfigured to achieve a more flexible semantic network that induces different reasoning paths in LLMs. These new pathways show a consistent pattern of responses that allows agreement on a single response. The results demonstrate that our proposal allows to discover and compensate for LLMs' semantic knowledge gaps in NLI, achieving significant improvements in accuracy, exceeding 10% for some models, and in particular for the non-entailment class. It is essential to note that LLMs need structured knowledge and not just more data to bridge reasoning gaps. Our hybrid approach directs attention to overlooked word relationships, allowing models to synthesize missing information. We believe that the future lies not in increasing model size, but in creating a semantic scafolding that mimics the flexibility of human thinking. Hopefully, our proposal will enable the development of more robust agents and interpretable reasoning, guiding AI toward reliable language understanding.
2026-09-23 04:00:00 · 大模型,AI应用,具身智能,Agent智能体,推理思考,搜索RAG,办公效率,Transformer,招聘HR,论文
AI 资讯

Optimizing Denoising Trajectories in dLLMs: A Lightweight Evolutionary Heuristic Approach

arXiv cs.CLarXiv:2609.26052v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to conventional Auto-Regressive (AR) Large Language Models (LLMs). By leveraging bidirectional attention and parallel decoding, dLLMs enable more efficient generation. However, they require a carefully designed denoising scheduler at inference time (absent during training) whose choice significantly impacts generation quality. While confidence-based heuristic schedulers have shown strong empirical performance, they suffer from two critical failure modes: EOS Overflow and Proximal Bias. Through in-depth analysis of the Transformer's attention patterns, we reveal that these failures stem from certain positions assigning disproportionately high attention weights to invalid tokens (e.g., [MASK] and [EOS]), which produce misleading confidence signals. Building on this insight, empirical evidence shows that valid attention scores can provide complementary guidance to conventional confidence-based heuristics, yet no single metric consistently excels across all scenarios, implying that the optimal denoising trajectory is highly context-dependent. To address this problem, we propose a lightweight evolutionary heuristic scheduler optimized using the Covariance Matrix Adaptation Evolution Strategy (CMA-ES). Our scheduler dynamically integrates multiple heuristic features with a contextual mean-field embedding, while requiring only 393 trainable parameters. Evaluated on LLaDA and Dream across four reasoning and planning benchmarks, our method consistently outperforms strong baselines, including conventional heuristics, block auto-regressive methods, and recent State-Of-The-Art (SOTA) approaches. To the best of our knowledge, it represents the most parameter-efficient neural scheduler to date. Our code is available at https://github.com/RS2002/Evo-Denoise .
2026-09-23 04:00:00 · 大模型,AI应用,开源,推理思考,搜索RAG,办公效率,Transformer,扩散模型,模型评测,向量数据库,招聘HR,论文
AI 资讯

Navigating Taxonomic Expansions of Entity Sets Driven by Knowledge Bases

arXiv cs.AIarXiv:2512.16953v3 Announce Type: replace Abstract: Recognizing similarities among entities is central to both human cognition and computational intelligence. Within this broader landscape, Entity Set Expansion is one prominent task aimed at taking an initial set of (tuples of) entities and identifying additional ones that share relevant semantic properties with the former, potentially repeating the process to form increasingly broader sets. However, this ``linear'' approach does not unveil the richer ``taxonomic'' structures present in knowledge resources. A recent logic-based framework introduces the notion of an expansion graph: a rooted directed acyclic graph where each node represents a semantic generalization labeled by a logical formula, and edges encode strict semantic inclusion. This structure supports taxonomic expansions of entity sets driven by knowledge bases. Yet, the potentially large size of such graphs may make full materialization impractical in real-world scenarios. To overcome this, we formalize reasoning tasks that check whether two tuples belong to comparable, incomparable, or the same nodes in the graph. Our results show that these tasks are intractable in general, that they remain intractable when the number of input tuples is bounded, and that they become solvable in polynomial time when the descriptions of the entities are small as well. The bounds we establish are tight. This enables local, incremental navigation of expansion graphs, supporting practical applications without requiring full graph construction.
2026-09-23 04:00:00 · 推理思考,办公效率,扩散模型,强化学习,端侧AI,论文
AI 资讯

Destination Support Restoration for Finite-Set Multimodal Trajectory Prediction

arXiv cs.AIarXiv:2609.25942v1 Announce Type: cross Abstract: Robots operating around pedestrians often reason over a finite set of predicted human futures. Repeated online updates can concentrate this limited prediction budget on dominant destinations and leave plausible alternatives underrepresented or absent, removing those alternatives from the finite representation available to downstream decision making. We introduce Destination Support Restoration (DSR), a causal post-selection operator that repairs destination support without retraining the host predictor or increasing the maintained set size. At a repair step, DSR evaluates a temporary destination-stratified candidate bank from the observed prefix, converts candidate evidence into integer target counts, protects representatives of active modes, and reallocates redundant surplus hypotheses to deficient modes. The maintained and returned sets retain exactly $N$ hypotheses, and DSR replaces at most $\lceil\rho N\rceil$ entries. Protected representatives preserve current categorical support; lineage-aware particle filters also preserve surviving resampling ancestors. Each replacement reduces the allocation mismatch to the evidence-driven target by one. On the complete 3,719-trajectory Edinburgh protocol over three seeds, DSR reduces MIF weighted ADE and FDE by 13.36% and 13.30% at $N=64$. Paired integrations with CLiFF, PPT, causal GDTS, Social Informer, and PECNet improve both metrics in every evaluated pair. These results show that finite-set support allocation is a useful prediction-side control point when a fixed hypothesis set serves as the interface to downstream systems.
2026-09-23 04:00:00 · 具身智能,多模态,办公效率,强化学习,招聘HR,网络安全,论文
AI 资讯

Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

arXiv cs.AIarXiv:2609.25643v1 Announce Type: new Abstract: Large language models (LLMs) excel at reasoning when scaled to hundreds of billions of parameters, but small- and mid-scale models remain brittle reasoners even with knowledge distillation (KD). We present Ladders-of-Thought (LoT), a framework that improves reasoning by combining progressive question rewrites with a self-evolving curriculum. LoT automatically generates semantically faithful but easier variants of reasoning problems, organizes them into difficulty buckets using step-based measures, and employs a self-evolving bandit scheduler to allocate training adaptively. Evaluated on two reasoning domains, math and multi-hop reasoning, across 1-8B models from different families, LoT consistently improves over KD. It delivers large gains on arithmetic tasks (e.g., +32 percentage points on AddSub, +25pp on SVAMP), +2-8pp improvements on in-domain test splits, and strong though dataset-dependent benefits on multi-hop reasoning (e.g., +16pp on QASC, +25pp on StrategyQA). LoT also converges faster than staged curricula, highlighting the value of adaptive progression. These results show that progressive rewrites coupled with adaptive curricula provide a simple yet effective recipe for strengthening reasoning in smaller LLMs.
2026-09-23 04:00:00 · 大模型,推理思考,办公效率,扩散模型,微调蒸馏,论文
AI 资讯

From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought

arXiv cs.AIarXiv:2609.25366v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is only meaningful if written reasoning causally constrains the answer. We introduce continuation-based causal testing, an ablation-patch intervention that perturbs one reasoning step, truncates the chain, and forces the model to continue from the corrupted prefix. It measures how load-bearing a CoT is for the final answer, a behavioral notion distinct from mechanistic faithfulness. Across Gemma-2-9B-IT, Llama-3.1-8B-Instruct, and DeepSeek-R1-Distill-Qwen-7B on GSM8K, MMLU, and BIG-Bench Hard, CoT load-bearingness tracks model-relative task difficulty: on easy tasks models silently bypass their own reasoning; on hard tasks they follow corrupted steps and propagate errors. A matched 2x2 analysis shows task difficulty dominates perturbation type: error propagation rises 16x from GSM8K to BBH multistep arithmetic, and a variance partition over 28,584 continuations attributes 98.8% of explained deviance to task difficulty versus 0.8% to perturbation type. Reasoning-specific RL suppresses error propagation and compresses the gradient. A four-variant judge-sensitivity analysis and blind two-annotator study (n=500) show the error-propagation vs. non-propagation label is invariant to judge prompt, with perfect inter-annotator agreement (Cohen's kappa = 1.00). This gradient creates a structural problem for CoT-based oversight and AI safety monitoring: where the trace is easy to read it carries little signal, and where it matters errors propagate before a monitor can intervene. Linear probes on hidden states separate silent bypass, self-correction, and error propagation, but additive activation steering provides limited causal control, flipping only about 25% of error-propagation cases at best. Behavioral mode is readable but not reliably controllable.
2026-09-23 04:00:00 · 大模型,Meta,阿里巴巴,DeepSeek,推理思考,办公效率,扩散模型,微调蒸馏,模型评测,提示工程,论文

‘I have a big decision to make’: Trump had a ‘good meeting’ with Iranian officials warning he may ‘annihilate the Islamic Republic’

Fortune

President Donald Trump said U.S. and Iranian officials met on Tuesday, shortly after threatening in his address to the U.N. General Assembly that he may “annihilate” the Islamic Republic if a deal isn’t reached soon to end the nearly seven-month war.

Trump confirmed the talks in an exchange with reporters but offered few details about how the discussions transpired or the substance of the exchange..

“It was a very good meeting,” Trump said.

Earlier Tuesday, Trump in his address to the General Assembly made the case that he’s shown fortitude while other leaders have equivocated as he defended his decision to start a war with Iran.

But Trump also suggested that he could be at a crossroads in the conflict even as he called on the Islamic Republic to fully reopen the Strait of Hormuz and return to the negotiating table to find a settlement to end the conflict.

“I have a big decision to make,” Trump said. “Will a deal be made with Iran that lets them rebuild and create a far greater country than it ever was before, maybe one of the greatest in the Middle East or even the world. Or do I annihilate the Islamic Republic and do it quickly, never giving them a chance to kill and destroy people and countries again?”

The president’s comments came in his wide-ranging address to the international body in which he also celebrated this year’s U.S. military operation to oust Venezuelan President Nicolás Maduro from power, a new agreement to expand the U.S. military presence in Greenland, and pushed back against calls for greater regulation of the artificial intelligence industry.

But at the heart of the address was a fulsome defense of the war and pushback on the notion that Iran has shown a measure of resilience.

“The terrorist regime is behaving badly, not because they are strong and confident, but because they are weak and desperate,” Trump said.

Trump also accused Iran of stalling to make a deal to fully reopen the Strait of Hormuz and negotiate an endgame to the war until after the midterm elections in the United States. He insisted the tactic would have no effect on how he approaches the conflict.

Most of Iran’s delegation walked out of the chamber as Trump railed against the Islamic Republic.

Trump’s speech began a busy day of diplomacy

After his speech, Trump gathered with Danish Prime Minister Mette Frederiksen and Greenlandic Prime Minister Jens-Frederik Nielsen on the sidelines of the General Assembly to sign a new security agreement.

The deal, announced last week, came together after Trump repeatedly threatened to “take over” Greenland, a territory of NATO-ally Denmark, by military force if necessary. Trump said Tuesday the U.S. will immediately begin the process of developing a larger military presence in Greenland including “building two very major military bases.”

Under the deal, Greenland remains Denmark’s territory.

Trump also held talks with new British Prime Minister Andy Burnham. The leaders said they discussed migration, the conflict in Iran, the Russia-Ukraine war and more.

Trump also met with Japanese Prime Minister Sanae Takaichi and Ukrainian President Volodymyr Zelenskyy, and will hold a more informal meeting with Venezuela’s acting President Delcy Rodríguez, according to the White House.

He’s also scheduled to meet with leaders and senior officials from Gulf Cooperation Council nations — Bahrain, Kuwait, Oman, Qatar, Saudi Arabia, and United Arab Emirates — to discuss the way ahead with the war in Iran and growing concern about recent Iran-backed rebel attacks on Saudi Arabia. Trump also plans to take part in an event on the Shield of Americas, an initiative Trump launched with Latin American countries earlier this year focused on combatting violent cartels.

Iran war looms large at General Assembly

The visit to the U.N. comes when Trump’s foreign policy approach — particularly his decision to start a war with Iran — has become a drag on his fellow Republicans’ hopes of retaining control of both chambers of Congress during the coming midterm elections. Nevertheless, Trump made the case in his speech that the world is better off for his leadership.

“As president of the most extraordinary nation in history,” Trump said. “I have no interest in leaving dangers to grow for another day. I do not believe in letting problems fester.”

In last year’s address at the U.N., Trump claimed the U.S. had “demolished” Iran’s nuclear capabilities after he ordered the June 2025 bombardment of three key facilities.

He returned to the General Assembly nearing the seven-month mark of a new war against Iran that he insists is necessary to address a nuclear threat posed by Iran.

The war has grown more complicated even though strikes by the United States and Israel have killed much of Iran’s top leadership and destroyed its air force and navy, while a U.S. blockade of Iranian oil has left Tehran in economic shambles.

The energy market contin

2026-09-22 18:59:12 · AI应用,Google,搜索RAG,办公效率,招聘HR,榜单评测
继续滚动加载更多…