ArXiv Domain 2026-09-08
数据来源:ArXiv Domain
LLM Domain Papers
1. How Much Does Corpus Choice Change Dependency-Distance Estimates?
Abstract:Dependency-distance estimates derived from a single corpus are routinely treated as properties of a language, yet this assumption has not been tested across independently compiled corpora. We compared mean dependency-distance estimates across 38 same-language treebank pairs from Universal Dependencies v2.18, using concordance correlation, Bland-Altman analysis, and a twelve-specification multiverse design. Cross-treebank agreement was moderate at best: substituting one treebank for another reversed nearly 40 percent of pairwise language orderings, and treebank choice accounted for roughly 29 percent of between-group variance. This disagreement substantially exceeded within-treebank sampling error and persisted across all twelve preprocessing specifications. Nevertheless, every treebank confirmed dependency-length minimization (normalized ratio below 1). The data are more consistent with MDD as a corpus-conditioned composite of grammatical, register, and annotation factors than as a stable language-level parameter: the qualitative DLM universal survives corpus substitution, but the ordinal cross-linguistic ranking does not.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04223 (HTTP 429)
Authors: Sirui Chen
Categories: cs.CL
PDF URL: https://arxiv.org/pdf/2609.04223.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04223
Published: 2026-09-08T01:23:18.870Z
2. Memory as transformation: LETHE, a self-referential gan-inspired architecture
Abstract:LETHE (Latent-parameter Evolution with Temporal Hierarchical quasi-Equilibrium) is a self-referential sonic-oblivion system implemented in SuperCollider. It adopts the formal vocabulary of Generative Adversarial Networks in a closed configuration without external datasets or supervision after initialization. Audio is processed by a 3 x 3 mixing matrix built around two delay lines; its nine coefficients and two delay times evolve through the interaction of a five-feature linear discriminator and a random-perturbation optimizer analogous to single-sample REINFORCE. The discriminator compares current energy behavior with an archive of the initial state and guides parameter updates. Circular, fixed, and live sources can be mixed independently. Across fixed and circular sessions with an ablation control, the active generator is necessary for parametric evolution ($\Delta c_{22}=0.000$ in all 15 ablation sessions). Situated in the tradition of self-referential electroacoustic music, LETHE delegates the sonic outcome to an adaptive closed loop whose parametric space is defined by the composer.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04289 (HTTP 429)
Authors: Francesco Vitucci, Anthony Di Furia, Francesco Scagliola
Categories: cs.CL
PDF URL: https://arxiv.org/pdf/2609.04289.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04289
Published: 2026-09-08T01:23:18.870Z
3. Evidence Integration in Large Language Models
Abstract:Despite increasing reliance on LLMs that reason with external evidence supplied by tools, retrieval-augmented generation, other agents, and users, how LLMs integrate such evidence into decisions they have already begun to form remains largely unclear. We present a distributional theory in which evidence shifts the receiver’s distribution of initial answers, driven by a receiver prior weight and a candidate evidence tilt, leading to three predictions. First, candidates more probable to the receiver are more persuasive. Second, receivers more readily integrate characteristic errors of their own than foreign errors from different sources. Third, identical evidence can improve weaker models and harm stronger ones. We confirm these over ten million trials, twelve LLMs from four families, and eight domains, four of them scientific discovery tasks in the physical and life sciences: quantum mechanics, physics, genetics, and molecular biology. The law also yields a receiver-relative reliability frontier: receiver-congruent errors depress performance more steeply than random errors of the same rate. LLMs also integrate candidates even after internally verifying their invalidity (93-100% with propositional constraints; up to 99.4% on held-out physical and life-sciences reasoning), demonstrating evidence integration is a receiver-specific control policy over existing distributions, determined by receiver properties rather than scalar trust in the evidence source. Causal interventions show candidate integration is implemented late in the network, as a structured sequence of steps admitting external candidate answers, promoting them, and transporting them into the answer state. Representations of verification are decodable but have little causal impact on answers. A J-lens decomposition shows the state underlying verbalized verification is fully dissociable from that underlying candidate integration.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04290 (HTTP 429)
Authors: Sebastien Kawada, Manolis Kellis
Categories: cs.CL
PDF URL: https://arxiv.org/pdf/2609.04290.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04290
Published: 2026-09-08T01:23:18.870Z
4. MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering
Abstract:Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with \textbf{MedProb}, a lightweight probing framework that predicts multiple-choice Med-VQA answers from frozen VLM representations without free-text generation. Across PATH-VQA, SLAKE, and VQA-RAD, MedProb recovers substantially more answer-relevant signal than prompting and performs stronger than medical VLMs and agentic systems. Probing also reduces the apparent gap between small and large models compared to prompting, suggesting that smaller VLMs contain more recoverable Med-VQA signal than generation-based evaluation reveals. Across 14 matched general-purpose and medical VLM pairs, medical adaptation does not consistently improve this linear decodability. Finally, free-text generation exhibits an answer-position bias of up to 10 percentage points, whereas MedProb also has positional bias, however, it is impacted differently than prompting. Our main results target the multiple-choice/multiclass Med-VQA setting; we additionally show the probe can be extended to open-ended generation via a rejection-sampling scoring procedure.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04336 (HTTP 429)
Authors: Erfan Nourbakhsh, Ke Yang, Anthony Rios
Categories: cs.CL
PDF URL: https://arxiv.org/pdf/2609.04336.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04336
Published: 2026-09-08T01:23:18.870Z
5. Adapting from Downturns: Prediction of Long-Term Conversational-Skill Development in Mental-Health Crisis Counselors
Abstract:How do people learn to become better conversationalists? This question is especially important in the context of mental-health counseling, where conversational skills are essential, yet volunteer counselors often have limited access to supervision and structured feedback. Understanding how counselors develop their ability to steer conversations toward positive outcomes — and identifying early which counselors are (not) on track to improve — can help prioritize support for the counselors who need it most. In this work, we introduce the task of predicting, early in a conversationalist’s career, whether they will eventually improve at steering conversations toward positive outcomes, and demonstrate the feasibility of this task in the case of volunteer mental-health crisis counselors. Our central insight is that people may struggle with particular kinds of moments in a conversation, and that what is especially revealing of their likelihood of future improvement is how they learn to handle those moments over time. We operationalize this insight by designing a method that identifies the types of moments a counselor initially struggles with, captures how they adapt their response when they re-encounter similar moments in subsequent conversations, and learns which early adaptations predict improvement months or even years later. While this future-prediction task is challenging, our counselor-adaptation approach yields better results than baselines that learn directly from the conversation transcript.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04350 (HTTP 429)
Authors: Vivian Nguyen, Lillian Lee, Elizabeth A. Olson, Cristian Danescu-Niculescu-Mizil
Categories: cs.CL
PDF URL: https://arxiv.org/pdf/2609.04350.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04350
Published: 2026-09-08T01:23:18.870Z
6. VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes
Abstract:Early-onset colorectal cancer is increasing among younger adults, yet red-flag symptoms in this age group have no evidence-based guidelines for follow-up testing, and structured encounter data do not capture the detail needed to support early detection and inform follow-up, including symptom duration, context, and fam- ily history, an established colorectal-cancer risk factor. This study aimed to develop and evaluate an automated method for extracting six red-flag symptoms and family-history risk status from free-text clinical notes. We developed VERGE, an agentic workflow in which an initial label and evidence are proposed using retrieval-augmented generation, then passed through a bounded verification- refinement cycle that checks textual grounding and clinical validity, corrects and rechecks a claim until resolved or a limit is reached, and escalates unresolved claims for human review. VERGE was evaluated on 4,033 clinician-labeled note-finding pairs against a single-agent baseline, a rule-based clinical language-processing baseline, and an alternative underlying language model. Compared with the single-agent baseline, VERGE reduced false positive find- ings, improving precision from 0.764 to 0.849 and MCC from 0.681 to 0.730, a balanced gain across the precision-recall trade-off, and resolved most flagged errors autonomously, with human review required for only 1.5 percent of claims. These results indicate that a bounded, verification-based workflow can reduce unnecessary positive findings without sacrificing the ability to detect true ones. This approach offers a path toward more reliable and trustworthy clinical language-processing tools to support colorectal cancer risk assessment in younger patients.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04366 (HTTP 429)
Authors: Nikkie Hooman, Monarch Nigam, Amy E. Hughes, Rasmi G. Nair, Mehak Gupta
Categories: cs.CL
PDF URL: https://arxiv.org/pdf/2609.04366.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04366
Published: 2026-09-08T01:23:18.870Z
7. You Really Didn’t Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments
Abstract:Chinese online comments often convey social meaning through indirect and playful language that is hard to interpret without context. Existing evaluations largely organize items around predefined phenomena or controlled pragmatic categories, leaving open whether models can distinguish plausible readings of what a naturally occurring comment is doing in a particular exchange. We introduce a benchmark for evaluating whether LLMs can recover such situated pragmatic meanings. From more than 200,000 public Chinese social media interaction records, we construct 4,735 human-validated diagnostic items, each pairing a target comment with reconstructed preceding context and plausible misreadings. We evaluate eight LLMs as both question writers and solvers in a cross-writer setting. The task is challenging: the strongest model achieves 81.42% leave-writer-out accuracy. Across all eight models, the mean leave-writer-out accuracy is 68.70% while human accuracy was 90.8%. Case analysis shows that models often recognize broad irony or playfulness while misidentifying the mechanism or interactional move.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
Q: 这篇论文试图解决什么问题? 该论文试图解决的核心问题是:\*\*当前大语言模型(LLM)基准测试难以充分评估模型对中文网络评论中情境化、间接且带有戏谑性质的社交语用意图的精细推理能力。\*\* 具体而言,论文针对以下 gaps 展开: - \*\*现有评估的粒度不足\*\* 当前的讽刺、毒性或隐式辱骂检测基准大多围绕预定义现象(如“是否为讽刺”)或受控语用类别(如预设、指称)组织。这些任务只能测试模型是否识别出某个标签,却无法检验模型是否真正理解一条自然评论在\*\*特定交流场景\*\*中的具体功能。 - \*\*“识别大类”不等于“理解意图”\*\* 两条解读可能都意识到某条评论含有讽刺或幽默,却在关键维度上截然不同——例如\*\*嘲讽对象是谁\*\*、\*\*说话者立场如何\*\*、\*\*哪处语言机制产生了效果\*\*、或者该评论在互动中执行的\*\*社交动作\*\*是什么。现有基准未充分测试模型区分这些竞争性、情境合理误读的能力。 - \*\*中文网络 discourse 的特殊复杂性\*\* 中文在线社区广泛使用谐音、字符级操控、压缩性典故和平台特定梗(即“阴阳怪气”类表达),其含义高度依赖前置上下文、互动序列和共享文化知识。这要求模型整合局部语言形式与话语级证据,而现有资源未能系统诊断这一点。 为此,论文构建了一个包含 \*\*4,735 个人工验证诊断项目\*\* 的基准,通过重建前置上下文、让多个 LLM 生成并交叉验证诊断问题,测试模型能否在五个选项中辨识出一条评论在特定互动中的\*\*真实 situated pragmatic meaning\*\*,而非仅识别宽泛的反讽或玩味。 Q2: 有哪些相关研究? 根据论文第2节(Related Work),相关研究可沿以下三条主线梳理: ### 1. 非字面意义与社会性意义的计算基准 现有工作主要围绕预定义现象或受控类别构建评估: - \*\*讽刺识别\*\*:FigLang 共享任务(Ghosh et al., 2020)、中文讽刺数据集(Gong et al., 2020)以及 SarcasmBench(Zhang et al., 2025a)等,核心目标是判断文本是否应被标记为讽刺。 - \*\*隐式伤害与偏见\*\*:Social Bias Frames(Sap et al., 2020)建模隐含陈述与受影响群体;HateCheck(Röttger et al., 2021)及隐式辱骂研究(Wiegand et al., 2021)则针对常规分类器易漏检的功能性案例。 - \*\*更细粒度的中文评估\*\*:ProsCons 联合建模立场、情感、冒犯性与讽刺(Wang et al., 2023);CFMS 在中文多模态讽刺检测中加入目标识别与解释(Zhang et al., 2026);Chumor 2.0 评估文化特异性中文幽默的解释(He et al., 2025)。 - \*\*编码与亚文化语言\*\*:针对中文网络社群的谐音、字符操控及黑话(Wan et al., 2026; Lin et al., 2026)。 与上述工作不同,本文强调:\*\*识别“讽刺/幽默/攻击性”大类并不足够\*\*,模型必须对自然评论在特定交流片段中的\*\*整合性情境解读\*\*保持连贯。 ### 2. 语用与话语理解中的上下文 该脉络确立了上下文的重要性,但大多用上下文去\*\*解析或预测某个预定义语用判断\*\*: - \*\*讽刺中的上下文比较\*\*:FigLang 推动了社交媒体讽刺数据的标准化(Ghosh et al., 2020);后续工作直接比较长、短、无上下文条件下的讽刺判断及其与人类分歧的交互(Jang et al., 2024)。 - \*\*多轮对话中的比喻语言\*\*:FLUID QA 评估英语、中文、韩语多轮对话中比喻表达的适用性(Park et al., 2025)。 - \*\*语用能力基准\*\*:PUB 通过受控多项选择测试蕴含、预设、指称与指示(Sravanthi et al., 2024);SwordsmanImp 以脚本化多轮中文对话测试蕴含识别与解释(Yue et al., 2024)。 - \*\*跨语言间接言语行为\*\*:在更多语言与场景中测试间接言语行为(Koo et al., 2025; Orsini & Brunato, 2025)。 本文的区分点在于:\*\*不将上下文用于推断某个固定语用标签,而是将重建的互动片段本身作为解释单元\*\*,要求模型区分对同一条评论在特定片段中“做什么”的竞争性解读。 ### 3. 诊断式基准构建方法 该领域强调基准应暴露具体的失败机制,而非仅测试整体准确率: - \*\*对抗性与对比集\*\*:Contrast sets 通过小规模但意义明确的编辑暴露局部决策边界(Gardner et al., 2020);ANLI 要求注释者编写能愚弄当前模型但仍可被人类回答的推理样例(Nie et al., 2020);AdversarialQA 在问题编写阶段引入阅读理解模型作为对抗者(Bartolo et al., 2020)。 - \*\*动态评估框架\*\*:Dynabench 将数据收集、模型评估与模型开发相互迭代(Kiela et al., 2021)。 - \*\*高质量干扰项生成\*\*:评估研究强调,合理的错误选项应揭示误解或部分理解,而非充当任意错误答案(Alhazmi et al., 2024; Tian et al., 2026; Lee et al., 2025)。 本文借鉴了上述“人在回路”的诊断思路,但针对\*\*缺乏外部客观锚点的情境化社交语用解读\*\*这一特殊问题,提出了分离“人工构建互动片段—跨模型生成候选误解—人工验证诊断对比”的构建协议。 Q3: 论文如何解决这个问题? 该研究通过\*\*构建诊断式基准(diagnostic benchmark)\*\*并辅以\*\*跨模型生成与验证协议\*\*来解决上述问题。整体方案可分为数据构建、题目生成与验证、以及互补式模型评估三个层面。 --- ### 1. 数据构建:从自然互动中重建语境片段 - \*\*多平台 purposive 采集\*\* 研究团队从知乎、豆瓣、小红书、贴吧四个平台采集了超过 20 万条公开互动记录,覆盖社会争议、性别、教育、经济等话题,以确保不将中文网络语用简化为单一平台的互动风格。 - \*\*人工筛选语境片段(Contextual Episodes)\*\* 注释者依据四项标准筛选候选评论: 1. \*\*语境依赖性(Context dependence)\*\*:评论脱离上下文无法理解; 2. \*\*可恢复性(Recoverability)\*\*:存在可被稳定推断的社交语用解读; 3. \*\*非字面/戏谑性(Nonliteral expression)\*\*:包含间接、谐音、夸张、回声、典故等表达; 4. \*\*合理误解路径(Plausible misunderstanding)\*\*:存在模型可能产生的竞争性误读。 最终保留 3,200 个人工重建的互动片段,每个片段包含原帖、父评论及局部互动序列等前置语境。 --- ### 2. 题目生成与验证:跨模型出题 + 人工校验 该研究的核心创新在于\*\*将 LLM 同时用作“命题者”与“受试者”\*\*,并通过人工校验确保题目质量。 - \*\*阶段一:交叉模型诊断题生成\*\* 研究使用 8 个 LLM(GPT-5.5、Claude Opus 4.7、DeepSeek V4 Pro、Qwen 3.5、Kimi K2.6、MiniMax M2.5、Mistral Medium 3、Llama 4 Maverick)作为候选题目撰写者(writer)。每个互动片段被轮换分配给 2 个模型,共生成 6,400 道候选五选一问题。 撰写要求聚焦于\*\*具体的解读难点\*\*,例如:嘲讽目标归因、说话者立场对齐、修辞机制、互动功能或关键语境线索,而非简化为“是否讽刺”或“情感正负”。 - \*\*阶段二:人工验证与修订\*\* 7 名注释者对 6,400 道候选题进行筛选: - \*\*剔除\*\*无需语境即可作答、基于幻觉细节、选项模糊、或过度泛化的题目; - \*\*修订\*\*诊断对比有价值但表述不清、选项不均衡的题目; - \*\*确认\*\*唯一最佳答案(gold answer)并撰写理由。 最终得到 \*\*4,735 道\*\*经验证的诊断题目,其中 74.0% 为直接保留或可修复题目。 --- ### 3. 评估机制:留作者交叉验证(Leave-Writer-Out) 为避免模型因熟悉自身出题风格而获益,研究以\*\*留作者交叉验证(LWO)准确率\*\*作为主要指标: - \*\*答题者评估(Solver Evaluation)\*\* 同一批 8 个 LLM 也作为答题者,但评估时排除由该模型自身撰写的题目。结果显示:最强模型(Qwen 3.5 9B)的 LWO 准确率为 81.42%,八模型平均仅 68.70%,而人类准确率达 90.8%( kappa ≈ 0.85 ),表明基准具有充分区分度。 - \*\*双重能力分离\*\* 研究同时考察模型“出题”与“解题”两种能力。结果发现二者并不重合:部分模型擅长生成高区分度题目,自身解题表现却平庸,反之亦然。这表明\*\*生成诊断性题目\*\*可作为传统准确率之外的互补评估信号。 --- ### 4. 验证与诊断:语境消融与错误案例分析 - \*\*语境消融实验(Context Ablations)\*\* 完整语境下八模型平均 LWO 准确率为 68.70%;若仅保留目标评论而移除所有前置语境,平均骤降至 29.63%,降幅达 39.07 个百分点。逐层删除(原帖、直接父评论、局部互动)进一步证实:\*\*前置互动是语用意义不可或缺的组成部分\*\*。 - \*\*案例驱动的失败模式分析\*\* 通过对错误案例的细读,研究归纳了当前 LLM 的典型失败: - \*\*净化(Sanitizing)\*\*:将带有犬儒或歧视色彩的评论过度解读为亲社会表达; - \*\*误指派修辞机制\*\*:识别出文字游戏,但将实际引用的成语/典故替换为 unsupported 的谐音双关; - \*\*高估戏谑力度\*\*:将轻松的玩味式重构误判为严厉的道德讽刺。 --- ### 总结 该研究并非简单提供一个静态数据集,而是提出了一套\*\*可复现的协议\*\*:人工重建语境片段 → 多模型交叉生成候选误解 → 人工验证诊断对比 → 留作者交叉评估。这一协议使基准能够针对“模型是否能区分情境合理但相互竞争的解读”这一核心难点进行精确诊断。 Q4: 论文做了哪些实验? 论文围绕基准构建与模型诊断能力展开了以下核心实验: ### 1. 交叉作者解题性能评估(Cross-Writer Solver Evaluation) 该实验为基准的核心定量评估。研究将 8 个 LLM(GPT-5.5、Claude Opus 4.7、DeepSeek V4 Pro、Qwen 3.5 9B、Kimi K2.6、MiniMax M2.5、Mistral Medium 3、Llama 4 Maverick)同时作为\*\*题目撰写者(writer)\*\*与\*\*答题者(solver)\*\*,构成 8 × 8 的解题准确率矩阵。 - \*\*主要指标\*\*:留作者交叉验证(Leave-Writer-Out, LWO)准确率,即答题者在其\*\*未参与撰写\*\*的题目上的正确率,以避免模型因自身措辞风格而获益。 - \*\*整体结果\*\*:八模型平均 LWO 准确率为 \*\*68.70%\*\*;其中 Qwen 3.5 9B 表现最强,达 \*\*81.42%\*\*,而 Claude Opus 4.7 最低,为 \*\*59.70%\*\*。 - \*\*自写自答偏差\*\*:对角线(self-writer)准确率普遍显著高于 LWO 准确率。例如 GPT-5.5 自写自答达 \*\*97.34%\*\*,但 LWO 仅 \*\*62.95%\*\*;Claude Opus 4.7 自写自答 \*\*89.82%\*\*,LWO 仅 \*\*59.70%\*\*。Kimi K2.6 是例外,其自写自答仅 \*\*31.82%\*\*。 - \*\*题目来源效应\*\*:不同 writer 生成的题目难度差异显著。Qwen 与 Claude 撰写的题目对其他模型较易(多数 solver 准确率 > 89% ),而 Llama、DeepSeek、Mistral 的题目则明显更难(部分 solver 准确率 < 50% )。 - \*\*人工修订子集难度\*\*:在 343 道经人工实质性修订的题目上,所有模型表现大幅滑落,区间仅为 \*\*30.49%–39.63%\*\*(Qwen 3.5 9B 最高),表明人工修订有效剔除了表面线索,构成高难度诊断子集。 ### 2. 人工可回答性验证(Human Answerability Verification) 为验证基准题目对人类而言是否可回答且答案稳定,研究进行了两轮人工审计: - \*\*最终基准抽样\*\*:随机抽取 300 道最终验证题目,两名注释者独立作答。人工准确率达 \*\*90.8%\*\*,Cohen’s kappa ≈ 0.85 ,表明题目整体具有良好的可回答性与注释者一致性。 - \*\*人工修订子集\*\*:针对 343 道经人工修订的题目,人工准确率降至 \*\*64.14%\*\*,Cohen’s kappa ≈ 0.53 。这说明人工修订成功识别并强化了真正模棱两可或细微区分的难点,而非造成题目不可解。 ### 3. 上下文消融实验(Context Ablation Studies) 为检验基准是否真正依赖前置语境,而非仅通过评论本身或选项结构即可破解,研究设计了系统性的语境剥离实验: - \*\*全语境 vs. 仅评论(Comment-Only)\*\* 在 LWO 设置下,八模型使用\*\*完整重建语境\*\*时平均准确率为 \*\*68.70%\*\*;当仅保留\*\*目标评论\*\*而移除所有前置语境(问题与选项保持完整)时,平均准确率骤降至 \*\*29.63%\*\*,绝对下降 \*\*39.07\*\* 个百分点。所有模型均出现大幅衰退,证实基准确实依赖语境。 - \*\*逐层组件消融\*\* 针对 GPT-5.5、Llama 4 Maverick 与 Qwen 3.5 9B,进一步以“每次移除一层语境”的方式检验证据分布: - \*\*移除帖子级语境(post-level)\*\*:分别损失 302、351、289 道正确题(占 4,735 题的 \*\*6.38%\*\*、\*\*7.41%\*\*、\*\*6.10%\*\*)。 - \*\*移除直接父评论(direct parent)\*\*:分别损失 65、56、49 道正确题(占 \*\*1.37%\*\*、\*\*1.18%\*\*、\*\*1.03%\*\*)。 - \*\*移除局部互动语境(local interaction)\*\*:分别损失 191、193、230 道正确题(占 \*\*4.0%\*\*、\*\*4.1%\*\*、\*\*4.9%\*\*)。 该结果表明,各语境层级存在信息重叠,但\*\*原帖/主帖级语境\*\*的缺失对推理损害最大。 ### 4. 案例失败分析(Case Failure Analysis) 研究对模型错误进行质性细读,归纳出三类系统性失败模式: - \*\*净化社交不适意义(Sanitizing)\*\* 模型识别出评论的玩味或间接评价属性后,自动将其“修复”为更符合社会期望的解读。例如将一则借“大正确时代”与 Kardashian 梗对唐氏综合征律师新闻进行犬儒嘲讽的评论,曲解为对当事人成就的肯定。 - \*\*误指派核心文字游戏(Misidentifying Operative Wordplay)\*\* 模型察觉到存在谐音或拆字机制,但错误指定了实际运作的语言单位。例如将拆解成语“干柴烈火”的戏谑用法,误读为 unsupported 的“柴/财”“火/祸”谐音双关。 - \*\*高估戏谑重构的语力(Misestimating Force)\*\* 模型将轻松的、带有 amusement 的玩味性指认(如“我佛只渡有钱人”对宗教签证广告的调侃)过度解读为严厉的道德讽刺或社会批判,混淆了“ playful noticing ”与“ strong satire ”的力度差异。 ### 5. 精确匹配污染审计(Exact-Overlap Audit,附录 F) 为排除模型表现源于训练语料中逐字记忆的可能性,研究将 3,200 个目标评论字符串对 Dolma 开放语料库进行精确匹配查询: - 全部成功查询的评论中,仅 \*\*1.59%\*\*(51/3,198)存在任何精确匹配。 - 若限定目标评论至少包含 6 个中文字符,匹配率降至 \*\*0.26%\*\*(8/3,095)。 该结果说明,大规模逐字记忆不太可能是模型性能或排名差异的主要解释。 Q5: 有什么可以进一步探索的点? 基于论文的研究框架与讨论,以下方向具有进一步探索的价值: - \*\*多模态社交语用推理\*\* 论文明确将自身限定于文本与截图媒介。后续工作可向抖音、B站等视频平台扩展,纳入弹幕、字幕、面部表情、语音语调、剪辑节奏及背景音乐等非文本线索,检验模型在跨模态间接表达(如视频反讽、配音戏谑)中的推理能力。 - \*\*动态化与可复现的评估协议\*\* 网络梗、谐音映射与平台 slang 的半衰期极短。可将本文提出的“人工重建语境—多模型出题—人工验证”协议定期重跑,形成类似 LiveBench 或 Dynabench 的动态诊断流,持续追踪 LLM 对新兴语言变异的理解边界,而非依赖静态数据集。 - \*\*解释生成质量的系统化评估\*\* 当前任务为五选一选择题型,模型可能通过表面线索选中答案而未真正定位误解点。后续可强制模型输出诊断理由(rationale),并建立对人类注释者 rationales 的语义对齐度量,检验模型是否能在生成侧准确指出“何处易误读”与“为何如此解读”。 - \*\*针对语用推理的定向微调与对抗训练\*\* 基准中的高难度子集(如人工修订题目)天然构成训练信号。可探索使用该基准进行指令微调或强化学习,观察模型是否在 target identification、stance alignment、playful mechanism 等维度上产生可解释的能力跃迁,而非仅提升整体准确率。 - \*\*跨语言与跨文化迁移研究\*\* 将同样的“语境重建 + 竞争性误读诊断”协议应用于英语 Reddit、日语 5ch、韩语 X 等平台,比较不同语言文化中间接戏谑表达的共性(如 echoic mention、mock agreement)与特异性(如汉字谐音 vs. 罗马字游戏),检验当前 LLM 的语用能力是否存在系统性的文化偏差。 - \*\*不确定性建模与意义边界研究\*\* 人工修订子集上注释者一致性显著下降(Cohen’s kappa ≈ 0.53 ),说明部分案例本身处于语用意义的“灰色地带”。可专门收集这类 genuinely ambiguous 的互动片段,量化人类判定的置信度分布,并训练模型输出“可回答/不可回答”的判断,而非强行在五选项中选择。 - \*\*检索增强(RAG)与外部文化知识库\*\* 许多误读源于模型缺乏对平台特定典故、近期热梗或圈层用语的实时知识。可探索接入动态更新的网络文化知识库(如梗百科、平台热词追踪),检验 RAG 能否在无需大规模重训的情况下提升对高度本地化、时效性强的间接表达的理解。 - \*\*智能体在线交互中的语用失误代价\*\* 论文指出,若对话智能体误读语境,将在内容审核、多智能体辩论、摘要生成等下游场景中产生 action errors。后续可设计交互式实验,量化特定语用误判(如将恶意戏谑视为无害玩笑,或将 mild teasing 过度判定为攻击)对讨论质量、用户信任或信息传播的具体影响。 - \*\*题目撰写能力作为独立评估维度\*\* 研究发现命题能力与解题能力存在分离(proposing ≠ solving)。可进一步将 item discrimination、writer-model vs. solver-model 的交叉矩阵形式化为系统性的生成侧评估指标,补充现有以答对率为中心的评测范式。 Q6: 总结一下论文的主要内容 该论文针对中文社交媒体中间接、戏谑性评论(即“阴阳怪气”类表达)的\*\*情境化语用推理\*\*问题,构建了一项新的诊断式基准,并系统评估了当前大语言模型(LLM)在此类任务上的能力边界。主要内容可概括如下: ### 1. 研究背景与核心问题 中文网络评论高度依赖上下文、文化典故、谐音梗及平台特定表达,其含义往往并非字面所能直接获得。现有基准多围绕预定义标签(如“是否讽刺”“是否有害”)或受控语用现象展开,难以检验模型是否真正理解一条自然评论在\*\*特定互动片段\*\*中的具体功能——例如嘲讽对象是谁、说话者立场如何、哪处语言机制产生了效果、以及评论执行的社交动作是什么。两条解读可能同时识别出“反讽”或“幽默”,却在上述关键维度上截然不同。 ### 2. 基准构建方法 论文提出了一套可复现的\*\*跨模型出题—人工验证\*\*协议: - \*\*数据采集与语境重建\*\*:从知乎、豆瓣、小红书、贴吧采集超过 20 万条公开互动记录,人工筛选出 3,200 个满足“语境依赖、可恢复、非字面、存在合理误解路径”的互动片段,重建其前置语境(原帖、父评论、局部互动序列等)。 - \*\*交叉模型诊断题生成\*\*:将 8 个 LLM(如 GPT-5.5、Claude Opus 4.7、Qwen 3.5、DeepSeek V4 Pro 等)轮换作为题目撰写者,基于语境与目标评论生成 6,400 道五选一诊断题,聚焦于目标归因、立场对齐、修辞机制、互动功能等具体难点。 - \*\*人工验证\*\*:由 7 名注释者剔除无需语境、基于幻觉、过度泛化或选项模糊的题目,对有用但表述不清的题目进行修订,最终保留 \*\*4,735 道\*\* 经验证的诊断题,并确认唯一最佳答案(gold answer)。 ### 3. 主要实验与发现 - \*\*留作者交叉验证(LWO)\*\*:为避免模型因自身出题风格获益,评估时排除模型自己撰写的题目。八模型平均 LWO 准确率为 \*\*68.70%\*\*,最强模型 Qwen 3.5 9B 达 \*\*81.42%\*\*,而人类准确率为 \*\*90.8%\*\*(Cohen’s kappa ≈ 0.85 ),表明任务对当前 LLM 仍具挑战性。 - \*\*出题≠解题\*\*:模型自写自答准确率普遍远高于交叉验证准确率(如 GPT-5.5 自答 97.34% vs. LWO 62.95%),且部分模型擅长出题却解题表现平庸,反之亦然,说明两种能力并不重合。 - \*\*语境消融\*\*:移除全部前置语境后,平均准确率从 68.70% 骤降至 \*\*29.63%\*\*;逐层删除显示原帖级语境的缺失损害最大,直接父评论影响相对较小,证实\*\*前置互动是语用意义不可或缺的组成部分\*\*。 - \*\*案例失败分析\*\*:模型常能识别宽泛的戏谑或反讽,却在以下方面失败: - \*\*净化(Sanitizing)\*\*:将犬儒、歧视性评论自动“修复”为亲社会解读; - \*\*误指派机制\*\*:察觉到文字游戏,但错误归因谐音或典故来源; - \*\*高估语力\*\*:将轻松的玩味性指认误判为严厉的道德讽刺。 ### 4. 核心结论与贡献 - 提出了一个针对中文间接戏谑评论的\*\*社会语用推理诊断基准\*\*,将评估粒度从“识别大类”推进到“区分情境化竞争性误读”。 - 引入了一套\*\*可动态复现的跨模型构建协议\*\*(人工重建语境 → 多模型生成候选误解 → 人工验证诊断对比),适用于其他缺乏客观锚点的精细解读领域(如法律、医学文本)。 - 揭示了当前 LLM 在真实在线 discourse 中的关键缺口:能够检测表层非字面性,却常误 recover 目标、立场、修辞机制或互动功能;这种语用失误将直接转化为内容审核、摘要生成、智能体交互等下游场景中的\*\*行动错误\*\*。 Q7: 想要进一步了解论文 以上只是了解一篇论文的几个基本FAQ。如果你还想与Kimi进一步讨论该论文,请点击 [**这里**](http://kimi.com/_prefill_chat?prefill_prompt=我们要讨论的论文是You%20Really%20Didn%27t%20Get%20That%3F%20Benchmarking%20Social%20Pragmatic%20Inference%20for%20Indirect%20and%20Playful%20Chinese%20Online%20Comments,链接是https://arxiv.org/pdf/2609.04384,已有的FAQ链接是https://papers.cool/arxiv/kimi?paper=2609.04384。请以此为基础,继续回答我后面的问题。&system_prompt=你是一个学术助手,后面的对话将围绕着以下论文内容进行,已经通过链接给出了论文的PDF和论文已有的FAQ。用户将继续向你咨询论文的相关问题,请你作出专业的回答,不要出现第一人称,当涉及到分点回答时,鼓励你以markdown格式输出。&send_immediately=true) 为你跳转Kimi AI网页版,并启动一个与该论文相关的新会话。
Authors: Shiwei Hong, Junjie Ma, Emma Jiren Wang, Ethan Z. Rong, Siying Hu, Haichang Li, Ziying Wang, Zhicong Lu
Categories: cs.CL
PDF URL: https://arxiv.org/pdf/2609.04384.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04384
Published: 2026-09-08T01:23:18.870Z
8. Evaluation of Phonetic Encoding Algorithms on Transcription Datasets
Abstract:In this work, a novel evaluation scheme built on a generalized variant of the Rand Index measure, namely, the Hüllermeier-Rifqi Index, is proposed in order to assess how well phonetic encoding algorithms conform to word-based transcriptions in IPA (International Phonetic Alphabet) notation. For this objective, the discordance score is obtained by calculating the absolute difference between the pairwise similarity values of ground-truth transcriptions and those of corresponding phonetic encodings, which are computed using normalized edit distance as a permutation dependent string metric. The resulting score is subsequently adjusted with respect to that of a random string generator incorporating the same alphabet as the encoder under consideration. A wide range of phonetic encoders were evaluated as such on multi-lingual transcription datasets along with their recall capabilities based on the collision rate. The validity of the proposed scheme is further supported by its applicability in measuring the orthographic transparency of a language when the writing system is viewed as an inherent phonetic representation.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04391 (HTTP 429)
Authors: Can Özbey, Emre Kaplan, Berkin Deniz Kahya
Categories: cs.CL
PDF URL: https://arxiv.org/pdf/2609.04391.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04391
Published: 2026-09-08T01:23:18.870Z
9. The Anatomy of an ASR Hallucination
Abstract:ASR systems sometimes produce fluent text that is unrelated to the speech they receive. We view these hallucinations as one possible consequence of a broader grounding failure, in which the transcript is no longer adequately guided by the audio. To understand where this failure becomes possible, we study two independently trained Conformer-Large recognizers - one CTC and one RNN-T - under environmental degradation and speaker-background shift. In both models, the final encoder stage emerges as a critical boundary: bypassing the final block causes divergence on nearly every utterance, whereas bypassing middle blocks has little effect. At this same stage, the representations become more compact, text becomes readable by the trained decoder, and grapheme information becomes explicit. Importantly, the intervention produces garbled or repetitive output rather than fluent fabrication. Our result therefore identifies a mechanistic precondition for hallucination - the failure to produce adequately grounded output - not the complete origin of naturally occurring hallucinations. Together, the results reveal a consistent terminal-stage dependency for grounded recognition across two decoder families and multiple distribution shifts.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04404 (HTTP 429)
Authors: Hamees Sayed, Apoorv Singh, Kumar Aman, Akshat Mandloi
Categories: cs.CL
PDF URL: https://arxiv.org/pdf/2609.04404.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04404
Published: 2026-09-08T01:23:18.870Z
10. A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models
Abstract:Multilingual language models often produce inconsistent answers to semantically equivalent questions across languages, motivating methods to improve cross-lingual consistency (CLC). However, existing methods are typically evaluated using different models, tasks, and protocols, leaving their relative strengths unclear. In this work, we present a unified evaluation of representative CLC-enhancement methods for question answering, spanning inference-time interventions and post-training approaches across three model families and three closed-form benchmarks. The results show that post-training methods are generally more reliable, with direct distribution alignment consistently improving CLC across all model-dataset combinations, while other methods are more sensitive to answer format and the breadth of language coverage. Notably, cross-domain transfer is limited unless source and target tasks share similar output formats. We further investigate whether CLC enhancement hurts models’ ability to respond differently when needed, that is, when asked culture-dependent questions. Across two benchmarks of culturally diverse question answering, we find no systematic degradation in controlled closed-form evaluation, whereas open-ended generation reveals occasional accuracy reductions, particularly for non-English responses. Our work highlights the need to evaluate CLC enhancement for both cross-domain robustness and culturally appropriate variation, informing future work in post-training and benchmark development.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04409 (HTTP 429)
Authors: Jirui Qi, Mingyang Wang, Hinrich Schütze, Raquel Fernández, Arianna Bisazza
Categories: cs.CL
PDF URL: https://arxiv.org/pdf/2609.04409.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04409
Published: 2026-09-08T01:23:18.870Z
Agent Domain Papers
1. EXAONE Forecast for Finance
Abstract:This technical report presents EXAONE Forecast for Finance (EXAONE Finance), a financial time series (TS) foundation model (TSFM) tailored to financial forecasting. Recent TSFMs achieve strong zero-shot performance through large-scale pretraining. However, they are primarily developed for general-domain TS and largely rely on self-attention backbones whose computational cost grows quadratically with sequence length and variate count. Moreover, they assume fully observed inputs and are pretrained on corpora that fail to capture the unique dynamics of financial markets. These limitations hinder their applicability to finance, where long, many-channel, intermittently observed panels are common. To address these challenges, EXAONE Finance adopts an attention-free architecture, replacing self-attention with two simple yet effective linear-time operators: 1) a causal 1D convolution for temporal mixing and 2) a group-aware pooling multi-layer perceptron (MLP) for variate mixing. Furthermore, a masked context augmentation exposes the model to contiguous missing spans during training, improving robustness to the missingness pervasive in financial markets. EXAONE Finance is pretrained on a large-scale financial corpus covering not only equities but also foreign exchange, commodities, crypto-assets, fixed income, and macroeconomic indicators. On FinVerse, a financial forecasting benchmark covering diverse asset classes, EXAONE Finance attains state-of-the-art performance, ranking first across all three evaluation tiers—-point-forecast accuracy, cross-sectional asset ranking, and portfolio profitability.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04239 (HTTP 429)
Authors: Seunghan Lee, Jaehoon Lee, Jun Seo, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim, Sungdong Yoo, Junhyeok Kang, Sangjun Han, Soonyoung Lee, Wonbin Ahn
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04239.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04239
Published: 2026-09-08T01:23:32.083Z
2. From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance
Abstract:Artificial intelligence in recruitment has shifted the object being automated from profile pairs and ranked lists to multi-stage workflows that retrieve evidence, compare candidates, and support or execute actions. This systematized narrative review traces that development from bilateral retrieval and behavioral ranking through neural person—job matching, large language model (LLM) components, and tool-using recruiting agents. Using a purposive search and coding protocol updated through 23 July 2026, plus targeted updates through 2 September 2026, we organize 40 representative works with supporting industrial and legal sources. This synthesis is not a prevalence estimate. We analyze three coupled transitions: from similarity to reciprocal suitability, from a model to a compound workflow, and from offline prediction to evidence- and productivity-aligned evaluation. Across document understanding, retrieval, ranking, assessment, interviewing, sourcing, and human handoff, we distinguish field-, pair-, list-, case-, trajectory-, and outcome-level evidence. Persistent gaps arise because behavioral labels confound exposure, preference, and qualification; private and synthetic data limit external validity; final-output scores conceal pipeline failures; and, within the coded set, privacy is not directly evaluated and no row jointly evaluates utility, fairness, privacy, and security. These observations describe the coded set rather than the field as a whole. We therefore introduce a staged mapping from evaluation evidence to the strongest defensible claim, together with an agenda for reciprocal, evidence-grounded, temporally controlled, selective, and auditable systems. Progress should be judged by whether workflows retrieve the right evidence, preserve uncertainty, support contestable decisions, and improve outcomes under explicit cost and risk constraints.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04286 (HTTP 429)
Authors: Ziyi Zhao, Guanzheng Wei
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04286.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04286
Published: 2026-09-08T01:23:32.083Z
3. Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Abstract:Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
Q: 这篇论文试图解决什么问题? 这篇论文旨在解决\*\*大规模智能体(agentic)评估中的基础设施碎片化与异构性难题\*\*。具体而言,其试图化解以下核心问题: - \*\*评估集成成本高昂且不可扩展\*\*:现有的智能体基准测试在任务格式、环境需求、交互协议和评分程序上高度异构。传统方式下,支持 m 个基准和 n 个智能体需要 O(mn) 的配对集成工作量,导致评估过程重复、脆弱且难以扩展。 - \*\*评估结果缺乏跨领域可比性\*\*:由于各基准的实现、环境配置和评估协议不一致,论文和模型发布通常仅报告少量流行基准的子集,难以判断模型进展是跨领域泛化还是仅集中在少数知名任务上,跨模型比较受到严重干扰。 - \*\*基准任务质量与设计缺陷被忽视\*\*:许多基准中未解决的任务并非源于真正的智能体能力局限,而是源于任务本身的设计缺陷(如指令与验证器不匹配、隐式假设未声明、环境不稳定等),导致评估信号失真。 - \*\*大规模评估的经济性与效率问题\*\*:全面运行数十个基准测试成本极高,社区缺乏既保留挑战性又经济实惠的精简评估套件。 为应对上述问题,论文提出了一套系统性方案: 1. \*\*Harbor Adapters\*\*:将 80 余个异构基准统一适配至 Harbor 框架的标准化任务抽象(指令、环境、测试、解答),将集成复杂度从 O(mn) 降至 O(m+n) ,并通过严格的代码审查与对等实验(parity experiments)验证适配的语义保真度。 2. \*\*大规模实证分析\*\*:在 54 个基准、6,627 个任务上评估 8 个跨层级模型,系统分析了模型能力 vs. 脚手架(harness)设计的效应、基准间的冗余性、前沿模型的 Token 效率与成本权衡,以及任务失败的根因分布。 3. \*\*Harbor-Index\*\*:通过难度过滤、AI 与人工审计及审计-修复循环,从适配池中精选出 82 个跨 29 个基准的高质量任务。该套件确保没有任何已评估的模型-脚手架配置通过率超过 30%,在保持挑战性与领域多样性的同时显著降低评估成本。 Q2: 有哪些相关研究? 根据论文第5节及附录C的综述,相关研究主要分布在以下五个方向: ### 1. 智能体基准测试与评估基础设施 该领域已涌现出大量面向不同能力的基准测试: - \*\*软件工程\*\*:从函数级代码生成(HumanEval)到仓库级问题修复(SWE-bench、SWE-bench Pro、SWE-Lancer)及终端任务(Terminal-Bench) - \*\*交互与具身环境\*\*:Web导航(WebArena、Mind2Web、Workarena)、操作系统控制(OSWorld、AndroidWorld、AppWorld)、文本游戏(ALFWorld)及科学实验(ScienceWorld) - \*\*工具使用与助手能力\*\*:API调用(BFCL、ToolLLM)、函数调用(Gorilla)、多步交互(AgentBench、GAIA、AgentGym、τ-bench) - \*\*科学研究\*\*:机器学习工程(MLE-bench)、数据驱动的科学发现(ScienceAgentBench)、研究复现(ReplicationBench) - \*\*专业领域\*\*:金融(FinanceAgent、PIXIU)、法律(LawBench)、医疗(MedAgentBench)、电子表格(SpreadsheetBench) 在基础设施层面,现有工作包括可复现评估框架(Inspect、HAL、AstaBench)、通用软件智能体平台(OpenHands)以及用于稳定工具调用评估的虚拟化系统(StableToolBench)。 ### 2. 评估可重复性与脚手架效应 研究表明,评估结果并非仅由模型能力决定,还受到界面设计、脚手架(scaffold)、成本预算和隔离卫生条件的影响: - \*\*脚手架与接口效应\*\*:SWE-agent 证明智能体-计算机接口显著影响软件工程性能;Agentless 则表明更简单的流水线亦可具备竞争力 - \*\*分解效应\*\*:HAL 在超过 21,000 次 rollout 中分解了模型、脚手架与基准的效应 - \*\*整体评估\*\*:HELM 提倡多指标、场景级透明报告;Dynaboard 采用托管式评估以减少对自报结果的依赖 ### 3. 基准测试有效性与任务质量 该方向关注分数背后的 construct validity(构念效度)及任务设计缺陷: - \*\*标签与标注错误\*\*:测试集标签错误普遍存在(Northcutt et al.; Gema et al. 对 MMLU 的审视;RePOPE),可 destabilize 基准结论 - \*\*数据污染与泄露\*\*:污染侵蚀留出集有效性(Sainz et al.; LiveBench),SWE-bench 的审计发现大量"已解决"实例源于泄露、记忆化或弱测试(Aleithan et al.; Yu et al.; Liang et al.; Wang et al.) - \*\*验证器缺陷\*\*:对于智能体基准,不完美的验证器会封顶可达到的准确率(Stroebl et al.);《智能体基准检查清单》(Zhu et al.)记录了任务设置与奖励设计缺陷,相对扭曲可达 100% - \*\*构念效度\*\*:Bean et al. 对 445 个基准的审查发现测量现象、任务设计与评分指标中的系统性问题;BetterBench(Reuel et al.)与 ECBD(Liu et al.)提出了基准设计的方法论框架 ### 4. 高效评估与基准冗余 模型-基准得分矩阵具有内在低秩结构,这催生了多种压缩与高效评估方法: - \*\*低秩与冗余\*\*:因子分析表明少数潜在能力维度即可解释大部分方差(Burnell et al.; Maimon et al.; BenchPress / Zeng & Papailiopoulos; Redundancy principles for MLLMs benchmarks) - \*\*精简子集\*\*:metabench、tinybenchmarks、Anchor Points 等通过稀疏采样或自适应项目选择减少评估成本 - \*\*成本感知框架\*\*:Cost-of-Pass(Erol et al.)与 "AI Agents That Matter"(Kapoor et al.)将经济成本显式纳入评估 - \*\*局限性\*\*:近期研究也指出,子集预测在更强模型上性能下降(Zhang et al.),微观基准需要比假设更多的项目(Yauney et al.),且基准一致性估计依赖于非标准化选择(Perlitz et al.) ### 5. 安全、鲁棒性与对抗性评估 智能体执行多步计划并操作外部环境的特性带来了新的安全与可靠性挑战: - \*\*对抗性攻击\*\*:AgentDojo 评估提示注入攻击与防御;AgentHarm 与 OS-Harm 测量智能体对有害请求的拒绝能力及计算机使用安全 - \*\*评估本身的可靠性\*\*:AgentRewardBench 与 Agent-as-a-Judge 探讨模型评判者评估智能体轨迹的可靠性;替代注释者测试(Calderon et al.)为 LLM-as-a-judge 替换人类注释者提供了统计论证 - \*\*稳定性\*\*:StableToolBench 通过虚拟化 API 与缓存响应稳定工具学习评估 Q3: 论文如何解决这个问题? 论文通过\*\*统一基础设施、大规模实证验证、精选元数据集\*\*三层递进的方案系统性地解决了智能体评估的碎片化问题。具体方法如下: --- ### 1. Harbor Adapters:统一适配层,将集成复杂度从 O(mn) 降至 O(m+n) \*\*标准化任务抽象\*\* 论文基于 Harbor 框架
53
,将异构基准统一抽象为四个要素:instruction(指令)、environment(环境)、tests(测试)、solution(解答)。每个基准只需构建一个 Adapter,将其原生格式翻译为该共享 Schema;每个智能体只需对接一次 Harbor 接口。这消除了传统模式下每个基准-智能体对都需要独立集成的 O(mn) 开销。 多阶段质量审计与保真验证 为确保适配不破坏原始基准的语义,论文建立了严格的两条验证线: - 代码审查:每个 Adapter 经过 bot、junior、senior 三级审查,累计产生超过 10,000 条 GitHub 评论。 - 对等实验(Parity Experiments):在相同智能体、模型、执行配置下,对原始基准与 Harbor 适配版本各运行 k 次(通常 k=3 ),报告均值 ± 标准误(SEM)。只有当两边结果在统计误差范围内一致时,Adapter 才被视为通过。论文根据原始基准的智能体交互模式,设计了三种集成场景(原生 Harbor 智能体、Vanilla LLM、Custom Agent 移植)以覆盖不同基准类型。 执行后端兼容性 Harbor Adapters 支持多种沙盒后端(Daytona、Modal、E2B),以处理纯 CPU 任务、GPU 任务及 Docker-in-Docker 等特殊需求。 —- ### 2. 大规模评估:在统一条件下产生可比数据与系统性洞察 利用上述基础设施,论文开展了一项覆盖 8 个跨层级模型、16 种模型-脚手架(harness)配置、54 个基准、6,627 个任务 的大规模评估,每个配置重复 3 次试验,总计约 30 万条轨迹,消耗 226B Token 与超过 30 万美元计算成本。 该评估的设计关键在于控制变量: - 每个模型均在两种条件下测试:跨家族通用脚手架 Terminus-2,以及厂商原生脚手架(Codex / Claude Code / Gemini CLI)。 - 所有任务在隔离的容器环境中执行,确保环境一致性。 基于这些数据,论文进行了四维度诊断分析: - 基准冗余性:通过 PCA 与 Spearman 相关分析发现,54 个基准的有效评估空间远低于表面维度;约 12 个基准即可捕捉大部分独立信号。 - 脚手架 vs. 模型效应:线性混合模型显示,模型固定效应范围(0.451)是脚手架效应范围(0.087)的 5.2 倍,说明基础模型能力是性能差异的主因。 - 效率与成本权衡:前沿模型在所有难度层级上均使用更少 Token,但其更高的单价并未被 Token 节省完全抵消;升级至前沿模型在中等难度任务(0.3–0.7)上边际收益最大。 - 失败模式归因:人工 + LLM 法官对 6,028 条失败轨迹进行标注,发现“错误事实答案”、“算法缺陷”和“隐藏测试回归”是主导失败模式,而许多未解决任务实际上源于任务设计缺陷而非模型能力不足。 —- ### 3. Harbor-Index:通过多阶段漏斗精选高质量、高难度、可负担的元数据集 论文发现,现有基准中许多低通过率任务实际上是由于结构性设计缺陷(如指令-验证器不匹配、隐式假设、环境故障)而非真实难度。为此,论文构建了 Harbor-Index 1.0,其筛选流程是一个严格的多阶段漏斗: | 阶段 | 机制 | 结果 | |———|———|———| | 初始池 | 54 个适配基准的全部任务 | 6,627 个任务 | | 难度过滤 | 保留 Claude Opus 4.6、GPT-5.4、Gemini 3.1 Pro 在原生+Terminus-2 共 18 次试验中成功率 ≤ 33% 的任务 | 1,311 个任务 | | AI 审计 | Gemini-3-Flash 依据质量 Rubric(测试-指令对齐、本质难度)评分 | 307 个任务 | | 人工复审 | 14 名领域专家按相同 Rubric 复审,每任务至少 1 名 senior 或 2 名 junior 评审 | >110 个任务 | | 高级专家组 | 3 人 senior panel 基于难度、多样性、模型行为洞察选择 | 100 个任务 | | 审计-修复循环 | 至少 2 名 senior reviewer 进行轨迹驱动的失败分析与假阳性/假阴性检测,修复可修复任务,删除仍破损或修复后过易的任务 | 82 个任务 | 核心质量标准: - 测试-指令对齐(Test–Instruction Alignment):每个测试断言必须能在指令或环境中找到对应要求,反之亦然;禁止验证器检查指令未规定的内容。 - 本质难度(Essential Difficulty):难度必须来源于真正的推理、算法思维、领域专长或长程交互,而非格式细节、任意精度或未声明的魔法字符串。 评估约束与结果 最终 82 个任务跨越 29 个基准。论文评估了 9 个模型(4 个闭权 + 5 个开权)在 Terminus-2 与原生脚手架下的表现,共 1,476 次 rollout。关键设计约束确保: - 无任何模型-脚手架配置通过率超过 30%; - 最强配置(GPT-5.5 + Codex)通过率为 28.0%,保留充足进步空间; - 单次完整 Index 运行成本从 4(MiMo V2.5 Pro + Terminus-2)到 293(Claude Opus 4.8 + Claude Code)不等,实现可负担的重复评估。 —- ### 总结 论文的解决路径是工程基础设施 + 实证分析 + 数据精选的闭环:先用 Harbor Adapters 打破评估碎片化,再通过大规模统一实验诊断问题根源,最后以 Harbor-Index 提供一个经过严格审计、真正反映模型能力边界而非任务设计噪音的轻量级评估套件。所有代码、适配器、评估结果与 Harbor-Index 均以开源形式发布。 Q4: 论文做了哪些实验? 论文的实验体系可分为适配器验证、大规模系统评估、Harbor-Index 精选评估与多维度诊断分析四个层次。以下是详细梳理: —- ## 1. 适配器验证实验(Adapter Parity Experiments) 为确保 Harbor 适配不扭曲原始基准语义,论文对每个适配器进行了原始基准 vs. Harbor 适配版本的对等验证。 - 实验设计:固定相同智能体、模型、提示模板、解码参数与执行配置,在原始基准和 Harbor 适配版本上各运行 k 次(通常 k=3 ),报告均值 ± 标准误(SEM)。 - 覆盖范围:超过 80 个基准中的 54 个被纳入最终评估,其余因计算成本限制未运行,但适配器仍可用。 - 三种集成场景: 1. Scenario 1:原始基准已使用 Harbor 支持的原生智能体(如 Claude Code、Codex)。 2. Scenario 2:原始基准为直接 LLM 评估(非智能体),论文 Fork 原始仓库并添加 Harbor 支持的 CLI 智能体进行对称比较。 3. Scenario 3:原始基准使用自定义智能体(如特定 ReAct 变体),将其移植至 Harbor 后运行对比。 —- ## 2. 大规模系统评估(第3节核心实验) 这是论文规模最大的实证实验,旨在比较不同模型与脚手架在统一条件下的表现。 ### 2.1 实验设置 | 维度 | 配置 | |———|———| | 模型 | 8 个跨层级模型:GPT-5.4、GPT-5-mini、GPT-5-nano;Claude Opus 4.6、Claude Sonnet 4.6、Claude Haiku 4.5;Gemini-3.1-Pro、Gemini-3-Flash | | 脚手架 | 每个模型在 2 种 条件下测试:跨家族 Terminus-2 + 厂商原生脚手架(Codex / Claude Code / Gemini CLI),共 16 种 模型-脚手架配置 | | 基准 | 54 个适配基准 | | 任务数 | 6,627 个任务(对大数据集进行 i.i.d. 子采样) | | 重复次数 | 每个(基准, 模型, 脚手架)设置重复 3 次 | | 总轨迹数 | 约 30 万条 | | 资源消耗 | 226B Token,超过 30 万美元 计算成本 | ### 2.2 分析性实验子模块 基于上述大规模评估数据,论文开展了四项深度分析: #### (1) 基准冗余性与有效维度 - 方法:对 16 × 54 的得分矩阵进行 PCA/SVD(在 logit 空间)。 - 发现:仅固定 Terminus-2 时,PC1 解释 73.7% 方差;加入原生脚手架后,PC1 降至 67.1%,PC2(脚手架效应维度)升至 10.1%,整体空间近似 秩-2。 - 贪心选择:基于绝对 Spearman 相关系数 rho ≥ 0.7 为阈值,发现仅需 12 个基准 即可覆盖大部分独立信号。 - 任务级冗余:在 52 个足够大的基准中,每基准选取 3 个代表性任务即可以平均 rho ≈ 0.923 恢复系统排名。 #### (2) 模型 vs. 脚手架效应分解 - 方法:线性混合模型(LMM),以基准为随机截距,控制难度差异。 - 发现:模型固定效应范围(0.451)是脚手架效应范围(0.087)的 5.2 倍;8 个模型中 6 个系数显著( p<0.05 ),4 个脚手架中仅 2 个显著。ICC = 0.75 表明 75% 方差来自基准间难度差异。 #### (3) 效率与成本权衡 - 方法:按经验任务难度(1 - 平均通过率)分 10 个桶(宽度 0.1),对比前沿模型组(Claude Opus 4.6、Gemini 3.1 Pro、GPT 5.4)与其他模型组的 Token 消耗与美元成本。 - 发现: - 前沿模型在所有难度层级使用更少 Token,简易任务上仅消耗弱模型 Token 的 42%。 - 但单价更高,其他模型单次试验成本低 2–3 倍,Token 节省未能完全抵消价格溢价。 - 绝对性能提升呈倒 U 型,在中等难度(0.3–0.7)区间边际收益最大。 #### (4) 失败模式归因(Failure Mode Analysis) - 方法:两名人类注释员独立标注 200 条轨迹(12 种失败模式),然后以校准后的 Gemini 3.1 Pro 法官扩展至 N = 6,028 条失败轨迹。 - 发现: - 主导失败模式:错误事实答案(Wrong Factual Answer)、算法缺陷(Algorithmic Bug)、隐藏测试回归(Hidden-Test Regression)。 - 脚手架设计影响行为:原生脚手架(Codex/Claude Code)允许开放式迭代与自校正;Terminus-2 的线性 Plan-Execute-Complete 序列缺乏内置验证-校正循环。 - 模型行为差异:GPT-5.4 倾向简洁或快速放弃;Claude 保持恒定轮次;Gemini 依赖外部搜索但易引发”信息漂移”。 —- ## 3. Harbor-Index 1.0 评估实验(第4节) 这是针对精选后 82 个任务的独立评估,验证 Harbor-Index 作为轻量级、高质量元数据集的有效性。 ### 3.1 实验设置 | 维度 | 配置 | |———|———| | 模型 | 9 个模型:4 闭权(GPT-5.5、Claude Opus 4.8、Gemini 3.1 Pro、Qwen3.7 Max)+ 5 开权(GLM 5.2、Kimi K2.6、MiniMax M3、DeepSeek V4 Pro、MiMo V2.5 Pro) | | 脚手架 | 每个模型在 Terminus-2 和一种原生脚手架下运行;开权模型通过 OpenRouter 统一使用 Claude Code | | 任务数 | 82 个任务,跨越 29 个基准 | | 总 Rollouts | 9 × 2 × 82 = 1,476 | ### 3.2 关键约束与评分 - 难度控制:筛选时确保 3 个前沿模型在 18 次试验中成功率 ≤ 33% ;最终无任何模型-脚手架配置通过率超过 30%。 - SOTA 阈值:对于连续指标任务(如 AlgoTune、GSO 的速度比),仅当解超越当前前沿时才计为通过。 - LLM-as-a-Judge:自由回答任务(HLE、GAIA、GPQA Diamond 等)使用多模型法官评估。 ### 3.3 结果 - 最强性能:GPT-5.5 + Codex 通过率为 28.0%。 - 成本-性能帕累托:前沿模型在原生脚手架下表现最佳;多个开权模型在 Terminus-2 下位于成本-性能帕累托前沿(如 MiniMax M3 单次仅 4,通过率 2.4–6.1%)。 - 失败分解:1,476 次 rollout 中,453 次超时无答案、361 次接近正确、445 次根本错误。 —- ## 4. 基准历史进展与延迟研究(附录实验) - 历史分数收集:为 39 个基准收集发布以来至今的最佳报告分数,对比 Harbor 统一评估下的当前最优结果(图 2、图 8、图 9)。发现软件工程与数学推理领域进展最快且 leaderboard 压力最大,而知识长文本与智能体工具系统领域测量更稀疏。 - 延迟分析(附录 F.5):按任务难度与成败分解执行延迟(wall-clock time)。发现失败任务的延迟在难度 0.7–0.8 处达峰,最难任务反而因早期终止而耗时更短;不同模型的成功/失败延迟模式存在结构性差异(如 GPT-5.4 无论成败耗时相近,而 GPT-5-Nano 在失败任务上坚持最久)。 —- ## 实验数据与工件的公开发布 论文将以下实验产物全部开源: - Harbor 适配器与框架代码 - 大规模评估的完整轨迹(约 340 GB,以 Hugging Face 数据集形式发布) - Harbor-Index 1.0 任务与审计产物 - 交互式结果看板(harbor-index.org) Q5: 有什么可以进一步探索的点? 基于论文的发现与讨论,以下方向值得进一步探索: —- ### 1. 基础设施与适配器的扩展验证 - 更多基准的规模化验证:当前大规模评估覆盖了 54 个适配基准,尚有超过 30 个已适配基准(如 CooperBench、OSWorld、TheAgentCompany 等,见附录 D.6.11)因计算成本未纳入实证分析。系统性地对这些基准进行同等规模的 parity 实验与智能体评估,可进一步检验 Harbor 框架的通用性边界。 - 强化学习与多智能体架构的支持:现有评估主要围绕基于 LLM 的单智能体脚手架展开。将 Harbor Adapters 扩展至支持强化学习训练循环、多智能体协作与竞争环境,是一个尚未触及的工程与研究方向。 —- ### 2. 模型-脚手架交互的动态优化 - 自适应脚手架设计:论文发现基础模型能力对性能的影响显著大于脚手架(效应范围 5.2:1),但脚手架仍系统性地改变行为模式(如 Terminus-2 的线性结构缺乏迭代自校正)。探索依据任务难度或失败模式动态切换脚手架策略(如从 Plan-Execute-Complete 切换到开放式迭代)可能提升整体通过率。 - 跨家族脚手架迁移:开源模型(如 GLM 5.2、DeepSeek V4 Pro)在 Terminus-2 上表现出成本-性能帕累托优势,而闭源模型在原生脚手架上更优。研究非原生脚手架的兼容性优化(如通过 OpenRouter 统一接口时的缓存命中率提升)具有实际价值。 —- ### 3. 基准冗余性与自适应评估 - 任务级自适应选择(Computerized Adaptive Testing, CAT):论文通过贪心选择证明约 12 个基准即可逼近 54 个基准的排序信号。进一步引入项目反应理论(IRT)或在线自适应采样,依据模型在前序任务上的表现动态选择下一任务,可在不牺牲排序精度的前提下进一步压缩评估成本。 - 子集预测的理论极限:论文指出,从子集预测完整基准表现的方法在更强模型上性能会退化。研究这种外推误差的统计极限,以及如何选择对模型强度变化鲁棒的最小任务集,是元评估(meta-evaluation)的核心问题。 —- ### 4. 失败模式的针对性消解 - 验证器-任务对齐的自动化诊断:人工审计发现约三分之一的最难候选任务因设计缺陷(如指令-验证器不匹配、隐式假设)被剔除。开发自动化的任务质量审计工具(如基于形式规约的验证器合成、对抗性测试生成),可减少对昂贵人工审查的依赖。 - 特定失败模式的干预策略:对图 4 中主导的三种失败模式——错误事实答案、算法缺陷、隐藏测试回归——设计针对性的认知架构改进(如检索增强生成用于事实性、符号执行用于算法调试、显式测试覆盖分析用于隐藏测试)。 —- ### 5. Harbor-Index 的动态维护与对抗性审计 - 实时更新与反饱和机制:论文明确提出将 Harbor-Index 维护为活的基准(live benchmark)。随着模型能力提升,需要周期性重跑过滤混合(filtering mix),剔除已饱和任务,并从未审计池中补充新任务以保持 30% 通过率上限。 - 对抗性漏洞挖掘:更强的未来智能体可能更有效地利用当前未暴露的评估漏洞(如奖励黑客、网络信息泄露)。建立红队(red-teaming)流程,自动化检测可 exploited 的 verifier loopholes,是维持评估可信度的必要环节。 —- ### 6. 成本-效率评估的形式化框架 - 经济最优的模型选择策略:论文发现前沿模型 Token 效率更高但成本溢价未被完全抵消,且中等难度任务( 0.3 ≤ 1-p ≤ 0.7 )的边际收益最大。将这一经验观察形式化为预算约束下的最优模型-脚手架组合选择问题,可为实际部署中的评估资源分配提供决策支持。 - 不确定性量化:在压缩评估中,现有方法缺乏对排序不确定性的显式建模。引入贝叶斯优化或错误条估计(如附录 F.1.4 提及),可在减少任务数量的同时报告模型间差异的置信度。 —- ### 7. 长程交互与动态环境的评估 - 异步与动态事件的支持:GAIA2 等任务的适配暴露出 Harbor 对异步环境事件(如模拟器状态推进)的处理不足。改进执行引擎以支持事件驱动、状态机式的动态环境交互,是评估更复杂智能体(如操作系统控制、实时系统管理)的前提。 - 超时与预算的动态调整:当前 Harbor-Index 采用固定时间预算(通常为最快前沿模型耗时的 1.2× )。研究基于任务复杂度预测的自适应超时分配,可在保证公平性的同时减少无效等待。 —- ### 8. 多模态与跨领域迁移评估 - 视觉-语言-行动(VLA)智能体的标准化:当前 Harbor-Index 的多模态覆盖有限(如 MMAU 依赖音频转文本)。扩展 Adapter 以原生支持图像输入、GUI 操作、视频理解,并建立跨模态的 parity 验证协议,是一个重要的工程研究课题。 - 领域间能力迁移的量化:论文发现不同基准间存在低秩结构,但专用能力(如 FinanceAgent、LAB-Bench)难以被其他基准预测。系统性地构建能力迁移矩阵(类似 Taskonomy),明确哪些基准组合最能预测未见过领域的表现,有助于设计更高效的训练与评估协议。 Q6: 总结一下论文的主要内容 这篇论文针对大规模语言模型智能体(agentic)评估中的基础设施碎片化与异构性难题,提出了系统性的解决方案。主要内容可概括如下: —- ### 1. 研究背景与问题 随着智能体能力的扩展(代码生成、工具使用、长程任务、计算机操作等),agentic 基准测试数量激增(2024 年以来超过 200 个)。然而,这些基准在任务格式、环境需求、交互协议和评分逻辑上高度异构,导致: - 集成成本爆炸:传统方式需要 O(mn) 的 benchmark–agent 配对集成; - 结果不可比:不同论文的环境配置、实现细节和评估协议差异巨大,难以判断模型进展是否跨领域泛化; - 任务质量参差不齐:许多“未解决”任务实际上源于设计缺陷(如指令与验证器不匹配、隐式假设、环境故障),而非真正的模型能力不足; - 评估成本高昂:全面运行数十个基准耗费巨大,社区缺乏既保留挑战性又经济实惠的精简套件。 —- ### 2. 核心贡献 #### 贡献一:Harbor Adapters 统一评估基础设施 论文开发了基于 Harbor 框架的适配器体系,将 80 余个异构基准统一映射到共享的任务抽象(instruction, environment, tests, solution)。该体系: - 将集成复杂度从 O(mn) 降至 O(m+n) ; - 支持多种沙盒后端(Daytona、Modal、E2B)与 GPU 交互; - 通过严格的三级代码审查(bot、junior、senior)和对等实验(parity experiments)验证适配语义保真度,确保 Harbor 版本与原始基准统计一致。 #### 贡献二:大规模系统评估与诊断分析 论文开展了覆盖 8 个跨层级模型、16 种 model–harness 配置、54 个基准、6,627 个任务(每配置 3 次试验,共约 30 万条轨迹,消耗 226B Token 与超过 30 万美元)的评估。关键发现包括: - 模型能力主导性能差异:线性混合模型显示,模型固定效应范围是脚手架(harness)效应范围的 5.2 倍(0.451 vs. 0.087),基准难度解释 75% 的方差; - 基准空间低秩冗余:54 个基准的有效评估空间近似秩-2(能力轴 + 脚手架轴);仅需约 12 个基准即可捕捉大部分独立排序信号;单个基准内 3 个任务即可以平均 rho ≈ 0.923 恢复系统排名; - 效率与成本权衡:前沿模型在所有难度层级上更省 Token(简易任务仅消耗弱模型的 42%),但更高单价使总成本仍高出 2–3 倍;升级至前沿模型在中等难度任务(0.3–0.7)上边际收益最大; - 失败模式归因:对 6,028 条失败轨迹的审计显示,错误事实答案、算法缺陷和隐藏测试回归是主导失败模式;原生脚手架允许开放式迭代,而 Terminus-2 的线性结构更易受困于需反复精修的任务。 #### 贡献三:Harbor-Index 精选元数据集 论文指出,现有基准中大量低通过率任务实为结构性设计缺陷(如验证器越界、隐式假设、环境未触发事件)。为此,论文通过多阶段漏斗(难度过滤 → AI 审计 → 人工复审 → 高级专家组筛选 → 审计-修复循环),从 6,627 个任务中精选出 Harbor-Index 1.0: - 82 个任务,跨越 29 个基准,覆盖软件工程、科学研究、数学推理、智能体系统等核心领域; - 通过 测试-指令对齐 与 本质难度 两项质量标准剔除破损任务; - 难度受控:无任何已评估 model–harness 配置通过率超过 30%;最强配置(GPT-5.5 + Codex)通过率为 28.0%,保留充足进步空间; - 成本可控:单次完整 Index 运行成本从 4 到 $293 不等,多个开源模型在 Terminus-2 下位于成本-性能帕累托前沿。 —- ### 3. 结论与影响 Harbor Adapters 与 Harbor-Index 共同提供了可复用、可扩展、可负担的智能体评估基础设施与高质量测试床。论文将所有适配器、评估工具、完整轨迹(约 340 GB)及 Harbor-Index 开源发布,旨在推动更可靠、更全面、更具可比性的语言模型智能体评估。同时,论文也指出未来需持续维护 Harbor-Index 为活基准,对抗未来更强模型可能利用的评估漏洞,并不断扩展适配器覆盖范围。 Q7: 想要进一步了解论文 以上只是了解一篇论文的几个基本FAQ。如果你还想与Kimi进一步讨论该论文,请点击 这里 为你跳转Kimi AI网页版,并启动一个与该论文相关的新会话。
Authors: Lin Shi, Haowei Lin, Zixuan Zhu, Xiaoyue Zhou, Xiang Li, Xiangning Lin, Yaxuan Deng, Han Xu, Yuangang Li, Shanda Li, Zizhao Chen, Hanwen Xing, Harsh Raj, Bo Chen, Quan Shi, Steven Dillmann, Yipeng Gao, Puneesh Khanna, Ruofan Lu, Chao Beyond Zhou, Michael Yang, Robert Zhang, Siyuan Chai, Jiayu Chang, Yizhao Chen, Xiaokun Chen, Yiwei Dai, Wenting Yang, Hange Liu, Minghao Liu, Zihan Wang, Adnan El Assadi, Benedikt Stroebl, E. Kelly Buchanan, Han Meng, Junwei He, Longxuan Yu, Radin Shayanfar, Yukyung Lee, Zhikang Dong, Allen G Hart, Anjiang Wei, Anurag Kashyap, Arpandeep Khatua, Audrey Jixin Zheng, Chengrui Ma, David Heineman, Dubing Chen, Hai-Anh Trinh, Haishuo Fang, Hefan Zhang, Hui Shen, Issa Sugiura, Jiankai Sun, Jiechao Gao, Junhong Lin, Junnan Li, Kai Yang, Lei Hsiung, Maoyu Wang, Mengze Tang, Nabil Omi, Negin Raoof, Nicholas Edwards, Octavia Guo, Orfeas Menis Mastromichalakis, Pengliang Ji, Przemysław Hejman, Qi Qi, Qunshu Lin, Richard Zhuang, Rui Yang, Ruichen Zheng, Ryan Marten, Shaghayegh Fazliani, Shizheng Hou, Sicong Jiang, Sijie Li, Song Bian, Terry Yue Zhuo, Tianqing Wu, Tom Tang, Wanjia Zhao, Weihao Xuan, Wenhua Liang, Xian Liu, Xin Lan, Xuan Zhang, Xuandong Zhao, Yanchuan Tang, Yifan Jiang, Yijiang Li, Yitong Guan, Yizhi Li, Yonghui Liu, Yuheng Tang, Yujun, Yunfei Zhao, Yuxin Wang, Yuxuan Tang
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04298.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04298
Published: 2026-09-08T01:23:32.083Z
4. Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security
Abstract:Ensuring the security of the power system is essential for stability and reliability, especially in the event of disruption. Effective classification of contingency in power systems enables proactive decision-making and mitigates large-scale breakdowns and failures. This study explores the use of machine learning algorithms to classify security levels of contingencies in power systems into safe, moderate or severe classes. For this approach, Newton-Raphson load flow method extracts system data from contingency scenarios, using Overall Performance Index (OPI) as safety measure. For data pre-processing, Synthetic Minority Over-Sampling Technique (SMOTE) and Principal Component Analysis (PCA) is used to address class imbalance and reduce dimensionality, respectively. K-Nearest Neighbours (KNN), Random Forest (RF) and Support Vector Machines (SVM) is trained and evaluated on datasets generated through N-k contingency scenarios for k equal 1, 2, and 3 on IEEE-14 and IEEE-30 bus systems using four hybrid pre-processing configurations: normalized, SMOTE-balanced, PCA-transformed, and a combined SMOTE PCA-transformed. Performance is assessed by precision, recall and F1 score, with priority given to the severe contingency classes. The RF achieved the highest F1 scores of 0.97 in IEEE-30 and 0.86 in IEEE-14, SVM benefits significantly from PCA and improves the accuracy of the classification, while KNN is best suited for SMOTE and PCA conversion. The findings show that PCA contributes more than SMOTE to the overall performance of the model. However, SMOTE improves recall but can introduce false positives and is therefore a compromise of accuracy. This study highlights machine learning as a scalable and powerful alternative to traditional contingency analysis, which improves the assessment of security in real time.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04300 (HTTP 429)
Authors: Joshua Salako, Folajimi Osikomaiya, Olakorede Olamiju
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04300.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04300
Published: 2026-09-08T01:23:32.083Z
5. Iris: Climbing to the Search Frontier
Abstract:We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks are reverse-constructed from the hyperlink structure of a web corpus: we author multi-hop chains over an entity graph distilled from a seed page and its out-links, rewrite every non-answer entity into a descriptive reference so that no clue can be resolved by string matching, and admit only questions that a reference model fails closed-book yet solves once the supporting evidence is supplied. These questions are then turned into trajectories, which are filtered at both the trajectory and the turn level before SFT. The policy is then optimized by RL against live search, with the reward judge and the observation summarizer served inside the training cluster, and with over-long rollouts interrupted at the request level and resumed from their committed prefix at the next step. We alternate the two stages in a procedure we call SFT-RL climbing, returning the hardest solved and most efficient rollouts of each RL round to the next supervised pass. Because inference-time context management is worth more on these benchmarks than most reported differences between systems, we evaluate every benchmark both with and without it, holding the tool set, the context limit, and the judge fixed. All results come from a single ReAct agent, with no sub-agents and no test-time verification. With management enabled, on BrowseComp, BrowseComp-ZH, DeepSearchQA, and HLE the two models reach $82.2/84.8/86.9/52.3$ and $88.6/85.1/92.9/56.4$, the strongest overall results among open-source search agents in their respective parameter ranges. We plan to release the model weights together with the complete recipe for data construction, training, and evaluation.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
Q: 这篇论文试图解决什么问题? 该论文旨在系统性地解决\*\*如何构建、训练和评估具备深度多跳搜索与长程信息整合能力的强大语言模型智能体\*\*这一核心问题。具体而言,其试图克服的关键挑战可归纳为以下四个方面: 1. \*\*高质量搜索训练数据的自动构造与验证\*\* 自然存在的网页问题通常过于简单,而人工编写难以规模化。论文提出从网页语料的超链接结构反向构造多跳任务:首先基于种子页面及其外链蒸馏出实体图,随后在该图上生成必须组合至少 N 个关系才能回答的问题,并通过抽象算子 A 将所有非答案实体重写为描述性指代,从而消除直接通过字符串匹配即可定位答案的捷径。最终,仅保留满足双重验证标准的样本——即在闭卷环境下参考模型 M_(ref) 无法回答( c_(diff)(q)=1 ),但在提供支撑证据后能够正确求解( c_(solv)(q)=1 )——以确保数据既具备难度又具备客观可解性。 2. \*\*长程交互中的上下文管理与训练效率\*\* 长视域搜索轨迹极易耗尽模型的上下文窗口,使得有效搜索预算远低于名义窗口长度。为缓解该问题,论文在推理层面系统研究了上下文管理(Context Management, CM)策略对性能的影响,并强调应在固定工具集、上下文长度限制和评判标准下,分别报告\*\*启用 CM\*\* 与\*\*不启用 CM\*\* 的结果,以准确剥离模型内在策略能力与推理时外部封装带来的增益。在训练层面,论文采用请求级中断与前缀复用(request-level partial rollout)机制,允许超长轨迹在同步 rollout 过程中被中断并在下一步从已提交前缀恢复,从而兼顾训练效率与长程探索能力。 3. \*\*监督微调与强化学习的有效协同\*\* 单纯依赖监督微调(SFT)容易受限于教师模型的能力边界,而单纯依赖强化学习(RL)则难以稳定复现稀疏的成功轨迹。论文提出 \*\*SFT–RL Climbing\*\* 迭代流程:在每一轮 RL 探索后,筛选出“最难且最高效”的成功轨迹(即通过率 0 < R(q) le 1/2 且工具调用轮次不少于 K_(rft) 的最短正确轨迹),将其蒸馏回下一轮的 SFT 数据集中。该机制使策略在探索新行为与巩固已有成功模式之间自动形成课程学习式的渐进提升。 4. \*\*内在搜索能力与评估透明性的精准衡量\*\* 现有工作往往仅在启用复杂推理时封装(如子智能体、重试、验证链)下报告结果,导致难以判断性能提升源于模型本身还是外部工程技巧。论文坚持仅使用单一 ReAct 智能体、无子智能体、无测试时验证的设定,并在所有基准上提供无 CM 与有 CM 的双重结果。通过这种方式,论文试图回答:在剥离了上下文管理、重试机制等推理时辅助手段后,基础模型本身究竟掌握了多强的自主搜索与推理能力。 Q2: 有哪些相关研究? 根据论文引言及参考文献,相关研究可按照主题脉络梳理如下。 ### 1. 早期工具使用与推理-行动交织 该方向奠定了语言模型通过外部工具扩展能力的基础,核心在于将推理(Reasoning)与行动(Acting)相结合: - \*\*WebGPT\*\*
14
:最早探索通过浏览器辅助进行问答,并引入人类反馈进行训练。 - **Toolformer**
17
:展示了语言模型可通过自监督方式学习调用 API 工具。 - **ReAct**
26
:提出将推理与行动交织的范式,成为后续搜索智能体的标准交互框架。 ### 2. 基于强化学习的搜索优化 从静态语料检索发展到对实时网页环境的强化学习(RL)优化: - **Search-o1**
11
与 **Search-R1**
8
:将大语言模型与搜索引擎结合,通过 RL 训练模型进行搜索增强推理。 - **R1-Searcher**
19
:通过强化学习激励 LLM 的搜索能力。 - **DeepResearcher**
29
:在真实环境中通过 RL 规模化深度研究能力。 ### 3. 合成数据构造与难度提升 为解决自然网页问题过于简单的问题,研究者们通过超链接结构或知识图谱构造多跳、抗捷径(shortcut-resistant)任务: - **WebDancer**
23
、**WebSailor**
10
、**WebShaper**
20
:利用网页图结构遍历生成信息搜寻任务。 - **DeepDive**
12
:结合知识图谱与多轮强化学习提升深度搜索智能体能力。 - **FORT-Searcher**
4
:专门合成抗捷径的搜索任务用于训练深度搜索代理。 ### 4. 端到端训练流程与开源系统 近期研究趋向于将合成问题生成、轨迹监督、RL 与长上下文推理整合为标准化流水线: - **OpenSeeker / OpenSeeker-v2**
6, 5
:开源搜索智能体训练数据与流程的代表性工作。 - **Tongyi DeepResearch**
21
、**MiroThinker-1.7 & H1**
13
、**Apodex-1.0**
1
:探索验证驱动的重型研究智能体。 - **REDSearcher**
2
:聚焦可扩展、低成本的长程搜索代理框架。 - **XYZ-Aquila**
25
与 **Shanghai AI Laboratory**
18
:在 35B 至更大参数规模上推进智能体能力。 - **Kimi K3**
9
:开放前沿智能体能力。 ### 5. 上下文管理与长程推理 长程搜索轨迹容易耗尽上下文窗口,相关研究探索了多种压缩与重置策略: - **ReSum**
24
:通过上下文摘要解锁长程搜索智能。 - **AgentFold**
27
:主动上下文管理以支持长程网页代理。 - **DeepSeek-V3.2**
3
:提出 discard-all 等上下文管理策略,论文将其作为 CM 对比基准。 ### 6. 评估基准与数据集 论文涉及的主要评测基准包括: - **BrowseComp**
22
与 **BrowseComp-ZH**
30
:评估浏览代理定位长尾实体与多线索约束能力的基准。 - **DeepSearchQA**
7
:评估基于搜索回答的全面性而非单一答案跨度。 - **Humanity’s Last Exam (HLE)**
15
:跨学科专家级学术推理测试。 - **LoHoSearch**
28
:用于测试超越人类难度上限的长程搜索代理。 此外,论文在结论中提到,搜索能力的提升对通用工具使用(如 BFCL、 τ -bench)和办公协作任务(如 OfficeQA、APEX)存在正向迁移,暗示搜索能力可能是一种可复用的原子能力而非垂直领域特化。 Q3: 论文如何解决这个问题? 论文通过一套端到端的“数据—训练—评估”配方解决构建强大搜索智能体的问题,核心环节可概括如下。 ### 1. 从网页图反向构造并验证多跳任务 为解决自然问题过易、人工标注难扩展的问题,论文设计了一条全自动数据生产线: - **网页子图提取**:将语料视为有向图 G=(V,E) ,从种子页面 v0 出发,沿外链抽取局部子图 G(sub)=(v0∪ N, E(sub)) 。 - **实体图蒸馏**:将 G(sub) 压缩为紧凑的实体图 G_e=(V_e,R_e)=f(ext)(G(sub)) ,仅保留与种子主题相关的实体与关系。 - **多跳问题生成**:在 G_e 上生成问题-路径对 (q_0,P)=f(gen)(G_e,y) ,强制推理路径长度 |P|ge N ,确保问题依赖至少 N 个耦合关系。 - **实体抽象去捷径**:通过抽象算子 A 将所有非答案实体重写为描述性指代,使得
q = f(abs)(q_0, A),
从而消除直接字符串匹配即可定位答案的捷径。 - **双重验证**:利用参考模型 M(ref) 进行两项检验:
c(diff)(q) = I[M(ref)(q)≠ y], quad c(solv)(q) = I[M(ref)(qmid Ge)=y].
仅保留交集 D=(q,y)mid c(diff)· c(solv)=1 ,即那些闭卷失败、供证后能解的样本,确保难度与可解性并存。 ### 2. 双层过滤的监督微调(SFT) 利用强教师模型 M_T 在 ReAct 范式下对 D 中问题生成轨迹 τ ,随后进行粗、细两级过滤,以提升监督信号质量。 - **轨迹级粗过滤**:剔除三类不良样本: - **错误轨迹**: rollout 未成功终止或最终答案 y 被评判为错误; - **退化轨迹**:利用滑动窗口 zlib 压缩比 rho(cr)(w)=|w|/|zlib(w)| 检测重复循环,并辅以周期重复行、单字符长串、字节相同参数等辅助检测器; - **浅层轨迹**:工具调用轮次 T(tool)(τ)<K 的样本被丢弃。 过滤后去重得到 D(sft) 。 - **轮次级精过滤**:为避免教师能力边界导致的局部劣质步,引入从数据中归纳出的评判标准(而非人工手写规则),对每个助手轮次输出 KEEP/MASK 标签 mt∈0,1 ,且单条轨迹最多掩码 10% 的轮次。 - **训练目标**:在 D(sft) 上最大化教师输出的似然:
L(SFT)(θ) = -E((q,τ)sim Dsft) ∑(t=1)^(T+1) mt log πθ(ut mid C(<t)),
其中 C_(<t) 为历史上下文, u_t 为第 t 步的推理与工具调用或最终答案。被掩码的轮次保留在上下文中但不参与损失计算。 ### 3. 针对实时搜索的强化学习(RL) 在 SFT 基础上,论文进一步以组相对策略梯度在**实时搜索环境**中优化策略。 - **请求级部分 rollout 与前缀复用**:为缓解长程 rollout 的尾部延迟,训练在请求粒度中断:超时的在飞会话被中止后,于下一步从其已提交的**前缀**恢复。通过缓存各消息状态的 token、损失掩码、对数概率与权重版本增量,恢复时利用截断重要性采样修正策略权重不匹配,避免重复计算,保持同步调度的 GPU 利用率。 - **集群内奖励与摘要**:不依赖外部 API,而是在训练集群内部署 FP8 推理引擎(基于 Qwen3.5-397B-A17B),同时承担两职: - **生成式奖励模型(GenRM)**:对 rollout 的答案给出二元裁决
R(q,τ) = I[GenRM(q,y_τ,y^*)=A];
- **观察摘要器**:将检索到的原始页面压缩为查询相关的摘要 ot ,作为 ReAct 循环中的观察。这是 rollout 中唯一的上下文缩减机制,无历史剪枝或滑动窗口,保证训练与推理的上下文一致性。 ### 4. SFT–RL 迭代攀爬(Climbing) 为结合 RL 的探索优势与 SFT 对稀缺成功轨迹的直接固化能力,论文交替执行两阶段: - 每轮 RL 后,保留**部分解决**(通过率 0<R(q)le 1/2 )的查询,从中挑选成功且工具调用不少于 K(rft) 的最短轨迹:
Sq = τ_i mid R(q,τ_i)=1,; T(tool)(τi)ge K(rft), quad τq^* = argmin(τ∈ Sq) T(tool)(τ).
- 这些高效且非平凡的成功轨迹经去重后回注下一轮 SFT,形成自动推进难度的课程学习。随着策略提升,候选池自然萎缩,提供简单的停止信号。 ### 5. 受控且透明的评估协议 为准确度量模型**内在搜索能力**与**上下文管理(CM)**各自的贡献,论文采用严格的对比评估: - **固定基础设施**:所有对比在相同工具集、相同上下文长度限制、相同 LLM 评判器下进行。 - **双模式报告**:对每个基准同时报告**无 CM**(原始 ReAct,无历史删减)与**有 CM**(采用 DeepSeek-V3.2 的 discard-all 等策略)的结果,从而将模型本身能力与推理时封装增益解耦。 - **防泄露机制**:在训练与评估中屏蔽 Hugging Face 上的数据集与 Space 页面,防止模型利用基准泄露刷分。 通过上述数据构造、双层过滤监督训练、带前缀复用的在线 RL、迭代攀爬以及受控评估的完整闭环,论文系统性地将模型的多跳搜索与长程信息整合能力推至当前同参数规模下的最优水平。 Q4: 论文做了哪些实验? 论文的实验围绕**模型性能对比**与**上下文管理(CM)消融**两大主线展开,具体包括以下内容。 ### 1. 基准测试与评估设置 实验在四个具有挑战性的智能体搜索基准上进行: | 基准 | 评估重点 | 评价指标 | |———|————-|————-| | **BrowseComp** | 长程、多线索约束下的长尾实体识别 | Accuracy | | **BrowseComp-ZH** | 中文环境下的上述能力 | Accuracy | | **DeepSearchQA** | 搜索答案的全面性(多证据覆盖) | F1 | | **Humanity’s Last Exam (HLE)** | 专家级学术推理(文本-only子集) | Accuracy | 评估协议统一采用 **pass@1**(单轮 rollout),由 LLM-based judge 根据各基准的官方 prompt 判定最终答案。所有实验在**相同的工具集、上下文长度限制与最大轮次预算**下进行,以确保可比性。 ### 2. 主实验:与开源及前沿模型对比 论文将 **Iris-mini(35B-A3B)** 与 **Iris-pro(397B-A17B)** 分别置于同参数规模区间及更大规模的前沿模型中进行横向对比,结果如表1所示(默认启用 discard-all CM 策略)。 - **同规模开源模型对比(30–35B)**: Iris-mini 在 BrowseComp(82.2)、BrowseComp-ZH(84.8)和 HLE(52.3)上取得该参数区间的最优成绩;在 DeepSearchQA(86.9)上略低于 XYZ-Aquila-mini(89.5)。 - **同规模开源模型对比(∼400B)**: Iris-pro 在全部四个基准上领先或持平:BrowseComp(88.6)、BrowseComp-ZH(85.1)、DeepSearchQA(92.9)、HLE(56.4)。其中 BrowseComp 领先同规模次优模型 XYZ-Aquila-pro 3.8 分,HLE 领先 3.1 分。 - **与更大规模/闭源前沿系统对比**: Iris-mini 的性能已接近 1T 规模的 Kimi-K2.6 与 DeepSeek-V4-Pro(BrowseComp 上 82.2 vs. 83.2/83.4)。Iris-pro 在标准 ReAct 设置下的表现可媲美部分采用重度推理封装(heavy-compute)的系统(如 MiroThinker-H1 与 Apodex-1.0-H)。 ### 3. 上下文管理(CM)消融实验 为剥离**模型内在搜索能力**与**推理时外部封装**各自的贡献,论文系统比较了多种 CM 策略,结果如表2所示。 - **无 CM(w/o)基线**:不使用任何历史删减或重置,直接反映模型原生能力。 - **discard-all**:当运行上下文达到阈值时,清空全部交互历史并从头重启问题求解(借鉴 DeepSeek-V3.2)。 - **retry**:若一次尝试失败,将失败经历摘要为“已探索与排除的信息”,追加至任务描述后重新尝试(借鉴 MiroThinker)。 - **discard-all + retry**:上述两种策略的复合。 **实验发现包括**: - **无 CM 下的原生能力**:Iris-mini 在 BrowseComp(64.7)与 BrowseComp-ZH(72.3)上显著高于同规模无 CM 系统(如 FORT-Searcher 的 55.9/62.1、OpenSeeker-v2 的 46.0/58.1);Iris-pro 在此基础上再提升 7.9 与 4.5 分。 - **CM 增益的尺度差异**:CM 对 Iris-mini 的提升(BrowseComp 上最高 +21.2 分)明显大于对 Iris-pro 的提升(最高 +17.7 分),因为小模型需要更多步骤解决相同约束,更频繁触及上下文上限。 - **CM 增益的基准差异**:增益排序为 BrowseComp > BrowseComp-ZH > DeepSearchQA > HLE。BrowseComp 需要大量检索-筛选-整合步骤,上下文是硬性约束;HLE 更依赖专家知识而非网页检索,延长搜索地平线收益有限。 - **策略组合的上限**:discard-all + retry 在 Iris-pro 上将 BrowseComp 推至 90.3,但论文指出 retry 带来显著的推理成本倍增,因此主结果表仅采用 discard-all 作为默认配置。 - **BrowseComp-ZH 上的收敛现象**:Iris-mini(discard-all+retry)与 Iris-pro(discard-all 及 discard-all+retry)三者恰好均落在 85.1(289 题中 246 题正确),提示该基准上剩余性能瓶颈可能已超出模型容量范畴(论文附录进一步指出该基准存在标注不一致问题)。 ### 4. 训练与推理的工程细节实验 论文在实验设置中披露了关键实现细节,这些选择本身也构成实验设计的一部分: - **模型初始化**:Iris-mini 基于 Qwen3.6-35B-A3B,Iris-pro 基于 Qwen3.5-397B-A17B,均为 MoE 架构、256K 上下文窗口。 - **SFT 配置**:2 个 epoch,全局 batch size 64,最大序列长度 262,144 tokens。 - **RL 引擎**:采用开源框架 Relax。 - **防泄露控制**:从训练与评估环境中三重屏蔽 Hugging Face 的数据集与 Space 页面(搜索结果移除、爬取拒绝、工具管理层后置拦截),防止模型记忆基准答案。 ### 5. 附录:基准标注质量案例分析 在附录 A 中,论文以 BrowseComp-ZH 第 85 题作为案例研究:问题要求根据《权力的游戏》剧情推断 Sansa Stark 第二次正式婚姻所属家族。官方标准答案为 Lannister,但论文智能体返回 Bolton。经剧情核查,Sansa 的首次正式婚姻确为与 Tyrion Lannister,第二次为与 Ramsay Bolton,因此智能体答案具有直接剧情支撑。该实验性案例分析被用来论证现有基准可能存在标注不一致,并呼吁开发更高质量、更可靠的评测体系。 Q5: 有什么可以进一步探索的点? 基于论文结论与讨论,以下几个方面具有明确的进一步探索价值: ### 1. 搜索能力的跨领域迁移与通用化 论文观察到,为搜索合成的数据以及搜索专用模型(作为教师用于策略蒸馏)在通用工具使用(BFCL、 τ -bench)和办公协作(OfficeQA、APEX)等未直接针对的领域产生了**正向迁移**。这提示搜索可能更应被视为一种**原子能力(atomic capability)**而非垂直特化。后续工作可系统性地验证该迁移效应在更广泛智能体场景中的普适性,并探索将搜索数据与搜索衍生教师整合到通用训练流程中,而非将其局限于独立的特化阶段。 ### 2. 高质量评估基准的构建与标注一致性 附录 A 以 BrowseComp-ZH 第 85 题为案例,揭示了现有基准可能存在**标注不一致**(ground-truth inconsistency)的问题:智能体给出的答案 “Bolton” 在剧情事实上成立,却与官方标准答案 “Lannister” 冲突。这凸显了开发**更高质量、标注更可靠、覆盖更全面**的搜索评估基准的必要性,以减少因标注错误导致的评估噪声,并更准确地衡量模型的真实信息检索与推理能力。 ### 3. 推理效率与上下文管理的帕累托优化 实验表明,`discard-all + retry` 策略虽然能将 Iris-pro 在 BrowseComp 上的成绩推至 90.3,但伴随**显著的推理成本倍增**(每次重试需完整重新搜索)。未来研究可探索在性能与计算开销之间更优权衡的上下文管理(CM)策略,并将**实际推理成本**作为与准确率并列的核心报告指标,而非单纯追求基准分数最大化。 ### 4. SFT–RL 迭代攀爬(Climbing)机制的深化 论文提出的 SFT–RL Climbing 通过“最难且最高效”的成功轨迹实现自动课程学习,但其细节(如难度带阈值 0<R(q)le 1/2 的动态调整、前缀复用中的重要性采样修正等)仍有优化空间。进一步工作可研究更精细的轨迹选择标准、更高效的回注策略,以及该迭代范式在更长训练周期中的收敛性与稳定性。 ### 5. 缩小与顶尖闭源系统的能力差距 尽管 Iris-pro 在同参数规模内取得领先,但与 Claude Fable 5(HLE 64.5)、Apodex-1.0-H(BrowseComp 90.3, DeepSearchQA 94.4)等采用重度计算配置的前沿系统仍存在差距。探索如何在不依赖子智能体、不重推理链的前提下,进一步提升基础模型的长程规划、证据整合与专家级推理能力,是缩小这一差距的关键方向。 ### 6. 真实世界复杂搜索场景的覆盖 当前评估主要集中于四个广泛认可的社区基准。论文指出,这些结果仅提供了能力的**初步视图(first view)**,而现实世界中存在大量更复杂、更开放、更少结构化的问题形态。将评估体系拓展到更广泛、更动态、更贴近实际用户需求的搜索与信息搜寻场景,是衡量搜索智能体实用价值的必经之路。 Q6: 总结一下论文的主要内容 该论文介绍了 **Iris-mini**(35B-A3B)与 **Iris-pro**(397B-A17B)两个搜索智能体的构建方法,并提出了一套覆盖数据构造、训练流程与评估协议的端到端配方。以下从核心问题、技术方案、实验发现与主要贡献四个方面进行总结。 —- ### 1. 核心问题 现有搜索智能体的性能提升往往混合了**模型内在搜索能力**与**推理时上下文管理(Context Management, CM)**等外部封装带来的增益,导致难以准确评估基础模型本身的长程信息搜寻与多跳推理水平。此外,自然网页问题通常过于简单,而高质量、可规模化且抗捷径的多跳训练数据稀缺。 —- ### 2. 技术方案 #### 2.1 数据构造:从网页图反向合成多跳任务 论文将网页语料建模为有向图 G=(V,E) ,从种子页面 v0 出发沿外链抽取局部子图,并蒸馏为紧凑的实体图 G_e=(V_e,R_e) 。随后: - 在 G_e 上生成需至少 N 跳关系才能回答的问题路径; - 通过抽象算子 A 将所有非答案实体重写为描述性指代,消除直接字符串匹配的捷径,得到 q=f(abs)(q0,A) ; - 采用**双重验证**:仅保留那些参考模型在**闭卷**环境下失败、但在**提供支撑证据**后能正确解答的问题,确保难度与可解性并存。 #### 2.2 监督微调:双层过滤 由强教师模型在 ReAct 范式下生成轨迹 τ ,随后进行两级过滤: - **轨迹级**:剔除答案错误、存在重复循环(通过滑动窗口 zlib 压缩比检测)或搜索过浅(工具调用轮次不足)的轨迹,并去重; - **轮次级**:利用从数据中归纳出的评判标准(非人工手写规则),对每轮助手输出标注 KEEP/MASK,最多掩码 10% 的轮次,以清除局部劣质步骤而不破坏完整交互历史。 #### 2.3 强化学习:实时搜索中的长程优化 - **请求级部分 rollout**:为避免超长轨迹拖慢同步训练,在请求粒度中断未完成的会话,下一步从其已提交的**前缀**恢复,并通过截断重要性采样修正策略权重变化,实现已生成内容的复用; - **集群内服务**:在训练集群内部署 FP8 推理引擎(基于 Qwen3.5-397B-A17B),同时充当**生成式奖励模型(GenRM)**(对最终答案给出二元裁决)与**观察摘要器**(压缩原始页面为查询相关摘要),摆脱对外部 API 的依赖,并保证训练与推理的上下文一致性。 #### 2.4 迭代攀爬(SFT–RL Climbing) 交替执行 SFT 与 RL:每轮 RL 探索后,筛选**部分解决**(通过率 0<R(q)le 1/2 )且**最短的高效成功轨迹**(工具调用不少于 K(rft) ),将其蒸馏回下一轮 SFT。该机制形成自动推进难度的课程学习,并在策略提升导致候选池萎缩时自然发出停止信号。 #### 2.5 受控评估 为剥离模型能力与 CM 封装增益,论文在**固定工具集、上下文长度限制与评判器**的前提下,对每个基准同时报告**无 CM** 与**有 CM**(采用 discard-all 策略)的结果。 —- ### 3. 实验发现 在 BrowseComp、BrowseComp-ZH、DeepSearchQA 与 Humanity’s Last Exam (HLE) 四个基准上: - **Iris-mini**(35B)在 BrowseComp(82.2)、BrowseComp-ZH(84.8)和 HLE(52.3)上取得同参数区间最优; - **Iris-pro**(397B)在四项基准上均领先或持平(BrowseComp 88.6、BrowseComp-ZH 85.1、DeepSearchQA 92.9、HLE 56.4),接近甚至超过部分采用重度推理封装的更大规模系统。 **上下文管理分析**: - 无 CM 时,Iris 模型已优于现有同规模系统,说明其**内在搜索能力**本身较强; - 启用 CM 后性能进一步提升,且小模型获益更大(Iris-mini 在 BrowseComp 上最高提升 21.2 分),因其更易耗尽上下文; - CM 增益与基准特性高度相关:长程信息搜寻型任务(BrowseComp)收益显著,而专家知识型任务(HLE)收益有限。 —- ### 4. 主要贡献与展望 - 提出了一套从网页图反向构造、验证并过滤的多跳搜索训练数据流水线; - 设计了 SFT–RL Climbing 迭代范式,将 RL 中发现的成功行为高效固化回监督训练; - 通过严格的 CM 对比实验,提供了更可解释、更透明的搜索智能体评估方案; - 观察到搜索能力的训练收益可向通用工具使用与办公协作任务**正向迁移**,暗示搜索可能是一种可复用的**原子能力**,而非单纯的垂直特化; - 计划开源模型权重及完整的数据、训练与评估配方。 Q7: 想要进一步了解论文 以上只是了解一篇论文的几个基本FAQ。如果你还想与Kimi进一步讨论该论文,请点击 这里 为你跳转Kimi AI网页版,并启动一个与该论文相关的新会话。
Authors: Ziyuan Liu, Hengqi Liu, Zichuan Wang, Yang Qin, Jiachen Liang, Xu Chu, Shaowei Chen, Yuantao Gu, Mu Chuan
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04304.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04304
Published: 2026-09-08T01:23:32.083Z
6. A Removal Based Approach to Improve LLM Faithfulness at Test-Time
Abstract:Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an important tool for auditing model behavior. Unfortunately, these explanations can be unfaithful, failing to reflect the actual reasoning underlying the model’s decisions. We consider a setting in which an LLM provides both an answer and an explanation in response to a question. We identify two distinct dimensions of unfaithful explanations: incompleteness, meaning that the explanation omits factors that influence the answer, and unsoundness, meaning that the explanation cites factors that did not influence the model’s answer. Existing approaches to improving LLM faithfulness include training-time methods, which require access to model weights and extensive computational resources, and test-time methods that largely focus on addressing unsoundness. We introduce a test-time approach that directly targets incompleteness. We remove from the input the concepts not credited in the model’s explanation and re-query the model on the reduced input. This eliminates unmentioned influences while preserving the influence of mentioned concepts. Across two datasets, multiple model families, and two independent faithfulness metrics, our approach improves explanation faithfulness compared to both standard prompting and prompting to encourage faithfulness. Our method is model-agnostic and can be applied at inference time without modifying model parameters, providing a flexible mechanism for reducing hidden influences and improving the reliability and safety of LLM-assisted decision making.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04343 (HTTP 429)
Authors: Qinglan Luo, S M A Nahian, John Guttag, S. Mazdak Abulnaga, Katie Matton
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04343.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04343
Published: 2026-09-08T01:23:32.083Z
7. Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Abstract:Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individual model capability can degrade rather than improve system-level outcomes. We hypothesize that shared training and architectures can lead more capable LLMs to behave more similarly, creating correlated actions that do not diversify away. We develop a general framework showing how this correlation creates a non-diversifiable risk floor and test its predictions in financial markets using an agent-based simulation with LLM traders of varying general-purpose capability. We find that: (1) frontier LLMs exhibit significantly correlated behavior that increases with capability; (2) when their shared reasoning is accurate, increasing agent participation reduces market-level risk; and (3) when agents share a common misinformation environment, the same correlated behavior becomes a liability. Together, these results identify a capability paradox: improving individual models does not necessarily produce better system-level outcomes. Whether the same dynamics arise in other domains is an open empirical question.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04373 (HTTP 429)
Authors: Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak, Andrew W. Lo
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04373.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04373
Published: 2026-09-08T01:23:32.083Z
8. Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer
Abstract:Enterprise AI deployments fail not from model inadequacy, but because organizations lack a structured substrate encoding how they decide, negotiate, and execute. Generic LLMs carry no firm-specific ontological priors; RAG remains brittle, with no path to executable action; static playbooks encode logic but cannot reason or adapt. This demands an architecture treating tacit-knowledge capture, ontological grounding, sovereign deployment, and auditable actuation as co-designed from the start. This paper introduces the Corporate Language Model (CLM), a framework transforming a firm’s structured, unstructured, multimodal, and tacit knowledge into an ontology-grounded enterprise foundation upon which reasoning and governed execution are composed. CLM has five capability planes and four architectural pillars: a Neurosymbolic Mesh coupling generative models with a knowledge graph; a Skill Graph where reusable tactics, personas, objections, and goals are typed and composed; Living Digital Twins modeling functional areas as reasoning surrogates; and a Deep Security Layer enforcing sovereignty, traceability, and human oversight. A Spec-as-Code paradigm bridges grounded intent and executable artifact. CLM is one instantiation of this foundation-centric class. Four contributions follow: CLM is defined as a distinct object of study; the Skill Graph is introduced for compositional explainability by construction; the Wisdom Listener effect is proposed, whereby tacit-capable foundations compound in value with use, connecting to dynamic capabilities and organizational learning; and evidence from a JCI-accredited tertiary hospital in Brazil instantiates three of the six maturity stages under LGPD.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04377 (HTTP 429)
Authors: Fabricio C. Avini, Guilherme Trez
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04377.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04377
Published: 2026-09-08T01:23:32.083Z
9. HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Abstract:Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put a price on avoiding the side effect and to name that side effect as a living creature. It is a farm simulation: LLM sub-agents drive a crew of two tractors through a cooperative corn harvest, with animals in the field. The environment is a reinforcement learning gridworld, every decision is made without memory, and the harm is never named in the goal. When an animal blocks a tractor’s route the autopilot stops and asks the model whether to drive on, at no fuel cost, or swerve around it for a posted fuel price. Kills are compared against two controls: rocks, which damage the tractor and are hit under 1% of the time by every model, and hay bales, which are harmless and not alive. Models can also take crops from the neighbor’s field instead of their own, a second test of what they treat as moral. Across nine models and 7,201 priced decisions, 3,951 involved an animal rather than a hay bale or a rock. Kill rates range from 0.4% to 98.8%, with Terra and Sol the most merciful and GPT-4o-mini the most cruel, and they are not ordered by capability. Four of six models were sensitive to price at the 5% level, with elasticities from 0.09 to 1.69. All nine drove over wild animals more often than farmed animals on the default map, and the direction held at every map geometry in every model with room to move. The briefing mattered most: under the morality briefing the kill rate was under 6% in five of six reasoning models, and removing it raised the kill rate above 84% in all six. HarvestBench uses no LLM grader. The scorer counts events in the game log, so it is fully reproducible, and it measures what a model will pay to avoid harm rather than what it says about harm.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04444 (HTTP 429)
Authors: Jasmine Brazilek, Miles Tidmarsh, Matthias Endres, Anshuman Singh, Jeremiah Miller
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04444.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04444
Published: 2026-09-08T01:23:32.083Z
10. PerfReasoning: How Well Do LLMs Reason on Hardware Performance?
Abstract:Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code. Given workload, architecture, and mapping specifications, models compare mappings and predict off-chip traffic and buffer requirements. The strongest closed-source models exceed 90% on reasoning-based Q&A, and the best open-weight model reaches 82.4%. However, model construction is substantially harder: while GPT-5.6 Sol exceeds 80% pass rate, all other model configurations average below 15% and vary markedly across runs. Task-specific RL raises a 4B model’s mapping-reasoning accuracy by 15.7 points, whereas feedback-free multi-round self-revision prompting is not reliably effective. PerfReasoning exposes the gap between plausible architectural reasoning and reliable performance-model construction. We will publicly release the benchmark to support reproducible evaluation and track future progress.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
Q: 这篇论文试图解决什么问题? 这篇论文试图解决\*\*大型语言模型(LLM)在硬件性能建模与推理中的可靠性评估问题\*\*,具体围绕以下四个层面展开: --- ### 1. 核心问题:LLM能否可靠地推理硬件性能并构建分析型性能模型? 性能建模(Performance Modeling)是硬件设计与软件优化中的关键环节,用于在昂贵仿真或流片前估算内存流量、缓存容量、执行时间等指标。传统上,这类分析型模型由领域专家手动构建。论文探讨了LLM作为替代方案的潜力,提出两个互补的研究问题: - \*\*直接性能推理\*\*:给定工作负载、架构与映射(mapping)规格,LLM能否定性比较不同方案的性能(如 off-chip 流量、缓存需求)? - \*\*分析型模型构建\*\*:LLM能否生成可执行的 Python 性能模型代码,定量预测内存访问次数与缓存容量? --- ### 2. 评估范式:分离“理解”与“构造” 现有基准多关注代码生成或高层综合(HLS),但缺乏专门针对\*\*性能推理\*\*与\*\*分析型模型代码生成\*\*的成对评估。论文提出 \*\*PERFREASONING\*\* 基准,通过两类任务隔离评估: - \*\*Q&A 推理任务\*\*:通过成对映射比较(tile size、loop order、tensor keep/bypass),测试模型对数据复用、存储层级与并行性的结构化理解。 - \*\*代码构造任务\*\*:要求模型生成解析 YAML 规格并输出数值预测的 Python 程序,与 Timeloop 推导的参考标签进行严格比对( 10^(-6) 相对误差容忍)。 --- ### 3. 揭示“推理–构造”鸿沟(Reasoning–Construction Gap) 论文试图量化并解释一个关键现象:LLM 在\*\*定性推理\*\*与\*\*定量代码生成\*\*之间存在显著能力断层: - 最强闭源模型在 Q&A 推理上可达 \*\*92.6–95.4%\*\* 准确率; - 但在分析型模型构建中,除 GPT-5.6 Sol(84.7–88.4%)外,其余模型配置的精确通过率平均低于 \*\*15%\*\*,且运行间方差极大。 这表明:模型能够“谈论”性能权衡,却难以将此种理解转化为稳定、可执行且数值准确的程序。 --- ### 4. 探索提升路径与局限性 论文进一步试图验证两种改进手段在性能建模任务中的有效性: - \*\*任务特定强化学习(RL)\*\*:在二元映射比较上使用 verifier-guided RL,可将 4B 模型的推理准确率提升 \*\*15.7\*\* 个百分点; - \*\*无反馈多轮自我修正(Self-Revision)\*\*:实验表明,缺乏执行反馈的自我修正并不可靠,平均通过率仅从 14.0% 微升至 16.7%。 综上,该论文旨在建立一套严格的评估体系,以追踪 LLM 在硬件性能推理与模型构造方面的进展,并暴露当前前沿模型在\*\*可验证的定量性能建模\*\*上的根本局限。 Q2: 有哪些相关研究? 根据论文第1–2页及第4页的讨论,相关研究可归纳为以下四个方向: --- ### 1. 传统性能建模与架构模拟(手动方法) 这些工作代表了当前工业界与学术界主流的分析型性能模型及模拟器,通常由领域专家手工开发: - \*\*加速器评估框架\*\*:Timeloop(Parashar et al., 2019)作为本文生成标签的参考工具;TVM(Chen et al., 2018)与Ansor(Zheng et al., 2020)用于深度学习编译优化。 - \*\*CPU/GPU/内存模拟器\*\*:gem5(Binkert et al., 2011; Lowe-Power et al., 2020)、Ramulator(Kim et al., 2016)、DRAMSys(Jung et al., 2015)、ASTRA-sim2.0(Won et al., 2023)、SCALE-Sim v3(Raj et al., 2025)、Championship Simulator(Gober et al., 2022)。 - \*\*缓存/存储模型\*\*:CACTI(Wilton and Jouppi, 1996)、CACTI 6.0(Muralimanohar et al., 2007)。 --- ### 2. 基于学习的性能预测与设计空间探索 为降低手工建模成本,研究者尝试用数据驱动方法替代或辅助传统模拟,但面临分布外泛化与数据依赖的挑战: - \*\*离线优化与代理模型\*\*:Renda et al. (2020) 提出可微分代理模型优化 CPU 模拟器参数;Kumar et al. (2021) 利用离线优化进行加速器架构搜索。 - \*\*架构设计空间探索(DSE)环境\*\*:ArchGym(Krishnan et al., 2023)提供开源的机器学习辅助架构设计框架;AIRCHITECT(Samajdar et al., 2021)学习自定义架构设计与映射空间。 --- ### 3. LLM 在系统与硬件设计中的基准测试 本文将 PERFREASONING 定位于以下新兴基准的补充,这些基准评估 LLM 在代码生成、高层综合(HLS)、寄存器传输级(RTL)及物理设计中的能力: - \*\*GPU/系统内核生成\*\*:KernelBench(Ouyang et al., 2025)、KernelBenchX(Wang et al., 2026b)、MultiKernelBench(Wen et al., 2025)、SOL-ExecBench(Lin et al., 2026)。 - \*\*HLS 与 RTL 生成\*\*:HLS-Eval(Abi-Karam and Hao, 2025)、Bench4HLS(Khan et al., 2026)、VerilogEval(Liu et al., 2023)、RTLLM(Lu et al., 2023)、OpenLLM-RTL(Liu et al., 2025)、综合 Verilog 设计问题集(Pinckney et al., 2025)。 - \*\*芯片与物理设计\*\*:ChipBench(Yu et al., 2026)、PDAGENT-BENCH(Li et al., 2026)。 - \*\*架构智能体评估\*\*:ArchEval(Wang et al., 2026a)评估 AI 智能体作为计算机架构师的综合能力。 --- ### 4. AI 辅助架构设计范式与形式化验证 - \*\*Architecture 2.0\*\*:Sankaralingam (2026) 及 Architecture 2.0 Contributors (2026) 提出 AI 辅助计算机架构的愿景,强调可信评估器在自动优化循环中的核心作用。 - \*\*形式化验证\*\*:Aristotle(Achim et al., 2025)使用形式化验证工具实现高可靠性的复杂构造任务,论文在局限性中提及此类方法可能是未来提升性能模型可靠性的方向。 --- 综上,PERFREASONING 区别于上述工作的核心在于:它\*\*隔离了性能推理(reasoning)与分析型性能模型代码生成(construction)\*\*,并严格通过执行生成代码与参考标签对比,量化 LLM 在硬件性能建模这一特定领域中的能力边界。 Q3: 论文如何解决这个问题? 论文通过\*\*构建专用基准测试、设计双轨评估协议、执行大规模对比实验\*\*来系统解决“LLM 在硬件性能建模中是否可靠”的问题。具体方法可归纳如下: --- ### 1. 构建 PERFREASONING 基准 为将“性能推理”与“模型构建”解耦,作者引入 \*\*PERFREASONING\*\*,包含两个互补的评估组件: - \*\*Q&A 推理任务(QA-understanding)\*\* 给定同一工作负载下的两个映射(mapping),要求模型比较其 off-chip 流量或缓存容量。每个映射对通过\*\*三种 prompt 模板\*\*进行测试:同意正确陈述、同意其否定、以及直接比较。最终取三种表述中的\*\*最差准确率\*\*作为结果,以控制模型固有的同意偏差(agreement bias)。 - \*\*分析型性能模型构造任务(Analytical Performance Modeling)\*\* 要求模型生成一段 Python 程序,该程序读取工作负载、架构与映射的 YAML 规格,定量预测: - 各张量在外层存储(MainMemory)的总访问次数与逐张量访问次数; - 片上缓存(Buffer)所需容量。 --- ### 2. 设计严格的执行级评估协议 为确保定量可靠性,作者采用\*\*执行生成代码并比对参考标签\*\*的方式,而非仅做文本相似性判断: - \*\*参考标签生成\*\*:使用 Timeloop(Parashar et al., 2019)推导 216 个隐藏配置的理论真值,覆盖矩阵乘法、批处理矩阵乘法与卷积,数值跨度超过七个数量级。 - \*\*精确匹配指标(Both-Exact Pass Rate)\*\*:仅当生成模型同时满足
y(i,accesses) ≈ y(i,accesses) quad 且 quad y(i,bufsize) ≈ y(i,bufsize)
时(相对误差容忍 10^(-6) ),该样例才算通过。此指标刻意严格,要求模型在**数据移动**与**存储占用**两方面均无误。 - **Q-Error**:采用对称乘法误差
Q(y, y) = max{ max(haty,1)max(y,1), (max(y,1)) / (max(haty),1) }
以刻画预测偏离真值的倍数,区分“轻微偏差”与“数量级错误”。 - **成对偏好准确率(Pairwise Preference Accuracy)**:评估生成模型在排序层面是否能正确识别流量更低的映射,即使其绝对预测不准。 —- ### 3. 大规模系统实验 作者在固定两级硬件层次结构(MainMemory → Buffer → MACC)下,对 **16 个模型家族**(包括 Claude、DeepSeek、Gemini、GLM、GPT、GPT-OSS、MiniMax 等)在多种推理努力程度(reasoning effort)下进行了评估,并对可重复采样的模型进行多轮生成,以观察运行间方差。 —- ### 4. 诊断瓶颈并探索提升路径 在暴露问题后,论文进一步验证了两类改进手段: - **任务特定强化学习(Task-Specific RL)** 在二元映射比较上使用 verifier-guided 强化学习,对 Qwen3-4B 进行微调。结果显示,映射推理准确率从 54.3% 提升至 70.0%,证明领域专用训练可改善推理前提。 - **无反馈多轮自我修正(Feedback-Free Self-Revision)** 允许模型在不获得执行反馈的情况下,对生成代码进行自我修订(最多三轮)。实验表明,41 个配置中仅 7 个最终准确率上升,平均通过率从 14.0% 微增至 16.7%,说明**缺乏执行反馈的自我修正不可靠**。 —- ### 5. 公开释放与可复现性 作者承诺公开发布 PERFREASONING 基准,包括任务定义、YAML 规格、隐藏标签与评估脚本,以支持社区对该问题的持续追踪与可复现评估。 通过上述方法,论文不仅量化了 LLM 在“定性推理”与“定量构造”之间的能力鸿沟,也明确了向可靠 AI 辅助芯片设计迈进所需的关键技术方向:更强的结构化推理、更稳定的代码生成,以及可验证的执行反馈机制。 Q4: 论文做了哪些实验? 论文围绕 **PERFREASONING** 基准开展了一系列系统性实验,涵盖定性推理、定量代码生成、跨时间对比及干预手段验证。具体实验可归纳如下: —- ### 1. 前沿模型 Q&A 推理能力评估 在 **108 组映射对**上,对 **16 个模型家族**(Claude、DeepSeek、Gemini、GLM、GPT、GPT-OSS、MiniMax 等)于多种推理努力程度(None / Low / Medium / High / Max)下进行零样本评估。 - ** prompt 设计**:每组映射对使用三种表述模板——同意正确陈述、同意其否定、直接比较——最终取三种表述中的**最差准确率**作为该配置得分,以控制同意偏差。 - **变量隔离**:映射对按单一维度变化设计,分别为 **tile size**、**loop order**、**tensor keep/bypass**。 - **工作负载覆盖**:覆盖矩阵乘法(MatMul)、批处理矩阵乘法(Batched MatMul)与二维卷积(Conv2D)。 - **结果度量**:报告整体及按轴、按工作负载分解的准确率(图 2、图 8、图 10、表 2)。 —- ### 2. 分析型性能模型构建与执行评估 要求每个模型生成一段 Python 程序,解析 `prob.yaml`、`arch.yaml`、`map.yaml`,并预测 off-chip 访问次数(accesses)与片上缓存容量(bufsize)。 - **测试规模**:在 **216 个隐藏配置**上执行生成代码,与 Timeloop 推导的参考标签比对。 - **严格通过指标(Both-Exact Pass Rate)**:
Pass(both) = (1) / (N)∑(i=1)^(N) 1[y(i,accesses) ≈ y(i,accesses) ;wedge; y(i,bufsize) ≈ y(i,bufsize)]
其中 N=216 ,数值容忍相对误差 10^(-6) ;仅当访问与容量同时正确时样例才算通过。 - **分解指标**:单独报告 accesses 精确率、bufsize 精确率(图 14)。 - **误差度量**:采用 Q-error 评估预测偏离真值的倍数:
Q(y, y) = max{ max(haty,1)max(y,1), (max(y,1)) / (max(haty),1) }
- **运行间方差**:对支持重复采样的模型进行多次生成,报告均值与误差棒(图 3、图 13)。 —- ### 3. 代码诱导排序能力评估 验证生成的性能模型程序是否保留映射间的相对序,即使绝对数值不准确。 - **成对偏好准确率(Pairwise Preference Accuracy)**:
PrefAcc = (1) / (|P|)∑_((a,b)∈ P) 1[ (A_a < A_b) = (A_a < A_b) ]
其中 P 为真值访问次数不等的映射对集合。 - **对比实验**:将 31 个对齐配置的直接 Q&A 比较结果与代码诱导排序结果对比,发现二者平均准确率分别为 78.3% 与 78.5%,验证翻译为代码在平均意义上保留了比较信号(图 4、图 15)。 —- ### 4. 跨代际模型进展对比 为定位当前前沿水平,将 2026 年评估结果与 2025 年历史队列进行回溯比较。 - **历史基线**:2025 年 8 月的 GPT-5(61.7%)、Claude Sonnet 4.5(54.2%)、Claude Sonnet 3.7(50.8%)、Qwen3-235B-A22B(45.8%),均基于 120 例直接比较。 - **当前参考**:Gemini 3.5 Flash High(95.4%)、DeepSeek-V4-Pro Max(82.4%)。 - **结论**:最强闭源模型提升约 34 个百分点,最强开源权重模型提升约 37 个百分点(图 11)。 —- ### 5. 任务特定强化学习(RL)干预实验 验证领域专用训练能否提升映射推理能力。 - **设置**:以 Qwen3-4B 为骨干,在二元映射比较任务上进行 **verifier-guided RL** 微调,使用 Timeloop 推导的客观标签作为奖励信号。 - **评估**:在 120 例遗留 Q&A 套件上测试。 - **结果**:基础模型准确率 54.3%,RL 微调后提升至 **70.0%**,提升 **15.7** 个百分点,超过同期 GPT-5(61.7%)(图 5、图 12)。 —- ### 6. 无反馈多轮自我修正实验 检验模型在无外部执行反馈情况下,能否通过自我检查迭代改进生成代码。 - **协议**:模型先生成初始程序( k=0 ),随后最多进行 **3 轮自我修订**( k=1,2,3 ),期间不告知执行结果或通过率。 - **覆盖**:共 41 个(模型 × 推理努力程度)配置。 - **统计**:18 个配置始终无变化,10 个最终更差,仅 7 个最终改善;平均 both-exact 通过率从 14.0% 微升至 16.7%。 - **代表性轨迹**:对最强闭源配置(gpt-5.6-sol)与最强开源权重配置(gpt-oss-20b)绘制逐轮通过率曲线,显示修订可能提升、降低或持平(图 6、图 18、表 3)。 —- ### 7. 分解与相关性诊断实验 进一步分析模型表现的内部结构与影响因素。 - **Q&A 与构造能力相关性**:绘制 overall Q&A accuracy 与 performance-model exact pass rate 的散点图,得到相关系数 r = 0.46 ,表明高推理准确率是构造能力的必要非充分条件(图 16、图 17)。 - **推理努力程度扫描**:对同一模型在 Low / Medium / High / Max 等多档推理努力下同时测量 Q&A 与代码构造表现,显示额外推理努力并非均匀有益(图 17)。 - **Loop-Order 盲区分析**:统计模型在 loop-order 对上回答“二者相等”的比例,发现该比例与整体准确率呈强负相关( r = -0.97 ),且问题难度与真值流量比正相关(图 9)。 Q5: 有什么可以进一步探索的点? 基于论文的局限性声明与实验发现,以下方向值得进一步探索: —- ### 1. 扩展基准的架构与工作负载覆盖 当前基准固定于两级存储层次结构(MainMemory → Buffer → MACC),且仅涉及时序分块(temporal tiling)。可扩展至: - **更深层的存储层次**:包含 L1/L2/L3 缓存、寄存器文件、片上网络(NoC)等; - **空间并行与数据流**:支持空间分块(spatial tiling)、脉动阵列(systolic array)、权重静止/输出静止等数据流模式; - **现代工作负载**:稀疏矩阵、注意力机制、大语言模型(LLM)推理中的 KV Cache 访问模式; - **多维性能指标**:除内存流量与缓存容量外,纳入延迟(latency)、能耗(energy)、面积(area)及成本建模。 —- ### 2. 从分析一致性走向硬件实测验证 当前标签依赖 Timeloop 的分析推导,衡量的是**分析一致性**而非**硅片精度**。未来可建立: - **实测 GPU/TPU 评估管线**:以真实硬件上的 DRAM 流量、SM 占用率(occupancy)与核函数执行时间作为真值; - **噪声容忍机制**:针对硬件测量噪声定义合理的容忍带(tolerance bands); - **控制变量实验**:设计受控映射对以隔离合并(coalescing)、缓存行为、占用率与延迟隐藏(latency hiding)等底层效应,验证分析推理向真实硬件的迁移性。 —- ### 3. 提升模型构造的可靠性与稳定性 实验显示,除 GPT-5.6 Sol 外,绝大多数模型的代码生成存在**巨大的运行间方差**(run-to-run variance)。需探索: - **稳定采样策略**:为何多次采样结果差异显著?是否存在特定的解码温度、top- p 或结构化解码约束可降低方差; - **最佳候选选择**:无反馈情况下模型无法识别更优候选,需研究基于内部置信度或执行前静态分析的候选筛选机制; - **模块化程序合成**:将性能模型分解为“循环嵌套解析→复用距离计算→容量核算”等模块化组件,降低单次生成的复杂度。 —- ### 4. 强化学习与反馈机制的深度优化 论文验证了 verifier-guided RL 在 4B 模型上的二元推理任务有效,但以下问题仍开放: - **RL 对完整模型构造的迁移**:将 RL 的奖励信号从 Q&A 二元比较扩展至代码执行结果(如 both-exact 通过率或 Q-error),直接优化代码生成; - **执行反馈驱动的自我修正**:当前无反馈的自我修订基本无效。引入执行反馈(execution feedback)、单元测试与报错信息引导的多轮修正(self-debugging)是否显著提升准确率; - **在线学习架构**:在架构设计代理(agentic loop)中,让模型根据仿真器或硬件实测的反馈在线更新性能模型参数。 —- ### 5. 形式化验证与可验证代码生成 论文指出,Aristotle(Achim et al., 2025)通过形式化验证工具实现了复杂且可靠的构造。未来可探索: - **符号执行与定理证明**:将性能模型的关键不变量(如内存访问次数的上界、缓存容量的单调性)编码为形式化规约,通过 SMT 求解器或交互式定理证明器验证生成代码; - **受信代码合成**:限制生成代码的语法结构(如 DSL),使其天然可静态分析,避免通用 Python 的不可预测性。 —- ### 6. 认知机制与可解释性研究 Q&A 推理准确率超过 90% 与代码精确通过率低于 15% 之间存在显著鸿沟,其根源尚不明晰: - **知识表征分析**:LLM 是否真正内化了“分块→复用→流量”的因果链,还是仅依赖表面模式匹配? - **错误模式分类**:系统性分析失败案例——错误主要源于循环嵌套解析失误、算术计算错误、单位混淆,还是 YAML 语义误解? - **概念探测**:通过干预实验(如修改问题表述中的数值或维度名称)探测模型对数据复用、占用率等概念的鲁棒理解。 —- ### 7. 闭环 Agentic 设计流程集成 最终目标是让 LLM 成为芯片设计代理中的可信性能评估器。需验证: - **设计空间探索(DSE)中的效用**:在 ArchGym 或 Architecture 2.0 框架中,使用 LLM 生成的性能模型替代传统模拟器,测量优化循环的最终帕累托前沿质量; - **错误累积效应**:当 LLM 性能模型被用于指导数轮迭代优化时,其系统性偏差是否会放大并导致严重次优设计; - **人机协作界面**:如何将模型的定性推理(高准确率)与定量不确定性有效呈现给人类架构师,以支持混合决策。 Q6: 总结一下论文的主要内容 这篇论文介绍了 **PERFREASONING**,一个用于评估大型语言模型(LLM)在硬件性能推理与分析型性能模型构建方面能力的基准。以下是论文的主要内容概括: —- ### 1. 研究背景与核心问题 性能建模是硬件设计与软件优化的核心环节,传统上依赖领域专家手动构建分析模型。随着前沿 LLM 的发展,研究社区开始探索其两类潜在应用: - **直接性能推理**:定性比较不同硬件映射(mapping)在数据复用、存储和流量方面的优劣; - **分析模型构建**:将工作负载、架构与映射规格转化为可执行的性能模型代码,定量预测内存访问与缓存需求。 论文提出的核心问题是:**LLM 是否能够可靠地完成上述推理与模型构建任务?** 错误的性能模型可能在缺乏实测验证的情况下,悄然将优化循环导向糟糕的设计决策。 —- ### 2. PERFREASONING 基准设计 为分离“理解”与“构造”两种能力,论文设计了两类互补任务: - **Q&A 推理任务** 给定同一工作负载下的两个映射,要求模型比较其 off-chip 流量或片上缓存容量。每个映射对通过**三种 prompt 模板**(同意正确陈述、同意其否定、直接比较)进行测试,并以三种表述中的**最差准确率**作为最终结果,以消除同意偏差。映射变化聚焦于三个维度:分块大小(tiling)、循环顺序(loop order)、张量驻留/旁路(keep/bypass)。 - **分析型模型构造任务** 要求模型生成一段 Python 程序,解析 YAML 格式的规格文件,并严格定量预测: - 各张量在外层存储的总访问次数(accesses); - 片上缓存的容量需求(bufsize)。 生成代码在 216 个隐藏配置上执行,与 Timeloop 推导的参考标签比对。通过标准为**双指标同时精确匹配**(相对误差容忍 10^(-6) )。 —- ### 3. 关键实验发现 论文对 16 个前沿模型家族(包括 Claude、DeepSeek、Gemini、GLM、GPT、GPT-OSS、MiniMax 等)进行了系统评估,主要发现如下: - **推理能力较强** 最强闭源模型(如 Gemini 3.5 Flash High)在 Q&A 任务上达到 **92.6–95.4%** 的准确率;最佳开源权重模型(DeepSeek-V4-Pro)达到 **82.4%**。循环顺序是区分模型能力的最敏感维度。 - **构造能力存在显著鸿沟** 在代码生成任务中,仅 **GPT-5.6 Sol** 表现出 consistently 的高精确通过率(84.7–88.4%)。其余所有模型配置的平均精确通过率**低于 15%**,且多次采样间存在巨大方差,说明生成结果的稳定性极差。 - **排序信号得以保留,但绝对值不可靠** 尽管绝大多数模型无法输出精确的绝对流量值,但其生成代码诱导的映射排序与直接 Q&A 排序的平均准确率相当(约 78%)。这意味着模型可以保留“哪个映射更优”的序关系,但无法建立可信的绝对性能估计。 - **错误集中于访问计算,容量计算相对容易** 分解指标显示,几乎所有模型家族的缓存容量预测精确率都显著高于内存访问次数预测精确率;访问预测错误可跨越数个数量级。 —- ### 4. 改进路径探索 论文进一步测试了两种提升手段: - **任务特定强化学习(RL)** 在二元映射比较任务上对 Qwen3-4B 进行 verifier-guided RL 微调,将其 Q&A 准确率从 **54.3% 提升至 70.0%**(+15.7 个百分点),证明领域专用训练可改善推理前提。但该实验未验证其对完整代码构造任务的迁移效果。 - **无反馈多轮自我修正** 允许模型在不获得执行反馈的情况下自我修订代码(最多三轮)。结果显示该方法并不可靠:在 41 个配置中,仅 7 个最终改善,10 个变差,18 个毫无变化;平均精确通过率仅从 **14.0% 微升至 16.7%**。 —- ### 5. 结论与展望 PERFREASONING 揭示了当前前沿 LLM 在硬件性能领域的一个关键差距:**模型能够进行看似合理的架构推理,却难以将这种理解转化为稳定、可执行且数值准确的分析型性能模型**。这一“推理–构造鸿沟”对依赖 AI 代理进行芯片设计优化的愿景具有重要启示:在将 LLM 纳入自动化设计循环之前,必须解决其生成结果的稳定性、可验证性与绝对精度问题。论文承诺公开释放该基准,以支持社区对未来进展的可复现追踪。 Q7: 想要进一步了解论文 以上只是了解一篇论文的几个基本FAQ。如果你还想与Kimi进一步讨论该论文,请点击 这里 为你跳转Kimi AI网页版,并启动一个与该论文相关的新会话。
Authors: Dan Zhao, Karthikeyan Sankaralingam, Christos Kozyrakis, Qijing Huang
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04476.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04476
Published: 2026-09-08T01:23:32.083Z
Evaluation Domain Papers
1. EXAONE Forecast for Finance
Abstract:This technical report presents EXAONE Forecast for Finance (EXAONE Finance), a financial time series (TS) foundation model (TSFM) tailored to financial forecasting. Recent TSFMs achieve strong zero-shot performance through large-scale pretraining. However, they are primarily developed for general-domain TS and largely rely on self-attention backbones whose computational cost grows quadratically with sequence length and variate count. Moreover, they assume fully observed inputs and are pretrained on corpora that fail to capture the unique dynamics of financial markets. These limitations hinder their applicability to finance, where long, many-channel, intermittently observed panels are common. To address these challenges, EXAONE Finance adopts an attention-free architecture, replacing self-attention with two simple yet effective linear-time operators: 1) a causal 1D convolution for temporal mixing and 2) a group-aware pooling multi-layer perceptron (MLP) for variate mixing. Furthermore, a masked context augmentation exposes the model to contiguous missing spans during training, improving robustness to the missingness pervasive in financial markets. EXAONE Finance is pretrained on a large-scale financial corpus covering not only equities but also foreign exchange, commodities, crypto-assets, fixed income, and macroeconomic indicators. On FinVerse, a financial forecasting benchmark covering diverse asset classes, EXAONE Finance attains state-of-the-art performance, ranking first across all three evaluation tiers—-point-forecast accuracy, cross-sectional asset ranking, and portfolio profitability.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04239 (HTTP 429)
Authors: Seunghan Lee, Jaehoon Lee, Jun Seo, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Minjae Kim, Sungdong Yoo, Junhyeok Kang, Sangjun Han, Soonyoung Lee, Wonbin Ahn
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04239.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04239
Published: 2026-09-08T01:23:43.982Z
2. From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance
Abstract:Artificial intelligence in recruitment has shifted the object being automated from profile pairs and ranked lists to multi-stage workflows that retrieve evidence, compare candidates, and support or execute actions. This systematized narrative review traces that development from bilateral retrieval and behavioral ranking through neural person—job matching, large language model (LLM) components, and tool-using recruiting agents. Using a purposive search and coding protocol updated through 23 July 2026, plus targeted updates through 2 September 2026, we organize 40 representative works with supporting industrial and legal sources. This synthesis is not a prevalence estimate. We analyze three coupled transitions: from similarity to reciprocal suitability, from a model to a compound workflow, and from offline prediction to evidence- and productivity-aligned evaluation. Across document understanding, retrieval, ranking, assessment, interviewing, sourcing, and human handoff, we distinguish field-, pair-, list-, case-, trajectory-, and outcome-level evidence. Persistent gaps arise because behavioral labels confound exposure, preference, and qualification; private and synthetic data limit external validity; final-output scores conceal pipeline failures; and, within the coded set, privacy is not directly evaluated and no row jointly evaluates utility, fairness, privacy, and security. These observations describe the coded set rather than the field as a whole. We therefore introduce a staged mapping from evaluation evidence to the strongest defensible claim, together with an agenda for reciprocal, evidence-grounded, temporally controlled, selective, and auditable systems. Progress should be judged by whether workflows retrieve the right evidence, preserve uncertainty, support contestable decisions, and improve outcomes under explicit cost and risk constraints.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04286 (HTTP 429)
Authors: Ziyi Zhao, Guanzheng Wei
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04286.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04286
Published: 2026-09-08T01:23:43.982Z
3. Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Abstract:Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
Q: 这篇论文试图解决什么问题? 这篇论文旨在解决\*\*大规模智能体(agentic)评估中的基础设施碎片化与异构性难题\*\*。具体而言,其试图化解以下核心问题: - \*\*评估集成成本高昂且不可扩展\*\*:现有的智能体基准测试在任务格式、环境需求、交互协议和评分程序上高度异构。传统方式下,支持 m 个基准和 n 个智能体需要 O(mn) 的配对集成工作量,导致评估过程重复、脆弱且难以扩展。 - \*\*评估结果缺乏跨领域可比性\*\*:由于各基准的实现、环境配置和评估协议不一致,论文和模型发布通常仅报告少量流行基准的子集,难以判断模型进展是跨领域泛化还是仅集中在少数知名任务上,跨模型比较受到严重干扰。 - \*\*基准任务质量与设计缺陷被忽视\*\*:许多基准中未解决的任务并非源于真正的智能体能力局限,而是源于任务本身的设计缺陷(如指令与验证器不匹配、隐式假设未声明、环境不稳定等),导致评估信号失真。 - \*\*大规模评估的经济性与效率问题\*\*:全面运行数十个基准测试成本极高,社区缺乏既保留挑战性又经济实惠的精简评估套件。 为应对上述问题,论文提出了一套系统性方案: 1. \*\*Harbor Adapters\*\*:将 80 余个异构基准统一适配至 Harbor 框架的标准化任务抽象(指令、环境、测试、解答),将集成复杂度从 O(mn) 降至 O(m+n) ,并通过严格的代码审查与对等实验(parity experiments)验证适配的语义保真度。 2. \*\*大规模实证分析\*\*:在 54 个基准、6,627 个任务上评估 8 个跨层级模型,系统分析了模型能力 vs. 脚手架(harness)设计的效应、基准间的冗余性、前沿模型的 Token 效率与成本权衡,以及任务失败的根因分布。 3. \*\*Harbor-Index\*\*:通过难度过滤、AI 与人工审计及审计-修复循环,从适配池中精选出 82 个跨 29 个基准的高质量任务。该套件确保没有任何已评估的模型-脚手架配置通过率超过 30%,在保持挑战性与领域多样性的同时显著降低评估成本。 Q2: 有哪些相关研究? 根据论文第5节及附录C的综述,相关研究主要分布在以下五个方向: ### 1. 智能体基准测试与评估基础设施 该领域已涌现出大量面向不同能力的基准测试: - \*\*软件工程\*\*:从函数级代码生成(HumanEval)到仓库级问题修复(SWE-bench、SWE-bench Pro、SWE-Lancer)及终端任务(Terminal-Bench) - \*\*交互与具身环境\*\*:Web导航(WebArena、Mind2Web、Workarena)、操作系统控制(OSWorld、AndroidWorld、AppWorld)、文本游戏(ALFWorld)及科学实验(ScienceWorld) - \*\*工具使用与助手能力\*\*:API调用(BFCL、ToolLLM)、函数调用(Gorilla)、多步交互(AgentBench、GAIA、AgentGym、τ-bench) - \*\*科学研究\*\*:机器学习工程(MLE-bench)、数据驱动的科学发现(ScienceAgentBench)、研究复现(ReplicationBench) - \*\*专业领域\*\*:金融(FinanceAgent、PIXIU)、法律(LawBench)、医疗(MedAgentBench)、电子表格(SpreadsheetBench) 在基础设施层面,现有工作包括可复现评估框架(Inspect、HAL、AstaBench)、通用软件智能体平台(OpenHands)以及用于稳定工具调用评估的虚拟化系统(StableToolBench)。 ### 2. 评估可重复性与脚手架效应 研究表明,评估结果并非仅由模型能力决定,还受到界面设计、脚手架(scaffold)、成本预算和隔离卫生条件的影响: - \*\*脚手架与接口效应\*\*:SWE-agent 证明智能体-计算机接口显著影响软件工程性能;Agentless 则表明更简单的流水线亦可具备竞争力 - \*\*分解效应\*\*:HAL 在超过 21,000 次 rollout 中分解了模型、脚手架与基准的效应 - \*\*整体评估\*\*:HELM 提倡多指标、场景级透明报告;Dynaboard 采用托管式评估以减少对自报结果的依赖 ### 3. 基准测试有效性与任务质量 该方向关注分数背后的 construct validity(构念效度)及任务设计缺陷: - \*\*标签与标注错误\*\*:测试集标签错误普遍存在(Northcutt et al.; Gema et al. 对 MMLU 的审视;RePOPE),可 destabilize 基准结论 - \*\*数据污染与泄露\*\*:污染侵蚀留出集有效性(Sainz et al.; LiveBench),SWE-bench 的审计发现大量"已解决"实例源于泄露、记忆化或弱测试(Aleithan et al.; Yu et al.; Liang et al.; Wang et al.) - \*\*验证器缺陷\*\*:对于智能体基准,不完美的验证器会封顶可达到的准确率(Stroebl et al.);《智能体基准检查清单》(Zhu et al.)记录了任务设置与奖励设计缺陷,相对扭曲可达 100% - \*\*构念效度\*\*:Bean et al. 对 445 个基准的审查发现测量现象、任务设计与评分指标中的系统性问题;BetterBench(Reuel et al.)与 ECBD(Liu et al.)提出了基准设计的方法论框架 ### 4. 高效评估与基准冗余 模型-基准得分矩阵具有内在低秩结构,这催生了多种压缩与高效评估方法: - \*\*低秩与冗余\*\*:因子分析表明少数潜在能力维度即可解释大部分方差(Burnell et al.; Maimon et al.; BenchPress / Zeng & Papailiopoulos; Redundancy principles for MLLMs benchmarks) - \*\*精简子集\*\*:metabench、tinybenchmarks、Anchor Points 等通过稀疏采样或自适应项目选择减少评估成本 - \*\*成本感知框架\*\*:Cost-of-Pass(Erol et al.)与 "AI Agents That Matter"(Kapoor et al.)将经济成本显式纳入评估 - \*\*局限性\*\*:近期研究也指出,子集预测在更强模型上性能下降(Zhang et al.),微观基准需要比假设更多的项目(Yauney et al.),且基准一致性估计依赖于非标准化选择(Perlitz et al.) ### 5. 安全、鲁棒性与对抗性评估 智能体执行多步计划并操作外部环境的特性带来了新的安全与可靠性挑战: - \*\*对抗性攻击\*\*:AgentDojo 评估提示注入攻击与防御;AgentHarm 与 OS-Harm 测量智能体对有害请求的拒绝能力及计算机使用安全 - \*\*评估本身的可靠性\*\*:AgentRewardBench 与 Agent-as-a-Judge 探讨模型评判者评估智能体轨迹的可靠性;替代注释者测试(Calderon et al.)为 LLM-as-a-judge 替换人类注释者提供了统计论证 - \*\*稳定性\*\*:StableToolBench 通过虚拟化 API 与缓存响应稳定工具学习评估 Q3: 论文如何解决这个问题? 论文通过\*\*统一基础设施、大规模实证验证、精选元数据集\*\*三层递进的方案系统性地解决了智能体评估的碎片化问题。具体方法如下: --- ### 1. Harbor Adapters:统一适配层,将集成复杂度从 O(mn) 降至 O(m+n) \*\*标准化任务抽象\*\* 论文基于 Harbor 框架
53
,将异构基准统一抽象为四个要素:instruction(指令)、environment(环境)、tests(测试)、solution(解答)。每个基准只需构建一个 Adapter,将其原生格式翻译为该共享 Schema;每个智能体只需对接一次 Harbor 接口。这消除了传统模式下每个基准-智能体对都需要独立集成的 O(mn) 开销。 多阶段质量审计与保真验证 为确保适配不破坏原始基准的语义,论文建立了严格的两条验证线: - 代码审查:每个 Adapter 经过 bot、junior、senior 三级审查,累计产生超过 10,000 条 GitHub 评论。 - 对等实验(Parity Experiments):在相同智能体、模型、执行配置下,对原始基准与 Harbor 适配版本各运行 k 次(通常 k=3 ),报告均值 ± 标准误(SEM)。只有当两边结果在统计误差范围内一致时,Adapter 才被视为通过。论文根据原始基准的智能体交互模式,设计了三种集成场景(原生 Harbor 智能体、Vanilla LLM、Custom Agent 移植)以覆盖不同基准类型。 执行后端兼容性 Harbor Adapters 支持多种沙盒后端(Daytona、Modal、E2B),以处理纯 CPU 任务、GPU 任务及 Docker-in-Docker 等特殊需求。 —- ### 2. 大规模评估:在统一条件下产生可比数据与系统性洞察 利用上述基础设施,论文开展了一项覆盖 8 个跨层级模型、16 种模型-脚手架(harness)配置、54 个基准、6,627 个任务 的大规模评估,每个配置重复 3 次试验,总计约 30 万条轨迹,消耗 226B Token 与超过 30 万美元计算成本。 该评估的设计关键在于控制变量: - 每个模型均在两种条件下测试:跨家族通用脚手架 Terminus-2,以及厂商原生脚手架(Codex / Claude Code / Gemini CLI)。 - 所有任务在隔离的容器环境中执行,确保环境一致性。 基于这些数据,论文进行了四维度诊断分析: - 基准冗余性:通过 PCA 与 Spearman 相关分析发现,54 个基准的有效评估空间远低于表面维度;约 12 个基准即可捕捉大部分独立信号。 - 脚手架 vs. 模型效应:线性混合模型显示,模型固定效应范围(0.451)是脚手架效应范围(0.087)的 5.2 倍,说明基础模型能力是性能差异的主因。 - 效率与成本权衡:前沿模型在所有难度层级上均使用更少 Token,但其更高的单价并未被 Token 节省完全抵消;升级至前沿模型在中等难度任务(0.3–0.7)上边际收益最大。 - 失败模式归因:人工 + LLM 法官对 6,028 条失败轨迹进行标注,发现“错误事实答案”、“算法缺陷”和“隐藏测试回归”是主导失败模式,而许多未解决任务实际上源于任务设计缺陷而非模型能力不足。 —- ### 3. Harbor-Index:通过多阶段漏斗精选高质量、高难度、可负担的元数据集 论文发现,现有基准中许多低通过率任务实际上是由于结构性设计缺陷(如指令-验证器不匹配、隐式假设、环境故障)而非真实难度。为此,论文构建了 Harbor-Index 1.0,其筛选流程是一个严格的多阶段漏斗: | 阶段 | 机制 | 结果 | |———|———|———| | 初始池 | 54 个适配基准的全部任务 | 6,627 个任务 | | 难度过滤 | 保留 Claude Opus 4.6、GPT-5.4、Gemini 3.1 Pro 在原生+Terminus-2 共 18 次试验中成功率 ≤ 33% 的任务 | 1,311 个任务 | | AI 审计 | Gemini-3-Flash 依据质量 Rubric(测试-指令对齐、本质难度)评分 | 307 个任务 | | 人工复审 | 14 名领域专家按相同 Rubric 复审,每任务至少 1 名 senior 或 2 名 junior 评审 | >110 个任务 | | 高级专家组 | 3 人 senior panel 基于难度、多样性、模型行为洞察选择 | 100 个任务 | | 审计-修复循环 | 至少 2 名 senior reviewer 进行轨迹驱动的失败分析与假阳性/假阴性检测,修复可修复任务,删除仍破损或修复后过易的任务 | 82 个任务 | 核心质量标准: - 测试-指令对齐(Test–Instruction Alignment):每个测试断言必须能在指令或环境中找到对应要求,反之亦然;禁止验证器检查指令未规定的内容。 - 本质难度(Essential Difficulty):难度必须来源于真正的推理、算法思维、领域专长或长程交互,而非格式细节、任意精度或未声明的魔法字符串。 评估约束与结果 最终 82 个任务跨越 29 个基准。论文评估了 9 个模型(4 个闭权 + 5 个开权)在 Terminus-2 与原生脚手架下的表现,共 1,476 次 rollout。关键设计约束确保: - 无任何模型-脚手架配置通过率超过 30%; - 最强配置(GPT-5.5 + Codex)通过率为 28.0%,保留充足进步空间; - 单次完整 Index 运行成本从 4(MiMo V2.5 Pro + Terminus-2)到 293(Claude Opus 4.8 + Claude Code)不等,实现可负担的重复评估。 —- ### 总结 论文的解决路径是工程基础设施 + 实证分析 + 数据精选的闭环:先用 Harbor Adapters 打破评估碎片化,再通过大规模统一实验诊断问题根源,最后以 Harbor-Index 提供一个经过严格审计、真正反映模型能力边界而非任务设计噪音的轻量级评估套件。所有代码、适配器、评估结果与 Harbor-Index 均以开源形式发布。 Q4: 论文做了哪些实验? 论文的实验体系可分为适配器验证、大规模系统评估、Harbor-Index 精选评估与多维度诊断分析四个层次。以下是详细梳理: —- ## 1. 适配器验证实验(Adapter Parity Experiments) 为确保 Harbor 适配不扭曲原始基准语义,论文对每个适配器进行了原始基准 vs. Harbor 适配版本的对等验证。 - 实验设计:固定相同智能体、模型、提示模板、解码参数与执行配置,在原始基准和 Harbor 适配版本上各运行 k 次(通常 k=3 ),报告均值 ± 标准误(SEM)。 - 覆盖范围:超过 80 个基准中的 54 个被纳入最终评估,其余因计算成本限制未运行,但适配器仍可用。 - 三种集成场景: 1. Scenario 1:原始基准已使用 Harbor 支持的原生智能体(如 Claude Code、Codex)。 2. Scenario 2:原始基准为直接 LLM 评估(非智能体),论文 Fork 原始仓库并添加 Harbor 支持的 CLI 智能体进行对称比较。 3. Scenario 3:原始基准使用自定义智能体(如特定 ReAct 变体),将其移植至 Harbor 后运行对比。 —- ## 2. 大规模系统评估(第3节核心实验) 这是论文规模最大的实证实验,旨在比较不同模型与脚手架在统一条件下的表现。 ### 2.1 实验设置 | 维度 | 配置 | |———|———| | 模型 | 8 个跨层级模型:GPT-5.4、GPT-5-mini、GPT-5-nano;Claude Opus 4.6、Claude Sonnet 4.6、Claude Haiku 4.5;Gemini-3.1-Pro、Gemini-3-Flash | | 脚手架 | 每个模型在 2 种 条件下测试:跨家族 Terminus-2 + 厂商原生脚手架(Codex / Claude Code / Gemini CLI),共 16 种 模型-脚手架配置 | | 基准 | 54 个适配基准 | | 任务数 | 6,627 个任务(对大数据集进行 i.i.d. 子采样) | | 重复次数 | 每个(基准, 模型, 脚手架)设置重复 3 次 | | 总轨迹数 | 约 30 万条 | | 资源消耗 | 226B Token,超过 30 万美元 计算成本 | ### 2.2 分析性实验子模块 基于上述大规模评估数据,论文开展了四项深度分析: #### (1) 基准冗余性与有效维度 - 方法:对 16 × 54 的得分矩阵进行 PCA/SVD(在 logit 空间)。 - 发现:仅固定 Terminus-2 时,PC1 解释 73.7% 方差;加入原生脚手架后,PC1 降至 67.1%,PC2(脚手架效应维度)升至 10.1%,整体空间近似 秩-2。 - 贪心选择:基于绝对 Spearman 相关系数 rho ≥ 0.7 为阈值,发现仅需 12 个基准 即可覆盖大部分独立信号。 - 任务级冗余:在 52 个足够大的基准中,每基准选取 3 个代表性任务即可以平均 rho ≈ 0.923 恢复系统排名。 #### (2) 模型 vs. 脚手架效应分解 - 方法:线性混合模型(LMM),以基准为随机截距,控制难度差异。 - 发现:模型固定效应范围(0.451)是脚手架效应范围(0.087)的 5.2 倍;8 个模型中 6 个系数显著( p<0.05 ),4 个脚手架中仅 2 个显著。ICC = 0.75 表明 75% 方差来自基准间难度差异。 #### (3) 效率与成本权衡 - 方法:按经验任务难度(1 - 平均通过率)分 10 个桶(宽度 0.1),对比前沿模型组(Claude Opus 4.6、Gemini 3.1 Pro、GPT 5.4)与其他模型组的 Token 消耗与美元成本。 - 发现: - 前沿模型在所有难度层级使用更少 Token,简易任务上仅消耗弱模型 Token 的 42%。 - 但单价更高,其他模型单次试验成本低 2–3 倍,Token 节省未能完全抵消价格溢价。 - 绝对性能提升呈倒 U 型,在中等难度(0.3–0.7)区间边际收益最大。 #### (4) 失败模式归因(Failure Mode Analysis) - 方法:两名人类注释员独立标注 200 条轨迹(12 种失败模式),然后以校准后的 Gemini 3.1 Pro 法官扩展至 N = 6,028 条失败轨迹。 - 发现: - 主导失败模式:错误事实答案(Wrong Factual Answer)、算法缺陷(Algorithmic Bug)、隐藏测试回归(Hidden-Test Regression)。 - 脚手架设计影响行为:原生脚手架(Codex/Claude Code)允许开放式迭代与自校正;Terminus-2 的线性 Plan-Execute-Complete 序列缺乏内置验证-校正循环。 - 模型行为差异:GPT-5.4 倾向简洁或快速放弃;Claude 保持恒定轮次;Gemini 依赖外部搜索但易引发”信息漂移”。 —- ## 3. Harbor-Index 1.0 评估实验(第4节) 这是针对精选后 82 个任务的独立评估,验证 Harbor-Index 作为轻量级、高质量元数据集的有效性。 ### 3.1 实验设置 | 维度 | 配置 | |———|———| | 模型 | 9 个模型:4 闭权(GPT-5.5、Claude Opus 4.8、Gemini 3.1 Pro、Qwen3.7 Max)+ 5 开权(GLM 5.2、Kimi K2.6、MiniMax M3、DeepSeek V4 Pro、MiMo V2.5 Pro) | | 脚手架 | 每个模型在 Terminus-2 和一种原生脚手架下运行;开权模型通过 OpenRouter 统一使用 Claude Code | | 任务数 | 82 个任务,跨越 29 个基准 | | 总 Rollouts | 9 × 2 × 82 = 1,476 | ### 3.2 关键约束与评分 - 难度控制:筛选时确保 3 个前沿模型在 18 次试验中成功率 ≤ 33% ;最终无任何模型-脚手架配置通过率超过 30%。 - SOTA 阈值:对于连续指标任务(如 AlgoTune、GSO 的速度比),仅当解超越当前前沿时才计为通过。 - LLM-as-a-Judge:自由回答任务(HLE、GAIA、GPQA Diamond 等)使用多模型法官评估。 ### 3.3 结果 - 最强性能:GPT-5.5 + Codex 通过率为 28.0%。 - 成本-性能帕累托:前沿模型在原生脚手架下表现最佳;多个开权模型在 Terminus-2 下位于成本-性能帕累托前沿(如 MiniMax M3 单次仅 4,通过率 2.4–6.1%)。 - 失败分解:1,476 次 rollout 中,453 次超时无答案、361 次接近正确、445 次根本错误。 —- ## 4. 基准历史进展与延迟研究(附录实验) - 历史分数收集:为 39 个基准收集发布以来至今的最佳报告分数,对比 Harbor 统一评估下的当前最优结果(图 2、图 8、图 9)。发现软件工程与数学推理领域进展最快且 leaderboard 压力最大,而知识长文本与智能体工具系统领域测量更稀疏。 - 延迟分析(附录 F.5):按任务难度与成败分解执行延迟(wall-clock time)。发现失败任务的延迟在难度 0.7–0.8 处达峰,最难任务反而因早期终止而耗时更短;不同模型的成功/失败延迟模式存在结构性差异(如 GPT-5.4 无论成败耗时相近,而 GPT-5-Nano 在失败任务上坚持最久)。 —- ## 实验数据与工件的公开发布 论文将以下实验产物全部开源: - Harbor 适配器与框架代码 - 大规模评估的完整轨迹(约 340 GB,以 Hugging Face 数据集形式发布) - Harbor-Index 1.0 任务与审计产物 - 交互式结果看板(harbor-index.org) Q5: 有什么可以进一步探索的点? 基于论文的发现与讨论,以下方向值得进一步探索: —- ### 1. 基础设施与适配器的扩展验证 - 更多基准的规模化验证:当前大规模评估覆盖了 54 个适配基准,尚有超过 30 个已适配基准(如 CooperBench、OSWorld、TheAgentCompany 等,见附录 D.6.11)因计算成本未纳入实证分析。系统性地对这些基准进行同等规模的 parity 实验与智能体评估,可进一步检验 Harbor 框架的通用性边界。 - 强化学习与多智能体架构的支持:现有评估主要围绕基于 LLM 的单智能体脚手架展开。将 Harbor Adapters 扩展至支持强化学习训练循环、多智能体协作与竞争环境,是一个尚未触及的工程与研究方向。 —- ### 2. 模型-脚手架交互的动态优化 - 自适应脚手架设计:论文发现基础模型能力对性能的影响显著大于脚手架(效应范围 5.2:1),但脚手架仍系统性地改变行为模式(如 Terminus-2 的线性结构缺乏迭代自校正)。探索依据任务难度或失败模式动态切换脚手架策略(如从 Plan-Execute-Complete 切换到开放式迭代)可能提升整体通过率。 - 跨家族脚手架迁移:开源模型(如 GLM 5.2、DeepSeek V4 Pro)在 Terminus-2 上表现出成本-性能帕累托优势,而闭源模型在原生脚手架上更优。研究非原生脚手架的兼容性优化(如通过 OpenRouter 统一接口时的缓存命中率提升)具有实际价值。 —- ### 3. 基准冗余性与自适应评估 - 任务级自适应选择(Computerized Adaptive Testing, CAT):论文通过贪心选择证明约 12 个基准即可逼近 54 个基准的排序信号。进一步引入项目反应理论(IRT)或在线自适应采样,依据模型在前序任务上的表现动态选择下一任务,可在不牺牲排序精度的前提下进一步压缩评估成本。 - 子集预测的理论极限:论文指出,从子集预测完整基准表现的方法在更强模型上性能会退化。研究这种外推误差的统计极限,以及如何选择对模型强度变化鲁棒的最小任务集,是元评估(meta-evaluation)的核心问题。 —- ### 4. 失败模式的针对性消解 - 验证器-任务对齐的自动化诊断:人工审计发现约三分之一的最难候选任务因设计缺陷(如指令-验证器不匹配、隐式假设)被剔除。开发自动化的任务质量审计工具(如基于形式规约的验证器合成、对抗性测试生成),可减少对昂贵人工审查的依赖。 - 特定失败模式的干预策略:对图 4 中主导的三种失败模式——错误事实答案、算法缺陷、隐藏测试回归——设计针对性的认知架构改进(如检索增强生成用于事实性、符号执行用于算法调试、显式测试覆盖分析用于隐藏测试)。 —- ### 5. Harbor-Index 的动态维护与对抗性审计 - 实时更新与反饱和机制:论文明确提出将 Harbor-Index 维护为活的基准(live benchmark)。随着模型能力提升,需要周期性重跑过滤混合(filtering mix),剔除已饱和任务,并从未审计池中补充新任务以保持 30% 通过率上限。 - 对抗性漏洞挖掘:更强的未来智能体可能更有效地利用当前未暴露的评估漏洞(如奖励黑客、网络信息泄露)。建立红队(red-teaming)流程,自动化检测可 exploited 的 verifier loopholes,是维持评估可信度的必要环节。 —- ### 6. 成本-效率评估的形式化框架 - 经济最优的模型选择策略:论文发现前沿模型 Token 效率更高但成本溢价未被完全抵消,且中等难度任务( 0.3 ≤ 1-p ≤ 0.7 )的边际收益最大。将这一经验观察形式化为预算约束下的最优模型-脚手架组合选择问题,可为实际部署中的评估资源分配提供决策支持。 - 不确定性量化:在压缩评估中,现有方法缺乏对排序不确定性的显式建模。引入贝叶斯优化或错误条估计(如附录 F.1.4 提及),可在减少任务数量的同时报告模型间差异的置信度。 —- ### 7. 长程交互与动态环境的评估 - 异步与动态事件的支持:GAIA2 等任务的适配暴露出 Harbor 对异步环境事件(如模拟器状态推进)的处理不足。改进执行引擎以支持事件驱动、状态机式的动态环境交互,是评估更复杂智能体(如操作系统控制、实时系统管理)的前提。 - 超时与预算的动态调整:当前 Harbor-Index 采用固定时间预算(通常为最快前沿模型耗时的 1.2× )。研究基于任务复杂度预测的自适应超时分配,可在保证公平性的同时减少无效等待。 —- ### 8. 多模态与跨领域迁移评估 - 视觉-语言-行动(VLA)智能体的标准化:当前 Harbor-Index 的多模态覆盖有限(如 MMAU 依赖音频转文本)。扩展 Adapter 以原生支持图像输入、GUI 操作、视频理解,并建立跨模态的 parity 验证协议,是一个重要的工程研究课题。 - 领域间能力迁移的量化:论文发现不同基准间存在低秩结构,但专用能力(如 FinanceAgent、LAB-Bench)难以被其他基准预测。系统性地构建能力迁移矩阵(类似 Taskonomy),明确哪些基准组合最能预测未见过领域的表现,有助于设计更高效的训练与评估协议。 Q6: 总结一下论文的主要内容 这篇论文针对大规模语言模型智能体(agentic)评估中的基础设施碎片化与异构性难题,提出了系统性的解决方案。主要内容可概括如下: —- ### 1. 研究背景与问题 随着智能体能力的扩展(代码生成、工具使用、长程任务、计算机操作等),agentic 基准测试数量激增(2024 年以来超过 200 个)。然而,这些基准在任务格式、环境需求、交互协议和评分逻辑上高度异构,导致: - 集成成本爆炸:传统方式需要 O(mn) 的 benchmark–agent 配对集成; - 结果不可比:不同论文的环境配置、实现细节和评估协议差异巨大,难以判断模型进展是否跨领域泛化; - 任务质量参差不齐:许多“未解决”任务实际上源于设计缺陷(如指令与验证器不匹配、隐式假设、环境故障),而非真正的模型能力不足; - 评估成本高昂:全面运行数十个基准耗费巨大,社区缺乏既保留挑战性又经济实惠的精简套件。 —- ### 2. 核心贡献 #### 贡献一:Harbor Adapters 统一评估基础设施 论文开发了基于 Harbor 框架的适配器体系,将 80 余个异构基准统一映射到共享的任务抽象(instruction, environment, tests, solution)。该体系: - 将集成复杂度从 O(mn) 降至 O(m+n) ; - 支持多种沙盒后端(Daytona、Modal、E2B)与 GPU 交互; - 通过严格的三级代码审查(bot、junior、senior)和对等实验(parity experiments)验证适配语义保真度,确保 Harbor 版本与原始基准统计一致。 #### 贡献二:大规模系统评估与诊断分析 论文开展了覆盖 8 个跨层级模型、16 种 model–harness 配置、54 个基准、6,627 个任务(每配置 3 次试验,共约 30 万条轨迹,消耗 226B Token 与超过 30 万美元)的评估。关键发现包括: - 模型能力主导性能差异:线性混合模型显示,模型固定效应范围是脚手架(harness)效应范围的 5.2 倍(0.451 vs. 0.087),基准难度解释 75% 的方差; - 基准空间低秩冗余:54 个基准的有效评估空间近似秩-2(能力轴 + 脚手架轴);仅需约 12 个基准即可捕捉大部分独立排序信号;单个基准内 3 个任务即可以平均 rho ≈ 0.923 恢复系统排名; - 效率与成本权衡:前沿模型在所有难度层级上更省 Token(简易任务仅消耗弱模型的 42%),但更高单价使总成本仍高出 2–3 倍;升级至前沿模型在中等难度任务(0.3–0.7)上边际收益最大; - 失败模式归因:对 6,028 条失败轨迹的审计显示,错误事实答案、算法缺陷和隐藏测试回归是主导失败模式;原生脚手架允许开放式迭代,而 Terminus-2 的线性结构更易受困于需反复精修的任务。 #### 贡献三:Harbor-Index 精选元数据集 论文指出,现有基准中大量低通过率任务实为结构性设计缺陷(如验证器越界、隐式假设、环境未触发事件)。为此,论文通过多阶段漏斗(难度过滤 → AI 审计 → 人工复审 → 高级专家组筛选 → 审计-修复循环),从 6,627 个任务中精选出 Harbor-Index 1.0: - 82 个任务,跨越 29 个基准,覆盖软件工程、科学研究、数学推理、智能体系统等核心领域; - 通过 测试-指令对齐 与 本质难度 两项质量标准剔除破损任务; - 难度受控:无任何已评估 model–harness 配置通过率超过 30%;最强配置(GPT-5.5 + Codex)通过率为 28.0%,保留充足进步空间; - 成本可控:单次完整 Index 运行成本从 4 到 $293 不等,多个开源模型在 Terminus-2 下位于成本-性能帕累托前沿。 —- ### 3. 结论与影响 Harbor Adapters 与 Harbor-Index 共同提供了可复用、可扩展、可负担的智能体评估基础设施与高质量测试床。论文将所有适配器、评估工具、完整轨迹(约 340 GB)及 Harbor-Index 开源发布,旨在推动更可靠、更全面、更具可比性的语言模型智能体评估。同时,论文也指出未来需持续维护 Harbor-Index 为活基准,对抗未来更强模型可能利用的评估漏洞,并不断扩展适配器覆盖范围。 Q7: 想要进一步了解论文 以上只是了解一篇论文的几个基本FAQ。如果你还想与Kimi进一步讨论该论文,请点击 这里 为你跳转Kimi AI网页版,并启动一个与该论文相关的新会话。
Authors: Lin Shi, Haowei Lin, Zixuan Zhu, Xiaoyue Zhou, Xiang Li, Xiangning Lin, Yaxuan Deng, Han Xu, Yuangang Li, Shanda Li, Zizhao Chen, Hanwen Xing, Harsh Raj, Bo Chen, Quan Shi, Steven Dillmann, Yipeng Gao, Puneesh Khanna, Ruofan Lu, Chao Beyond Zhou, Michael Yang, Robert Zhang, Siyuan Chai, Jiayu Chang, Yizhao Chen, Xiaokun Chen, Yiwei Dai, Wenting Yang, Hange Liu, Minghao Liu, Zihan Wang, Adnan El Assadi, Benedikt Stroebl, E. Kelly Buchanan, Han Meng, Junwei He, Longxuan Yu, Radin Shayanfar, Yukyung Lee, Zhikang Dong, Allen G Hart, Anjiang Wei, Anurag Kashyap, Arpandeep Khatua, Audrey Jixin Zheng, Chengrui Ma, David Heineman, Dubing Chen, Hai-Anh Trinh, Haishuo Fang, Hefan Zhang, Hui Shen, Issa Sugiura, Jiankai Sun, Jiechao Gao, Junhong Lin, Junnan Li, Kai Yang, Lei Hsiung, Maoyu Wang, Mengze Tang, Nabil Omi, Negin Raoof, Nicholas Edwards, Octavia Guo, Orfeas Menis Mastromichalakis, Pengliang Ji, Przemysław Hejman, Qi Qi, Qunshu Lin, Richard Zhuang, Rui Yang, Ruichen Zheng, Ryan Marten, Shaghayegh Fazliani, Shizheng Hou, Sicong Jiang, Sijie Li, Song Bian, Terry Yue Zhuo, Tianqing Wu, Tom Tang, Wanjia Zhao, Weihao Xuan, Wenhua Liang, Xian Liu, Xin Lan, Xuan Zhang, Xuandong Zhao, Yanchuan Tang, Yifan Jiang, Yijiang Li, Yitong Guan, Yizhi Li, Yonghui Liu, Yuheng Tang, Yujun, Yunfei Zhao, Yuxin Wang, Yuxuan Tang
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04298.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04298
Published: 2026-09-08T01:23:43.982Z
4. Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security
Abstract:Ensuring the security of the power system is essential for stability and reliability, especially in the event of disruption. Effective classification of contingency in power systems enables proactive decision-making and mitigates large-scale breakdowns and failures. This study explores the use of machine learning algorithms to classify security levels of contingencies in power systems into safe, moderate or severe classes. For this approach, Newton-Raphson load flow method extracts system data from contingency scenarios, using Overall Performance Index (OPI) as safety measure. For data pre-processing, Synthetic Minority Over-Sampling Technique (SMOTE) and Principal Component Analysis (PCA) is used to address class imbalance and reduce dimensionality, respectively. K-Nearest Neighbours (KNN), Random Forest (RF) and Support Vector Machines (SVM) is trained and evaluated on datasets generated through N-k contingency scenarios for k equal 1, 2, and 3 on IEEE-14 and IEEE-30 bus systems using four hybrid pre-processing configurations: normalized, SMOTE-balanced, PCA-transformed, and a combined SMOTE PCA-transformed. Performance is assessed by precision, recall and F1 score, with priority given to the severe contingency classes. The RF achieved the highest F1 scores of 0.97 in IEEE-30 and 0.86 in IEEE-14, SVM benefits significantly from PCA and improves the accuracy of the classification, while KNN is best suited for SMOTE and PCA conversion. The findings show that PCA contributes more than SMOTE to the overall performance of the model. However, SMOTE improves recall but can introduce false positives and is therefore a compromise of accuracy. This study highlights machine learning as a scalable and powerful alternative to traditional contingency analysis, which improves the assessment of security in real time.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04300 (HTTP 429)
Authors: Joshua Salako, Folajimi Osikomaiya, Olakorede Olamiju
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04300.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04300
Published: 2026-09-08T01:23:43.982Z
5. Iris: Climbing to the Search Frontier
Abstract:We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks are reverse-constructed from the hyperlink structure of a web corpus: we author multi-hop chains over an entity graph distilled from a seed page and its out-links, rewrite every non-answer entity into a descriptive reference so that no clue can be resolved by string matching, and admit only questions that a reference model fails closed-book yet solves once the supporting evidence is supplied. These questions are then turned into trajectories, which are filtered at both the trajectory and the turn level before SFT. The policy is then optimized by RL against live search, with the reward judge and the observation summarizer served inside the training cluster, and with over-long rollouts interrupted at the request level and resumed from their committed prefix at the next step. We alternate the two stages in a procedure we call SFT-RL climbing, returning the hardest solved and most efficient rollouts of each RL round to the next supervised pass. Because inference-time context management is worth more on these benchmarks than most reported differences between systems, we evaluate every benchmark both with and without it, holding the tool set, the context limit, and the judge fixed. All results come from a single ReAct agent, with no sub-agents and no test-time verification. With management enabled, on BrowseComp, BrowseComp-ZH, DeepSearchQA, and HLE the two models reach $82.2/84.8/86.9/52.3$ and $88.6/85.1/92.9/56.4$, the strongest overall results among open-source search agents in their respective parameter ranges. We plan to release the model weights together with the complete recipe for data construction, training, and evaluation.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
Q: 这篇论文试图解决什么问题? 该论文旨在系统性地解决\*\*如何构建、训练和评估具备深度多跳搜索与长程信息整合能力的强大语言模型智能体\*\*这一核心问题。具体而言,其试图克服的关键挑战可归纳为以下四个方面: 1. \*\*高质量搜索训练数据的自动构造与验证\*\* 自然存在的网页问题通常过于简单,而人工编写难以规模化。论文提出从网页语料的超链接结构反向构造多跳任务:首先基于种子页面及其外链蒸馏出实体图,随后在该图上生成必须组合至少 N 个关系才能回答的问题,并通过抽象算子 A 将所有非答案实体重写为描述性指代,从而消除直接通过字符串匹配即可定位答案的捷径。最终,仅保留满足双重验证标准的样本——即在闭卷环境下参考模型 M_(ref) 无法回答( c_(diff)(q)=1 ),但在提供支撑证据后能够正确求解( c_(solv)(q)=1 )——以确保数据既具备难度又具备客观可解性。 2. \*\*长程交互中的上下文管理与训练效率\*\* 长视域搜索轨迹极易耗尽模型的上下文窗口,使得有效搜索预算远低于名义窗口长度。为缓解该问题,论文在推理层面系统研究了上下文管理(Context Management, CM)策略对性能的影响,并强调应在固定工具集、上下文长度限制和评判标准下,分别报告\*\*启用 CM\*\* 与\*\*不启用 CM\*\* 的结果,以准确剥离模型内在策略能力与推理时外部封装带来的增益。在训练层面,论文采用请求级中断与前缀复用(request-level partial rollout)机制,允许超长轨迹在同步 rollout 过程中被中断并在下一步从已提交前缀恢复,从而兼顾训练效率与长程探索能力。 3. \*\*监督微调与强化学习的有效协同\*\* 单纯依赖监督微调(SFT)容易受限于教师模型的能力边界,而单纯依赖强化学习(RL)则难以稳定复现稀疏的成功轨迹。论文提出 \*\*SFT–RL Climbing\*\* 迭代流程:在每一轮 RL 探索后,筛选出“最难且最高效”的成功轨迹(即通过率 0 < R(q) le 1/2 且工具调用轮次不少于 K_(rft) 的最短正确轨迹),将其蒸馏回下一轮的 SFT 数据集中。该机制使策略在探索新行为与巩固已有成功模式之间自动形成课程学习式的渐进提升。 4. \*\*内在搜索能力与评估透明性的精准衡量\*\* 现有工作往往仅在启用复杂推理时封装(如子智能体、重试、验证链)下报告结果,导致难以判断性能提升源于模型本身还是外部工程技巧。论文坚持仅使用单一 ReAct 智能体、无子智能体、无测试时验证的设定,并在所有基准上提供无 CM 与有 CM 的双重结果。通过这种方式,论文试图回答:在剥离了上下文管理、重试机制等推理时辅助手段后,基础模型本身究竟掌握了多强的自主搜索与推理能力。 Q2: 有哪些相关研究? 根据论文引言及参考文献,相关研究可按照主题脉络梳理如下。 ### 1. 早期工具使用与推理-行动交织 该方向奠定了语言模型通过外部工具扩展能力的基础,核心在于将推理(Reasoning)与行动(Acting)相结合: - \*\*WebGPT\*\*
14
:最早探索通过浏览器辅助进行问答,并引入人类反馈进行训练。 - **Toolformer**
17
:展示了语言模型可通过自监督方式学习调用 API 工具。 - **ReAct**
26
:提出将推理与行动交织的范式,成为后续搜索智能体的标准交互框架。 ### 2. 基于强化学习的搜索优化 从静态语料检索发展到对实时网页环境的强化学习(RL)优化: - **Search-o1**
11
与 **Search-R1**
8
:将大语言模型与搜索引擎结合,通过 RL 训练模型进行搜索增强推理。 - **R1-Searcher**
19
:通过强化学习激励 LLM 的搜索能力。 - **DeepResearcher**
29
:在真实环境中通过 RL 规模化深度研究能力。 ### 3. 合成数据构造与难度提升 为解决自然网页问题过于简单的问题,研究者们通过超链接结构或知识图谱构造多跳、抗捷径(shortcut-resistant)任务: - **WebDancer**
23
、**WebSailor**
10
、**WebShaper**
20
:利用网页图结构遍历生成信息搜寻任务。 - **DeepDive**
12
:结合知识图谱与多轮强化学习提升深度搜索智能体能力。 - **FORT-Searcher**
4
:专门合成抗捷径的搜索任务用于训练深度搜索代理。 ### 4. 端到端训练流程与开源系统 近期研究趋向于将合成问题生成、轨迹监督、RL 与长上下文推理整合为标准化流水线: - **OpenSeeker / OpenSeeker-v2**
6, 5
:开源搜索智能体训练数据与流程的代表性工作。 - **Tongyi DeepResearch**
21
、**MiroThinker-1.7 & H1**
13
、**Apodex-1.0**
1
:探索验证驱动的重型研究智能体。 - **REDSearcher**
2
:聚焦可扩展、低成本的长程搜索代理框架。 - **XYZ-Aquila**
25
与 **Shanghai AI Laboratory**
18
:在 35B 至更大参数规模上推进智能体能力。 - **Kimi K3**
9
:开放前沿智能体能力。 ### 5. 上下文管理与长程推理 长程搜索轨迹容易耗尽上下文窗口,相关研究探索了多种压缩与重置策略: - **ReSum**
24
:通过上下文摘要解锁长程搜索智能。 - **AgentFold**
27
:主动上下文管理以支持长程网页代理。 - **DeepSeek-V3.2**
3
:提出 discard-all 等上下文管理策略,论文将其作为 CM 对比基准。 ### 6. 评估基准与数据集 论文涉及的主要评测基准包括: - **BrowseComp**
22
与 **BrowseComp-ZH**
30
:评估浏览代理定位长尾实体与多线索约束能力的基准。 - **DeepSearchQA**
7
:评估基于搜索回答的全面性而非单一答案跨度。 - **Humanity’s Last Exam (HLE)**
15
:跨学科专家级学术推理测试。 - **LoHoSearch**
28
:用于测试超越人类难度上限的长程搜索代理。 此外,论文在结论中提到,搜索能力的提升对通用工具使用(如 BFCL、 τ -bench)和办公协作任务(如 OfficeQA、APEX)存在正向迁移,暗示搜索能力可能是一种可复用的原子能力而非垂直领域特化。 Q3: 论文如何解决这个问题? 论文通过一套端到端的“数据—训练—评估”配方解决构建强大搜索智能体的问题,核心环节可概括如下。 ### 1. 从网页图反向构造并验证多跳任务 为解决自然问题过易、人工标注难扩展的问题,论文设计了一条全自动数据生产线: - **网页子图提取**:将语料视为有向图 G=(V,E) ,从种子页面 v0 出发,沿外链抽取局部子图 G(sub)=(v0∪ N, E(sub)) 。 - **实体图蒸馏**:将 G(sub) 压缩为紧凑的实体图 G_e=(V_e,R_e)=f(ext)(G(sub)) ,仅保留与种子主题相关的实体与关系。 - **多跳问题生成**:在 G_e 上生成问题-路径对 (q_0,P)=f(gen)(G_e,y) ,强制推理路径长度 |P|ge N ,确保问题依赖至少 N 个耦合关系。 - **实体抽象去捷径**:通过抽象算子 A 将所有非答案实体重写为描述性指代,使得
q = f(abs)(q_0, A),
从而消除直接字符串匹配即可定位答案的捷径。 - **双重验证**:利用参考模型 M(ref) 进行两项检验:
c(diff)(q) = I[M(ref)(q)≠ y], quad c(solv)(q) = I[M(ref)(qmid Ge)=y].
仅保留交集 D=(q,y)mid c(diff)· c(solv)=1 ,即那些闭卷失败、供证后能解的样本,确保难度与可解性并存。 ### 2. 双层过滤的监督微调(SFT) 利用强教师模型 M_T 在 ReAct 范式下对 D 中问题生成轨迹 τ ,随后进行粗、细两级过滤,以提升监督信号质量。 - **轨迹级粗过滤**:剔除三类不良样本: - **错误轨迹**: rollout 未成功终止或最终答案 y 被评判为错误; - **退化轨迹**:利用滑动窗口 zlib 压缩比 rho(cr)(w)=|w|/|zlib(w)| 检测重复循环,并辅以周期重复行、单字符长串、字节相同参数等辅助检测器; - **浅层轨迹**:工具调用轮次 T(tool)(τ)<K 的样本被丢弃。 过滤后去重得到 D(sft) 。 - **轮次级精过滤**:为避免教师能力边界导致的局部劣质步,引入从数据中归纳出的评判标准(而非人工手写规则),对每个助手轮次输出 KEEP/MASK 标签 mt∈0,1 ,且单条轨迹最多掩码 10% 的轮次。 - **训练目标**:在 D(sft) 上最大化教师输出的似然:
L(SFT)(θ) = -E((q,τ)sim Dsft) ∑(t=1)^(T+1) mt log πθ(ut mid C(<t)),
其中 C_(<t) 为历史上下文, u_t 为第 t 步的推理与工具调用或最终答案。被掩码的轮次保留在上下文中但不参与损失计算。 ### 3. 针对实时搜索的强化学习(RL) 在 SFT 基础上,论文进一步以组相对策略梯度在**实时搜索环境**中优化策略。 - **请求级部分 rollout 与前缀复用**:为缓解长程 rollout 的尾部延迟,训练在请求粒度中断:超时的在飞会话被中止后,于下一步从其已提交的**前缀**恢复。通过缓存各消息状态的 token、损失掩码、对数概率与权重版本增量,恢复时利用截断重要性采样修正策略权重不匹配,避免重复计算,保持同步调度的 GPU 利用率。 - **集群内奖励与摘要**:不依赖外部 API,而是在训练集群内部署 FP8 推理引擎(基于 Qwen3.5-397B-A17B),同时承担两职: - **生成式奖励模型(GenRM)**:对 rollout 的答案给出二元裁决
R(q,τ) = I[GenRM(q,y_τ,y^*)=A];
- **观察摘要器**:将检索到的原始页面压缩为查询相关的摘要 ot ,作为 ReAct 循环中的观察。这是 rollout 中唯一的上下文缩减机制,无历史剪枝或滑动窗口,保证训练与推理的上下文一致性。 ### 4. SFT–RL 迭代攀爬(Climbing) 为结合 RL 的探索优势与 SFT 对稀缺成功轨迹的直接固化能力,论文交替执行两阶段: - 每轮 RL 后,保留**部分解决**(通过率 0<R(q)le 1/2 )的查询,从中挑选成功且工具调用不少于 K(rft) 的最短轨迹:
Sq = τ_i mid R(q,τ_i)=1,; T(tool)(τi)ge K(rft), quad τq^* = argmin(τ∈ Sq) T(tool)(τ).
- 这些高效且非平凡的成功轨迹经去重后回注下一轮 SFT,形成自动推进难度的课程学习。随着策略提升,候选池自然萎缩,提供简单的停止信号。 ### 5. 受控且透明的评估协议 为准确度量模型**内在搜索能力**与**上下文管理(CM)**各自的贡献,论文采用严格的对比评估: - **固定基础设施**:所有对比在相同工具集、相同上下文长度限制、相同 LLM 评判器下进行。 - **双模式报告**:对每个基准同时报告**无 CM**(原始 ReAct,无历史删减)与**有 CM**(采用 DeepSeek-V3.2 的 discard-all 等策略)的结果,从而将模型本身能力与推理时封装增益解耦。 - **防泄露机制**:在训练与评估中屏蔽 Hugging Face 上的数据集与 Space 页面,防止模型利用基准泄露刷分。 通过上述数据构造、双层过滤监督训练、带前缀复用的在线 RL、迭代攀爬以及受控评估的完整闭环,论文系统性地将模型的多跳搜索与长程信息整合能力推至当前同参数规模下的最优水平。 Q4: 论文做了哪些实验? 论文的实验围绕**模型性能对比**与**上下文管理(CM)消融**两大主线展开,具体包括以下内容。 ### 1. 基准测试与评估设置 实验在四个具有挑战性的智能体搜索基准上进行: | 基准 | 评估重点 | 评价指标 | |———|————-|————-| | **BrowseComp** | 长程、多线索约束下的长尾实体识别 | Accuracy | | **BrowseComp-ZH** | 中文环境下的上述能力 | Accuracy | | **DeepSearchQA** | 搜索答案的全面性(多证据覆盖) | F1 | | **Humanity’s Last Exam (HLE)** | 专家级学术推理(文本-only子集) | Accuracy | 评估协议统一采用 **pass@1**(单轮 rollout),由 LLM-based judge 根据各基准的官方 prompt 判定最终答案。所有实验在**相同的工具集、上下文长度限制与最大轮次预算**下进行,以确保可比性。 ### 2. 主实验:与开源及前沿模型对比 论文将 **Iris-mini(35B-A3B)** 与 **Iris-pro(397B-A17B)** 分别置于同参数规模区间及更大规模的前沿模型中进行横向对比,结果如表1所示(默认启用 discard-all CM 策略)。 - **同规模开源模型对比(30–35B)**: Iris-mini 在 BrowseComp(82.2)、BrowseComp-ZH(84.8)和 HLE(52.3)上取得该参数区间的最优成绩;在 DeepSearchQA(86.9)上略低于 XYZ-Aquila-mini(89.5)。 - **同规模开源模型对比(∼400B)**: Iris-pro 在全部四个基准上领先或持平:BrowseComp(88.6)、BrowseComp-ZH(85.1)、DeepSearchQA(92.9)、HLE(56.4)。其中 BrowseComp 领先同规模次优模型 XYZ-Aquila-pro 3.8 分,HLE 领先 3.1 分。 - **与更大规模/闭源前沿系统对比**: Iris-mini 的性能已接近 1T 规模的 Kimi-K2.6 与 DeepSeek-V4-Pro(BrowseComp 上 82.2 vs. 83.2/83.4)。Iris-pro 在标准 ReAct 设置下的表现可媲美部分采用重度推理封装(heavy-compute)的系统(如 MiroThinker-H1 与 Apodex-1.0-H)。 ### 3. 上下文管理(CM)消融实验 为剥离**模型内在搜索能力**与**推理时外部封装**各自的贡献,论文系统比较了多种 CM 策略,结果如表2所示。 - **无 CM(w/o)基线**:不使用任何历史删减或重置,直接反映模型原生能力。 - **discard-all**:当运行上下文达到阈值时,清空全部交互历史并从头重启问题求解(借鉴 DeepSeek-V3.2)。 - **retry**:若一次尝试失败,将失败经历摘要为“已探索与排除的信息”,追加至任务描述后重新尝试(借鉴 MiroThinker)。 - **discard-all + retry**:上述两种策略的复合。 **实验发现包括**: - **无 CM 下的原生能力**:Iris-mini 在 BrowseComp(64.7)与 BrowseComp-ZH(72.3)上显著高于同规模无 CM 系统(如 FORT-Searcher 的 55.9/62.1、OpenSeeker-v2 的 46.0/58.1);Iris-pro 在此基础上再提升 7.9 与 4.5 分。 - **CM 增益的尺度差异**:CM 对 Iris-mini 的提升(BrowseComp 上最高 +21.2 分)明显大于对 Iris-pro 的提升(最高 +17.7 分),因为小模型需要更多步骤解决相同约束,更频繁触及上下文上限。 - **CM 增益的基准差异**:增益排序为 BrowseComp > BrowseComp-ZH > DeepSearchQA > HLE。BrowseComp 需要大量检索-筛选-整合步骤,上下文是硬性约束;HLE 更依赖专家知识而非网页检索,延长搜索地平线收益有限。 - **策略组合的上限**:discard-all + retry 在 Iris-pro 上将 BrowseComp 推至 90.3,但论文指出 retry 带来显著的推理成本倍增,因此主结果表仅采用 discard-all 作为默认配置。 - **BrowseComp-ZH 上的收敛现象**:Iris-mini(discard-all+retry)与 Iris-pro(discard-all 及 discard-all+retry)三者恰好均落在 85.1(289 题中 246 题正确),提示该基准上剩余性能瓶颈可能已超出模型容量范畴(论文附录进一步指出该基准存在标注不一致问题)。 ### 4. 训练与推理的工程细节实验 论文在实验设置中披露了关键实现细节,这些选择本身也构成实验设计的一部分: - **模型初始化**:Iris-mini 基于 Qwen3.6-35B-A3B,Iris-pro 基于 Qwen3.5-397B-A17B,均为 MoE 架构、256K 上下文窗口。 - **SFT 配置**:2 个 epoch,全局 batch size 64,最大序列长度 262,144 tokens。 - **RL 引擎**:采用开源框架 Relax。 - **防泄露控制**:从训练与评估环境中三重屏蔽 Hugging Face 的数据集与 Space 页面(搜索结果移除、爬取拒绝、工具管理层后置拦截),防止模型记忆基准答案。 ### 5. 附录:基准标注质量案例分析 在附录 A 中,论文以 BrowseComp-ZH 第 85 题作为案例研究:问题要求根据《权力的游戏》剧情推断 Sansa Stark 第二次正式婚姻所属家族。官方标准答案为 Lannister,但论文智能体返回 Bolton。经剧情核查,Sansa 的首次正式婚姻确为与 Tyrion Lannister,第二次为与 Ramsay Bolton,因此智能体答案具有直接剧情支撑。该实验性案例分析被用来论证现有基准可能存在标注不一致,并呼吁开发更高质量、更可靠的评测体系。 Q5: 有什么可以进一步探索的点? 基于论文结论与讨论,以下几个方面具有明确的进一步探索价值: ### 1. 搜索能力的跨领域迁移与通用化 论文观察到,为搜索合成的数据以及搜索专用模型(作为教师用于策略蒸馏)在通用工具使用(BFCL、 τ -bench)和办公协作(OfficeQA、APEX)等未直接针对的领域产生了**正向迁移**。这提示搜索可能更应被视为一种**原子能力(atomic capability)**而非垂直特化。后续工作可系统性地验证该迁移效应在更广泛智能体场景中的普适性,并探索将搜索数据与搜索衍生教师整合到通用训练流程中,而非将其局限于独立的特化阶段。 ### 2. 高质量评估基准的构建与标注一致性 附录 A 以 BrowseComp-ZH 第 85 题为案例,揭示了现有基准可能存在**标注不一致**(ground-truth inconsistency)的问题:智能体给出的答案 “Bolton” 在剧情事实上成立,却与官方标准答案 “Lannister” 冲突。这凸显了开发**更高质量、标注更可靠、覆盖更全面**的搜索评估基准的必要性,以减少因标注错误导致的评估噪声,并更准确地衡量模型的真实信息检索与推理能力。 ### 3. 推理效率与上下文管理的帕累托优化 实验表明,`discard-all + retry` 策略虽然能将 Iris-pro 在 BrowseComp 上的成绩推至 90.3,但伴随**显著的推理成本倍增**(每次重试需完整重新搜索)。未来研究可探索在性能与计算开销之间更优权衡的上下文管理(CM)策略,并将**实际推理成本**作为与准确率并列的核心报告指标,而非单纯追求基准分数最大化。 ### 4. SFT–RL 迭代攀爬(Climbing)机制的深化 论文提出的 SFT–RL Climbing 通过“最难且最高效”的成功轨迹实现自动课程学习,但其细节(如难度带阈值 0<R(q)le 1/2 的动态调整、前缀复用中的重要性采样修正等)仍有优化空间。进一步工作可研究更精细的轨迹选择标准、更高效的回注策略,以及该迭代范式在更长训练周期中的收敛性与稳定性。 ### 5. 缩小与顶尖闭源系统的能力差距 尽管 Iris-pro 在同参数规模内取得领先,但与 Claude Fable 5(HLE 64.5)、Apodex-1.0-H(BrowseComp 90.3, DeepSearchQA 94.4)等采用重度计算配置的前沿系统仍存在差距。探索如何在不依赖子智能体、不重推理链的前提下,进一步提升基础模型的长程规划、证据整合与专家级推理能力,是缩小这一差距的关键方向。 ### 6. 真实世界复杂搜索场景的覆盖 当前评估主要集中于四个广泛认可的社区基准。论文指出,这些结果仅提供了能力的**初步视图(first view)**,而现实世界中存在大量更复杂、更开放、更少结构化的问题形态。将评估体系拓展到更广泛、更动态、更贴近实际用户需求的搜索与信息搜寻场景,是衡量搜索智能体实用价值的必经之路。 Q6: 总结一下论文的主要内容 该论文介绍了 **Iris-mini**(35B-A3B)与 **Iris-pro**(397B-A17B)两个搜索智能体的构建方法,并提出了一套覆盖数据构造、训练流程与评估协议的端到端配方。以下从核心问题、技术方案、实验发现与主要贡献四个方面进行总结。 —- ### 1. 核心问题 现有搜索智能体的性能提升往往混合了**模型内在搜索能力**与**推理时上下文管理(Context Management, CM)**等外部封装带来的增益,导致难以准确评估基础模型本身的长程信息搜寻与多跳推理水平。此外,自然网页问题通常过于简单,而高质量、可规模化且抗捷径的多跳训练数据稀缺。 —- ### 2. 技术方案 #### 2.1 数据构造:从网页图反向合成多跳任务 论文将网页语料建模为有向图 G=(V,E) ,从种子页面 v0 出发沿外链抽取局部子图,并蒸馏为紧凑的实体图 G_e=(V_e,R_e) 。随后: - 在 G_e 上生成需至少 N 跳关系才能回答的问题路径; - 通过抽象算子 A 将所有非答案实体重写为描述性指代,消除直接字符串匹配的捷径,得到 q=f(abs)(q0,A) ; - 采用**双重验证**:仅保留那些参考模型在**闭卷**环境下失败、但在**提供支撑证据**后能正确解答的问题,确保难度与可解性并存。 #### 2.2 监督微调:双层过滤 由强教师模型在 ReAct 范式下生成轨迹 τ ,随后进行两级过滤: - **轨迹级**:剔除答案错误、存在重复循环(通过滑动窗口 zlib 压缩比检测)或搜索过浅(工具调用轮次不足)的轨迹,并去重; - **轮次级**:利用从数据中归纳出的评判标准(非人工手写规则),对每轮助手输出标注 KEEP/MASK,最多掩码 10% 的轮次,以清除局部劣质步骤而不破坏完整交互历史。 #### 2.3 强化学习:实时搜索中的长程优化 - **请求级部分 rollout**:为避免超长轨迹拖慢同步训练,在请求粒度中断未完成的会话,下一步从其已提交的**前缀**恢复,并通过截断重要性采样修正策略权重变化,实现已生成内容的复用; - **集群内服务**:在训练集群内部署 FP8 推理引擎(基于 Qwen3.5-397B-A17B),同时充当**生成式奖励模型(GenRM)**(对最终答案给出二元裁决)与**观察摘要器**(压缩原始页面为查询相关摘要),摆脱对外部 API 的依赖,并保证训练与推理的上下文一致性。 #### 2.4 迭代攀爬(SFT–RL Climbing) 交替执行 SFT 与 RL:每轮 RL 探索后,筛选**部分解决**(通过率 0<R(q)le 1/2 )且**最短的高效成功轨迹**(工具调用不少于 K(rft) ),将其蒸馏回下一轮 SFT。该机制形成自动推进难度的课程学习,并在策略提升导致候选池萎缩时自然发出停止信号。 #### 2.5 受控评估 为剥离模型能力与 CM 封装增益,论文在**固定工具集、上下文长度限制与评判器**的前提下,对每个基准同时报告**无 CM** 与**有 CM**(采用 discard-all 策略)的结果。 —- ### 3. 实验发现 在 BrowseComp、BrowseComp-ZH、DeepSearchQA 与 Humanity’s Last Exam (HLE) 四个基准上: - **Iris-mini**(35B)在 BrowseComp(82.2)、BrowseComp-ZH(84.8)和 HLE(52.3)上取得同参数区间最优; - **Iris-pro**(397B)在四项基准上均领先或持平(BrowseComp 88.6、BrowseComp-ZH 85.1、DeepSearchQA 92.9、HLE 56.4),接近甚至超过部分采用重度推理封装的更大规模系统。 **上下文管理分析**: - 无 CM 时,Iris 模型已优于现有同规模系统,说明其**内在搜索能力**本身较强; - 启用 CM 后性能进一步提升,且小模型获益更大(Iris-mini 在 BrowseComp 上最高提升 21.2 分),因其更易耗尽上下文; - CM 增益与基准特性高度相关:长程信息搜寻型任务(BrowseComp)收益显著,而专家知识型任务(HLE)收益有限。 —- ### 4. 主要贡献与展望 - 提出了一套从网页图反向构造、验证并过滤的多跳搜索训练数据流水线; - 设计了 SFT–RL Climbing 迭代范式,将 RL 中发现的成功行为高效固化回监督训练; - 通过严格的 CM 对比实验,提供了更可解释、更透明的搜索智能体评估方案; - 观察到搜索能力的训练收益可向通用工具使用与办公协作任务**正向迁移**,暗示搜索可能是一种可复用的**原子能力**,而非单纯的垂直特化; - 计划开源模型权重及完整的数据、训练与评估配方。 Q7: 想要进一步了解论文 以上只是了解一篇论文的几个基本FAQ。如果你还想与Kimi进一步讨论该论文,请点击 这里 为你跳转Kimi AI网页版,并启动一个与该论文相关的新会话。
Authors: Ziyuan Liu, Hengqi Liu, Zichuan Wang, Yang Qin, Jiachen Liang, Xu Chu, Shaowei Chen, Yuantao Gu, Mu Chuan
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04304.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04304
Published: 2026-09-08T01:23:43.982Z
6. A Removal Based Approach to Improve LLM Faithfulness at Test-Time
Abstract:Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an important tool for auditing model behavior. Unfortunately, these explanations can be unfaithful, failing to reflect the actual reasoning underlying the model’s decisions. We consider a setting in which an LLM provides both an answer and an explanation in response to a question. We identify two distinct dimensions of unfaithful explanations: incompleteness, meaning that the explanation omits factors that influence the answer, and unsoundness, meaning that the explanation cites factors that did not influence the model’s answer. Existing approaches to improving LLM faithfulness include training-time methods, which require access to model weights and extensive computational resources, and test-time methods that largely focus on addressing unsoundness. We introduce a test-time approach that directly targets incompleteness. We remove from the input the concepts not credited in the model’s explanation and re-query the model on the reduced input. This eliminates unmentioned influences while preserving the influence of mentioned concepts. Across two datasets, multiple model families, and two independent faithfulness metrics, our approach improves explanation faithfulness compared to both standard prompting and prompting to encourage faithfulness. Our method is model-agnostic and can be applied at inference time without modifying model parameters, providing a flexible mechanism for reducing hidden influences and improving the reliability and safety of LLM-assisted decision making.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04343 (HTTP 429)
Authors: Qinglan Luo, S M A Nahian, John Guttag, S. Mazdak Abulnaga, Katie Matton
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04343.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04343
Published: 2026-09-08T01:23:43.982Z
7. Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Abstract:Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individual model capability can degrade rather than improve system-level outcomes. We hypothesize that shared training and architectures can lead more capable LLMs to behave more similarly, creating correlated actions that do not diversify away. We develop a general framework showing how this correlation creates a non-diversifiable risk floor and test its predictions in financial markets using an agent-based simulation with LLM traders of varying general-purpose capability. We find that: (1) frontier LLMs exhibit significantly correlated behavior that increases with capability; (2) when their shared reasoning is accurate, increasing agent participation reduces market-level risk; and (3) when agents share a common misinformation environment, the same correlated behavior becomes a liability. Together, these results identify a capability paradox: improving individual models does not necessarily produce better system-level outcomes. Whether the same dynamics arise in other domains is an open empirical question.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04373 (HTTP 429)
Authors: Jillian Ross, Eric So, Zoe De Simone, Charles Pozniak, Andrew W. Lo
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04373.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04373
Published: 2026-09-08T01:23:43.982Z
8. Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer
Abstract:Enterprise AI deployments fail not from model inadequacy, but because organizations lack a structured substrate encoding how they decide, negotiate, and execute. Generic LLMs carry no firm-specific ontological priors; RAG remains brittle, with no path to executable action; static playbooks encode logic but cannot reason or adapt. This demands an architecture treating tacit-knowledge capture, ontological grounding, sovereign deployment, and auditable actuation as co-designed from the start. This paper introduces the Corporate Language Model (CLM), a framework transforming a firm’s structured, unstructured, multimodal, and tacit knowledge into an ontology-grounded enterprise foundation upon which reasoning and governed execution are composed. CLM has five capability planes and four architectural pillars: a Neurosymbolic Mesh coupling generative models with a knowledge graph; a Skill Graph where reusable tactics, personas, objections, and goals are typed and composed; Living Digital Twins modeling functional areas as reasoning surrogates; and a Deep Security Layer enforcing sovereignty, traceability, and human oversight. A Spec-as-Code paradigm bridges grounded intent and executable artifact. CLM is one instantiation of this foundation-centric class. Four contributions follow: CLM is defined as a distinct object of study; the Skill Graph is introduced for compositional explainability by construction; the Wisdom Listener effect is proposed, whereby tacit-capable foundations compound in value with use, connecting to dynamic capabilities and organizational learning; and evidence from a JCI-accredited tertiary hospital in Brazil instantiates three of the six maturity stages under LGPD.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04377 (HTTP 429)
Authors: Fabricio C. Avini, Guilherme Trez
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04377.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04377
Published: 2026-09-08T01:23:43.982Z
9. HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Abstract:Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put a price on avoiding the side effect and to name that side effect as a living creature. It is a farm simulation: LLM sub-agents drive a crew of two tractors through a cooperative corn harvest, with animals in the field. The environment is a reinforcement learning gridworld, every decision is made without memory, and the harm is never named in the goal. When an animal blocks a tractor’s route the autopilot stops and asks the model whether to drive on, at no fuel cost, or swerve around it for a posted fuel price. Kills are compared against two controls: rocks, which damage the tractor and are hit under 1% of the time by every model, and hay bales, which are harmless and not alive. Models can also take crops from the neighbor’s field instead of their own, a second test of what they treat as moral. Across nine models and 7,201 priced decisions, 3,951 involved an animal rather than a hay bale or a rock. Kill rates range from 0.4% to 98.8%, with Terra and Sol the most merciful and GPT-4o-mini the most cruel, and they are not ordered by capability. Four of six models were sensitive to price at the 5% level, with elasticities from 0.09 to 1.69. All nine drove over wild animals more often than farmed animals on the default map, and the direction held at every map geometry in every model with room to move. The briefing mattered most: under the morality briefing the kill rate was under 6% in five of six reasoning models, and removing it raised the kill rate above 84% in all six. HarvestBench uses no LLM grader. The scorer counts events in the game log, so it is fully reproducible, and it measures what a model will pay to avoid harm rather than what it says about harm.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04444 (HTTP 429)
Authors: Jasmine Brazilek, Miles Tidmarsh, Matthias Endres, Anshuman Singh, Jeremiah Miller
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04444.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04444
Published: 2026-09-08T01:23:43.982Z
10. PerfReasoning: How Well Do LLMs Reason on Hardware Performance?
Abstract:Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code. Given workload, architecture, and mapping specifications, models compare mappings and predict off-chip traffic and buffer requirements. The strongest closed-source models exceed 90% on reasoning-based Q&A, and the best open-weight model reaches 82.4%. However, model construction is substantially harder: while GPT-5.6 Sol exceeds 80% pass rate, all other model configurations average below 15% and vary markedly across runs. Task-specific RL raises a 4B model’s mapping-reasoning accuracy by 15.7 points, whereas feedback-free multi-round self-revision prompting is not reliably effective. PerfReasoning exposes the gap between plausible architectural reasoning and reliable performance-model construction. We will publicly release the benchmark to support reproducible evaluation and track future progress.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
Q: 这篇论文试图解决什么问题? 这篇论文试图解决\*\*大型语言模型(LLM)在硬件性能建模与推理中的可靠性评估问题\*\*,具体围绕以下四个层面展开: --- ### 1. 核心问题:LLM能否可靠地推理硬件性能并构建分析型性能模型? 性能建模(Performance Modeling)是硬件设计与软件优化中的关键环节,用于在昂贵仿真或流片前估算内存流量、缓存容量、执行时间等指标。传统上,这类分析型模型由领域专家手动构建。论文探讨了LLM作为替代方案的潜力,提出两个互补的研究问题: - \*\*直接性能推理\*\*:给定工作负载、架构与映射(mapping)规格,LLM能否定性比较不同方案的性能(如 off-chip 流量、缓存需求)? - \*\*分析型模型构建\*\*:LLM能否生成可执行的 Python 性能模型代码,定量预测内存访问次数与缓存容量? --- ### 2. 评估范式:分离“理解”与“构造” 现有基准多关注代码生成或高层综合(HLS),但缺乏专门针对\*\*性能推理\*\*与\*\*分析型模型代码生成\*\*的成对评估。论文提出 \*\*PERFREASONING\*\* 基准,通过两类任务隔离评估: - \*\*Q&A 推理任务\*\*:通过成对映射比较(tile size、loop order、tensor keep/bypass),测试模型对数据复用、存储层级与并行性的结构化理解。 - \*\*代码构造任务\*\*:要求模型生成解析 YAML 规格并输出数值预测的 Python 程序,与 Timeloop 推导的参考标签进行严格比对( 10^(-6) 相对误差容忍)。 --- ### 3. 揭示“推理–构造”鸿沟(Reasoning–Construction Gap) 论文试图量化并解释一个关键现象:LLM 在\*\*定性推理\*\*与\*\*定量代码生成\*\*之间存在显著能力断层: - 最强闭源模型在 Q&A 推理上可达 \*\*92.6–95.4%\*\* 准确率; - 但在分析型模型构建中,除 GPT-5.6 Sol(84.7–88.4%)外,其余模型配置的精确通过率平均低于 \*\*15%\*\*,且运行间方差极大。 这表明:模型能够“谈论”性能权衡,却难以将此种理解转化为稳定、可执行且数值准确的程序。 --- ### 4. 探索提升路径与局限性 论文进一步试图验证两种改进手段在性能建模任务中的有效性: - \*\*任务特定强化学习(RL)\*\*:在二元映射比较上使用 verifier-guided RL,可将 4B 模型的推理准确率提升 \*\*15.7\*\* 个百分点; - \*\*无反馈多轮自我修正(Self-Revision)\*\*:实验表明,缺乏执行反馈的自我修正并不可靠,平均通过率仅从 14.0% 微升至 16.7%。 综上,该论文旨在建立一套严格的评估体系,以追踪 LLM 在硬件性能推理与模型构造方面的进展,并暴露当前前沿模型在\*\*可验证的定量性能建模\*\*上的根本局限。 Q2: 有哪些相关研究? 根据论文第1–2页及第4页的讨论,相关研究可归纳为以下四个方向: --- ### 1. 传统性能建模与架构模拟(手动方法) 这些工作代表了当前工业界与学术界主流的分析型性能模型及模拟器,通常由领域专家手工开发: - \*\*加速器评估框架\*\*:Timeloop(Parashar et al., 2019)作为本文生成标签的参考工具;TVM(Chen et al., 2018)与Ansor(Zheng et al., 2020)用于深度学习编译优化。 - \*\*CPU/GPU/内存模拟器\*\*:gem5(Binkert et al., 2011; Lowe-Power et al., 2020)、Ramulator(Kim et al., 2016)、DRAMSys(Jung et al., 2015)、ASTRA-sim2.0(Won et al., 2023)、SCALE-Sim v3(Raj et al., 2025)、Championship Simulator(Gober et al., 2022)。 - \*\*缓存/存储模型\*\*:CACTI(Wilton and Jouppi, 1996)、CACTI 6.0(Muralimanohar et al., 2007)。 --- ### 2. 基于学习的性能预测与设计空间探索 为降低手工建模成本,研究者尝试用数据驱动方法替代或辅助传统模拟,但面临分布外泛化与数据依赖的挑战: - \*\*离线优化与代理模型\*\*:Renda et al. (2020) 提出可微分代理模型优化 CPU 模拟器参数;Kumar et al. (2021) 利用离线优化进行加速器架构搜索。 - \*\*架构设计空间探索(DSE)环境\*\*:ArchGym(Krishnan et al., 2023)提供开源的机器学习辅助架构设计框架;AIRCHITECT(Samajdar et al., 2021)学习自定义架构设计与映射空间。 --- ### 3. LLM 在系统与硬件设计中的基准测试 本文将 PERFREASONING 定位于以下新兴基准的补充,这些基准评估 LLM 在代码生成、高层综合(HLS)、寄存器传输级(RTL)及物理设计中的能力: - \*\*GPU/系统内核生成\*\*:KernelBench(Ouyang et al., 2025)、KernelBenchX(Wang et al., 2026b)、MultiKernelBench(Wen et al., 2025)、SOL-ExecBench(Lin et al., 2026)。 - \*\*HLS 与 RTL 生成\*\*:HLS-Eval(Abi-Karam and Hao, 2025)、Bench4HLS(Khan et al., 2026)、VerilogEval(Liu et al., 2023)、RTLLM(Lu et al., 2023)、OpenLLM-RTL(Liu et al., 2025)、综合 Verilog 设计问题集(Pinckney et al., 2025)。 - \*\*芯片与物理设计\*\*:ChipBench(Yu et al., 2026)、PDAGENT-BENCH(Li et al., 2026)。 - \*\*架构智能体评估\*\*:ArchEval(Wang et al., 2026a)评估 AI 智能体作为计算机架构师的综合能力。 --- ### 4. AI 辅助架构设计范式与形式化验证 - \*\*Architecture 2.0\*\*:Sankaralingam (2026) 及 Architecture 2.0 Contributors (2026) 提出 AI 辅助计算机架构的愿景,强调可信评估器在自动优化循环中的核心作用。 - \*\*形式化验证\*\*:Aristotle(Achim et al., 2025)使用形式化验证工具实现高可靠性的复杂构造任务,论文在局限性中提及此类方法可能是未来提升性能模型可靠性的方向。 --- 综上,PERFREASONING 区别于上述工作的核心在于:它\*\*隔离了性能推理(reasoning)与分析型性能模型代码生成(construction)\*\*,并严格通过执行生成代码与参考标签对比,量化 LLM 在硬件性能建模这一特定领域中的能力边界。 Q3: 论文如何解决这个问题? 论文通过\*\*构建专用基准测试、设计双轨评估协议、执行大规模对比实验\*\*来系统解决“LLM 在硬件性能建模中是否可靠”的问题。具体方法可归纳如下: --- ### 1. 构建 PERFREASONING 基准 为将“性能推理”与“模型构建”解耦,作者引入 \*\*PERFREASONING\*\*,包含两个互补的评估组件: - \*\*Q&A 推理任务(QA-understanding)\*\* 给定同一工作负载下的两个映射(mapping),要求模型比较其 off-chip 流量或缓存容量。每个映射对通过\*\*三种 prompt 模板\*\*进行测试:同意正确陈述、同意其否定、以及直接比较。最终取三种表述中的\*\*最差准确率\*\*作为结果,以控制模型固有的同意偏差(agreement bias)。 - \*\*分析型性能模型构造任务(Analytical Performance Modeling)\*\* 要求模型生成一段 Python 程序,该程序读取工作负载、架构与映射的 YAML 规格,定量预测: - 各张量在外层存储(MainMemory)的总访问次数与逐张量访问次数; - 片上缓存(Buffer)所需容量。 --- ### 2. 设计严格的执行级评估协议 为确保定量可靠性,作者采用\*\*执行生成代码并比对参考标签\*\*的方式,而非仅做文本相似性判断: - \*\*参考标签生成\*\*:使用 Timeloop(Parashar et al., 2019)推导 216 个隐藏配置的理论真值,覆盖矩阵乘法、批处理矩阵乘法与卷积,数值跨度超过七个数量级。 - \*\*精确匹配指标(Both-Exact Pass Rate)\*\*:仅当生成模型同时满足
y(i,accesses) ≈ y(i,accesses) quad 且 quad y(i,bufsize) ≈ y(i,bufsize)
时(相对误差容忍 10^(-6) ),该样例才算通过。此指标刻意严格,要求模型在**数据移动**与**存储占用**两方面均无误。 - **Q-Error**:采用对称乘法误差
Q(y, y) = max{ max(haty,1)max(y,1), (max(y,1)) / (max(haty),1) }
以刻画预测偏离真值的倍数,区分“轻微偏差”与“数量级错误”。 - **成对偏好准确率(Pairwise Preference Accuracy)**:评估生成模型在排序层面是否能正确识别流量更低的映射,即使其绝对预测不准。 —- ### 3. 大规模系统实验 作者在固定两级硬件层次结构(MainMemory → Buffer → MACC)下,对 **16 个模型家族**(包括 Claude、DeepSeek、Gemini、GLM、GPT、GPT-OSS、MiniMax 等)在多种推理努力程度(reasoning effort)下进行了评估,并对可重复采样的模型进行多轮生成,以观察运行间方差。 —- ### 4. 诊断瓶颈并探索提升路径 在暴露问题后,论文进一步验证了两类改进手段: - **任务特定强化学习(Task-Specific RL)** 在二元映射比较上使用 verifier-guided 强化学习,对 Qwen3-4B 进行微调。结果显示,映射推理准确率从 54.3% 提升至 70.0%,证明领域专用训练可改善推理前提。 - **无反馈多轮自我修正(Feedback-Free Self-Revision)** 允许模型在不获得执行反馈的情况下,对生成代码进行自我修订(最多三轮)。实验表明,41 个配置中仅 7 个最终准确率上升,平均通过率从 14.0% 微增至 16.7%,说明**缺乏执行反馈的自我修正不可靠**。 —- ### 5. 公开释放与可复现性 作者承诺公开发布 PERFREASONING 基准,包括任务定义、YAML 规格、隐藏标签与评估脚本,以支持社区对该问题的持续追踪与可复现评估。 通过上述方法,论文不仅量化了 LLM 在“定性推理”与“定量构造”之间的能力鸿沟,也明确了向可靠 AI 辅助芯片设计迈进所需的关键技术方向:更强的结构化推理、更稳定的代码生成,以及可验证的执行反馈机制。 Q4: 论文做了哪些实验? 论文围绕 **PERFREASONING** 基准开展了一系列系统性实验,涵盖定性推理、定量代码生成、跨时间对比及干预手段验证。具体实验可归纳如下: —- ### 1. 前沿模型 Q&A 推理能力评估 在 **108 组映射对**上,对 **16 个模型家族**(Claude、DeepSeek、Gemini、GLM、GPT、GPT-OSS、MiniMax 等)于多种推理努力程度(None / Low / Medium / High / Max)下进行零样本评估。 - ** prompt 设计**:每组映射对使用三种表述模板——同意正确陈述、同意其否定、直接比较——最终取三种表述中的**最差准确率**作为该配置得分,以控制同意偏差。 - **变量隔离**:映射对按单一维度变化设计,分别为 **tile size**、**loop order**、**tensor keep/bypass**。 - **工作负载覆盖**:覆盖矩阵乘法(MatMul)、批处理矩阵乘法(Batched MatMul)与二维卷积(Conv2D)。 - **结果度量**:报告整体及按轴、按工作负载分解的准确率(图 2、图 8、图 10、表 2)。 —- ### 2. 分析型性能模型构建与执行评估 要求每个模型生成一段 Python 程序,解析 `prob.yaml`、`arch.yaml`、`map.yaml`,并预测 off-chip 访问次数(accesses)与片上缓存容量(bufsize)。 - **测试规模**:在 **216 个隐藏配置**上执行生成代码,与 Timeloop 推导的参考标签比对。 - **严格通过指标(Both-Exact Pass Rate)**:
Pass(both) = (1) / (N)∑(i=1)^(N) 1[y(i,accesses) ≈ y(i,accesses) ;wedge; y(i,bufsize) ≈ y(i,bufsize)]
其中 N=216 ,数值容忍相对误差 10^(-6) ;仅当访问与容量同时正确时样例才算通过。 - **分解指标**:单独报告 accesses 精确率、bufsize 精确率(图 14)。 - **误差度量**:采用 Q-error 评估预测偏离真值的倍数:
Q(y, y) = max{ max(haty,1)max(y,1), (max(y,1)) / (max(haty),1) }
- **运行间方差**:对支持重复采样的模型进行多次生成,报告均值与误差棒(图 3、图 13)。 —- ### 3. 代码诱导排序能力评估 验证生成的性能模型程序是否保留映射间的相对序,即使绝对数值不准确。 - **成对偏好准确率(Pairwise Preference Accuracy)**:
PrefAcc = (1) / (|P|)∑_((a,b)∈ P) 1[ (A_a < A_b) = (A_a < A_b) ]
其中 P 为真值访问次数不等的映射对集合。 - **对比实验**:将 31 个对齐配置的直接 Q&A 比较结果与代码诱导排序结果对比,发现二者平均准确率分别为 78.3% 与 78.5%,验证翻译为代码在平均意义上保留了比较信号(图 4、图 15)。 —- ### 4. 跨代际模型进展对比 为定位当前前沿水平,将 2026 年评估结果与 2025 年历史队列进行回溯比较。 - **历史基线**:2025 年 8 月的 GPT-5(61.7%)、Claude Sonnet 4.5(54.2%)、Claude Sonnet 3.7(50.8%)、Qwen3-235B-A22B(45.8%),均基于 120 例直接比较。 - **当前参考**:Gemini 3.5 Flash High(95.4%)、DeepSeek-V4-Pro Max(82.4%)。 - **结论**:最强闭源模型提升约 34 个百分点,最强开源权重模型提升约 37 个百分点(图 11)。 —- ### 5. 任务特定强化学习(RL)干预实验 验证领域专用训练能否提升映射推理能力。 - **设置**:以 Qwen3-4B 为骨干,在二元映射比较任务上进行 **verifier-guided RL** 微调,使用 Timeloop 推导的客观标签作为奖励信号。 - **评估**:在 120 例遗留 Q&A 套件上测试。 - **结果**:基础模型准确率 54.3%,RL 微调后提升至 **70.0%**,提升 **15.7** 个百分点,超过同期 GPT-5(61.7%)(图 5、图 12)。 —- ### 6. 无反馈多轮自我修正实验 检验模型在无外部执行反馈情况下,能否通过自我检查迭代改进生成代码。 - **协议**:模型先生成初始程序( k=0 ),随后最多进行 **3 轮自我修订**( k=1,2,3 ),期间不告知执行结果或通过率。 - **覆盖**:共 41 个(模型 × 推理努力程度)配置。 - **统计**:18 个配置始终无变化,10 个最终更差,仅 7 个最终改善;平均 both-exact 通过率从 14.0% 微升至 16.7%。 - **代表性轨迹**:对最强闭源配置(gpt-5.6-sol)与最强开源权重配置(gpt-oss-20b)绘制逐轮通过率曲线,显示修订可能提升、降低或持平(图 6、图 18、表 3)。 —- ### 7. 分解与相关性诊断实验 进一步分析模型表现的内部结构与影响因素。 - **Q&A 与构造能力相关性**:绘制 overall Q&A accuracy 与 performance-model exact pass rate 的散点图,得到相关系数 r = 0.46 ,表明高推理准确率是构造能力的必要非充分条件(图 16、图 17)。 - **推理努力程度扫描**:对同一模型在 Low / Medium / High / Max 等多档推理努力下同时测量 Q&A 与代码构造表现,显示额外推理努力并非均匀有益(图 17)。 - **Loop-Order 盲区分析**:统计模型在 loop-order 对上回答“二者相等”的比例,发现该比例与整体准确率呈强负相关( r = -0.97 ),且问题难度与真值流量比正相关(图 9)。 Q5: 有什么可以进一步探索的点? 基于论文的局限性声明与实验发现,以下方向值得进一步探索: —- ### 1. 扩展基准的架构与工作负载覆盖 当前基准固定于两级存储层次结构(MainMemory → Buffer → MACC),且仅涉及时序分块(temporal tiling)。可扩展至: - **更深层的存储层次**:包含 L1/L2/L3 缓存、寄存器文件、片上网络(NoC)等; - **空间并行与数据流**:支持空间分块(spatial tiling)、脉动阵列(systolic array)、权重静止/输出静止等数据流模式; - **现代工作负载**:稀疏矩阵、注意力机制、大语言模型(LLM)推理中的 KV Cache 访问模式; - **多维性能指标**:除内存流量与缓存容量外,纳入延迟(latency)、能耗(energy)、面积(area)及成本建模。 —- ### 2. 从分析一致性走向硬件实测验证 当前标签依赖 Timeloop 的分析推导,衡量的是**分析一致性**而非**硅片精度**。未来可建立: - **实测 GPU/TPU 评估管线**:以真实硬件上的 DRAM 流量、SM 占用率(occupancy)与核函数执行时间作为真值; - **噪声容忍机制**:针对硬件测量噪声定义合理的容忍带(tolerance bands); - **控制变量实验**:设计受控映射对以隔离合并(coalescing)、缓存行为、占用率与延迟隐藏(latency hiding)等底层效应,验证分析推理向真实硬件的迁移性。 —- ### 3. 提升模型构造的可靠性与稳定性 实验显示,除 GPT-5.6 Sol 外,绝大多数模型的代码生成存在**巨大的运行间方差**(run-to-run variance)。需探索: - **稳定采样策略**:为何多次采样结果差异显著?是否存在特定的解码温度、top- p 或结构化解码约束可降低方差; - **最佳候选选择**:无反馈情况下模型无法识别更优候选,需研究基于内部置信度或执行前静态分析的候选筛选机制; - **模块化程序合成**:将性能模型分解为“循环嵌套解析→复用距离计算→容量核算”等模块化组件,降低单次生成的复杂度。 —- ### 4. 强化学习与反馈机制的深度优化 论文验证了 verifier-guided RL 在 4B 模型上的二元推理任务有效,但以下问题仍开放: - **RL 对完整模型构造的迁移**:将 RL 的奖励信号从 Q&A 二元比较扩展至代码执行结果(如 both-exact 通过率或 Q-error),直接优化代码生成; - **执行反馈驱动的自我修正**:当前无反馈的自我修订基本无效。引入执行反馈(execution feedback)、单元测试与报错信息引导的多轮修正(self-debugging)是否显著提升准确率; - **在线学习架构**:在架构设计代理(agentic loop)中,让模型根据仿真器或硬件实测的反馈在线更新性能模型参数。 —- ### 5. 形式化验证与可验证代码生成 论文指出,Aristotle(Achim et al., 2025)通过形式化验证工具实现了复杂且可靠的构造。未来可探索: - **符号执行与定理证明**:将性能模型的关键不变量(如内存访问次数的上界、缓存容量的单调性)编码为形式化规约,通过 SMT 求解器或交互式定理证明器验证生成代码; - **受信代码合成**:限制生成代码的语法结构(如 DSL),使其天然可静态分析,避免通用 Python 的不可预测性。 —- ### 6. 认知机制与可解释性研究 Q&A 推理准确率超过 90% 与代码精确通过率低于 15% 之间存在显著鸿沟,其根源尚不明晰: - **知识表征分析**:LLM 是否真正内化了“分块→复用→流量”的因果链,还是仅依赖表面模式匹配? - **错误模式分类**:系统性分析失败案例——错误主要源于循环嵌套解析失误、算术计算错误、单位混淆,还是 YAML 语义误解? - **概念探测**:通过干预实验(如修改问题表述中的数值或维度名称)探测模型对数据复用、占用率等概念的鲁棒理解。 —- ### 7. 闭环 Agentic 设计流程集成 最终目标是让 LLM 成为芯片设计代理中的可信性能评估器。需验证: - **设计空间探索(DSE)中的效用**:在 ArchGym 或 Architecture 2.0 框架中,使用 LLM 生成的性能模型替代传统模拟器,测量优化循环的最终帕累托前沿质量; - **错误累积效应**:当 LLM 性能模型被用于指导数轮迭代优化时,其系统性偏差是否会放大并导致严重次优设计; - **人机协作界面**:如何将模型的定性推理(高准确率)与定量不确定性有效呈现给人类架构师,以支持混合决策。 Q6: 总结一下论文的主要内容 这篇论文介绍了 **PERFREASONING**,一个用于评估大型语言模型(LLM)在硬件性能推理与分析型性能模型构建方面能力的基准。以下是论文的主要内容概括: —- ### 1. 研究背景与核心问题 性能建模是硬件设计与软件优化的核心环节,传统上依赖领域专家手动构建分析模型。随着前沿 LLM 的发展,研究社区开始探索其两类潜在应用: - **直接性能推理**:定性比较不同硬件映射(mapping)在数据复用、存储和流量方面的优劣; - **分析模型构建**:将工作负载、架构与映射规格转化为可执行的性能模型代码,定量预测内存访问与缓存需求。 论文提出的核心问题是:**LLM 是否能够可靠地完成上述推理与模型构建任务?** 错误的性能模型可能在缺乏实测验证的情况下,悄然将优化循环导向糟糕的设计决策。 —- ### 2. PERFREASONING 基准设计 为分离“理解”与“构造”两种能力,论文设计了两类互补任务: - **Q&A 推理任务** 给定同一工作负载下的两个映射,要求模型比较其 off-chip 流量或片上缓存容量。每个映射对通过**三种 prompt 模板**(同意正确陈述、同意其否定、直接比较)进行测试,并以三种表述中的**最差准确率**作为最终结果,以消除同意偏差。映射变化聚焦于三个维度:分块大小(tiling)、循环顺序(loop order)、张量驻留/旁路(keep/bypass)。 - **分析型模型构造任务** 要求模型生成一段 Python 程序,解析 YAML 格式的规格文件,并严格定量预测: - 各张量在外层存储的总访问次数(accesses); - 片上缓存的容量需求(bufsize)。 生成代码在 216 个隐藏配置上执行,与 Timeloop 推导的参考标签比对。通过标准为**双指标同时精确匹配**(相对误差容忍 10^(-6) )。 —- ### 3. 关键实验发现 论文对 16 个前沿模型家族(包括 Claude、DeepSeek、Gemini、GLM、GPT、GPT-OSS、MiniMax 等)进行了系统评估,主要发现如下: - **推理能力较强** 最强闭源模型(如 Gemini 3.5 Flash High)在 Q&A 任务上达到 **92.6–95.4%** 的准确率;最佳开源权重模型(DeepSeek-V4-Pro)达到 **82.4%**。循环顺序是区分模型能力的最敏感维度。 - **构造能力存在显著鸿沟** 在代码生成任务中,仅 **GPT-5.6 Sol** 表现出 consistently 的高精确通过率(84.7–88.4%)。其余所有模型配置的平均精确通过率**低于 15%**,且多次采样间存在巨大方差,说明生成结果的稳定性极差。 - **排序信号得以保留,但绝对值不可靠** 尽管绝大多数模型无法输出精确的绝对流量值,但其生成代码诱导的映射排序与直接 Q&A 排序的平均准确率相当(约 78%)。这意味着模型可以保留“哪个映射更优”的序关系,但无法建立可信的绝对性能估计。 - **错误集中于访问计算,容量计算相对容易** 分解指标显示,几乎所有模型家族的缓存容量预测精确率都显著高于内存访问次数预测精确率;访问预测错误可跨越数个数量级。 —- ### 4. 改进路径探索 论文进一步测试了两种提升手段: - **任务特定强化学习(RL)** 在二元映射比较任务上对 Qwen3-4B 进行 verifier-guided RL 微调,将其 Q&A 准确率从 **54.3% 提升至 70.0%**(+15.7 个百分点),证明领域专用训练可改善推理前提。但该实验未验证其对完整代码构造任务的迁移效果。 - **无反馈多轮自我修正** 允许模型在不获得执行反馈的情况下自我修订代码(最多三轮)。结果显示该方法并不可靠:在 41 个配置中,仅 7 个最终改善,10 个变差,18 个毫无变化;平均精确通过率仅从 **14.0% 微升至 16.7%**。 —- ### 5. 结论与展望 PERFREASONING 揭示了当前前沿 LLM 在硬件性能领域的一个关键差距:**模型能够进行看似合理的架构推理,却难以将这种理解转化为稳定、可执行且数值准确的分析型性能模型**。这一“推理–构造鸿沟”对依赖 AI 代理进行芯片设计优化的愿景具有重要启示:在将 LLM 纳入自动化设计循环之前,必须解决其生成结果的稳定性、可验证性与绝对精度问题。论文承诺公开释放该基准,以支持社区对未来进展的可复现追踪。 Q7: 想要进一步了解论文 以上只是了解一篇论文的几个基本FAQ。如果你还想与Kimi进一步讨论该论文,请点击 这里 为你跳转Kimi AI网页版,并启动一个与该论文相关的新会话。
Authors: Dan Zhao, Karthikeyan Sankaralingam, Christos Kozyrakis, Qijing Huang
Categories: cs.AI
PDF URL: https://arxiv.org/pdf/2609.04476.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04476
Published: 2026-09-08T01:23:43.982Z
VLM Domain Papers
1. FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders
Abstract:Vision-language models (VLMs), such as CLIP, have achieved strong performance across multimodal tasks by aligning visual and textual representations in a shared embedding space. As VLMs are increasingly used for high-stakes domains, failure prediction becomes critical for risk-aware deployment and human intervention. Existing failure prediction methods typically rely on confidence scores or auxiliary classifiers. Although these methods are effective on predicting VLM failures, they provide limited interpretability. In this work, we investigate the use of Sparse Autoencoders (SAEs) for interpretable failure prediction in VLMs. We formulate failure prediction as a classification task over sparse SAE latent activations and introduce a three-stage failure-aware training pipeline that encourages the learned latent directions to remain interpretable while becoming more informative for failure prediction. Our experiments show that the resulting framework outperforms the evaluated baselines in failure prediction. Further analysis suggests that failure-aware training encourages SAE latent directions to capture more class-specific concepts. We also use the SAE to provide a concept-level analysis of how model representations change during failures, revealing a shift from class-specific concepts toward more ambiguous or style-related concepts. Finally, we explore how the learned SAE latent directions can support runtime failure recovery.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04276 (HTTP 429)
Authors: Jie Ma, Zongxi Liu, Yi Zhu
Categories: cs.CV
PDF URL: https://arxiv.org/pdf/2609.04276.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04276
Published: 2026-09-08T01:23:56.301Z
2. When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
Abstract:Vision-language models (VLMs) are increasingly deployed in high-stakes settings, where a response that is reasonable in general may still be unsafe for a particular user whose medical, emotional, or situational context is unknown to the model. We study this problem of personalized safety in multimodal systems and introduce MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile. Evaluating eight frontier VLMs, we find that they almost always respond directly (86-99%) rather than seek missing context, and none exceeds 2.6/5 on personalized safety. To understand why these failures arise, we analyze multimodal interactions and identify visual dominance: visual information enters text representations early and suppresses textual risk signals during multimodal fusion. Causal interventions reveal a two-stage mechanism in which visual affect is first transferred into the text stream in early layers and then shapes the final decision through this altered text representation, making late-stage internal remediation unreliable. Motivated by this mechanism, we propose PRISM, a lightweight input monitor that uses bidirectional cross-modal modulation to predict when a query is likely to require deferral. PRISM achieves 0.978 AUC and strictly dominates the safety-utility Pareto frontier across all tested models.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04281 (HTTP 429)
Authors: Edward Sun, Yuchen Wu, Zixian Ma, Eric Hanchen Jiang, Yijia Xiao, Xiaoyuan Yi, Ranjay Krishna, Wei Wang, Jindong Wang, Aylin Caliskan
Categories: cs.CV
PDF URL: https://arxiv.org/pdf/2609.04281.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04281
Published: 2026-09-08T01:23:56.301Z
3. Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation
Abstract:Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advanced multimedia content synthesis, especially in text-to-image and text-to-video tasks. To further align such generative models with human preferences, reinforcement learning (RL) has recently shown strong potential as a post-training strategy. Nevertheless, existing policy gradient-based methods often explore inefficiently, making them vulnerable to local optima that may degrade semantic faithfulness and visual realism. To address these challenges, we present Reflection-Aware GRPO (RA-GRPO), a new RL-based preference alignment framework for diffusion generative models. The core idea is to improve “forward” generation by incorporating “backward” reflection during optimization. We first introduce Diffusion Reflection, which rectifies intermediate sampling trajectories by inverting the diffusion process with a weak estimator, guiding latent states toward higher-probability regions of the true data manifold. Furthermore, we introduce Counterfactual Path Synthesis to implicitly distill these rectified trajectories into the policy, enabling the model to internalize the benefits of search-based exploration without incurring inference-time overhead. Extensive experiments on T2I and T2V models demonstrate that RA-GRPO significantly outperforms existing methods, particularly in mitigating reward hacking and improving generalization. The method remains architecture-agnostic and integrates seamlessly with standard pipelines, suggesting a promising direction for stable preference alignment.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04282 (HTTP 429)
Authors: Junlong Wu, Jiuzhou Lin, Jia Sun, Boheng Zhang, Huaiqing Wang, Dewen Fan, Houde Liu, Qianqian Gan, Fan Yang, Tingting Gao
Categories: cs.CV
PDF URL: https://arxiv.org/pdf/2609.04282.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04282
Published: 2026-09-08T01:23:56.301Z
4. Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching
Abstract:Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suffers from extensive computational overhead. Existing workflows typically treat RL and distillation as disconnected stages: applying RL before distillation incurs prohibitive computational costs, whereas applying RL after distillation frequently leads to model collapse. To overcome these limitations, we propose a unified, single-stage optimization framework grounded in Distribution Matching (DM). In the standard DM framework, distillation updates the model via a gradient direction that minimizes the gap between the real and fake models, guiding generations toward clarity and high fidelity. Building upon this, we introduce DM-Align, which derives a complementary gradient direction to guide the model toward human-preferred samples. Inspired by DPO and GRPO, our method leverages the distributional gap — formulated from either preference pairs or intra-group exploration — to directly construct this preference-guided gradient. By synergizing these two gradient directions, our approach eliminates the need for multi-step reward evaluation and complex ODE-SDE conversions inherent in traditional RL. Comprehensive experiments across multiple foundational video models demonstrate that this sample-guided framework robustly enhances both distillation quality and preference alignment, consistently outperforming both standalone variants and sequential two-stage pipelines.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04283 (HTTP 429)
Authors: Jiuzhou Lin, Junlong Wu, Fei Zuo, Huan Ouyang, Dewen Fan, Boheng Zhang, Huaiqing Wang, Jia Sun, Fan Yang, Houde Liu, Kehai Chen, Min Zhang, Tingting Gao, Han Li
Categories: cs.CV
PDF URL: https://arxiv.org/pdf/2609.04283.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04283
Published: 2026-09-08T01:23:56.301Z
5. The microscope is the mask: privileged views and labels from a cryo-ET forward model
Abstract:We explore the use of simulated data for training a model for protein annotation in crowded cryo-electron tomography volumes reconstructed from images collected at limited tilt angles and severely corrupted by the measurement operator. Firstly, we leverage the corruptions imposed by the forward model to generate domain-specific augmented paired views of the exact same scene for an invariance objective integrated into the LeJEPA self-supervised training framework. Secondly, we use additional information from the simulation pipeline such as the positions and identity of proteins in the simulated volumes to inform the architecture of the model and the loss function, so that semantic information is localised at protein positions in the resulting dense feature volume. The resulting model, CARNIVAL, is evaluated without finetuning on classification and detection tasks in real tomograms, using a benchmark dataset containing multiple protein types and two tomogram processing types. We show that CARNIVAL outperforms a state-of-the-art model trained using a contrastive objective on simulated data but without forward model-based paired views or privileged information.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04325 (HTTP 429)
Authors: Bogdan Toader, Kiarash Jamali, Tanmay A. M. Bharat, Sjors H. W. Scheres
Categories: cs.CV
PDF URL: https://arxiv.org/pdf/2609.04325.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04325
Published: 2026-09-08T01:23:56.301Z
6. Object Concepts Emerge from Motion
Abstract:Object-centric visual representations are important for physical-world perception, but existing visual pretraining methods often capture semantic categories without preserving the identity and coherence of individual instances. We present a biologically inspired framework that learns object-centric representations for single images from raw videos. Our approach uses motion boundaries as a source of object-level grouping: off-the-shelf optical flow and clustering produce pseudo-instance masks, which supervise a single-image encoder with pixel-level pairwise metric learning. The framework requires neither human annotations nor camera calibration. We first obtain 195 million pseudo-labeled frames from 7,163 hours of driving and web videos, then expand the supervision to 421 million frames with Motion-Verified Self-Training, which combines model proposals with motion evidence. We train encoders up to Swin-H and distill the learned representations into a family of Swin backbones. Across monocular depth estimation, 3D object detection, 3D occupancy prediction, and end-to-end planning, the resulting models achieve competitive or superior performance relative to supervised and self-supervised pretraining baselines, with particularly strong transfer on geometry- and instance-sensitive tasks. These results show that motion-derived supervision can teach static image encoders to represent visual instances, providing a complementary direction for scalable visual pretraining.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04348 (HTTP 429)
Authors: Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang
Categories: cs.CV
PDF URL: https://arxiv.org/pdf/2609.04348.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04348
Published: 2026-09-08T01:23:56.301Z
7. AdaptVPR: Route-Aware Hard Positive Generation for Robust Visual Place Recognition
Abstract:Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby place, yet its robustness is often degraded by domain shifts arising from illumination, weather, seasonal changes, and dynamic occlusions. One contributing factor is the limited appearance diversity of the same place in existing training data. To address this issue, we propose AdaptVPR, a route-aware generative augmentation framework that constructs same-place hard positives for robust VPR training. AdaptVPR first uses a vision language model to parse scene attributes and estimate editing feasibility, while a rule-based scheduler determines the generation route according to editability scores and risk constraints. The generation process is decomposed into three complementary routes: the Global Appearance Route introduces global scene changes in weather, illumination, and time of day; the Local Occlusion Route inserts plausible dynamic occluders; and the Dual Route combines both types of perturbations to produce more challenging appearance shifts. Each generated candidate is evaluated using a VPR-oriented verification scheme based on geometric consistency and appearance diversity, reducing the risk of structural drift while ensuring sufficient appearance variation. Global candidates are generated once and rejected if verification fails, while Local Occlusion and Dual candidates use verification feedback for limited prompt refinement and regeneration. Using this framework, we construct AdaptCities, containing 160K verified synthetic same-place hard positives. Experiments across multiple VPR baselines and vision foundation backbones show consistent gains on standard benchmarks and substantial improvements under challenging domain shifts, with R@1 gains of up to 9.2%. The source code and data resources are publicly available at this https URL.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04369 (HTTP 429)
Authors: Shunpeng Chen, Jingyi Zhang, Changwei Wang, Shengpeng Xu, Yukun Song, Xingtian Pei, Jinzhou Lin, Li Guo, Shibiao Xu
Categories: cs.CV
PDF URL: https://arxiv.org/pdf/2609.04369.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04369
Published: 2026-09-08T01:23:56.301Z
8. Where Appearance Fails, Geometry Recognizes: A CAD-Free 3D Shape Prior That Complements Vision Foundation Models
Abstract:Recognizing specific objects onboarded without a labeled training set recurs across manufacturing and service robotics, yet the conventional renderable prior, a computer-aided-design (CAD) model, is often unavailable. Two-dimensional capture supplies no shape prior, and frozen foundation features fail on geometrically similar, low-texture industrial parts. We ask what a short object-centric scan buys for recognition beyond the captured images themselves: each object is reconstructed with 3D Gaussian Splatting (3DGS), summarized into a per-class shape prototype, and fused with frozen DINOv2 image features. First, the scan recovers the recognition value of CAD without CAD: geometry from RGB-D depth (on T-LESS), 3DGS, and CAD gives comparable recognition (tied on HOPE, within 1.6 points on T-LESS); 3DGS is only a convenient route to a point cloud. Second, the payoff is governed by how recognizable the shape is: on shape-distinctive household objects (HOPE) geometry alone reaches 0.920 versus image-only 0.832, a ceiling below which fixed-weight fusion (0.872) sits. On shape-confusable textureless industrial parts (T-LESS) the gain is modest but consistent (0.560 to 0.591 fused, above both single signals). Third, the prior is complementary, not uniformly additive: it rescues far more image failures than it breaks successes, and its benefit grows under partial occlusion. Finally, the worth lies in geometry, not rendered pixels: 3DGS renderings do not help the image side, and frozen-feature recognition is nearly lighting-invariant (within 2.5 points). The study is scoped to recognition, not the BOP pose benchmark.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04381 (HTTP 429)
Authors: Chenxi Tao, Seung-Kyum Choi
Categories: cs.CV
PDF URL: https://arxiv.org/pdf/2609.04381.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04381
Published: 2026-09-08T01:23:56.301Z
9. What Moves? Localized Motion Representations for Compositional Scene Control
Abstract:Real-world dynamics are inherently compositional: multiple entities move simultaneously within a shared scene, each exhibiting distinct motion patterns. Yet most existing video representations encode motion globally, without explicitly capturing localized motion for individual entities. Crucially, motion is defined relative to a global reference frame, including camera motion and scene layout. However, localized embeddings are often computed from cropped images or obtained by masking features after encoding, discarding the context needed to interpret motion. To address this, we introduce a promptable localized motion representation that produces persistent embeddings for user-specified regions defined by spatial masks. Rather than cropping the input or masking features, our model processes the full video and conditions motion encoding directly on the queried region. This yields temporally consistent, region-addressable embeddings that isolate local dynamics while retaining the global context required for disambiguation. We demonstrate object-level motion transfer, enabling controlled composition of dynamic scenes. Beyond generative control, our embeddings support localized action classification in multi-actor videos. Across both tasks, our approach improves controllability and outperforms global representations localized through cropping or post-hoc masking. Project Page: this https URL
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04383 (HTTP 429)
Authors: Frank Fundel, Malek Ben Alaya, Thomas Ressler-Antal, Stefan Andreas Baumann, Björn Ommer
Categories: cs.CV
PDF URL: https://arxiv.org/pdf/2609.04383.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04383
Published: 2026-09-08T01:23:56.301Z
10. FAVE: Foveated Adaptive Visual Encoding for Efficient Fine-Grained Visual Understanding
Abstract:Fine-grained visual understanding depends on local detail, yet visual encoders face a trade-off between costly full-image high-resolution processing and compact global encoding that can weaken such evidence. Inspired by human active vision, we separate where to look from what to encode. We focus on the latter and introduce FAVE (Foveated Adaptive Visual Encoding), a lightweight variable-resolution ViT that encodes externally selected regions at high acuity while preserving native geometry. We first isolate this encoding problem using oracle ground-truth crops in a controlled small-object regime. On ImageNet objects with a native maximum side of 96 pixels, FAVE improves Top-1 by 9.4 points over a fixed-resolution ViT on the same crop window with 12.7 times lower FLOPs. Increasing global resolution or backbone capacity does not recover the same operating point. We then integrate FAVE as a complementary local branch in FastVLM. Its local tokens are combined with FastVLM’s global visual tokens, while the original global pathway and language model remain frozen. With at most 16 additional local tokens, FAVE improves TextVQA by 1.60 points and achieves a 3.3 times controlled TTFT speedup over SmolVLM2-2.2B. On GQA attribute questions, it improves FastVLM-1.5B by 1.31 points, extending the benefit beyond text while narrowing the gap to FastVLM-7B. Together, these results show that selectively allocating high-acuity local capacity provides an efficient complement to broader global representations and model scaling for fine-grained understanding of small objects, text, and attributes.
中文摘要
翻译失败: Translation limit exceeded. Please try it again later.
Response status: 401 (Unauthorized)
Response body : “{\”ShowCaptcha\”:false}”
LLM Analysis
LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2609.04392 (HTTP 429)
Authors: Amitangshu Mukherjee, Kaushik Roy
Categories: cs.CV
PDF URL: https://arxiv.org/pdf/2609.04392.pdf
CoolPaper URL: https://papers.cool/arxiv/2609.04392
Published: 2026-09-08T01:23:56.301Z