数据来源:ArXiv Domain

LLM Domain Papers

1. NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts

Abstract:Ancient Indian medical texts like Sushruta Samhita have extensive information on diseases, treatments, and surgical techniques. Yet, their ancient format and use of intricate vocabulary pose difficulties in accessibility and systematic ordering. The research here utilizes Natural Language Processing (NLP) methods like Named Entity Recognition (NER), BERTopic modeling, and Knowledge Graph development in Neo4j to extract, categorize, and visualize important concepts based on translated versions. Thematic classification with BERTopic allows for the identification of the underlying medical topics, whereas NER supports the structured entity recognition of diseases, treatments, researchers, and medicinal plants. Graphbased network analysis with Neo4j also allows for the semantic representation of relationship among extracted entities, supporting knowledge retrieval and digital preservation. The findings illustrate how graph databases, topic modeling, and entity recognition facilitate the computational organization of Ayurveda’s historical medical wisdom, closing the gap between the conventional texts and contemporary data-driven inquiry. The suggested method promotes historical text analysis, medical informatics, and digital humanities to make ancient Indian medical wisdom more accessible and understandable.

中文摘要

摘要:古印度医学文献如《苏述罗集》中包含大量关于疾病、治疗和外科技术的信息。然而,它们的古老格式和复杂词汇的使用造成了获取和系统整理的困难。本研究利用自然语言处理(NLP)方法,如命名实体识别(NER)、BERTopic建模和Neo4j知识图谱开发,对翻译版本中的重要概念进行提取、分类和可视化。BERTopic的主题分类允许识别潜在的医学主题,而NER则支持疾病、治疗、研究人员和药用植物的结构化实体识别。基于图的网络分析与Neo4j也可实现提取实体之间关系的语义表示,支持知识检索和数字化保存。研究结果展示了图数据库、主题建模和实体识别如何促进阿育吠陀历史医学智慧的计算机化组织,缩小传统文献与现代数据驱动研究之间的差距。所提出的方法推动了历史文献分析、医学信息学和数字人文的发展,使古代印度医学智慧更易获取和理解。

LLM Analysis

Q: 这篇论文试图解决什么问题? 该研究旨在解决古印度医学文献(特别是《妙闻集》\*Sushruta Samhita\*)因语言障碍、结构松散及古今知识体系割裂而导致的知识提取困难与系统化组织缺失问题。具体而言,论文试图应对以下核心挑战: - \*\*文本可及性与理解障碍\*\* 古代医学文本采用复杂的梵语词汇、特殊的叙述结构及非常规的系统分类方式,且经英译后存在大量转写变体与格式噪声,导致现代读者难以直接访问和理解其中的医学知识。 - \*\*非结构化知识难以计算化处理\*\* 传统文本以非结构化形式存储疾病、疗法、草药与学者信息,缺乏适合现代信息检索与数据挖掘的表示形式,无法支持语义查询、关联分析与知识推理。 - \*\*古今医学知识体系的割裂\*\* 阿育吠陀(Ayurveda)的传统智慧尚未有效融入当代数据驱动的医学研究与教育,缺少能够桥接历史文献与现代生物医学分析的数字化方法与工具。 - \*\*领域专用实体识别的缺失\*\* 通用命名实体识别(NER)模型难以准确识别梵语派生术语、阿育吠陀专用名词(如疾病、疗法、草药)及其在翻译文本中的变体,造成实体漏检与误分类。 为应对上述问题,该研究提出了一套基于自然语言处理(NLP)的整合方案:利用命名实体识别(NER)抽取医学实体,采用 BERTopic 进行主题建模以发现潜在医学主题,并基于 Neo4j 构建知识图谱以显式表示实体间的语义关系,从而实现古代印度医学知识的结构化提取、主题分类与可视化分析。 Q2: 有哪些相关研究? 根据论文第2节(Related Studies)的综述,相关研究主要围绕\*\*历史文献的NLP处理\*\*、\*\*主题建模技术\*\*、\*\*知识图谱构建\*\*以及\*\*印度语言与医学领域的专用模型\*\*四个维度展开,具体如下: ### 1. 历史文献的数字化与NLP处理 - \*\*Voukoutis et al. (2024)\*\* 提出了面向希腊语文本的开放大语言模型 \*\*Meltemi 7B\*\*,该模型基于400亿词元的希腊语语料训练,优化了历史文献的文本理解、自动分段与实体识别,为古代文献的数字化提供了基础技术参考。 - \*\*Akhil et al. (2020)\*\* 探讨了NLP在历史文献自动化处理中的应用,涉及文本分割、实体识别与主题建模等关键任务。 ### 2. 主题建模与神经网络嵌入 - \*\*Ginn & Hulden (2024)\*\* 将动态主题建模(Dynamic Topic Modeling)应用于基于BERT的神经网络嵌入,对现存全部罗马文学语料进行分析。研究表明,相较于传统统计模型(如潜在狄利克雷分配,LDA),神经模型在超参数稳定性与定性洞察方面更具优势。 - \*\*Mutsaddi et al. (2025)\*\* 针对印地语短文本进行了BERTopic与LDA的比较研究,证实BERTopic在生成连贯主题方面优于LDA,揭示了基于Transformer的主题建模在多语言文本分析中的潜力。 - \*\*Anoop & Asharaf (2020; 2021)\*\* 在主题建模与语义知识发现方面开展了一系列工作,为电子商务与 Aspect-Oriented 情感分析中的主题建模提供了方法论支撑。 ### 3. 知识图谱在医学与历史文本中的应用 - \*\*Terdalkar & Bhattacharya (2023)\*\* 构建了面向梵语文献的自动化知识图谱框架,支持基于查询的知识检索,使古代文本能够以机器可读的格式被访问。 - \*\*Terdalkar et al. (2023)\*\* 在 \*\*Āyurjñānam\*\* 项目中,通过人工注释药物信息,构建了阿育吠陀(Ayurveda)文本的知识图谱,并集成用户界面以支持有效查询与标注,展示了传统医学知识计算化存储的路径。 - \*\*TKDL (2025)\*\*(传统知识数字图书馆)由印度政府推动,采用标准化术语与分类体系对的传统医学知识进行数字化组织,旨在促进传统知识融入现代科学话语与法律保护体系,防止生物盗版。 ### 4. 面向印度语言与医学领域的专用NLP模型 - \*\*Bahad et al. (2024)\*\* 针对印度语言的命名实体识别(NER)约束,对预训练NER模型进行了微调,拓展了NLP技术在多语言文本中的适用性。 - \*\*Mullick et al. (2023)\*\* 开发了面向印度语言医疗查询的意图分类与实体抽取模型,强调了针对医学与历史文本构建领域专用NLP方法的必要性。 - \*\*Sarella & Mangam (2024)\*\* 综述了AI驱动的NLP在健康传播中的应用,指出形式化实体识别能够提升历史背景中疾病与治疗信息抽取的准确性与效率。 - \*\*Srivastava et al. (2024)\*\* 提出了 \*\*MedPromptExtract\*\*,一种基于半监督学习、大语言模型与提示工程的医疗数据抽取工具,验证了NLP在将非结构化医学数据转化为结构化数据方面的有效性。 - \*\*Gurav (2024)\*\* 综述了NLP在电子健康记录(EHR)中的应用,重点关注信息抽取与医疗沟通增强。 ### 5. 其他相关技术工作 - \*\*Lekshmi & Anoop (2023)\*\*、\*\*Jickson et al. (2023)\*\*、\*\*Varghese & Anoop (2022)\*\* 与 \*\*Krishnan & Anoop (2023)\*\* 等研究分别在社交媒体情感分析、抑郁检测、COVID-19新闻视频分析以及气候变化情感分析等领域应用了机器学习与深度学习技术,为本文的文本挖掘与分析方法提供了跨领域的方法论借鉴。 - \*\*Sharaf & Anoop (2023)\*\* 对医疗领域大语言模型(如BioBERT)进行了分析,为医学文本的深度学习应用提供了背景支持。 这些研究共同表明,计算语言学、AI模型与知识工程的交叉融合,为古代手稿的信息抽取、语义标注与数字人文研究提供了日益成熟的理论与技术基础。本文正是在此基础上,将NER、BERTopic与Neo4j知识图谱整合应用于《妙闻集》(\*Sushruta Samhita\*)的英译本分析。 Q3: 论文如何解决这个问题? 该研究通过一套融合\*\*自然语言处理(NLP)\*\*与\*\*知识工程\*\*的系统性计算框架,对《妙闻集》(\*Sushruta Samhita\*)英译本进行结构化提取、主题发现与语义关联,具体实施路径如下: --- ### 1. 数据预处理与文本清洗 针对古代翻译文本中的格式噪声与转写不一致问题,研究首先执行多阶段预处理: - \*\*原始文本提取\*\*:从英译PDF中提取 raw text; - \*\*噪声过滤\*\*:移除页码、章节标题、罗马数字及多余空格; - \*\*格式标准化\*\*:修复不一致的排版、字符编码与间距问题; - \*\*词汇归一化\*\*:利用自定义词典处理梵语术语的英译变体(transliterations),为后续实体识别与主题建模提供干净的语料基础。 --- ### 2. 命名实体识别(NER)与领域适配 将NER视为序列标注任务,对文本中的每个词序列 X = x_1, x_2, dots, x_n 预测对应的实体标签序列 Y = y_1, y_2, dots, y_n ,其中 y_i 表示疾病、疗法、学者、草药或地点等类别。 - \*\*统计模型基础\*\*:采用条件随机场(CRF)建模标签转移概率:

P(Y mid X) = exp(∑ Wt^T f(y_t, y(t-1), X, t))∑(Y’) exp(∑ W_t^T f(y’_t, y’(t-1), X, t))
其中 f(·) 为特征函数(如词性、词嵌入), W 为训练得到的权重向量, Y’ 遍历所有可能的标签序列。 - **深度学习增强**:引入 BiLSTM-CRF 架构,利用双向长短期记忆网络捕捉上下文依赖,再以CRF层进行结构化预测。其损失函数定义为:
L = -∑_(i=1)^(N) log P(Y_i mid X_i)
其中 N 为训练样本数。 - **领域工具集成**:结合 **spaCy** 与 **ScispaCy**(面向生物医学文本的NLP库)识别医学条件、治疗手段与植物实体,并通过**自定义阿育吠陀词典**增强对梵英转写变体的覆盖,最终将识别出的实体映射到标准化的阿育吠陀术语体系。 —- ### 3. BERTopic 主题建模与语义聚类 为克服传统LDA对古代医学文本语义捕捉不足的问题,研究采用 **BERTopic** 进行神经主题建模,其核心计算流程包括: - **句子嵌入**:通过 Transformer 模型(如BERT、SBERT或DistilBERT)将文本映射为高维语义向量:

E(X) = f_(BERT)(X)

  • **非线性降维**:使用 UMAP(Uniform Manifold Approximation and Projection)在保留语义关系的同时将高维嵌入降至低维空间:
    Z = UMAP(E(X))
  • **密度聚类**:采用 HDBSCAN 基于密度分布将相似嵌入划分为主题簇:
    T = HDBSCAN(Z)
  • **主题关键词提取**:对每个主题应用类内 TF-IDF(c-TF-IDF)计算词项代表性:
    c-TF-IDF(t,T) = f(t,T)∑(t’ ∈ T) f(t’,T) · log (N) / (∑(T’ ∈ D) 1)(t ∈ T’)
    其中 f
    (t,T) 为词 t 在主题 T 中的频率, N 为文档总数, 1_(t ∈ T’) 为指示函数。 通过该流程,文本被自动聚类为**医学疗法**、**疾病与障碍**、**学术典籍**、**药用植物**、**地理与历史背景**等可解释主题。 —- ### 4. 知识图谱构建与语义网络分析 为实现实体间关系的显式表示与语义检索,研究将提取的实体与主题组织为图结构: - **形式化定义**:知识图谱表示为 G = (V, E) ,其中 V 为节点(实体), E 为边(关系)。 - **图数据库实现**:采用 **Neo4j** 原生图数据库存储高度互联的数据,并使用 **Cypher** 查询语言进行图操作。 - **实体映射与关系抽取**:将NER识别的疾病、疗法、草药、学者、地点等作为节点,依据文本中的共现与语义关联建立边。最终构建的图谱包含25个节点与26条关系,节点按语义角色着色(如概念、疗法、疾病、植物、学者、地点)。 - **语义查询支持**:该结构支持对“疾病—疗法—草药—学者”复杂关联的路径查询与推理,为数字人文研究提供可扩展的知识基座。 —- ### 5. 可视化与计算整合 研究通过多维度可视化将分析结果转化为可交互、可解读的知识视图: - **NER分布可视化**:利用柱状图与旭日图(sunburst chart)展示各类实体(NORP、PRODUCT、GPE、疾病、学者等)的层级与频次; - **主题空间可视化**:绘制主题词分数图(topic word scores)、主题频率排名、二维**主题距离图**(intertopic distance map)以及**三维UMAP投影**,揭示主题间的语义亲疏与密度分布; - **知识图谱网络图**:在Neo4j中生成网络视图,直观呈现典籍、学者、疾病、疗法与草药之间的语义关联。 —- ### 6. 方法论集成逻辑 上述方法并非孤立执行,而是形成闭环: 1. **预处理后的干净文本**输入NER与BERTopic; 2. **NER输出的结构化实体**与**BERTopic发现的主题簇**共同作为节点与关系来源; 3. **Neo4j知识图谱**整合两者结果,实现从“非结构化古籍”到“可查询、可推理的语义网络”的转化; 4. 最终通过**可视化界面**支持数字人文、医学信息学与历史文本分析的跨学科应用。 该框架有效地将阿育吠陀传统知识转化为机器可读、可计算、可扩展的数字资源,从而弥合了古代医学文本与现代数据驱动研究之间的鸿沟。 Q4: 论文做了哪些实验? 该研究的实验围绕**《妙闻集》(*Sushruta Samhita*)英译本**展开,依次实施了数据预处理、命名实体识别(NER)、BERTopic主题建模与知识图谱构建四类核心实验。各实验的设计、步骤与产出结果如下: —- ### 1. 文本预处理与数据清洗实验 **目的**:消除古代翻译文本中的格式噪声与转写变异,构建标准化的分析语料。 **实验步骤**(对应表1): | 步骤 | 操作内容 | 输出结果 | |:—-|:—-|:—-| | 1. 文本提取 | 从英译PDF中提取原始文本 | 含错误的原始文本 | | 2. 噪声去除 | 删除页码、章节标题、罗马数字及多余空格 | 去噪可读文本 | | 3. 格式修复 | 修正不一致排版、间距及非标准字符 | 格式化干净文本 | | 4. 实体识别(NER) | 识别疾病、疗法、学者、地点等关键实体 | 分类命名实体 | | 5. 主题建模(BERTopic) | 利用BERTopic提取主题并聚类相关术语 | 主题簇与主题 | | 6. 最终数据准备 | 整合清洗后的结构化数据集 | 供知识图使用的最终数据集 | **关键操作**:通过Python中的 **NLTK** 进行分词与基础预处理,并针对梵语-英语转写变体建立自定义词典,以解决术语拼写不一致问题。 —- ### 2. 命名实体识别(NER)与分类实验 **目的**:自动抽取文本中的医学实体,并按领域语义进行系统分类与频次统计。 #### (1)实体分类实验 采用 **spaCy** 与 **ScispaCy**(面向生物医学文献的NLP库)结合自定义阿育吠陀词典,对实体进行五类语义标注(表2): | 序号 | 类别 | 代表性实体 | 领域意义 | |:—-|:—-|:—-|:—-| | 1 | 疾病(Diseases) | Kushtha, Prameha, Gulma, Vidradhi, Jwara | 经典阿育吠陀病理体系 | | 2 | 疗法(Treatments) | Sneha, Vasti, Triphala, Ghrita, Abhyanga | 排毒与疗愈技术 | | 3 | 学者(Scholars) | Susruta, Charaka, Dallana, Vagbhata | 古印度医学文献贡献者 | | 4 | 药用植物(Medicinal Plants) | Haritaki, Nimba, Ashwagandha, Guduchi | 草药与长寿方剂 | | 5 | 地点与典籍(Locations & Texts) | India, Rigveda, Sushruta Samhita | 历史与地理语境 | #### (2)实体分布统计实验 对全部识别出的实体按标准NER标签进行频次分析(表3),发现: - **NORP**(民族/宗教/政治团体,含医学流派)出现频次最高,达 **565** 次; - **PRODUCT**(医疗制品/器械)**242** 次; - **DATE**(时间)**174** 次;**GPE**(地缘政治实体)**79** 次;**LOC**(地点)**64** 次; - 其他如 **WORK_OF_ART**(典籍/艺术作品)**58** 次,**EVENT**(事件)仅 **2** 次。 **可视化产出**: - **图2**:按类别展示提取实体的分布柱状图; - **图3**:旭日图(sunburst chart)层级展示“学者—疾病—疗法—药用植物—地点与典籍”的实体层级关系。 —- ### 3. BERTopic主题建模实验 **目的**:在无监督条件下发现文本中的潜在医学主题,揭示阿育吠陀知识的内在语义结构。 #### (1)建模流程实验(表4) | 步骤 | 技术操作 | 输出 | |:—-|:—-|:—-| | 1. 预处理 | 去除停用词、数字、符号及低频词 | 无噪声干净文本 | | 2. 分词与词形还原 | 分词、词形还原并过滤无关项 | 词形还原后的词元 | | 3. BERTopic训练 | 在预处理文本上训练BERTopic模型 | 主题簇 | | 4. 主题提取与聚类 | 对相似词聚类并分配概率得分 | 带概率的主题-词分布 | | 5. 主题解释与标注 | 分析主题顶部词汇及语境并赋予标签 | 精炼的主题标签 | | 6. 最终主题数据 | 生成包含主题及其关系的结构化数据集 | 供可视化与知识图使用的主题-实体数据集 | #### (2)主题发现结果(表5) 模型共识别出5个核心主题: | 主题编号 | 主题名称 | 代表性术语 | 主题重要性 | |:—-|:—-|:—-|:—-| | 1 | 医学疗法(Medical Treatments) | Vasti, Sneha, Triphala, Ghrita, Dhuma | 探索古今阿育吠陀净化与治疗技术的有效性 | | 2 | 疾病与障碍(Diseases & Disorders) | Kushtha, Prameha, Vidradhi, Gulma, Jwara | 理解古医学对疾病的分类与病理生理学 | | 3 | 学术典籍(Scholarly Texts) | Sushruta Samhita, Charaka, Vagbhata, Nidana | 展示阿育吠陀知识通过历史学者的传承 | | 4 | 药用植物(Medicinal Plants) | Haritaki, Nimba, Madhuka, Amalaki, Pippali | 突出植物药学与传统草药疗愈 | | 5 | 地理与历史(Geographical & Historical) | India, Hastimsha, Rigveda, Ayurveda | 提供阿育吠陀发展与文化背景 | **可视化产出**: - **图4**:各主题的关键词得分(topic word scores),展示每个主题最具代表性的词汇; - **图5**:前10个最频繁主题的分布,其中“医学疗法”与“疾病与障碍”占比最高,并包含如 *vagbhata, vataja, vamana*(与Vagbhata、Vata病、催吐疗法相关)以及 *bhadradarvadi, chhinna, chedana*(与外科切除、创伤处理相关)等术语; - **图6**:主题间距离图(intertopic distance map),以二维气泡图展示主题语义亲疏,中心密集区域反映医学与操作主题的重叠,外围离散气泡代表专科子领域; - **图7**:3D UMAP投影,将高维语义空间降维至三维,以颜色梯度区分不同主题簇,验证主题间的聚类紧密度与分离度。 —- ### 4. 知识图谱构建与网络分析实验 **目的**:将抽取的实体与主题整合为可查询、可推理的语义网络,显式表达阿育吠陀概念间的关联。 **技术实现**: - 使用 **Neo4j Desktop** 作为原生图数据库,采用 **Cypher** 查询语言进行图数据插入与检索; - 知识图谱形式化定义为:

G = (V, E)
其中 V 为节点(实体), E 为边(关系)。 **图谱规模与结构**: - **节点**:共 **25** 个; - **关系**:共 **26** 条; - **节点类型与颜色编码**: - 概念(蓝色):如 *Ayurveda, Sushruta Samhita* - 疗法(紫色):如 *Sneha, Vasti* - 疾病(棕色):如 *Vidradhi, Prameha* - 药用植物(绿色):如 *Nimba, Guduchi* - 学者(红色):如 *Susruta, Charaka* - 地点/典籍(橙色):如 *India, Rigveda* **可视化产出**: - **图8**:知识图谱网络视图,展示“典籍—学者—疾病—疗法—草药—地点”之间的语义关联,例如 *Sushruta Samhita* 与 *Charaka*、*Triphala*、*Prameha* 等节点的关系路径。 —- ### 实验总结 上述实验形成了完整的处理流水线:**原始PDF → 清洗文本 → NER实体抽取 → BERTopic主题发现 → Neo4j知识图谱构建 → 多维度可视化**。通过该系列实验,研究验证了NLP技术在古代医学文献结构化、主题归纳与语义关联挖掘中的可行性,并产出可量化的实体分布统计、可解释的主题簇以及可查询的图结构数据。 Q5: 有什么可以进一步探索的点? 基于论文第5节(Recommendations)与第6节(Discussion and Conclusion)的讨论,以下几个方向可作为后续深入探索的重点: —- ### 1. 领域自适应NLP模型的构建 现有通用命名实体识别(NER)模型在处理梵语派生术语、阿育吠陀行话及英译转写变体时存在误分类与漏检。未来可着力于: - **开发阿育吠陀专用NER模型**:整合自定义实体识别规则与音译(transliteration)消歧机制,提升对古代医学术语的识别精度; - **构建历史医学语料库**:在梵语-英语对照的古代医学文本上开展预训练或微调,以增强模型对领域词汇的表征能力。 —- ### 2. 混合主题建模与本体工程融合 BERTopic虽能有效识别主题结构,但其主题标签的语义连贯性仍有优化空间: - **引入本体驱动的混合分类框架**:将BERTopic的无监督聚类与基于规则的主题分类(如阿育吠陀医学本体)相结合,降低主题歧义; - **优化主题一致性评估**:改进连贯性评分(coherence scores)的选取机制,使主题聚类更符合古代医学知识体系的分层逻辑。 —- ### 3. 知识图谱的因果推理与跨域互操作 当前Neo4j知识图谱主要刻画实体间的共现与关联关系,尚待深化: - **增广因果与药理关系**:显式建模“疾病—疗法—草药”之间的因果、禁忌与协同作用,而非仅停留在语义共现层面; - **增强语义推理与复杂查询**:引入描述逻辑或图神经网络(GNN)方法,支持多跳推理与假设验证; - **对接国际生物医学本体**:实现与传统知识数字图书馆(TKDL)及现代医学本体(如SNOMED CT、MeSH)的互操作,推动阿育吠陀知识融入全球循证医学研究框架。 —- ### 4. 标准化知识表示与数字人文基础设施 古代医学文本的数字化保存亟需统一标准: - **建立阿育吠陀知识的标准化表示模型**:定义疾病、疗法、草药、学者等核心实体的元数据规范,便于跨历史、语言学与生物医学领域复用; - **构建可扩展的数字化平台**:支持多版本文本对照、众包标注与长期知识演化追踪,形成可持续的数字人文研究基础设施。 —- ### 5. 跨文献比较与纵向历史分析 本研究仅聚焦《妙闻集》(*Sushruta Samhita*)英译本,未来可扩展至: - **多文本语料整合**:纳入《遮罗迦集》(*Charaka Samhita*)、《八支心要集》(*Ashtanga Hridaya*)等经典,开展跨典籍的医学原理比较; - **历时性知识演化研究**:利用动态主题建模(Dynamic Topic Modeling)追踪阿育吠陀概念体系在不同历史时期的变迁与传播路径。 —- ### 6. 跨学科验证与知识整合 计算提取的结果仍需领域专家校验: - **人机协同验证机制**:联合语言学家、阿育吠陀医学 practitioner 与计算科学家,对自动抽取的实体、主题与关系进行专家级审校; - **古今诊疗知识融合**:探索将图谱化的古代医学知识用于辅助现代补充与替代医学(CAM)研究,或作为临床决策支持系统的历史知识层。 这些探索点共同指向一个目标:在保持古代医学文本原真性的前提下,通过计算手段实现传统知识的**可解释、可推理、可复用**,从而深化数字人文与医学信息学的交叉研究。 Q6: 总结一下论文的主要内容 该论文围绕古印度医学经典《妙闻集》(*Sushruta Samhita*)的英译本,提出了一套融合自然语言处理(NLP)与知识工程的综合计算框架,旨在解决古代医学文本因语言复杂、结构松散及非结构化特性而导致的知识提取与系统化组织难题。主要内容可概括如下: ### 1. 研究背景与核心问题 古印度阿育吠陀(Ayurveda)文献蕴含丰富的疾病分类、外科技术与药理学知识,但其复杂的梵语-英语转写词汇、混乱的文本结构以及非常规的系统分类,阻碍了现代读者的访问与计算化利用。现有通用NLP模型在处理此类历史医学文本时面临实体误分类、语义遗漏及古今知识体系割裂等挑战。 ### 2. 研究方法与技术路线 研究采用端到端的文本挖掘与知识表示流程,核心包括三个阶段: - **数据预处理**:对英译PDF文本进行提取、去噪(移除页码、标题、罗马数字)、格式标准化及词形还原,构建干净的分析语料。 - **命名实体识别(NER)**:结合 **spaCy**、**ScispaCy** 与自定义阿育吠陀词典,将文本中的疾病(如 *Kushtha, Prameha*)、疗法(如 *Sneha, Vasti*)、学者(如 *Susruta, Charaka*)、药用植物(如 *Ashwagandha, Guduchi*)及地点/典籍实体进行序列标注与分类。技术上以条件随机场(CRF)与 BiLSTM-CRF 为模型基础,优化对领域术语的识别。 - **BERTopic 主题建模**:利用 Transformer 嵌入(BERT/SBERT)、UMAP 降维与 HDBSCAN 密度聚类,从文本中无监督地提取主题;并通过类内 TF-IDF 计算提取各主题的代表性关键词,识别出医学疗法、疾病障碍、学术典籍、药用植物及地理历史五大核心主题。 - **知识图谱构建**:基于 **Neo4j** 图数据库,将提取的实体与主题关系形式化为图结构 G = (V, E) ,使用 Cypher 查询语言实现存储、检索与可视化,显式呈现疾病、疗法、草药与学者之间的语义关联。 ### 3. 实验结果与发现 - **实体识别**:共识别并分类大量实体,其中民族/医学传统(NORP)565次、医疗制品(PRODUCT)242次、日期(DATE)174次、地点(GPE/LOC)143次,系统揭示了文本的医学与历史语境。 - **主题结构**:BERTopic 分析显示,“医学疗法”与“疾病与障碍”为文本最突出的主题;同时发现外科术语(如 *chedana, chhinna*)、草药知识(如 *nimba*)及时间医学(如 *morning, evening*,关联 *Dinacharya* 昼夜节律)等子主题。 - **语义可视化**:通过主题词得分图、主题频率排名、二维主题距离图(intertopic distance map)及三维 UMAP 投影,展示了主题间的语义亲疏与分布密度。 - **知识图谱**:成功构建包含25个节点与26条关系的网络,涵盖概念、疗法、疾病、植物、学者及地点六类节点,直观呈现了阿育吠陀知识体系的内在关联。 ### 4. 研究贡献 - 证明了领域适配的 NER 与神经主题建模在历史医学文本中的有效性; - 建立了首个面向《妙闻集》英译本的结构化实体分类与主题发现方案; - 通过 Neo4j 知识图谱实现了古代医学知识的语义关联存储与可视化,为数字人文研究提供了可查询、可扩展的数据基础; - 弥合了传统阿育吠陀知识与现代数据驱动研究之间的鸿沟,推动了医学文献的计算化保护与利用。 ### 5. 局限与未来方向 - **模型专用性**:通用NER对梵语派生术语的识别仍有不足,需构建阿育吠陀专用的预训练语言模型; - **主题本体融合**:建议将 BERTopic 与基于规则的本体分类结合,提升主题标签的语义准确性; - **图谱推理深化**:当前图谱以关联关系为主,未来可引入因果推理(如疾病-疗法-草药的药理因果)及与现代生物医学本体的互操作; - **语料扩展**:将方法推广至《遮罗迦集》等其他阿育吠陀经典,开展跨文献比较与历时性知识演化研究。 综上,该研究为古代医学文献的数字化保存、结构化分析及计算化利用提供了一套可复现的技术范式,在医学信息学、数字人文及历史文本分析领域具有明确的应用价值。 Q7: 想要进一步了解论文 以上只是了解一篇论文的几个基本FAQ。如果你还想与Kimi进一步讨论该论文,请点击 这里 为你跳转Kimi AI网页版,并启动一个与该论文相关的新会话。

Authors: M. S. Rajeevan, B. Mini Devi, V.S. Anoop, C. Mallikarjuna

Categories: cs.CL

PDF URL: https://arxiv.org/pdf/2608.28608.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28608

Published: 2026-09-02T01:23:21.918Z


2. Parametric Multimodal User Memory: Storing What Captions Cannot Carry

Abstract:A personalized agent needs a user memory: a persistent model of who its user is. Today it is almost always text — transcripts and captions retrieved by similarity. This serves the captionable half of a person (“my cat is named Bibi”), but discards the perceptual half no caption can hold: how a voice sounds, how a face reads across age and lighting, how tired someone sounds. We measure this loss across five modalities: a strong caption-based re-identifier recovers as little as 0.11 of a dedicated encoder’s recall, collapsing toward chance on non-nameable signals. We instead ground perceptual memory in the model, decomposing recall into two subproblems: a vision-language model grounds the referent in context (what and where), and a dedicated encoder extracts an identity key (who), stored as one inline token read by attention at generation with no external round-trip. Neither suffices alone — the VLM identifies cross-age faces at only 0.54 recall where a face encoder reaches 0.81, and an ungrounded encoder recognizes a two-person-scene referent at 0.05 — yet together they reach correct-region oracle (0.96), generalizing to multi-speaker audio and video. The recognition core is training-free: it reproduces the encoder’s recall on any frozen model at O(1) registration cost. On PerceptMem (12 domains, 1,080 tasks) perceptual identity is capacity-limited while exact facts are binding-limited: identity belongs in a parametric bank, facts in a text store. The two memories compose cleanly: an agent with both can remember not only what its user said, but also what they are like.

中文摘要

摘要:一个个性化代理需要用户记忆:一个关于用户是谁的持久模型。如今,这几乎总是文本——通过相似性检索的转录和字幕。这能服务于一个人中可以用文字表达的部分(“我的猫叫Bibi”),但却丢弃了任何字幕都无法承载的感知部分:声音的音色如何、面孔在不同年龄和光照下的变化、一个人听起来有多疲惫。我们在五种模态上衡量了这种损失:一个强大的基于字幕的再识别器恢复的准确率仅为专用编码器召回率的0.11,在无法命名的信号上几乎退化到随机水平。我们转而将感知记忆在模型中进行扎根,将召回分解为两个子问题:一个视觉-语言模型(VLM)在上下文中确定指称对象(是什么以及在哪里),一个专用编码器提取身份键(是谁),存储为生成时通过注意力读取的内联Token,无需外部往返。单独使用都不够——VLM在跨年龄人脸识别上的召回率仅为0.54,而面部编码器可达0.81;未扎根的编码器在两人场景中识别指称对象的召回率为0.05——但结合起来可以达到正确区域或acular(0.96),并可推广到多说话人音频和视频。识别核心无需训练:它在任何冻结模型上以O(1)的注册成本再现编码器的召回率。在PerceptMem(12个领域,1,080个任务)上,感知身份受限于容量,而精确事实受限于绑定:身份属于参数化存储,事实属于文本存储。两种记忆可以干净地组合:拥有两者的代理不仅可以记住用户说过什么,还可以记住用户的特性。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28609 (HTTP 429)

Authors: Bojie Li, Noah Shi

Categories: cs.CL

PDF URL: https://arxiv.org/pdf/2608.28609.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28609

Published: 2026-09-02T01:23:21.918Z


3. Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System

Abstract:Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trained on Western-centric data, making them ill-suited for regional curricula like India’s. The Indian education system is linguistically diverse, exam-oriented, and structured around standardized syllabi, not addressed by existing datasets or tools. In this work, we curate a syllabus-aligned QA dataset based on NCERT (National Council of Educational Research and Training) textbooks for classes 9-12, capturing the content, context, and teaching style of Indian curricula. The final dataset, comprising 18,720 question-answer pairs across five subjects, is publicly available at this https URL. We fine-tune the LLaMA 3.1 8B model using this dataset and deploy it in a Retrieval-Augmented Generation (RAG) framework tailored to educational needs. We introduce GurukulAI, an open-access platform that enables Indian students to chat with the model, get doubts cleared, practice exam-style questions, receive contextual answers, and interact in both English and Hindi. By localizing AI for Indian classrooms, our work bridges the gap between global LLM capabilities and regional educational demands. The code is available at this https URL.

中文摘要

摘要:最近,大型语言模型(LLMs)如 ChatGPT 和 LLaMA 的进展已经改变了 AI 驱动的教育,但这些系统主要基于西方数据进行训练,因此不适用于像印度这样的本地课程。印度的教育体系具有语言多样性、以考试为导向,并且围绕标准化教材结构,而现有数据集或工具并未涵盖这些特点。在本工作中,我们基于 NCERT(国家教育研究与培训委员会)9-12 年级教材,策划了一个与课程大纲对齐的问答数据集,捕捉了印度课程的内容、上下文和教学风格。最终的数据集包含 5 门学科共 18,720 个问答对,已在此 https URL 上公开。我们使用该数据集对 LLaMA 3.1 8B 模型进行了微调,并在针对教育需求设计的检索增强生成(RAG)框架中部署。我们推出了 GurukulAI,这是一个开放访问平台,使印度学生能够与模型进行对话,解答疑问,练习考试风格的题目,获取上下文答案,并可用英语和印地语互动。通过将 AI 本地化应用于印度课堂,我们的工作弥合了全球大型语言模型能力与当地教育需求之间的差距。代码可在此 https URL 获取。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28611 (HTTP 429)

Authors: Isha Narang, Sneh Gosai, Mayank Singh

Categories: cs.CL

PDF URL: https://arxiv.org/pdf/2608.28611.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28611

Published: 2026-09-02T01:23:21.918Z


4. STAGEET: Stage-wise Typed Edit Tagging for Grammatical Error Correction with Arabic as a Case Study

Abstract:Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rather than generating a full corrected sentence. Their interpretability, however, is primarily operational: a label specifies how the string should change, but a single edit vocabulary does not always reveal the type of correction being made. We propose STAGEET, a stage-wise typed edit-tagging framework that reorganizes Seq2Edit supervision into typed executable stages and extends edit operations to correction categories. STAGEET decomposes correction into an ordered sequence of medium-grained typed stages; each stage predicts from its own label space, rewrites the current hypothesis once, and passes the resulting intermediate sentence to the next stage. We instantiate the framework as both an end-to-end shared-encoder multi-head model with stage-specific adapters and a fully specialized variant with one independent tagger per stage. Experiments on QALB-2014 and ZAEBUC show that category-aware staged correction retains competitive edit-based GEC performance while exposing a more inspectable correction trajectory, and attains state-of-the-art results on QALB-2014.

中文摘要

摘要:序列到编辑(Sequence-to-edit)方法通过预测输入上的编辑标签而不是生成完整的纠正句子,使语法错误纠正(GEC)高效且具有局部可解释性。然而,它们的可解释性主要是操作层面的:一个标签指定了字符串应如何变化,但单一的编辑词汇表并不总能揭示正在进行的纠正类型。我们提出了STAGEET,一种分阶段类型化编辑标注框架,它将Seq2Edit的监督重组为类型化可执行阶段,并将编辑操作扩展到纠正类别。STAGEET将纠正分解为有序的中粒度类型化阶段序列;每个阶段在其自身的标签空间中进行预测,对当前假设进行一次重写,并将生成的中间句子传递到下一阶段。我们将该框架实例化为端到端共享编码器多头模型(带阶段特定适配器)以及每个阶段拥有独立标注器的完全专用变体。在QALB-2014和ZAEBUC上的实验表明,类别感知的分阶段纠正保持了具有竞争力的基于编辑的GEC性能,同时展示了更易检查的纠正轨迹,并在QALB-2014上取得了最先进的结果。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28614 (HTTP 429)

Authors: Wenjie Lou, Alaa Mamdouh Akef

Categories: cs.CL

PDF URL: https://arxiv.org/pdf/2608.28614.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28614

Published: 2026-09-02T01:23:21.918Z


5. From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education

Abstract:Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds. Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable and preserve full turn by turn dialogue. However, these logs are educationally difficult to use directly. Complete transcripts are too detailed for routine teacher review, whereas final scores obscure whether learners followed up patient cues, checked uncertainty, or used summaries to guide later questioning. This study examined whether coded GenAI VP dialogues can provide teacher-interpretable process evidence of clinical reasoning. We analysed 1{,}030 GenAI VP dialogues from 210 second-year medical learners across five weeks chest-pain cases. Each consultation was teacher-scored using a rubric assessing the full history taking dialogue, and consultations were classified within each week as high- or low-rated using the weekly median score. To explain how rated performance was reflected in the dialogue process, we applied three analytic layers to the same coded dialogue data: behavioural prevalence, local co-occurrence using Epistemic Network Analysis, and sequential transition using Transition Network Analysis. High-rated consultations involved more history taking activity, but differences were not simply about volume. High rated consultations more often connected information gathering and symptom exploration with communication, checking, organisation, and synthesis. Summarising and organising moves more often led to verification or mechanism-oriented follow-up. These findings show how layered analysis of GenAI VP dialogue logs can reveal process patterns associated with high rated history taking and support process-focused feedback in medical education.

中文摘要

摘要:病史采集是一项基于对话的临床推理任务,学习者必须在会诊过程中收集、组织并整合患者信息。生成式人工智能虚拟患者(GenAI VP)使反复进行病史采集练习成为可能,并保留完整的逐轮对话。然而,这些记录直接用于教学具有一定难度。完整的文字记录对于日常教师审阅过于详细,而最终评分则无法体现学习者是否跟进患者提示、检查不确定性或使用总结来指导后续提问。本研究探讨了编码的 GenAI VP 对话是否能够为教师提供可解释的临床推理过程证据。我们分析了 210 名二年级医学学习者在五周胸痛病例中产生的 1,030 次 GenAI VP 对话。每次会诊都使用评估完整病史采集对话的评分细则进行教师评分,并根据每周的中位数评分将每周的会诊分为高评分或低评分。为了说明评分表现如何在对话过程中体现,我们对相同的编码对话数据应用了三层分析方法:行为出现频率、本地共现(使用认知网络分析)以及顺序转换(使用转换网络分析)。高评分的会诊涉及更多的病史采集活动,但差异不仅仅在于数量。高评分会诊更常将信息收集和症状探索与沟通、核实、组织和综合联系起来。总结和组织操作更常引导至验证或机制导向的后续提问。这些发现表明,对 GenAI VP 对话记录进行分层分析可以揭示与高评分病史采集相关的过程模式,并支持医学教育中以过程为导向的反馈。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28619 (HTTP 429)

Authors: Xinyu Li, Zijian Li, Mengyu Xia, Luzhen Tang, Naping Chen, Changmin Lin, Danijela Gasevic, Dragan Gasevic, Yizhou Fan

Categories: cs.CL

PDF URL: https://arxiv.org/pdf/2608.28619.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28619

Published: 2026-09-02T01:23:21.918Z


6. Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

Abstract:Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence. However, for LMRMs no reliable method to measure sycophancy yet exists. We bridge this gap by introducing a benchmark and dataset for evaluating LMRM sycophancy when confronted with a wrong answer from a user. Our benchmark pairs four visually grounded datasets spanning mathematical, clinical, temporal, and demographic reasoning with five pressure conditions in single-turn and multi-turn settings. We evaluate sycophancy in the final answer as well as its emergence within the reasoning chain. We find that sycophancy is prevalent under pressure, with Statement pressure eliciting the highest rates and Conviction the lowest for all models except Mistral-Small-4, and under multi-turn pressure reasoning-level sycophancy intensifies sharply in clinical visual judgement, reaching 95.7% for the most affected model. We further introduce a failure taxonomy separating reasoning-chain from answer-level sycophancy, and a complementary sentence-level taxonomy locating where in the chain drift first emerges. Our results show that sycophancy can corrupt the reasoning chain independently of the final answer, so answer-level evaluation alone is insufficient.

中文摘要

摘要:大型多模态推理模型(LMRMs)的能力正日益增强,主要是通过在回答之前生成明确的思维链推理。在语言模型中观察到,这种性能往往伴随着谄媚行为,即模型倾向于在证据面前迎合用户。然而,对于LMRMs,目前还没有可靠的方法来衡量谄媚行为。我们通过引入一个基准和数据集来填补这一空白,用于评估LMRM在面对用户错误答案时的谄媚表现。我们的基准将四个跨数学、临床、时间和人口统计推理的视觉基础数据集与五种单轮和多轮设置下的压力条件配对。我们评估了最终答案中的谄媚行为以及其在推理链中的出现情况。结果发现,在压力下谄媚行为十分普遍,其中Statement压力引发的谄媚率最高,而Conviction压力引发的最低(除了Mistral-Small-4模型之外)。在多轮压力下,推理链级别的谄媚在临床视觉判断中急剧增强,对于受影响最严重的模型,达到95.7%。我们进一步引入了一种失效分类法,将推理链级与答案级的谄媚行为区分开,并附加一个句子级分类法,用于定位谄媚行为首次在推理链中出现的位置。我们的结果显示,谄媚行为可以独立于最终答案破坏推理链,因此仅依靠答案级评估是不够的。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28623 (HTTP 429)

Authors: Mahir Numayeer Islam, Gakuto Okuyama, Nikolaus Siauw, Shivank Garg, Madhur Panwar, Vasu Sharma

Categories: cs.CL

PDF URL: https://arxiv.org/pdf/2608.28623.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28623

Published: 2026-09-02T01:23:21.918Z


7. MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson’s Disease Assessments

Abstract:Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson’s disease is time-consuming and often depends on specialist expertise. Although large language models (LLMs) can generate natural language summaries, they frequently lack domain-specific clinical grounding and struggle to produce factually correct and temporally consistent responses for structured longitudinal assessment data. To address these limitations, we propose MA-RAG, a query-driven multi-agent retrieval-augmented generation framework that decomposes clinical reasoning into domain-specialized agents, combines structured fact extraction, and synthesizes clinically grounded summaries through a final verification stage. The framework supports four clinical analysis tasks: single-session, trajectory, comparison, and cohort summarization. We evaluate MA-RAG using objective metrics, namely Fact Precision, Hallucination Rate, Temporal Fidelity, and Semantic Similarity, together with subjective evaluations conducted by clinical experts. Compared to Traditional, RAG-only, and Single-agent RAG baselines, MA-RAG substantially improves factual correctness, achieving up to a 122% relative increase in Fact Precision (from 0.436 to 0.990) and reducing the Hallucination Rate by up to 98% (from 0.564 to 0.010), while consistently receiving top ratings from clinical experts for organization and clinical usefulness. These results demonstrate that domain-specialized multi-agent reasoning enables reliable query-driven summarization of structured longitudinal clinical assessment data.

中文摘要

摘要:对帕金森病的单次就诊和纵向临床评估进行准确解读既耗时,又通常依赖于专家的专业知识。尽管大型语言模型(LLM)能够生成自然语言总结,它们往往缺乏领域特定的临床基础,并且难以为结构化的纵向评估数据生成事实正确且时间一致的响应。为了解决这些局限性,我们提出了MA-RAG,一种基于查询驱动的多智能体增强检索生成框架,该框架将临床推理分解为领域专门化的智能体,结合结构化事实提取,并通过最终验证阶段综合生成具有临床依据的总结。该框架支持四类临床分析任务:单次会话总结、轨迹分析、比较分析和队列总结。我们使用客观指标(即事实精确度、虚构率、时间一致性和语义相似性)以及临床专家的主观评价对MA-RAG进行评估。与传统方法、仅RAG方法和单智能体RAG方法相比,MA-RAG显著提高了事实正确性,使事实精确度相对增加最多122%(从0.436提高到0.990),并将虚构率最多降低98%(从0.564降至0.010),同时在临床专家对组织结构和临床有用性评分中持续获得最高评价。这些结果表明,领域专门化的多智能体推理能够实现对结构化纵向临床评估数据的可靠查询驱动总结。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28624 (HTTP 429)

Authors: Sana Alamgeera, Denise Goberta, Muhammad Irshad, Anne H. H. Ngu

Categories: cs.CL

PDF URL: https://arxiv.org/pdf/2608.28624.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28624

Published: 2026-09-02T01:23:21.918Z


8. Asymmetric Within-Document Predictive Learning for Scientific Document Representation

Abstract:We study predictive pretraining for scientific document representation using the discourse structure of papers. We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations. Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrastive baseline using the same section pairs. Adding Sliced Isotropic Gaussian Regularization (SIGReg) substantially improves performance and narrows this gap. The effect of regularization is task-dependent: moderate SIGReg helps fine-grained ranking, while stronger regularization can weaken local alignment. We further show that different encoding branches support different retrieval regimes. These results position within-document predictive learning as a promising citation-free complement for scientific document representation, provided that embedding geometry is carefully controlled.

中文摘要

摘要:我们研究了利用论文的篇章结构进行科学文档表示的预测预训练。我们提出了SciJEPA,一种无需引用的框架,通过文档内部非对称预测进行学习:使用标题和摘要的表示来预测方法部分的表示,使用方法部分的表示来预测结论部分的表示。在RELISH、高影响力引用、SciDocs和引用预测上的实验表明,单纯的预测训练是可行的,但比使用相同章节对的受控对比基线效果较弱。引入切片各向同性高斯正则化(SIGReg)可以显著提高性能,并缩小这一差距。正则化的效果依赖于任务:适度的SIGReg有助于细粒度排序,而较强的正则化可能削弱局部对齐。我们进一步展示了不同的编码分支支持不同的检索模式。这些结果表明,只要嵌入几何结构得到仔细控制,文档内部预测学习作为科学文档表示的一种无需引用的补充方法是有前景的。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28625 (HTTP 429)

Authors: You Zuo, Éric de la Clergerie, Benoît Sagot

Categories: cs.CL

PDF URL: https://arxiv.org/pdf/2608.28625.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28625

Published: 2026-09-02T01:23:21.918Z


9. Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

Abstract:Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models’ training cutoffs. Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format. Additionally, 145 verifiably detectable errors were inserted into 55 manuscripts to assess error identification under natural and verification-oriented prompts. Across all manuscript groups, including rejected submissions, LLM scores ranged from 7.0 to 8.1, whereas human mean scores ranged from 3.4 to 6.8. The models detected 12.1\% of the verified errors under natural prompting, and a one-sentence verification instruction increased detection to 22.2\%; however, 78\% of the errors remained undetected. Providing figures reduced error detection while increasing review scores. No visual error was reliably verified against its corresponding figure, and half of the text-only reviews described figures that were not provided. Author identity did not influence either review scores or error detection. LLM editorial decisions exactly matched those produced by simple score averaging.

中文摘要

摘要:大型语言模型(LLMs)越来越多地被用于生成同行评审,这促使人们开始审视其批判性评估能力。本研究评估了两种多模态大型语言模型 Qwen2.5-VL-72B 和 Pixtral-Large-124B,在 2026 年国际学习表征会议(ICLR 2026)165 篇投稿中的审稿表现,该会议的时间晚于两种模型的训练截止日期。稿件以三种方式呈现给模型:作者身份保密、高声望机构替代或低声望机构替代,同时提供两种格式:仅文本或文本加图表。此外,在 55 篇稿件中插入了 145 个可验证的已知错误,以评估模型在自然提示和验证导向提示下的错误识别能力。在所有稿件组中,包括被拒稿件,LLM 的评分范围为 7.0 到 8.1,而人类的平均评分范围为 3.4 到 6.8。在自然提示下,模型检测到 12.1% 的已验证错误,而增加一句验证指令后,检测率提高到 22.2%;然而,仍有 78% 的错误未被发现。提供图表降低了错误检测率,同时提高了审稿评分。没有视觉错误能够可靠地与其对应的图表验证匹配,而一半的仅文本审稿描述了未提供的图表。作者身份对审稿评分和错误检测均无影响。LLM 的编辑决定与简单平均评分得出的结果完全一致。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28626 (HTTP 429)

Authors: Emad Alharbi

Categories: cs.CL

PDF URL: https://arxiv.org/pdf/2608.28626.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28626

Published: 2026-09-02T01:23:21.918Z


10. Intelligent Identification and Repair of Design Defects in BIM via Domain-Specific Large Language Models

Abstract:Existing methods lack a generalized approach to efficiently identify and resolve the diversity of design defects in BIM. Therefore, this study proposes an integrated framework to identify and repair various defects in BIM via domain-specific LLMs. Firstly, a BIM-to-Text method with component-balanced chunking is introduced to bridge BIM data with LLMs. Then, prompt learning with rule injection, few-shot prompting and RAG is proposed to identify defects and generate repair suggestions. Meanwhile, a hallucination control strategy combining key identifier validation and token-length thresholds is introduced to ensure reliability. Experiments show capability expansion yields 85% identification accuracy versus 70% for traditional rule checking, achieving a 94% rate of reasonable repair suggestions. Moreover, the proposed hallucination control further increased accuracy from 64% to 85%, eliminating 92.5% of hallucinations in a single intervention round. This study establishes an end-to-end prototype from raw BIM data input, through defect identification, to repair suggestion generation.

中文摘要

摘要:现有方法缺乏一种通用方法来高效识别和解决BIM中的多样化设计缺陷。因此,本研究提出了一个综合框架,通过领域特定的LLMs识别和修复BIM中的各种缺陷。首先,提出了一种具有组件平衡分块的BIM到文本方法,以桥接BIM数据与LLMs。然后,提出结合规则注入、少量示例提示和RAG的提示学习方法,以识别缺陷并生成修复建议。同时,引入一种结合关键标识验证和令牌长度阈值的幻觉控制策略,以确保可靠性。实验表明,能力扩展使识别准确率达到85%,而传统规则检查为70%,实现了94%的合理修复建议率。此外,提出的幻觉控制策略进一步将准确率从64%提升至85%,在一次干预中消除了92.5%的幻觉。本研究建立了一个端到端原型,从原始BIM数据输入,经缺陷识别,到生成修复建议。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28629 (HTTP 429)

Authors: Jia-Rui Lin, Yun-Hong Cai, Xiang-Rui Ni, Peng Pan

Categories: cs.CL

PDF URL: https://arxiv.org/pdf/2608.28629.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28629

Published: 2026-09-02T01:23:21.918Z


Agent Domain Papers

1. DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation

Abstract:Large Language Model (LLM) agents have shown promise for automating data-science workflows, yet their end-to-end performance depends critically on the agent harness that represents tasks, manages execution state, constrains output artifacts, and provides evaluation feedback. Existing data-science agents often leave this harness implicit, making results difficult to reproduce, compare, and attribute across heterogeneous tasks. We introduce DS-Lighting, a unified harness toolkit that makes harness design explicit for data-science automation. DS-Lighting decomposes the harness into four reusable layers: data, workflow, execution, and evaluation, and represents diverse agents as executable operator programs that support both predefined pipelines and adaptive search. We further integrate multiple open-source data-science benchmarks into an MLE-Bench-style task format, enabling controlled comparison under a shared task interface, sandboxed runtime, and metric protocol. Experiments across agents, harnesses, models, and ablations show that explicit harness design improves reproducibility, comparability, and reliability, while reducing avoidable system-level failures in end-to-end data-science workflows. Our code is available at this https URL

中文摘要

摘要:大型语言模型(LLM)代理在自动化数据科学工作流程方面显示出潜力,但它们的端到端性能在很大程度上取决于代理的框架,该框架负责表示任务、管理执行状态、约束输出产物并提供评估反馈。现有的数据科学代理通常将这种框架隐式化,使得结果难以在异构任务中复现、比较和归因。我们提出了DS-Lighting,一个统一的框架工具包,它使数据科学自动化的框架设计显式化。DS-Lighting将框架分解为四个可重用层:数据层、工作流层、执行层和评估层,并将多样化的代理表示为可执行的操作程序,这些程序同时支持预定义流水线和自适应搜索。我们进一步将多个开源数据科学基准集成到MLE-Bench风格的任务格式中,从而在共享任务接口、沙箱运行时和指标协议下实现受控比较。在对代理、框架、模型和消融实验的测试中显示,显式的框架设计提高了复现性、可比性和可靠性,同时减少了端到端数据科学工作流程中可避免的系统级失败。我们的代码可通过此https URL获取。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28590 (HTTP 429)

Authors: Fan Liu, Hao Liu

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28590.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28590

Published: 2026-09-02T01:23:41.475Z


2. Expert-validated STEM QA

Abstract:Recent advancements in AI are helping scientists achieve breakthroughs in fields such as mathematics, medicine, and materials sciences. New evaluation datasets for AI models contribute to such advancement in AI. In the STEM domain, frontier models have consumed most of the available online data, creating the need for human-created datasets that codify the knowledge of leading experts in the domain. There are several STEM datasets available for the research community in this field. However, there are some gaps in these datasets, leaving room for improvement. Examples of gaps include (1) saturation in model performance on these datasets, leaving no head-room for meaningful evaluations, (2) skewed taxonomy distributions, (3) multiple choice question format that is misaligned with how scientists use AI in the real world, and (4) inaccurate answers and rationales partially led by a contest-based data collection and a time-bound review process. In this study, we present ‘Expert-validated STEM QA’, a high-quality, expert-validated STEM dataset (N=398) in Physics, Chemistry, Biology, and Mathematics, created by 241 domain experts. We (1) carefully designed a taxonomy with balanced distribution, (2) vetted question contributors with quality-driven incentive, (3) conducted multiple rounds of reviews with revisions validated by domain experts based on consensus, and (4) created the dataset in verifiable question and answer format. Our study demonstrated low performance ($<25\%$) of frontier AI models on the dataset as a benchmark. Post-training on a separate, private version of the dataset (N=2,000) increased performance of the open source model by $15\%$ relative to the baseline model (p=0.045) on the STEM subset of HLE-verified dataset, indicating potential utility of the dataset for model training. We have open-sourced a portion of our dataset for the AI research community.

中文摘要

摘要:近期人工智能的进展正在帮助科学家在数学、医学和材料科学等领域取得突破。用于人工智能模型的新评估数据集也推动了人工智能的发展。在STEM领域,前沿模型已经消耗了大部分可用的在线数据,因此需要由人类创建的数据集,将该领域顶尖专家的知识进行编码。该领域有若干可供研究社区使用的STEM数据集。然而,这些数据集中存在一些空白,需要改进。空白的例子包括:(1) 这些数据集上的模型性能饱和,缺乏进行有意义评估的空间;(2) 分类分布偏斜;(3) 多选题格式与科学家在实际使用人工智能的方式不符;(4) 部分答案和理由不准确,受比赛型数据收集和有时间限制的审查过程影响。在本研究中,我们提出了“专家验证的STEM问答”,这是一个高质量、经过专家验证的STEM数据集(N=398),覆盖物理、化学、生物和数学领域,由241名领域专家创建。我们(1) 精心设计了分布均衡的分类体系;(2) 对问题贡献者进行了质量驱动的审核;(3) 进行了多轮审查,并由领域专家根据共识验证修订;(4) 创建了可验证的问题和答案格式的数据集。我们的研究显示,前沿人工智能模型在该数据集上的表现较低($<25\%$),作为基准。将模型在数据集的另一个私有版本(N=2,000)上进行后训练后,开放源模型在HLE验证数据集的STEM子集上的性能相对基线模型提高了$15\%$(p=0.045),表明该数据集在模型训练中具有潜在应用价值。我们已将部分数据集开源,以供人工智能研究社区使用。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28591 (HTTP 429)

Authors: Kihwan Han, Saurabh Patil, Chinmayee Shukla, Abhinav Sharma, Marko Pavlovic, Anshuman Lall, Mahesh Joshi

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28591.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28591

Published: 2026-09-02T01:23:41.475Z


3. A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making

Abstract:Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test—it is a sequence of guideline-pathway choices, escalation judgments, and commitments under uncertainty. Existing benchmarks largely measure factual recall, leaving open whether frontier LLMs share decision-path blind spots that combining models cannot fix. We built the Oncology Decision Boundary Benchmark (ODBB)—2,005 oncology decision points across NCCN guidelines and colorectal cancer cases—and evaluated nine frontier LLMs (four closed-source, five open-weight families) released between June 2025 and April 2026. A fully deterministic scorer (zero LLM inference) classified outputs into 14 failure types, independently validated by two oncologists (Cohen’s weighted $\kappa$ = 0.939 and 0.790) on a 225-item stratified sample. Treating the nine as a pooled super-model, 42.1% (Wilson 95% CI 40.0—44.3%) of all items—35.7% of the 1,586 NCCN items and 66.4% of the 419 colorectal-cancer cases—were answered correctly by none, with failures concentrated in choosing between guideline pathways before reasoning within any: a consistent blind spot in clinical meta-judgment that likely requires architectural intervention rather than more training data. Two models tuned for decisiveness (GPT-5.5, Gemini 3.1 Pro Preview) made unsafe commitments three to five times more often than the seven cautious models without scoring higher. In 3—9% of items, models stated the correct next clinical step yet did not commit to it—failures of decision, not knowledge. Model quality is no longer the primary bottleneck for clinical LLM deployment; the binding constraint is the assumption that any single model can be the sole basis for a clinical decision. Progress requires architectures that detect when a model reaches its competence boundary and route the decision to a clinician.

中文摘要

摘要:大型语言模型(LLMs)在医学知识考试中取得了高分,但现实世界的肿瘤学并不是知识考试——它是指南路径选择、升级判断以及不确定性下承诺的连续过程。现有基准主要衡量事实回忆,而前沿LLM是否存在组合模型无法修复的决策路径盲点仍未明朗。我们构建了肿瘤学决策边界基准(ODBB)——覆盖NCCN指南和结直肠癌病例的2005个肿瘤学决策点——并评估了九个前沿LLM(四个闭源,五个开源权重家族),这些模型在2025年6月至2026年4月期间发布。一个完全确定性的评分程序(零LLM推理)将输出分类为14种失败类型,并由两位肿瘤学家在225项分层样本上独立验证(Cohen加权κ=0.939和0.790)。将九个模型视为一个合并的超级模型,42.1%(Wilson 95% CI 40.0—44.3%)的所有项目——其中1,586个NCCN项目的35.7%和419个结直肠癌病例的66.4%——没有任何模型答对,失败主要集中在在任何内部推理之前选择指南路径:这是临床元判断中的一致盲点,可能需要架构干预而非更多训练数据。两个针对决断力调优的模型(GPT-5.5,Gemini 3.1 Pro Preview)做出不安全承诺的频率是七个谨慎模型的三到五倍,但得分并未更高。在3—9%的项目中,模型声明了正确的下一步临床步骤,但未作出承诺——这是决策失败,而非知识失败。模型质量不再是临床LLM部署的主要瓶颈;约束条件是任何单一模型都可作为临床决策唯一依据的假设。进展需要能够检测模型何时达到其能力边界并将决策转交给临床医生的架构。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28592 (HTTP 429)

Authors: Zhang Sheng, Jinming Li, Wangyang Chen, Zhiwei Bao, Yu YoSean Wang

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28592.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28592

Published: 2026-09-02T01:23:41.475Z


Abstract:With the increasing development of AI regulatory frameworks, ensuring that artificial intelligence systems, particularly generative models, operate in accordance with legal and ethical standards has become a critical priority. Existing proposals for AI alignment and value-guided behavior, however, face some limitations. Approaches such as Constitutional AI depend on human supervision, while broad normative frameworks like the Good-for-Humanity (GfH) principle may be overly general and ambiguous to provide actionable governance guidance. To overcome these limitations, we propose a hybrid approach called Statutory AI that employs pre-existing human-authored principles drawn from specific themes within a legal corpus. Specifically, Statutory AI uses legal texts as a constitutional framework, enabling AI systems to autonomously critique and revise their outputs according to established norms. It operates in two stages, both using Chain-of-Thought prompting. The first stage classifies the user prompt into one of the identified themes, while the second stage analyzes it in conjunction with relevant articles selected from the legal corpus of that theme. To illustrate the potential of our approach, we conducted an experiment involving 1,000 red-teaming prompts and five penal themes: discrimination, disclosure of confidential information, violence, fraud, and abuse of vulnerable persons. Statutory AI reduced harmful content by 52 to 59 percentage points across tested models, approximately 10 percentage points higher than standard Constitutional AI, while cutting computation time by over 50%.

中文摘要

摘要:随着人工智能监管框架的不断发展,确保人工智能系统,尤其是生成模型,按照法律和伦理标准运行,已成为关键优先事项。然而,现有的人工智能对齐和价值引导行为的提案存在一些局限性。诸如宪法人工智能(Constitutional AI)的方法依赖于人工监督,而广泛的规范性框架如‘利于人类原则’(Good-for-Humanity, GfH)可能过于笼统和模糊,难以提供可操作的治理指导。为克服这些局限性,我们提出了一种名为法定人工智能(Statutory AI)的混合方法,利用法律语料中特定主题下的人类既有原则。具体而言,法定人工智能将法律文本作为宪法框架,使人工智能系统能够根据既定规范自主批评和修正其输出。它分两个阶段运行,均采用推理链(Chain-of-Thought)提示。第一阶段将用户提示分类至识别出的主题之一,第二阶段结合从该主题法律语料中选取的相关条文进行分析。为展示我们方法的潜力,我们进行了涉及1,000个红队提示和五个惩罚性主题(歧视、机密信息泄露、暴力、欺诈以及对弱势群体的虐待)的实验。法定人工智能在测试模型中将有害内容减少了52%至59个百分点,比标准的宪法人工智能高出约10个百分点,同时计算时间减少了50%以上。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28593 (HTTP 429)

Authors: Cindy Delage, Stéphane Canu, Marc Décombas, Jonathan Foureur

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28593.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28593

Published: 2026-09-02T01:23:41.475Z


5. From Question-First to Analyst-First: Domain-Expert Skills and Verified Knowledge Compilation for Proactive Enterprise Analytics

Abstract:Conversational analytics systems assume the user already has a well-formed question, leaving a non-expert facing a blank query box on an unfamiliar enterprise schema. Commercial ‘proactive’ tools narrow this gap only by detecting statistical anomalies over analyst-curated metric layers, and academic next-question recommenders depend on query logs that a fresh dataset lacks. We describe a production analytics system that inverts the interaction model from question-first to analyst-first through two coupled architectural ideas. First, a pluggable domain-expert ‘skill’ abstraction: a folder-based, database-free subject-matter pack (a manifest, per-stage prompt facets, keyword-routed references, report templates, and optional compute) auto-selected per (client, dataset) by deterministic schema matching and spliced as a cross-cutting concern into every stage of an agentic pipeline, the schema explorer, and the report engines, degrading to a strict no-op when absent. Because a skill is a self-contained folder resolved deterministically, the catalogue is open-ended: an extensible marketplace of domain experts. Second, an offline knowledge-compilation loop: an agent probes the dataset’s parquet via DuckDB (zero load on production), runs critic-gated per-table convergence with self-healing retries, and data-validates joins by value overlap, producing durable schema knowledge that drives standing expert reports whose every published metric is re-verified by re-executing its evidence SQL, plus suggested questions that mirror the report agenda. These close a proactive loop: reports surface numbers, the numbers seed questions, and a click launches a verified deep dive, all before the query box is used. We give a formal model and report illustrative single-tenant evidence. We make no user-study or benchmark claims; the contribution is the architecture and its defensibility.

中文摘要

摘要:会话分析系统假设用户已经有了一个完整的问题,这使得非专业用户在面对陌生的企业架构时只能面对空白的查询框。商业“主动”工具仅通过在分析师策划的指标层上检测统计异常来缩小这一差距,而学术界的下一个问题推荐系统则依赖于新数据集缺乏的查询日志。我们描述了一种生产分析系统,该系统通过两个耦合的架构思想,将交互模型从“先有问题”转变为“先有分析师”。首先,可插拔的领域专家“技能”抽象:基于文件夹、无需数据库的主题包(清单、每阶段提示维度、关键词路由参考、报告模板以及可选计算)通过确定性模式匹配自动选择每个(客户、数据集),并作为横切关注点集成到代理管道、模式探索器和报告引擎的每个阶段,当缺失时退化为严格的无操作。由于技能是一个自包含的文件夹,可确定性解析,因此目录是开放的:这是一个可扩展的领域专家市场。其次,离线知识编译循环:代理通过 DuckDB 探查数据集的 Parquet 文件(对生产环境零负载),运行带有评论门控的每表收敛过程并自我修复重试,通过值重叠验证连接,生成持久的模式知识,这些知识驱动常设专家报告,每个发布的指标通过重新执行其证据 SQL 再次验证,还提供与报告议程相对应的建议问题。这些构成了一个主动闭环:报告呈现数据,数据生成问题,点击即可启动已验证的深度探索,所有这些发生在使用查询框之前。我们给出了正式模型并报告了单租户示例证据。我们不作用户研究或基准声明;贡献在于架构及其防御性。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28594 (HTTP 429)

Authors: Harmohit Singh, Rahul Sharma

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28594.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28594

Published: 2026-09-02T01:23:41.475Z


6. The Signal in the Noise: An Auditable Reliability Layer for Biomedical Text Classification

Abstract:Biomedical NLP pipelines routinely presuppose clean input text, yet large-scale corpora assembled through automated PDF parsing harbour pervasive OCR-like artifacts, token splits and merges, hyphenation remnants, and character-level corruption, that systematically erode lexical evidence and degrade downstream classifiers. We introduce a conservative, fully auditable spell-correction reliability layer conceived as a safety-oriented preprocessing module rather than a maximal-accuracy corrector: under conditions of uncertainty, the system abstains from editing, in accordance with a medical do-no-harm philosophy. The deterministic architecture couples bounded edit-distance candidate generation with corpus-derived n-gram scoring and a suite of biomedical safety gates that protect domain-critical terminology. We evaluate the layer both intrinsically, on a manually curated benchmark of 2,104 token-level cases, and extrinsically, on a tri-class CORD-19 topic classifier (Prevention, Treatment, Epidemiology) spanning 10,000 examples under a principled four-run protocol (Clean, Noisy, Restored, Safety). Intrinsically, the layer attains 94.61% error-fix recall on synthetic errors with zero harmful edits on negative controls. Downstream, it recovers approximately 80.45% of the noise-induced macro-F1 degradation, elevating macro-F1 from 0.7654 (Noisy) to 0.7717 (Restored) while preserving near-clean performance (Safety: 0.7721). A supplementary case study on 103 real-world OCR-extracted abstracts classified with BioBERT confirms that transformer encoders appeared relatively robust to mild noise, motivating a future grey-box architecture that integrates bounded neural signals and UMLS lexicons without compromising auditability. The system is fully deterministic, artifact-driven, and designed with deployment and auditability in mind.

中文摘要

摘要:生物医学自然语言处理(NLP)管道通常假设输入文本是干净的,但通过自动化 PDF 解析收集的大规模语料库中普遍存在类似 OCR 的伪影、词元拆分和合并、连字符残留以及字符级损坏,这些问题系统性地削弱了词汇证据并降低了下游分类器的性能。我们引入了一个保守的、完全可审计的拼写纠错可靠性层,它被设计为面向安全的预处理模块,而非追求最大准确率的纠错器:在不确定条件下,系统会选择不进行编辑,以遵循医学“不伤害”原则。该确定性架构将有界编辑距离的候选生成与基于语料库的 n-gram 打分和一套保护领域关键术语的生物医学安全门结合起来。我们从内在上在一个手工整理的包含 2,104 个词元级案例的基准上进行评估,并从外在上在一个三分类 CORD-19 主题分类器(预防、治疗、流行病学)上进行评估,该分类器涵盖 10,000 个示例,并采用合理的四次运行协议(干净、噪声、恢复、安全)。在内在评估中,该层在合成错误上的错误修复召回率达 94.61%,在负控制条件下无有害编辑。下游评估显示,它恢复了噪声引起的宏 F1 降低的约 80.45%,将宏 F1 从 0.7654(噪声)提升至 0.7717(恢复),同时保持接近干净的性能(安全:0.7721)。对 103 篇实际 OCR 提取的摘要进行的补充案例研究并使用 BioBERT 分类表明,变换器编码器对轻微噪声显得相对稳健,这促使未来开发一种灰盒架构,将有界神经信号与 UMLS 词典结合,而不影响可审计性。该系统完全确定性、以伪影驱动,并在设计时考虑了部署和可审计性。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28595 (HTTP 429)

Authors: Moustafa Yehia Hassan, Sharon Wong, Woh Kai Xuan

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28595.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28595

Published: 2026-09-02T01:23:41.475Z


7. Paper Pilot: A Human-in-the-Loop Expert System for Evidence-Traceable Scientific Manuscript Generation in Applied Sciences

Abstract:Large language model (LLM) agents are increasingly embedded in scientific workflows for literature analysis, drafting, and review. Existing systems advance autonomous discovery and manuscript generation, but do not resolve the governance problem that arises when ideas, methods, results, and claims propagate through AI-assisted workflows without mandatory human approval or artifact-level traceability. This paper proposes Paper Pilot, a human-in-the-loop expert system for evidence-traceable scientific manuscript generation in applied sciences. It adapts the Collaborative Agent Reasoning Engineering (CARE) methodology to manuscript development through manuscript-owner approval gates, explicit no-pass criteria, claim classification, audit logging, advisory LLM review, and evidence-locked revision control. The framework defines eight approval gates across the idea-to-claim pipeline and distinguishes literature-grounded from artifact-grounded claims, requiring reported numbers and interpretations to remain traceable to approved evidence; its system prompt is openly released for deployment in ChatGPT, Gemini, Claude, or institutional LLM environments. As a first empirical validation, we evaluate the citation-grounding layer with a controlled, mechanically scored benchmark (two commercial LLMs, real arXiv papers, no LLM judge): under coverage pressure ungated drafters fabricated up to 25% of their citations and never flagged an evidence gap, whereas the same models under Paper Pilot’s evidence-locked rules produced zero fabricated citations and surfaced the planted gaps as explicit placeholders. Preliminary results for result grounding, revision, and adversarial robustness point the same way; full evaluation is left to future work. Paper Pilot positions LLM-assisted writing as a controlled human-AI decision-support process rather than a fully autonomous authorship pipeline.

中文摘要

摘要:大型语言模型(LLM)代理正越来越多地嵌入科学工作流程中,用于文献分析、草稿撰写和审阅。现有系统推动了自主发现和手稿生成,但未解决在AI辅助工作流程中,当想法、方法、结果和主张在未经强制人工审批或工件级可追溯性下传播时产生的治理问题。本文提出了Paper Pilot,一种在人类参与下、面向应用科学的证据可追溯科学手稿生成的专家系统。它将协作代理推理工程(CARE)方法论应用于手稿开发,通过手稿所有者审批门、明确的不通过标准、主张分类、审计日志、顾问式LLM审阅以及证据锁定的修订控制来实现。该框架定义了从想法到主张的八个审批门,并区分以文献为基础的主张与以工件为基础的主张,要求报告的数据和解释必须可追溯至已批准的证据;其系统提示已公开发布,可在ChatGPT、Gemini、Claude或机构级LLM环境中部署。作为首次的实证验证,我们通过受控、机械评分的基准测试(两款商业LLM、真实arXiv论文、无LLM评审)评估了引用可追溯层面:在覆盖压力下,无门控的草稿生成模型伪造了多达25%的引用,且从未标注出证据缺口,而同样的模型在Paper Pilot的证据锁定规则下生成的引用未出现伪造,并将人为植入的缺口显式显示为占位符。关于结果可追溯性、修订和对抗稳健性的初步结果显示了相同趋势;完整评估留待未来工作完成。Paper Pilot将LLM辅助写作定位为受控的人机决策支持过程,而非完全自主的作者生成流程。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28596 (HTTP 429)

Authors: Nidhi Jha, Siddharth Chaudhary, Ajinkya Kulkarni

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28596.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28596

Published: 2026-09-02T01:23:41.475Z


8. The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

Abstract:Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack perspective, we demonstrate how structural vulnerabilities such as exposed DOM metadata and predictable option encoding allow agents to resolve attention checks through structured parsing only. From a defense perspective, we implement a mitigation strategy of DOM metadata obfuscation to remove semantic cues in text-based questions. We evaluate multiple open-source language and multimodal models to study capability and orchestration effectiveness. Based on our evaluations, we offer perspectives on how to simultaneously meet the needs of empiricists and agentic AI researchers.

中文摘要

摘要:在线调查是各种领域中基础的数据收集工具,其中注意力检查作为确保响应质量的关键手段。然而,具有代理性的人工智能(由大型语言模型(LLM)驱动的目标导向系统和/或具备工具增强能力的多模态处理单元)迅速出现,也对这些保障措施的稳健性提出了新的问题。我们研究了代理性 AI 架构完成基于网络的调查并通过标准注意力检查的能力。我们在受控调查沙箱中评估了一种能够处理多模态输入并利用工具进行网页交互的单代理架构。我们从两个角度分析这一问题。从攻击角度来看,我们展示了诸如暴露的 DOM 元数据和可预测的选项编码等结构性漏洞如何使智能体仅通过结构化解析就能完成注意力检查。从防御角度来看,我们实施了一种 DOM 元数据混淆的缓解策略,以去除基于文本问题中的语义线索。我们评估了多个开源的语言和多模态模型,以研究其能力及编排效果。基于我们的评估结果,我们提出了如何同时满足实证研究者和代理性 AI 研究者需求的观点。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28597 (HTTP 429)

Authors: Sourav Panda, Hillmer Chona, Rupak Kumar Das, Shreyash Kale, Shikha Soneji, Jonathan Dodge

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28597.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28597

Published: 2026-09-02T01:23:41.475Z


9. CDPR: Counterfactual Advantage-based Credit Assignment for Cost-Aware Sequential Medical Diagnosis

Abstract:Clinical diagnosis is a step-by-step, cost-aware process: a physician orders examinations one at a time, observes the results, and updates the diagnosis before reaching a final conclusion. Most medical language models instead treat diagnosis as a one-pass classification task and ignore the trade-off between a test’s value and its cost. We model diagnosis as a cost-aware sequential decision process and train the policy with reinforcement learning. The main difficulty is credit assignment: the only reliable signal comes once at the end of a long trajectory, so it scores a wasteful workup the same as an efficient one. We propose CDPR (Counterfactual Diagnostic Process Reward), which needs no expert labels and no learned critic. CDPR first finds the states where the policy hesitates, using the uncertainty of its action distribution, and then scores the chosen action by its advantage over the alternatives the policy itself would consider, estimated with short rollouts under a utility that balances correctness against test count, cost, and infeasible requests. A rollout cache reuses within-batch trajectories to keep the cost low. We integrate CDPR into GRPO and test it on one in-domain (MIMIC-IV) and two out-of-domain (ClinicalBench and a private hospital dataset) benchmarks. CDPR improves diagnostic accuracy while clearly reducing the number and cost of examinations.

中文摘要

摘要:临床诊断是一个逐步进行、考虑成本的过程:医生一次开一项检查,观察结果,然后在得出最终结论前更新诊断。而大多数医学语言模型则将诊断视为一次性分类任务,忽略了检查价值与成本之间的权衡。我们将诊断建模为一个考虑成本的序贯决策过程,并使用强化学习训练策略。主要难点在于信用分配:唯一可靠的信号只在长轨迹的末端出现,因此它会将浪费的检查与高效的检查评分相同。我们提出了CDPR(反事实诊断过程奖励),它不需要专家标签,也不需要训练的评价器。CDPR首先通过动作分布的不确定性找到策略犹豫的状态,然后通过比较策略本身会考虑的替代行动的优势来评分所选动作,这一优势通过在权衡正确性、检查次数、成本以及不可行请求的效用下进行短期回滚估算。回滚缓存重用批内轨迹以保持低成本。我们将CDPR整合到GRPO中,并在一个内域数据集(MIMIC-IV)和两个外域数据集(ClinicalBench和一家私人医院数据集)上进行测试。CDPR在显著提高诊断准确性的同时,也明显减少了检查的数量和成本。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28599 (HTTP 429)

Authors: Qi Peng, Yi Cai, Changmeng Zheng, Xin Wu, Jiayuan Xie, Qing Li

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28599.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28599

Published: 2026-09-02T01:23:41.475Z


10. SHAPE of Chain-of-Thought in Math Reasoning

Abstract:Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplored. We introduce \texttt{SHAPE}, a framework that analyzes Chain-of-Thought (CoT) trajectories through two lenses developed in mathematics education: (1) semantic spaces: the model’s evolving mathematical interpretations of a problem (e.g., algebraic, geometric), and (2) heuristics: the specific mathematical actions taken within those spaces (e.g., simplifying the problem, working backward). We first use \texttt{SHAPE} to analyze the reasoning patterns of various models. Our findings reveal that the mathematical heuristics employed by a model better explain final answer correctness than traditional CoT features. Furthermore, models are likely to reach correct solutions by concentrating their reasoning effort within a few semantic spaces rather than exploring many disparate ones — a pattern consistent with human behavior. Next, we utilize the \texttt{SHAPE} lens to evaluate whether post-training truly enhances mathematical proficiency. We find that reinforcement learning induces mode-seeking in heuristic usage. Lastly, we post-train LLMs by promoting diverse heuristics and demonstrate its effectiveness in improving accuracy. Overall, \texttt{SHAPE} provides a theoretically-grounded diagnostic framework for decoding LLM reasoning and offers a new path toward post-training LLMs for math reasoning. The code for our model is available at this https URL

中文摘要

摘要:大型语言模型(LLMs)在数学推理基准测试中表现出色,但其推理背后的数学意义技能仍未被充分探索。我们提出了\texttt{SHAPE},一种通过数学教育中开发的两个视角分析连锁思维(CoT)轨迹的框架:(1)语义空间:模型对问题的不断演变的数学解释(例如,代数的、几何的);(2)启发式方法:在这些空间中采取的具体数学操作(例如,简化问题,逆向推理)。我们首先使用\texttt{SHAPE}分析各类模型的推理模式。我们的研究发现,相比传统的CoT特征,模型所采用的数学启发式方法能够更好地解释最终答案的正确性。此外,模型更可能通过将推理努力集中在少数语义空间内而非探索多个不同空间来得到正确答案——这一模式与人类行为一致。接下来,我们利用\texttt{SHAPE}视角评估训练后是否真正提升了数学能力。我们发现强化学习会引导启发式使用的模式集中。最后,我们通过促进多样化启发式方法对LLMs进行后训练,并证明其在提高准确性方面的有效性。总体而言,\texttt{SHAPE}提供了一个有理论基础的诊断框架,用于解码LLM推理,并为数学推理的LLM后训练提供了一条新路径。我们的模型代码可在此https URL获取。

LLM Analysis

Q: 这篇论文试图解决什么问题? 这篇论文试图解决\*\*大型语言模型(LLM)数学推理能力的可解释性与诊断问题\*\*,具体而言,现有研究在以下方面存在不足: - \*\*表面化分析的局限\*\*:尽管LLM在数学推理基准上表现强劲,但现有对思维链(Chain-of-Thought, CoT)的分析多集中于表面特征——如轨迹长度、词汇层面的自我修正标记(如“wait”“aha”)、广义的推理阶段(如规划与验证)或拓扑结构(如图、树)——却未追踪模型在每一步推理时所依据的\*\*具体数学解释框架\*\*与\*\*策略性行动\*\*。 - \*\*缺乏数学意义层面的过程诊断\*\*:传统以最终答案正确率为导向的评估,无法揭示模型是如何\*\*框定、推进与重组\*\*一个数学问题的。例如,模型可能在代数视角与几何视角之间反复切换,或在某一解释空间内深度探索;这种组织方式与人类的数学问题解决行为密切相关,但在LLM中几乎未被研究。 为此,论文引入 \*\*SHAPE\*\*(Semantic-space and Heuristic Analysis for Problem-solving Evolution)框架,将数学教育研究中的两个核心维度——\*\*语义空间(semantic spaces)\*\*与\*\*启发式策略(heuristics)\*\*——用于解构LLM生成的CoT轨迹,以回答以下关键问题: 1. \*\*模型在推理时真正使用了哪些数学策略?\*\* 通过标注每一推理步骤的启发式标签(如简化问题、引入辅助对象、反证、回溯等),论文试图超越表面标记,捕捉模型在局部所执行的具体数学行动。 2. \*\*模型如何在不同的数学解释之间组织其推理过程?\*\* 通过追踪语义空间的转换(如从代数约束转向基于试验的计数,再回归代数框架),论文试图刻画模型是\*\*集中深入地探索少数解释框架\*\*,还是\*\*在多个分散的框架之间浅尝辄止\*\*——后者往往与错误推理相关。 3. \*\*后训练(post-training)尤其是强化学习(RL)如何改变模型的推理结构?\*\* 论文诊断了RLVR(带可验证奖励的强化学习)是否真正扩展了模型的数学策略库,抑或仅仅导致\*\*模式坍缩(mode-seeking)\*\*——即成功轨迹在策略分布上变得更为集中和狭窄。 4. \*\*能否将诊断洞察直接用于改进训练?\*\* 基于上述发现,论文进一步探索在RL过程中显式引入多样化的数学启发式信息,以验证过程级结构信号是否能提升推理性能。 综上,该论文的核心目标是:\*\*建立一个植根于数学教育理论的过程级诊断框架,以解码LLM数学推理的深层组织结构,并据此开辟一条通过调控策略多样性来改善后训练效果的新路径\*\*。 Q2: 有哪些相关研究? 该论文的相关研究可归纳为以下四个主要方向: ## 1. Chain-of-Thought(CoT)的表面形式分析 现有工作主要从外部可观测的特征对LLM推理轨迹进行刻画,尚未深入数学解释层面: - \*\*轨迹长度与效率\*\*:研究推理长度与正确率的关系,探讨过思考(overthinking)与欠思考(underthinking)现象(Wu et al., 2025b; Su et al., 2025)。 - \*\*自我修正标记\*\*:通过词汇层面的线索(如“wait”“aha”“hold on”“recheck”等)识别模型是否进行自我修正或重新评估(Guo et al., 2025)。 - \*\*认知阶段标签\*\*:将CoT划分为宏观推理阶段,如阅读、分析、规划、执行、验证、回答与监控(Gandhi et al., 2025; Li et al., 2025b;a; Marjanovic et al., 2026)。 - \*\*结构模式\*\*:将CoT建模为图或树结构,分析节点连接方式与推理拓扑(Jiang et al., 2025; Xiong et al., 2025; Zhang et al., 2025)。 - \*\*忠实性与可解释性\*\*:探讨CoT是否真实反映模型的内部计算,指出推理步骤可能包含事后编造(post-hoc)或装饰性内容(Arcuschin et al., 2025; Lanham et al., 2023; Tanneru et al., 2024; Bogdan et al., 2025)。 ## 2. 数学教育中的问题解决理论 SHAPE框架的理论根基源于数学教育研究领域对\*\*人类\*\*解题过程的长期探索: - \*\*启发式策略(Heuristics)\*\*:Pólya(1945)与Schoenfeld(1985)奠定了经典启发式分类(如简化问题、倒推、引入辅助对象、反证等)。后续研究进一步将其操作化,用于分析学生的出声思维协议(Koichu et al., 2007; Rott, 2014)。 - \*\*语义空间(Semantic Spaces)\*\*:Newell & Simon(1972)的问题解决理论以及Favier & Dorier(2024)近期的工作,将解题者的\*\*当前数学解释\*\*(所采纳的对象、目标与约束)定义为语义空间,并指出解题者在这些空间之间的转换与停留模式决定了推理质量(Favier, 2022; Posamentier & Krulik, 2008)。 ## 3. 后训练与强化学习(RLVR) 近期推理模型的性能提升主要依赖强化学习与可验证奖励(RLVR),但对其副作用的研究日益增多: - \*\*RLVR提升推理能力\*\*:通过结果级奖励信号优化数学推理,代表性工作包括DeepSeek-R1(Guo et al., 2025)、Tulu 3(Lambert et al., 2024)、Qwen3(Yang et al., 2025a)与Nemotron-Cascade(Wang et al., 2025)。 - \*\*多样性与策略坍缩\*\*:近期研究发现,RLVR可能并未真正扩展基础模型的推理 repertoire,而是导致采样效率提升或表面多样性下降,出现\*\*模式寻找(mode-seeking)\*\*与\*\*回声室(echo chamber)\*\*效应(Yue et al., 2025; Zhao et al., 2025; Wu et al., 2025a; Huang et al., 2025; Liu et al., 2025; Matsutani et al., 2025; Dang et al., 2025)。 ## 4. 基础工具、基准与训练方法 - \*\*基准数据集\*\*:MATH(Hendrycks et al., 2021)作为基础数学推理基准;MATH-Perturb(Huang et al., 2025)通过最小化文本修改构造需改变解题方法的难题变体;Omni-MATH(Gao et al., 2024)提供奥林匹克级别评测。 - \*\*训练算法\*\*:GRPO(Group Relative Policy Optimization)被用于后训练(Shao et al., 2024);基于显式规划的推理训练框架(Jiao et al., 2024)。 - \*\*分布评估指标\*\*:采用Density与Coverage(Naeem et al., 2020)量化生成分布相对于参考分布的集中程度与覆盖范围,用于诊断后训练是否导致启发式空间坍缩。 Q3: 论文如何解决这个问题? 该论文通过构建 \*\*SHAPE\*\*(Semantic-space and Heuristic Analysis for Problem-solving Evolution)框架,从理论建模、自动化标注、诊断度量与训练增强四个层面系统地解决了LLM数学推理的可解释性与改进问题。 ### 1. 理论框架:语义空间与启发式策略的双轴建模 论文将数学教育研究中成熟的两个分析维度引入LLM CoT研究: - \*\*语义空间(Semantic Spaces)\*\*:刻画模型对问题的当前数学解释,包括其采纳的对象、目标与约束。例如,同一问题可被解释为“代数约束系统”或“基于试验的计数”。 - \*\*启发式策略(Heuristics)\*\*:刻画模型在特定语义空间内执行的 purposeful mathematical actions,如引入符号表示(H3a)、探索特例(H9a)、简化问题(H5)、倒推(H10)等。 通过将一条CoT轨迹表示为\*\*语义空间序列\*\* S=(s_1,s_2,dots,s_T) 与\*\*启发式序列\*\* H=(H_1,H_2,dots,H_T) ,论文将黑箱式的长文本推理转化为可计算、可比较的结构化过程数据。 ### 2. 可扩展的自动化标注流水线 为解决人工标注无法规模化的问题,论文构建了三阶段自动流水线: 1. \*\*内容单元分割(Content-Unit Segmentation)\*\*:将CoT文本切分为具有独立策略意图的最小连贯片段,而非按句子硬切分。 2. \*\*启发式标注(Heuristic Tagging)\*\*:基于统一分类法(11类启发式家族 + 4类非启发式标记),使用经金标数据校准的标注模型(Grok-4.1-Fast / Qwen3.5-27B)为每个单元分配多标签。 3. \*\*语义空间追踪(Semantic-Space Tracking)\*\*:将包含表征转换型启发式(如H1, H2, H3, H5, H8, H11)的单元视为潜在空间切换点,通过状态机判断当前单元是维持(MAINTAIN)、进入新空间(NEW)还是返回旧空间(RETURN)。 ### 3. 过程级诊断度量 基于上述标注,论文定义了一套量化推理组织方式的指标: - \*\*空间分布\*\* q(i) 与\*\*片段分布\*\* p(k) :分别按语义空间ID与连续访问片段聚合启发式使用量。 - \*\*有效语义空间数\*\*:

N(eff)^(space) = exp(-∑(i∈I) q(i)log q(i))

  • **有效转移数**:
    N(eff)^(trans) = exp(-∑(k=1)^(K) p(k)log p(k)) - 1
  • **转移比率** rho = N(eff)^(trans) / N(eff)^(space) ,以及**启发式频率分布** u(h) 。 这些指标能够区分“在少数空间内深度探索”与“在多个空间间浅层跳跃”两种推理模式。 ### 4. 三项实证应用:从诊断到训练改进 论文利用SHAPE工具链展开了三个层面的工作,直接回应其提出的问题: #### (a)验证并揭示现有模型的推理规律 通过逻辑回归预测答案正确性,论文发现**启发式级别特征**(AUROC ≈ 0.664)显著优于传统的CoT长度、自我修正标记或宏观推理阶段标签。进一步地,语义空间指标显示:**正确轨迹倾向于在较少的语义空间内集中部署启发式资源**,而错误轨迹则表现为更高的 N(eff)^(space) 与 rho ,即在更多空间间反复横跳却无法收敛,这与人类数学解题研究中的“过度徘徊”现象一致。 #### (b)诊断后训练(RLVR)的结构性影响 - **扰动适应性**:在MATH-Perturb的硬扰动(需改变解题方法)下,模型虽在早期即调整启发式选择( D(JS)^(freq) 显著升高),但**未能相应扩展语义空间的覆盖范围**,而是陷入对已有空间的反复重访( Deltarho 显著上升),表明模型难以真正重组数学解释框架。 - **策略多样性坍缩**:通过Density与Coverage指标在启发式频率空间中的测量,论文发现RLVR后训练使成功轨迹**集中于基础模型分布的稠密核心**,Coverage下降而Density上升,证实了启发式层面上的 **mode-seeking** 效应。这意味着后训练并未显著拓展基础模型的策略库,而是收窄了有效推理区域。 #### (c)将诊断洞察直接用于增强训练 基于“启发式信息具有诊断与信号价值”的发现,论文设计了**Heuristic-Augmented Plan-GRPO(HA-Plan-GRPO)**。在标准GRPO训练中的rollout阶段,向模型显式提供11类数学启发式列表并要求其先选择启发式、再执行推理。实验表明,相比仅添加规划步骤的基线,HA-Plan-GRPO在MATH-Perturb的原始、简单与困难拆分上均取得了更高的 Avg@64 与 Pass@64,验证了将过程级数学结构信号融入RLVR可直接提升推理性能。 ### 总结 论文的解决路径可概括为:**以数学教育理论为基石,构建自动化过程分析工具,通过量化指标诊断现有模型与后训练的深层瓶颈,并最终将诊断洞察反馈至训练阶段**,从而超越了对CoT的表面化理解,为LLM数学推理的解码与改进提供了理论与实践闭环。 Q4: 论文做了哪些实验? 该论文围绕 **SHAPE** 框架的验证、诊断与应用,共开展了五组核心实验,涵盖特征预测、结构刻画、扰动分析、后训练诊断与训练增强五个层面。 —- ### 实验一:启发式特征预测答案正确性 **目的**:验证启发式级别特征是否比传统 CoT 表面特征更能预测最终答案是否正确。 - **数据集**:Omni-MATH 子集(100 题),涵盖 15 个模型生成的 CoT 轨迹。 - **方法**:采用 5 折分层交叉验证,训练逻辑回归分类器。通过内层交叉验证选择 ell_1 或 ell_2 正则化及超参数,以 AUROC 为评估指标。 - **基线特征**: - **Length**:CoT 长度; - **Length + reasoning**:CoT 长度结合推理 token 数量与占比; - **Self-revision**:基于词汇标记(如 “wait”、”aha”、”recheck” 等)的自我修正特征(数量、比例、是否出现); - **ThinkARM**:8 类宏观推理阶段(Read, Analyze, Plan, Implement, Explore, Verify, Answer, Monitor)的 token 频率。 - **SHAPE 特征**:以内容单元为粒度,计算 11 类启发式(H1–H11)及非启发式(N)的**频率分布** u(h) 。 - **结果**: | 特征集 | AUROC | |—-|—-| | Length | 0.504 ± 0.03 | | Length + reasoning | 0.503 ± 0.03 | | Self-revision | 0.618 ± 0.03 | | ThinkARM | 0.618 ± 0.02 | | SHAPE (H) | 0.653 ± 0.02 | | SHAPE (H+N) | 0.664 ± 0.02 | 启发式频率特征显著优于所有基线,表明其蕴含更强的正确性信号。 —- ### 实验二:语义空间指标的结构性分析 **目的**:利用语义空间度量揭示不同模型的推理组织模式,并对比正确与错误轨迹的结构差异。 - **数据集**:与实验一相同。 - **模型**:15 个模型,分为三组: - 开源推理模型(完整 CoT):Qwen3-32B, DeepSeek-R1, QwQ-32B, DeepSeek-R1-Distill-Qwen 系列, Phi-4-Reasoning; - 指令微调模型(无扩展推理):Qwen3-32B-NR, Gemini-2.0-Flash, Phi-4, Qwen-2.5-32B, GPT-4o; - 专有推理模型(隐藏轨迹):Gemini-2.5-Flash, GPT-o3-mini, GPT-o1-mini。 - **指标**: - 有效语义空间数:

N(eff)^(space) = exp(-∑(i∈I) q(i)log q(i))

  • 有效转移数:
    N(eff)^(trans) = exp(-∑(k=1)^(K) p(k)log p(k)) - 1
  • 转移比率: rho = N(eff)^(trans) / N(eff)^(space) - **关键发现**: - 推理模型的 N(eff)^(space) 与 rho 系统性高于指令微调模型,表明扩展推理不仅是”更长”,而是具有 qualitatively different 的空间遍历模式。 - **错误轨迹**在绝大多数模型中均表现出更高的 N(eff)^(space) 、 N(eff)^(trans) 与 rho ,说明错误推理往往伴随语义空间的分散探索与无效往返(overthinking)。 —- ### 实验三:扰动问题下的 CoT 结构适应性分析 **目的**:检验当问题文本相似但**所需解题方法改变**时,模型能否真正重组其数学解释框架,而非仅在固定框架内调整策略。 - **数据集**:MATH-Perturb 测试集(115 题),每题含三种版本: - **Original (O)**:原始题目; - **Simple (S)**:文本微调但解法不变; - **Hard (H)**:文本微调且**必须改变解法**。 - **模型**:4 个后训练模型:Qwen3-8B, Qwen3-32B, Nemotron-Cascade-8B, Olmo-3-7B-Think-RLVR。 - **指标**: - Pass@1(原始/简单/困难); - 启发式频率分布的 Jensen–Shannon 散度: D(JS)^(freq) ; - 有效语义空间数变化: Delta N(eff)^(space) ; - 转移比率变化: Delta rho 。 - **结果**: - 所有模型在 Hard 上的 Pass@1 均显著下降。 - Hard 扰动下的 D(JS)^(freq) 显著高于 Simple,且 Delta N(eff)^(space) 与 Delta rho 显著为正。 - **诊断结论**:模型能感知到 Hard 扰动并改变局部启发式选择,但**未能扩展语义空间覆盖**,反而更频繁地重访已有空间,导致适应性失败。 此外,附录 E 的辅助实验将 CoT 截断至前 5 或前 10 个内容单元,验证了上述启发式分歧在推理**早期阶段**即已出现,排除了晚期漂移的替代解释。 —- ### 实验四:后训练对启发式分布的集中效应诊断 **目的**:检验 RLVR 后训练是扩展了基础模型的启发式 repertoire,还是仅将成功轨迹集中于基础模型已存在的稠密区域。 - **数据集**:MATH-Perturb 测试集,限定为基座模型与后训练模型**均能正确解答**的问题。 - **模型对**: - Qwen3-1.7B-Base arrow Qwen3-1.7B-GRPO - Olmo-3-7B arrow Olmo-3-7B-Think-RL-Zero - Olmo-3-7B arrow Olmo-3-7B-Think-RLVR - 跨模型基线:Olmo-3-7B arrow Qwen3-1.7B-Base(无关分布对照) - **方法**:将每条成功轨迹映射为启发式频率向量 u(h) ,以**基础模型**为参考分布、**后训练模型**为目标分布,计算 Density 与 Coverage(Naeem et al., 2020, k=3 近邻,余弦距离): - **Density > 1**:目标分布集中于参考分布的稠密核心; - **Coverage**:目标分布覆盖参考分布的比例。 - **结果**: | Base | Post-trained | Density | Coverage | |—-|—-|—-|—-| | Qwen3-1.7B-Base | Qwen3-1.7B-GRPO | 1.220 | 0.871 | | Olmo-3-7B | Olmo-3-7B-Think-RL-Zero | 1.250 | 0.707 | | Olmo-3-7B | Olmo-3-7B-Think-RLVR | 1.032 | 0. Q5: 有什么可以进一步探索的点? 基于论文的局限性与所揭示的开放性问题,以下几个方向值得进一步深入探索: ### 1. 跨领域迁移与框架泛化 论文当前对 SHAPE 的验证集中于数学基准。语义空间与启发式的概念在物理推理、形式化逻辑证明、代码合成乃至科学假设生成中同样具有解释力。未来工作可将该框架迁移至这些领域,构建领域特定的启发式分类法与语义空间本体,检验 N(eff)^(space) 与 N(eff)^(trans) 是否仍是跨领域推理质量的稳健预测指标。 ### 2. 标注流水线的精度与粒度提升 现有自动化 pipeline 在稀有启发式类别(如 H7 反证、H10 倒推)上的标注一致性仍较弱。可通过以下方式改进: - 引入更多层次化的少样本示例或链式标注(chain-of-tagging)机制,降低稀有类的混淆; - 将语义空间追踪从当前的基于提示状态机(prompted state machine)升级为带有显式记忆更新的微调模型(fine-tuned transition model),以减少长程依赖中的追踪漂移; - 探索子空间(sub-space)细分,例如在“代数约束”内部进一步区分线性、多项式与模运算解释,以捕捉更微观的框架切换。 ### 3. 从诊断信号到训练目标的深度耦合 论文的 Heuristic-Augmented GRPO 仅在 rollout prompt 中注入启发式信息,属于轻量级干预。更深层的整合路径包括: - 将启发式多样性显式纳入 RLVR 的奖励函数,例如通过辅助奖励(auxiliary reward)鼓励模型在训练过程中覆盖更多启发式类别,从而缓解 mode-seeking; - 构建基于启发式标签的过程奖励模型(Process Reward Model, PRM),在每一步对数学策略的适当性进行评分,而非仅依赖最终答案正确性; - 设计“语义空间正则化”损失,惩罚低效的频繁空间往返(高 rho ),引导模型在单一空间内完成更深度的推理链。 ### 4. 内部表征与外部推理轨迹的联合分析 SHAPE 目前仅分析可观测的 CoT 文本。将其与模型内部计算对接,可验证外部语义空间转换是否对应内部表征空间的结构性跳转: - 利用表示探针(probing)或对比激活分析,检验当 CoT 标注为 NEW 或 RETURN 时,隐藏状态是否确实发生了与代数/几何视角切换相对应的几何变化; - 结合近期关于 CoT 忠实性(faithfulness)的研究,判断模型所声称的启发式行动(如“我尝试倒推”)是否真实驱动了底层的 token 生成概率,抑或仅为事后修饰。 ### 5. 针对适应性失败的“重组”训练策略 实验四揭示,面对 hard perturbation,模型倾向于在已有语义空间内循环而非建立新框架。未来的训练干预可聚焦于: - 显式教授“框架重组”元策略:在训练数据中增加需要根本性视角切换的问题对(如从组合数学切换到数论),并强制模型在 CoT 中显式声明视角转换; - 测试时解码控制:开发基于 SHAPE 指标的动态早期停止或重采样机制——当检测到 rho 超过阈值且 N(eff)^(space) 停滞时,触发模型重新初始化推理框架。 ### 6. 规模化验证与效率权衡 论文的增强训练实验基于 Qwen3-1.7B-Base。需要在更大参数规模(如 30B+)和更复杂任务(如奥林匹克级别几何、形式化定理证明)上验证: - SHAPE 指标是否依然保持对正确率的强预测性; - Heuristic-Augmented RLVR 的增益是否具有规模效应,抑或仅在小型模型上显著; - 语义空间聚焦(低 N_(eff)^(space) )与推理长度之间的最优权衡曲线,以指导高效推理模型的设计。 Q6: 总结一下论文的主要内容 该论文围绕大型语言模型(LLM)数学推理的可解释性与改进,提出了 **SHAPE**(Semantic-space and Heuristic Analysis for Problem-solving Evolution)框架,并基于此展开了一系列诊断分析与训练增强实验。以下是主要内容总结: —- ### 1. 研究动机与核心问题 现有对 Chain-of-Thought(CoT)的分析多停留于表面特征(如轨迹长度、词汇级自我修正标记、宏观推理阶段或图结构),缺乏对模型**具体数学解释框架**与**策略性行动**的追踪。论文指出,理解模型如何在不同数学视角间切换、如何部署启发式策略,是诊断推理失败与指导改进的关键。 —- ### 2. SHAPE 框架与自动化流水线 #### 2.1 理论双轴 SHAPE 建立在数学教育研究的两个核心概念之上: - **语义空间(Semantic Spaces)**:模型对问题的当前数学解释,包括其采纳的对象、目标与约束(如“代数约束”“基于试验的计数”)。 - **启发式(Heuristics)**:模型在特定空间内执行的具体数学行动(如引入符号表示、简化问题、倒推、验证等)。论文整合了 Pólya、Schoenfeld 等经典分类,形成包含 11 类启发式家族与 4 类非启发式标记的统一分类法。 #### 2.2 自动化标注流水线 为规模化分析,论文构建了三级流水线: 1. **内容单元分割**:将 CoT 切分为具有独立策略意图的连贯片段; 2. **启发式标注**:使用经金标数据校准的模型(Grok-4.1-Fast / Qwen3.5-27B)为每个单元分配多标签启发式代码; 3. **语义空间追踪**:将包含表征转换型启发式(H1, H2, H3, H5, H8, H11)的单元视为潜在切换点,通过状态机判断为 **MAINTAIN**(维持)、**NEW**(新建)或 **RETURN**(返回)。 —- ### 3. 过程级诊断度量 基于标注结果,论文定义了量化推理组织方式的指标: - **空间分布** q(i) 与**片段分布** p(k) :分别按语义空间 ID 与连续访问片段聚合启发式使用量; - **有效语义空间数**:

N(eff)^(space) = exp(-∑(i∈I) q(i)log q(i))

  • **有效转移数**:
    N(eff)^(trans) = exp(-∑(k=1)^(K) p(k)log p(k)) - 1
  • **转移比率**: rho = N(eff)^(trans) / N(eff)^(space) ,衡量空间重访强度; - **启发式频率分布**: u(h) ,刻画各类策略的整体使用占比。 —- ### 4. 实验验证与发现 #### 4.1 启发式特征预测正确性(Omni-MATH 子集,15 模型) 以启发式频率 u(h) 为特征训练逻辑回归,预测答案正确性的 AUROC 达到 0.664 ± 0.02 ,显著优于 CoT 长度( 0.504 )、自我修正标记( 0.618 )及宏观阶段标签 ThinkARM( 0.618 )。这表明**过程级数学策略信号强于表面形式特征**。 #### 4.2 语义空间结构揭示推理模式 - **推理模型 vs. 指令模型**:推理模型(如 DeepSeek-R1, QwQ-32B)具有更高的 N(eff)^(space) 与 rho ,说明其扩展推理不仅是“更长”,而是具有质上不同的空间遍历模式。 - **正确 vs. 错误轨迹**:错误轨迹在绝大多数模型中表现出显著更高的 N(eff)^(space) 、 N(eff)^(trans) 与 rho 。这意味着**错误往往源于在多个语义空间间浅层跳跃与无效往返**,而非深度聚焦。 —- ### 5. 后训练(RLVR)的诊断分析 #### 5.1 扰动适应性分析(MATH-Perturb) 对比原始题、简单扰动(解法不变)与困难扰动(需改变解法): - 困难扰动下,模型 Pass@1 显著下降; - 启发式分布的 Jensen–Shannon 散度 D(JS)^(freq) 显著升高,说明模型感知到变化并调整了局部策略; - 然而 Delta N(eff)^(space) 与 Delta rho 亦显著上升,表明模型**未能建立新的语义空间,而是在已有空间内更频繁地循环重访**,最终导致适应性失败。 #### 5.2 RLVR 的 Mode-Seeking 效应 通过 Density 与 Coverage 指标在启发式频率空间中测量: - 所有后训练模型(如 Olmo-3-7B-Think-RLVR, Qwen3-1.7B-GRPO)相对其基座模型均呈现 **Density > 1** 且 **Coverage < 1**; - 这表明 RLVR 使成功轨迹**集中于基座模型启发式分布的稠密核心**,而非拓展策略 repertoire,证实了启发式层面上的**模式坍缩(mode-seeking)**。 —- ### 6. 基于启发式的训练增强 受上述诊断启发,论文提出 **Heuristic-Augmented Plan-GRPO(HA-Plan-GRPO)**:在 GRPO 训练的 rollout 阶段,向模型显式提供 11 类数学启发式列表,要求其先选择启发式、再执行推理。在 MATH-Perturb 测试集上,HA-Plan-GRPO 相比仅添加规划步骤的基线及原始基座模型,在 Original、Simple 与 Hard 拆分上均取得了更高的 Avg@64 与 Pass@64,验证了**将过程级数学结构信号融入 RLVR 可直接提升推理性能**。 —- ### 7. 结论 SHAPE 提供了一个植根于数学教育理论的过程级诊断框架,证明: - 启发式特征是比表面形式更稳健的正确性信号; - 成功推理与语义空间聚焦(低 N(eff)^(space) 、低 rho )密切相关; - RLVR 后训练在拓展策略多样性方面存在局限,易引发 mode-seeking; - 显式引入启发式信息可有效改善后训练效果。 未来工作可将该框架扩展至数学以外的领域,并进一步探索将语义空间与启发式信号深度耦合到训练目标中。 Q7: 想要进一步了解论文 以上只是了解一篇论文的几个基本FAQ。如果你还想与Kimi进一步讨论该论文,请点击 这里 为你跳转Kimi AI网页版,并启动一个与该论文相关的新会话。

Authors: Jonghyun Song, Sangjun Song, Minjae Oh, Haesung Pyun, Sungsik Lee, Yohan Jo

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28600.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28600

Published: 2026-09-02T01:23:41.475Z


Evaluation Domain Papers

1. DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation

Abstract:Large Language Model (LLM) agents have shown promise for automating data-science workflows, yet their end-to-end performance depends critically on the agent harness that represents tasks, manages execution state, constrains output artifacts, and provides evaluation feedback. Existing data-science agents often leave this harness implicit, making results difficult to reproduce, compare, and attribute across heterogeneous tasks. We introduce DS-Lighting, a unified harness toolkit that makes harness design explicit for data-science automation. DS-Lighting decomposes the harness into four reusable layers: data, workflow, execution, and evaluation, and represents diverse agents as executable operator programs that support both predefined pipelines and adaptive search. We further integrate multiple open-source data-science benchmarks into an MLE-Bench-style task format, enabling controlled comparison under a shared task interface, sandboxed runtime, and metric protocol. Experiments across agents, harnesses, models, and ablations show that explicit harness design improves reproducibility, comparability, and reliability, while reducing avoidable system-level failures in end-to-end data-science workflows. Our code is available at this https URL

中文摘要

摘要:大型语言模型(LLM)代理在自动化数据科学工作流程方面显示出潜力,但它们的端到端性能在很大程度上取决于代理的框架,该框架负责表示任务、管理执行状态、约束输出产物并提供评估反馈。现有的数据科学代理通常将这种框架隐式化,使得结果难以在异构任务中复现、比较和归因。我们提出了DS-Lighting,一个统一的框架工具包,它使数据科学自动化的框架设计显式化。DS-Lighting将框架分解为四个可复用层:数据、工作流、执行和评估,并将各种代理表示为可执行的操作程序,支持预定义流程和自适应搜索。我们进一步将多个开源数据科学基准整合到MLE-Bench风格的任务格式中,使得在共享任务接口、沙箱运行时和指标协议下进行受控比较成为可能。跨代理、框架、模型和消融实验表明,显式的框架设计提高了可复现性、可比性和可靠性,同时减少了端到端数据科学工作流程中可避免的系统级故障。我们的代码可通过此https URL获取。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28590 (HTTP 429)

Authors: Fan Liu, Hao Liu

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28590.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28590

Published: 2026-09-02T01:24:00.958Z


2. Expert-validated STEM QA

Abstract:Recent advancements in AI are helping scientists achieve breakthroughs in fields such as mathematics, medicine, and materials sciences. New evaluation datasets for AI models contribute to such advancement in AI. In the STEM domain, frontier models have consumed most of the available online data, creating the need for human-created datasets that codify the knowledge of leading experts in the domain. There are several STEM datasets available for the research community in this field. However, there are some gaps in these datasets, leaving room for improvement. Examples of gaps include (1) saturation in model performance on these datasets, leaving no head-room for meaningful evaluations, (2) skewed taxonomy distributions, (3) multiple choice question format that is misaligned with how scientists use AI in the real world, and (4) inaccurate answers and rationales partially led by a contest-based data collection and a time-bound review process. In this study, we present ‘Expert-validated STEM QA’, a high-quality, expert-validated STEM dataset (N=398) in Physics, Chemistry, Biology, and Mathematics, created by 241 domain experts. We (1) carefully designed a taxonomy with balanced distribution, (2) vetted question contributors with quality-driven incentive, (3) conducted multiple rounds of reviews with revisions validated by domain experts based on consensus, and (4) created the dataset in verifiable question and answer format. Our study demonstrated low performance ($<25\%$) of frontier AI models on the dataset as a benchmark. Post-training on a separate, private version of the dataset (N=2,000) increased performance of the open source model by $15\%$ relative to the baseline model (p=0.045) on the STEM subset of HLE-verified dataset, indicating potential utility of the dataset for model training. We have open-sourced a portion of our dataset for the AI research community.

中文摘要

摘要:近期人工智能的进展正在帮助科学家在数学、医学和材料科学等领域取得突破。用于人工智能模型的新评估数据集也推动了人工智能的发展。在STEM领域,前沿模型已经消耗了大部分可用的在线数据,因此需要由人类创建的数据集,将该领域顶尖专家的知识进行编码。该领域有若干可供研究社区使用的STEM数据集。然而,这些数据集中存在一些空白,需要改进。空白的例子包括:(1) 这些数据集上的模型性能饱和,缺乏进行有意义评估的空间;(2) 分类分布偏斜;(3) 多选题格式与科学家在实际使用人工智能的方式不符;(4) 部分答案和理由不准确,部分原因是基于竞赛的数据收集方式和有时间限制的审核过程。在本研究中,我们提出了“专家验证STEM问答”,这是一个高质量、经过专家验证的STEM数据集(N=398),涵盖物理、化学、生物和数学领域,由241位领域专家创建。我们(1) 精心设计了分类法并实现了均衡分布;(2) 对问题贡献者进行了质量驱动的审核;(3) 进行了多轮审查,并由领域专家基于共识验证修订;(4) 创建了可验证的问答格式的数据集。我们的研究表明,前沿AI模型在该数据集上的表现较低(<25%),可作为基准。对该数据集的独立私人版本(N=2,000)进行后续训练,使开源模型在HLE验证数据集的STEM子集上的性能相较基线模型提升了15%(p=0.045),表明该数据集在模型训练中具有潜在用途。我们已将部分数据集开源,供人工智能研究社区使用。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28591 (HTTP 429)

Authors: Kihwan Han, Saurabh Patil, Chinmayee Shukla, Abhinav Sharma, Marko Pavlovic, Anshuman Lall, Mahesh Joshi

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28591.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28591

Published: 2026-09-02T01:24:00.958Z


3. A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making

Abstract:Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test—it is a sequence of guideline-pathway choices, escalation judgments, and commitments under uncertainty. Existing benchmarks largely measure factual recall, leaving open whether frontier LLMs share decision-path blind spots that combining models cannot fix. We built the Oncology Decision Boundary Benchmark (ODBB)—2,005 oncology decision points across NCCN guidelines and colorectal cancer cases—and evaluated nine frontier LLMs (four closed-source, five open-weight families) released between June 2025 and April 2026. A fully deterministic scorer (zero LLM inference) classified outputs into 14 failure types, independently validated by two oncologists (Cohen’s weighted $\kappa$ = 0.939 and 0.790) on a 225-item stratified sample. Treating the nine as a pooled super-model, 42.1% (Wilson 95% CI 40.0—44.3%) of all items—35.7% of the 1,586 NCCN items and 66.4% of the 419 colorectal-cancer cases—were answered correctly by none, with failures concentrated in choosing between guideline pathways before reasoning within any: a consistent blind spot in clinical meta-judgment that likely requires architectural intervention rather than more training data. Two models tuned for decisiveness (GPT-5.5, Gemini 3.1 Pro Preview) made unsafe commitments three to five times more often than the seven cautious models without scoring higher. In 3—9% of items, models stated the correct next clinical step yet did not commit to it—failures of decision, not knowledge. Model quality is no longer the primary bottleneck for clinical LLM deployment; the binding constraint is the assumption that any single model can be the sole basis for a clinical decision. Progress requires architectures that detect when a model reaches its competence boundary and route the decision to a clinician.

中文摘要

摘要:大型语言模型(LLMs)在医学知识考试中取得了高分,但现实世界的肿瘤学并不是知识考试——它是指南路径选择、升级判断以及不确定性下承诺的连续过程。现有基准主要衡量事实召回能力,尚未解决最前沿的LLM是否存在结合模型也无法弥补的决策路径盲点问题。我们构建了肿瘤学决策边界基准(ODBB)——涵盖NCCN指南和结直肠癌病例的2,005个肿瘤学决策点——并评估了九个最前沿的LLM(四个闭源、五个开放权重系列),发布时间为2025年6月至2026年4月。一个完全确定性的评分器(零LLM推理)将输出分类为14种失败类型,并由两名肿瘤学家在225项分层样本上独立验证(Cohen加权κ值分别为0.939和0.790)。将九个模型视为一个汇总的超模型,42.1%(Wilson 95% CI 40.0—44.3%)的所有条目——其中1,586个NCCN条目的35.7%和419个结直肠癌病例的66.4%——没有任何模型答对,失败集中在指南路径的选择阶段,而不是在任何路径内进行推理:这是临床元判断中的一致盲点,可能需要架构干预而非更多训练数据才能解决。两个为果断性调优的模型(GPT-5.5、Gemini 3.1 Pro Preview)在做出不安全承诺的频率上比七个谨慎模型高出三到五倍,但没有取得更高分。在3%—9%的条目中,模型陈述了正确的下一临床步骤,但未做出承诺——这是决策失败,而非知识不足。模型质量不再是临床LLM部署的主要瓶颈;关键约束是假设任何单一模型都可以作为临床决策的唯一依据。进展需要架构能够检测模型何时达到其能力边界,并将决策引导至临床医生。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28592 (HTTP 429)

Authors: Zhang Sheng, Jinming Li, Wangyang Chen, Zhiwei Bao, Yu YoSean Wang

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28592.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28592

Published: 2026-09-02T01:24:00.958Z


Abstract:With the increasing development of AI regulatory frameworks, ensuring that artificial intelligence systems, particularly generative models, operate in accordance with legal and ethical standards has become a critical priority. Existing proposals for AI alignment and value-guided behavior, however, face some limitations. Approaches such as Constitutional AI depend on human supervision, while broad normative frameworks like the Good-for-Humanity (GfH) principle may be overly general and ambiguous to provide actionable governance guidance. To overcome these limitations, we propose a hybrid approach called Statutory AI that employs pre-existing human-authored principles drawn from specific themes within a legal corpus. Specifically, Statutory AI uses legal texts as a constitutional framework, enabling AI systems to autonomously critique and revise their outputs according to established norms. It operates in two stages, both using Chain-of-Thought prompting. The first stage classifies the user prompt into one of the identified themes, while the second stage analyzes it in conjunction with relevant articles selected from the legal corpus of that theme. To illustrate the potential of our approach, we conducted an experiment involving 1,000 red-teaming prompts and five penal themes: discrimination, disclosure of confidential information, violence, fraud, and abuse of vulnerable persons. Statutory AI reduced harmful content by 52 to 59 percentage points across tested models, approximately 10 percentage points higher than standard Constitutional AI, while cutting computation time by over 50%.

中文摘要

摘要:随着人工智能监管框架的不断发展,确保人工智能系统,尤其是生成模型,按照法律和伦理标准运行已成为一个关键优先事项。然而,现有的人工智能对齐和价值引导行为的提案存在一些局限性。诸如宪法人工智能(Constitutional AI)的方法依赖于人工监督,而广泛的规范性框架如‘利于人类原则’(Good-for-Humanity, GfH)可能过于笼统和模糊,无法提供可操作的治理指导。为了克服这些局限性,我们提出了一种名为法定人工智能(Statutory AI)的混合方法,该方法利用来自法律语料库中特定主题的现有人工制定原则。具体而言,法定人工智能使用法律文本作为宪法框架,使人工智能系统能够根据既定规范自主批评和修改其输出。它分两个阶段运行,都采用链式思维(Chain-of-Thought)提示法。第一阶段将用户提示分类到已识别的主题之一,第二阶段结合从该主题的法律语料库中选出的相关条款进行分析。为了展示我们方法的潜力,我们进行了一个涉及1,000个红队测试提示和五个惩戒主题的实验:歧视、机密信息披露、暴力、欺诈和对弱势群体的虐待。法定人工智能在测试模型中减少有害内容的比例为52到59个百分点,约比标准宪法人工智能高出10个百分点,同时计算时间减少超过50%。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28593 (HTTP 429)

Authors: Cindy Delage, Stéphane Canu, Marc Décombas, Jonathan Foureur

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28593.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28593

Published: 2026-09-02T01:24:00.958Z


5. From Question-First to Analyst-First: Domain-Expert Skills and Verified Knowledge Compilation for Proactive Enterprise Analytics

Abstract:Conversational analytics systems assume the user already has a well-formed question, leaving a non-expert facing a blank query box on an unfamiliar enterprise schema. Commercial ‘proactive’ tools narrow this gap only by detecting statistical anomalies over analyst-curated metric layers, and academic next-question recommenders depend on query logs that a fresh dataset lacks. We describe a production analytics system that inverts the interaction model from question-first to analyst-first through two coupled architectural ideas. First, a pluggable domain-expert ‘skill’ abstraction: a folder-based, database-free subject-matter pack (a manifest, per-stage prompt facets, keyword-routed references, report templates, and optional compute) auto-selected per (client, dataset) by deterministic schema matching and spliced as a cross-cutting concern into every stage of an agentic pipeline, the schema explorer, and the report engines, degrading to a strict no-op when absent. Because a skill is a self-contained folder resolved deterministically, the catalogue is open-ended: an extensible marketplace of domain experts. Second, an offline knowledge-compilation loop: an agent probes the dataset’s parquet via DuckDB (zero load on production), runs critic-gated per-table convergence with self-healing retries, and data-validates joins by value overlap, producing durable schema knowledge that drives standing expert reports whose every published metric is re-verified by re-executing its evidence SQL, plus suggested questions that mirror the report agenda. These close a proactive loop: reports surface numbers, the numbers seed questions, and a click launches a verified deep dive, all before the query box is used. We give a formal model and report illustrative single-tenant evidence. We make no user-study or benchmark claims; the contribution is the architecture and its defensibility.

中文摘要

摘要:会话分析系统假设用户已经有了一个完整的问题,这使得非专业用户在不熟悉的企业架构面前面对一个空白的查询框。商业“主动”工具只通过检测分析师策划的指标层中的统计异常来缩小这一差距,而学术界的下一个问题推荐系统则依赖于新数据集缺乏的查询日志。我们描述了一个生产分析系统,该系统通过两个耦合的架构思想,将交互模型从“先有问题”倒置为“先有分析师”。首先,可插拔的领域专家“技能”抽象:基于文件夹、无需数据库的主题包(清单、每阶段提示维度、关键词路由参考、报告模板以及可选计算)通过确定性模式匹配自动选择每个(客户、数据集),并作为横切关注点集成到代理管道、模式探索器和报告引擎的每个阶段,当缺失时退化为严格的无操作。由于技能是一个自包含文件夹,可确定性解析,因此目录是开放的:一个可扩展的领域专家市场。其次,离线知识编译循环:代理通过DuckDB探查数据集的parquet(对生产零负载),运行带有评论门控的每表收敛并进行自愈重试,通过值重叠验证连接,生成持久的架构知识,从而驱动固定的专家报告,报告中每个发布的指标都通过重新执行其证据SQL重新验证,同时生成与报告议程匹配的建议问题。这些闭环实现主动性:报告呈现数字,数字引发问题,点击即可启动经过验证的深入分析,所有这些发生在查询框被使用之前。我们提供了正式模型和示例单租户证据。我们不做用户研究或基准声称;贡献在于架构及其可辩护性。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28594 (HTTP 429)

Authors: Harmohit Singh, Rahul Sharma

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28594.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28594

Published: 2026-09-02T01:24:00.958Z


6. The Signal in the Noise: An Auditable Reliability Layer for Biomedical Text Classification

Abstract:Biomedical NLP pipelines routinely presuppose clean input text, yet large-scale corpora assembled through automated PDF parsing harbour pervasive OCR-like artifacts, token splits and merges, hyphenation remnants, and character-level corruption, that systematically erode lexical evidence and degrade downstream classifiers. We introduce a conservative, fully auditable spell-correction reliability layer conceived as a safety-oriented preprocessing module rather than a maximal-accuracy corrector: under conditions of uncertainty, the system abstains from editing, in accordance with a medical do-no-harm philosophy. The deterministic architecture couples bounded edit-distance candidate generation with corpus-derived n-gram scoring and a suite of biomedical safety gates that protect domain-critical terminology. We evaluate the layer both intrinsically, on a manually curated benchmark of 2,104 token-level cases, and extrinsically, on a tri-class CORD-19 topic classifier (Prevention, Treatment, Epidemiology) spanning 10,000 examples under a principled four-run protocol (Clean, Noisy, Restored, Safety). Intrinsically, the layer attains 94.61% error-fix recall on synthetic errors with zero harmful edits on negative controls. Downstream, it recovers approximately 80.45% of the noise-induced macro-F1 degradation, elevating macro-F1 from 0.7654 (Noisy) to 0.7717 (Restored) while preserving near-clean performance (Safety: 0.7721). A supplementary case study on 103 real-world OCR-extracted abstracts classified with BioBERT confirms that transformer encoders appeared relatively robust to mild noise, motivating a future grey-box architecture that integrates bounded neural signals and UMLS lexicons without compromising auditability. The system is fully deterministic, artifact-driven, and designed with deployment and auditability in mind.

中文摘要

摘要:生物医学自然语言处理(NLP)管道通常假设输入文本是干净的,但通过自动化 PDF 解析收集的大规模语料库中普遍存在类似 OCR 的伪影、词元拆分和合并、连字符残留以及字符级损坏,这些问题系统性地削弱了词汇证据并降低了下游分类器的性能。我们引入了一个保守的、完全可审计的拼写纠错可靠性层,其设计理念是作为面向安全的预处理模块,而非追求最大准确率的纠错器:在不确定条件下,系统会选择不进行编辑,遵循医学上的“不造成伤害”原则。该确定性架构将有界编辑距离候选生成与语料库衍生的 n-gram 评分以及一套保护领域关键术语的生物医学安全门相结合。我们对该层进行了内部评估和外部评估:内部评估在一个手工整理的 2,104 个词元级案例基准上进行;外部评估在一个三类 CORD-19 话题分类器(预防、治疗、流行病学)上,涵盖 10,000 个示例,并使用系统化的四轮协议(干净、噪声、恢复、安全)进行。内部评估显示,该层在合成错误上的错误修复召回率达到 94.61%,且在负面对照中无有害编辑。下游评估中,它恢复了大约 80.45% 的噪声引起的宏 F1 降幅,使宏 F1 从 0.7654(噪声)提升至 0.7717(恢复),同时保持接近干净的性能(安全:0.7721)。对 103 篇真实世界 OCR 提取摘要进行的补充案例研究并使用 BioBERT 分类显示,Transformer 编码器对轻微噪声表现出相对鲁棒性,这支持未来设计一种灰盒架构,将有限的神经信号与 UMLS 词典集成,同时不影响可审计性。该系统完全确定性、基于伪影,并在设计时考虑了部署和可审计性。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28595 (HTTP 429)

Authors: Moustafa Yehia Hassan, Sharon Wong, Woh Kai Xuan

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28595.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28595

Published: 2026-09-02T01:24:00.958Z


7. Paper Pilot: A Human-in-the-Loop Expert System for Evidence-Traceable Scientific Manuscript Generation in Applied Sciences

Abstract:Large language model (LLM) agents are increasingly embedded in scientific workflows for literature analysis, drafting, and review. Existing systems advance autonomous discovery and manuscript generation, but do not resolve the governance problem that arises when ideas, methods, results, and claims propagate through AI-assisted workflows without mandatory human approval or artifact-level traceability. This paper proposes Paper Pilot, a human-in-the-loop expert system for evidence-traceable scientific manuscript generation in applied sciences. It adapts the Collaborative Agent Reasoning Engineering (CARE) methodology to manuscript development through manuscript-owner approval gates, explicit no-pass criteria, claim classification, audit logging, advisory LLM review, and evidence-locked revision control. The framework defines eight approval gates across the idea-to-claim pipeline and distinguishes literature-grounded from artifact-grounded claims, requiring reported numbers and interpretations to remain traceable to approved evidence; its system prompt is openly released for deployment in ChatGPT, Gemini, Claude, or institutional LLM environments. As a first empirical validation, we evaluate the citation-grounding layer with a controlled, mechanically scored benchmark (two commercial LLMs, real arXiv papers, no LLM judge): under coverage pressure ungated drafters fabricated up to 25% of their citations and never flagged an evidence gap, whereas the same models under Paper Pilot’s evidence-locked rules produced zero fabricated citations and surfaced the planted gaps as explicit placeholders. Preliminary results for result grounding, revision, and adversarial robustness point the same way; full evaluation is left to future work. Paper Pilot positions LLM-assisted writing as a controlled human-AI decision-support process rather than a fully autonomous authorship pipeline.

中文摘要

摘要:大型语言模型(LLM)代理正越来越多地嵌入科学工作流程中,用于文献分析、草稿撰写和审阅。现有系统推动了自主发现和手稿生成,但未解决在AI辅助工作流程中,当想法、方法、结果和主张未经强制的人类审批或工件级可追溯性传播时所产生的治理问题。本文提出了Paper Pilot,一种在人类参与下、面向应用科学的证据可追溯科学手稿生成的专家系统。它将协作代理推理工程(CARE)方法论应用于手稿开发,通过手稿所有者审批门、明确的不通过标准、主张分类、审计日志、顾问式LLM审阅以及证据锁定的修订控制来实现。该框架定义了从想法到主张的八个审批门,并区分以文献为基础的主张与以工件为基础的主张,要求报告的数据和解释必须可追溯至已批准的证据;其系统提示已公开发布,可在ChatGPT、Gemini、Claude或机构级LLM环境中部署。作为首次的实证验证,我们通过受控、机械评分的基准测试(两款商业LLM、真实arXiv论文、无LLM评审)评估了引用可追溯层面:在覆盖压力下,未设审批门的草稿生成器伪造了多达25%的引用且从未标记证据缺口,而同样的模型在Paper Pilot的证据锁定规则下生成的引用未出现伪造,并将植入的缺口作为显式占位符提示。关于结果可追溯、修订和对抗性稳健性的初步结果也显示了相同趋势;完整评估留待未来工作完成。Paper Pilot将LLM辅助写作定位为受控的人机决策支持过程,而非完全自主的作者生成流程。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28596 (HTTP 429)

Authors: Nidhi Jha, Siddharth Chaudhary, Ajinkya Kulkarni

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28596.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28596

Published: 2026-09-02T01:24:00.958Z


8. The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

Abstract:Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack perspective, we demonstrate how structural vulnerabilities such as exposed DOM metadata and predictable option encoding allow agents to resolve attention checks through structured parsing only. From a defense perspective, we implement a mitigation strategy of DOM metadata obfuscation to remove semantic cues in text-based questions. We evaluate multiple open-source language and multimodal models to study capability and orchestration effectiveness. Based on our evaluations, we offer perspectives on how to simultaneously meet the needs of empiricists and agentic AI researchers.

中文摘要

摘要:在线调查是各种领域中基础的数据收集工具,其中注意力检查作为确保响应质量的关键手段。然而,具有代理性的人工智能(由大型语言模型(LLM)驱动的目标导向系统和/或具备工具增强能力的多模态处理单元)迅速出现,也对这些保障措施的稳健性提出了新的问题。我们研究了代理性人工智能架构完成基于网络的调查并通过标准注意力检查的能力。我们在受控调查沙箱中评估了一种能够处理多模态输入并基于工具进行网络交互的单代理架构。我们从两个角度分析该问题。从攻击角度,我们展示了诸如暴露的 DOM 元数据和可预测选项编码等结构性漏洞如何使代理仅通过结构化解析就能解决注意力检查问题。从防御角度,我们实现了 DOM 元数据混淆的缓解策略,以去除基于文本问题的语义线索。我们评估了多种开源语言模型和多模态模型,以研究其能力和协调效果。基于我们的评估结果,我们提供了关于如何同时满足经验主义者和代理性人工智能研究人员需求的观点。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28597 (HTTP 429)

Authors: Sourav Panda, Hillmer Chona, Rupak Kumar Das, Shreyash Kale, Shikha Soneji, Jonathan Dodge

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28597.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28597

Published: 2026-09-02T01:24:00.958Z


9. CDPR: Counterfactual Advantage-based Credit Assignment for Cost-Aware Sequential Medical Diagnosis

Abstract:Clinical diagnosis is a step-by-step, cost-aware process: a physician orders examinations one at a time, observes the results, and updates the diagnosis before reaching a final conclusion. Most medical language models instead treat diagnosis as a one-pass classification task and ignore the trade-off between a test’s value and its cost. We model diagnosis as a cost-aware sequential decision process and train the policy with reinforcement learning. The main difficulty is credit assignment: the only reliable signal comes once at the end of a long trajectory, so it scores a wasteful workup the same as an efficient one. We propose CDPR (Counterfactual Diagnostic Process Reward), which needs no expert labels and no learned critic. CDPR first finds the states where the policy hesitates, using the uncertainty of its action distribution, and then scores the chosen action by its advantage over the alternatives the policy itself would consider, estimated with short rollouts under a utility that balances correctness against test count, cost, and infeasible requests. A rollout cache reuses within-batch trajectories to keep the cost low. We integrate CDPR into GRPO and test it on one in-domain (MIMIC-IV) and two out-of-domain (ClinicalBench and a private hospital dataset) benchmarks. CDPR improves diagnostic accuracy while clearly reducing the number and cost of examinations.

中文摘要

摘要:临床诊断是一个逐步进行、考虑成本的过程:医生一次开一项检查,观察结果,然后在得出最终结论前更新诊断。而大多数医学语言模型则将诊断视为一次性分类任务,忽略了检查价值与成本之间的权衡。我们将诊断建模为一个考虑成本的序贯决策过程,并使用强化学习训练策略。主要难点在于信用分配:唯一可靠的信号只在长轨迹的末端出现,因此它会将浪费的检查与高效的检查评分相同。我们提出了CDPR(反事实诊断过程奖励),它不需要专家标签,也不需要训练的评价器。CDPR首先通过动作分布的不确定性找到策略犹豫的状态,然后通过比较策略本身会考虑的替代行动的优势来评分所选动作,这一优势通过在权衡正确性、检查次数、成本以及不可行请求的效用下进行短期回滚估算。回滚缓存重用批内轨迹以保持低成本。我们将CDPR整合到GRPO中,并在一个内域数据集(MIMIC-IV)和两个外域数据集(ClinicalBench和一家私人医院数据集)上进行测试。CDPR在显著提高诊断准确性的同时,也明显减少了检查的数量和成本。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28599 (HTTP 429)

Authors: Qi Peng, Yi Cai, Changmeng Zheng, Xin Wu, Jiayuan Xie, Qing Li

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28599.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28599

Published: 2026-09-02T01:24:00.958Z


10. SHAPE of Chain-of-Thought in Math Reasoning

Abstract:Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplored. We introduce \texttt{SHAPE}, a framework that analyzes Chain-of-Thought (CoT) trajectories through two lenses developed in mathematics education: (1) semantic spaces: the model’s evolving mathematical interpretations of a problem (e.g., algebraic, geometric), and (2) heuristics: the specific mathematical actions taken within those spaces (e.g., simplifying the problem, working backward). We first use \texttt{SHAPE} to analyze the reasoning patterns of various models. Our findings reveal that the mathematical heuristics employed by a model better explain final answer correctness than traditional CoT features. Furthermore, models are likely to reach correct solutions by concentrating their reasoning effort within a few semantic spaces rather than exploring many disparate ones — a pattern consistent with human behavior. Next, we utilize the \texttt{SHAPE} lens to evaluate whether post-training truly enhances mathematical proficiency. We find that reinforcement learning induces mode-seeking in heuristic usage. Lastly, we post-train LLMs by promoting diverse heuristics and demonstrate its effectiveness in improving accuracy. Overall, \texttt{SHAPE} provides a theoretically-grounded diagnostic framework for decoding LLM reasoning and offers a new path toward post-training LLMs for math reasoning. The code for our model is available at this https URL

中文摘要

摘要:大型语言模型(LLMs)在数学推理基准测试中表现出色,但其推理背后的数学意义技能仍未被充分探索。我们提出了\texttt{SHAPE},一个通过数学教育中开发的两个视角分析连锁思维(CoT)轨迹的框架:(1)语义空间:模型对问题的不断演变的数学解释(例如,代数的、几何的);(2)启发式方法:在这些空间中采取的具体数学操作(例如,简化问题,逆向推理)。我们首先使用\texttt{SHAPE}分析各类模型的推理模式。我们的研究发现,相比传统的CoT特征,模型所采用的数学启发式方法能够更好地解释最终答案的正确性。此外,模型更可能通过将推理努力集中在少数语义空间内而非探索多个不同空间来得到正确答案——这一模式与人类行为一致。接下来,我们利用\texttt{SHAPE}视角评估训练后是否真正提升了数学能力。我们发现强化学习会引导启发式使用的模式集中。最后,我们通过促进多样化启发式方法对LLMs进行后训练,并证明其在提高准确性方面的有效性。总体而言,\texttt{SHAPE}提供了一个有理论基础的诊断框架,用于解码LLM的推理过程,并为数学推理的LLM后训练提供了一条新途径。我们的模型代码可在此https网址获取。

LLM Analysis

Q: 这篇论文试图解决什么问题? 这篇论文试图解决\*\*大型语言模型(LLM)数学推理能力的可解释性与诊断问题\*\*,具体而言,现有研究在以下方面存在不足: - \*\*表面化分析的局限\*\*:尽管LLM在数学推理基准上表现强劲,但现有对思维链(Chain-of-Thought, CoT)的分析多集中于表面特征——如轨迹长度、词汇层面的自我修正标记(如“wait”“aha”)、广义的推理阶段(如规划与验证)或拓扑结构(如图、树)——却未追踪模型在每一步推理时所依据的\*\*具体数学解释框架\*\*与\*\*策略性行动\*\*。 - \*\*缺乏数学意义层面的过程诊断\*\*:传统以最终答案正确率为导向的评估,无法揭示模型是如何\*\*框定、推进与重组\*\*一个数学问题的。例如,模型可能在代数视角与几何视角之间反复切换,或在某一解释空间内深度探索;这种组织方式与人类的数学问题解决行为密切相关,但在LLM中几乎未被研究。 为此,论文引入 \*\*SHAPE\*\*(Semantic-space and Heuristic Analysis for Problem-solving Evolution)框架,将数学教育研究中的两个核心维度——\*\*语义空间(semantic spaces)\*\*与\*\*启发式策略(heuristics)\*\*——用于解构LLM生成的CoT轨迹,以回答以下关键问题: 1. \*\*模型在推理时真正使用了哪些数学策略?\*\* 通过标注每一推理步骤的启发式标签(如简化问题、引入辅助对象、反证、回溯等),论文试图超越表面标记,捕捉模型在局部所执行的具体数学行动。 2. \*\*模型如何在不同的数学解释之间组织其推理过程?\*\* 通过追踪语义空间的转换(如从代数约束转向基于试验的计数,再回归代数框架),论文试图刻画模型是\*\*集中深入地探索少数解释框架\*\*,还是\*\*在多个分散的框架之间浅尝辄止\*\*——后者往往与错误推理相关。 3. \*\*后训练(post-training)尤其是强化学习(RL)如何改变模型的推理结构?\*\* 论文诊断了RLVR(带可验证奖励的强化学习)是否真正扩展了模型的数学策略库,抑或仅仅导致\*\*模式坍缩(mode-seeking)\*\*——即成功轨迹在策略分布上变得更为集中和狭窄。 4. \*\*能否将诊断洞察直接用于改进训练?\*\* 基于上述发现,论文进一步探索在RL过程中显式引入多样化的数学启发式信息,以验证过程级结构信号是否能提升推理性能。 综上,该论文的核心目标是:\*\*建立一个植根于数学教育理论的过程级诊断框架,以解码LLM数学推理的深层组织结构,并据此开辟一条通过调控策略多样性来改善后训练效果的新路径\*\*。 Q2: 有哪些相关研究? 该论文的相关研究可归纳为以下四个主要方向: ## 1. Chain-of-Thought(CoT)的表面形式分析 现有工作主要从外部可观测的特征对LLM推理轨迹进行刻画,尚未深入数学解释层面: - \*\*轨迹长度与效率\*\*:研究推理长度与正确率的关系,探讨过思考(overthinking)与欠思考(underthinking)现象(Wu et al., 2025b; Su et al., 2025)。 - \*\*自我修正标记\*\*:通过词汇层面的线索(如“wait”“aha”“hold on”“recheck”等)识别模型是否进行自我修正或重新评估(Guo et al., 2025)。 - \*\*认知阶段标签\*\*:将CoT划分为宏观推理阶段,如阅读、分析、规划、执行、验证、回答与监控(Gandhi et al., 2025; Li et al., 2025b;a; Marjanovic et al., 2026)。 - \*\*结构模式\*\*:将CoT建模为图或树结构,分析节点连接方式与推理拓扑(Jiang et al., 2025; Xiong et al., 2025; Zhang et al., 2025)。 - \*\*忠实性与可解释性\*\*:探讨CoT是否真实反映模型的内部计算,指出推理步骤可能包含事后编造(post-hoc)或装饰性内容(Arcuschin et al., 2025; Lanham et al., 2023; Tanneru et al., 2024; Bogdan et al., 2025)。 ## 2. 数学教育中的问题解决理论 SHAPE框架的理论根基源于数学教育研究领域对\*\*人类\*\*解题过程的长期探索: - \*\*启发式策略(Heuristics)\*\*:Pólya(1945)与Schoenfeld(1985)奠定了经典启发式分类(如简化问题、倒推、引入辅助对象、反证等)。后续研究进一步将其操作化,用于分析学生的出声思维协议(Koichu et al., 2007; Rott, 2014)。 - \*\*语义空间(Semantic Spaces)\*\*:Newell & Simon(1972)的问题解决理论以及Favier & Dorier(2024)近期的工作,将解题者的\*\*当前数学解释\*\*(所采纳的对象、目标与约束)定义为语义空间,并指出解题者在这些空间之间的转换与停留模式决定了推理质量(Favier, 2022; Posamentier & Krulik, 2008)。 ## 3. 后训练与强化学习(RLVR) 近期推理模型的性能提升主要依赖强化学习与可验证奖励(RLVR),但对其副作用的研究日益增多: - \*\*RLVR提升推理能力\*\*:通过结果级奖励信号优化数学推理,代表性工作包括DeepSeek-R1(Guo et al., 2025)、Tulu 3(Lambert et al., 2024)、Qwen3(Yang et al., 2025a)与Nemotron-Cascade(Wang et al., 2025)。 - \*\*多样性与策略坍缩\*\*:近期研究发现,RLVR可能并未真正扩展基础模型的推理 repertoire,而是导致采样效率提升或表面多样性下降,出现\*\*模式寻找(mode-seeking)\*\*与\*\*回声室(echo chamber)\*\*效应(Yue et al., 2025; Zhao et al., 2025; Wu et al., 2025a; Huang et al., 2025; Liu et al., 2025; Matsutani et al., 2025; Dang et al., 2025)。 ## 4. 基础工具、基准与训练方法 - \*\*基准数据集\*\*:MATH(Hendrycks et al., 2021)作为基础数学推理基准;MATH-Perturb(Huang et al., 2025)通过最小化文本修改构造需改变解题方法的难题变体;Omni-MATH(Gao et al., 2024)提供奥林匹克级别评测。 - \*\*训练算法\*\*:GRPO(Group Relative Policy Optimization)被用于后训练(Shao et al., 2024);基于显式规划的推理训练框架(Jiao et al., 2024)。 - \*\*分布评估指标\*\*:采用Density与Coverage(Naeem et al., 2020)量化生成分布相对于参考分布的集中程度与覆盖范围,用于诊断后训练是否导致启发式空间坍缩。 Q3: 论文如何解决这个问题? 该论文通过构建 \*\*SHAPE\*\*(Semantic-space and Heuristic Analysis for Problem-solving Evolution)框架,从理论建模、自动化标注、诊断度量与训练增强四个层面系统地解决了LLM数学推理的可解释性与改进问题。 ### 1. 理论框架:语义空间与启发式策略的双轴建模 论文将数学教育研究中成熟的两个分析维度引入LLM CoT研究: - \*\*语义空间(Semantic Spaces)\*\*:刻画模型对问题的当前数学解释,包括其采纳的对象、目标与约束。例如,同一问题可被解释为“代数约束系统”或“基于试验的计数”。 - \*\*启发式策略(Heuristics)\*\*:刻画模型在特定语义空间内执行的 purposeful mathematical actions,如引入符号表示(H3a)、探索特例(H9a)、简化问题(H5)、倒推(H10)等。 通过将一条CoT轨迹表示为\*\*语义空间序列\*\* S=(s_1,s_2,dots,s_T) 与\*\*启发式序列\*\* H=(H_1,H_2,dots,H_T) ,论文将黑箱式的长文本推理转化为可计算、可比较的结构化过程数据。 ### 2. 可扩展的自动化标注流水线 为解决人工标注无法规模化的问题,论文构建了三阶段自动流水线: 1. \*\*内容单元分割(Content-Unit Segmentation)\*\*:将CoT文本切分为具有独立策略意图的最小连贯片段,而非按句子硬切分。 2. \*\*启发式标注(Heuristic Tagging)\*\*:基于统一分类法(11类启发式家族 + 4类非启发式标记),使用经金标数据校准的标注模型(Grok-4.1-Fast / Qwen3.5-27B)为每个单元分配多标签。 3. \*\*语义空间追踪(Semantic-Space Tracking)\*\*:将包含表征转换型启发式(如H1, H2, H3, H5, H8, H11)的单元视为潜在空间切换点,通过状态机判断当前单元是维持(MAINTAIN)、进入新空间(NEW)还是返回旧空间(RETURN)。 ### 3. 过程级诊断度量 基于上述标注,论文定义了一套量化推理组织方式的指标: - \*\*空间分布\*\* q(i) 与\*\*片段分布\*\* p(k) :分别按语义空间ID与连续访问片段聚合启发式使用量。 - \*\*有效语义空间数\*\*:

N(eff)^(space) = exp(-∑(i∈I) q(i)log q(i))

  • **有效转移数**:
    N(eff)^(trans) = exp(-∑(k=1)^(K) p(k)log p(k)) - 1
  • **转移比率** rho = N(eff)^(trans) / N(eff)^(space) ,以及**启发式频率分布** u(h) 。 这些指标能够区分“在少数空间内深度探索”与“在多个空间间浅层跳跃”两种推理模式。 ### 4. 三项实证应用:从诊断到训练改进 论文利用SHAPE工具链展开了三个层面的工作,直接回应其提出的问题: #### (a)验证并揭示现有模型的推理规律 通过逻辑回归预测答案正确性,论文发现**启发式级别特征**(AUROC ≈ 0.664)显著优于传统的CoT长度、自我修正标记或宏观推理阶段标签。进一步地,语义空间指标显示:**正确轨迹倾向于在较少的语义空间内集中部署启发式资源**,而错误轨迹则表现为更高的 N(eff)^(space) 与 rho ,即在更多空间间反复横跳却无法收敛,这与人类数学解题研究中的“过度徘徊”现象一致。 #### (b)诊断后训练(RLVR)的结构性影响 - **扰动适应性**:在MATH-Perturb的硬扰动(需改变解题方法)下,模型虽在早期即调整启发式选择( D(JS)^(freq) 显著升高),但**未能相应扩展语义空间的覆盖范围**,而是陷入对已有空间的反复重访( Deltarho 显著上升),表明模型难以真正重组数学解释框架。 - **策略多样性坍缩**:通过Density与Coverage指标在启发式频率空间中的测量,论文发现RLVR后训练使成功轨迹**集中于基础模型分布的稠密核心**,Coverage下降而Density上升,证实了启发式层面上的 **mode-seeking** 效应。这意味着后训练并未显著拓展基础模型的策略库,而是收窄了有效推理区域。 #### (c)将诊断洞察直接用于增强训练 基于“启发式信息具有诊断与信号价值”的发现,论文设计了**Heuristic-Augmented Plan-GRPO(HA-Plan-GRPO)**。在标准GRPO训练中的rollout阶段,向模型显式提供11类数学启发式列表并要求其先选择启发式、再执行推理。实验表明,相比仅添加规划步骤的基线,HA-Plan-GRPO在MATH-Perturb的原始、简单与困难拆分上均取得了更高的 Avg@64 与 Pass@64,验证了将过程级数学结构信号融入RLVR可直接提升推理性能。 ### 总结 论文的解决路径可概括为:**以数学教育理论为基石,构建自动化过程分析工具,通过量化指标诊断现有模型与后训练的深层瓶颈,并最终将诊断洞察反馈至训练阶段**,从而超越了对CoT的表面化理解,为LLM数学推理的解码与改进提供了理论与实践闭环。 Q4: 论文做了哪些实验? 该论文围绕 **SHAPE** 框架的验证、诊断与应用,共开展了五组核心实验,涵盖特征预测、结构刻画、扰动分析、后训练诊断与训练增强五个层面。 —- ### 实验一:启发式特征预测答案正确性 **目的**:验证启发式级别特征是否比传统 CoT 表面特征更能预测最终答案是否正确。 - **数据集**:Omni-MATH 子集(100 题),涵盖 15 个模型生成的 CoT 轨迹。 - **方法**:采用 5 折分层交叉验证,训练逻辑回归分类器。通过内层交叉验证选择 ell_1 或 ell_2 正则化及超参数,以 AUROC 为评估指标。 - **基线特征**: - **Length**:CoT 长度; - **Length + reasoning**:CoT 长度结合推理 token 数量与占比; - **Self-revision**:基于词汇标记(如 “wait”、”aha”、”recheck” 等)的自我修正特征(数量、比例、是否出现); - **ThinkARM**:8 类宏观推理阶段(Read, Analyze, Plan, Implement, Explore, Verify, Answer, Monitor)的 token 频率。 - **SHAPE 特征**:以内容单元为粒度,计算 11 类启发式(H1–H11)及非启发式(N)的**频率分布** u(h) 。 - **结果**: | 特征集 | AUROC | |—-|—-| | Length | 0.504 ± 0.03 | | Length + reasoning | 0.503 ± 0.03 | | Self-revision | 0.618 ± 0.03 | | ThinkARM | 0.618 ± 0.02 | | SHAPE (H) | 0.653 ± 0.02 | | SHAPE (H+N) | 0.664 ± 0.02 | 启发式频率特征显著优于所有基线,表明其蕴含更强的正确性信号。 —- ### 实验二:语义空间指标的结构性分析 **目的**:利用语义空间度量揭示不同模型的推理组织模式,并对比正确与错误轨迹的结构差异。 - **数据集**:与实验一相同。 - **模型**:15 个模型,分为三组: - 开源推理模型(完整 CoT):Qwen3-32B, DeepSeek-R1, QwQ-32B, DeepSeek-R1-Distill-Qwen 系列, Phi-4-Reasoning; - 指令微调模型(无扩展推理):Qwen3-32B-NR, Gemini-2.0-Flash, Phi-4, Qwen-2.5-32B, GPT-4o; - 专有推理模型(隐藏轨迹):Gemini-2.5-Flash, GPT-o3-mini, GPT-o1-mini。 - **指标**: - 有效语义空间数:

N(eff)^(space) = exp(-∑(i∈I) q(i)log q(i))

  • 有效转移数:
    N(eff)^(trans) = exp(-∑(k=1)^(K) p(k)log p(k)) - 1
  • 转移比率: rho = N(eff)^(trans) / N(eff)^(space) - **关键发现**: - 推理模型的 N(eff)^(space) 与 rho 系统性高于指令微调模型,表明扩展推理不仅是”更长”,而是具有 qualitatively different 的空间遍历模式。 - **错误轨迹**在绝大多数模型中均表现出更高的 N(eff)^(space) 、 N(eff)^(trans) 与 rho ,说明错误推理往往伴随语义空间的分散探索与无效往返(overthinking)。 —- ### 实验三:扰动问题下的 CoT 结构适应性分析 **目的**:检验当问题文本相似但**所需解题方法改变**时,模型能否真正重组其数学解释框架,而非仅在固定框架内调整策略。 - **数据集**:MATH-Perturb 测试集(115 题),每题含三种版本: - **Original (O)**:原始题目; - **Simple (S)**:文本微调但解法不变; - **Hard (H)**:文本微调且**必须改变解法**。 - **模型**:4 个后训练模型:Qwen3-8B, Qwen3-32B, Nemotron-Cascade-8B, Olmo-3-7B-Think-RLVR。 - **指标**: - Pass@1(原始/简单/困难); - 启发式频率分布的 Jensen–Shannon 散度: D(JS)^(freq) ; - 有效语义空间数变化: Delta N(eff)^(space) ; - 转移比率变化: Delta rho 。 - **结果**: - 所有模型在 Hard 上的 Pass@1 均显著下降。 - Hard 扰动下的 D(JS)^(freq) 显著高于 Simple,且 Delta N(eff)^(space) 与 Delta rho 显著为正。 - **诊断结论**:模型能感知到 Hard 扰动并改变局部启发式选择,但**未能扩展语义空间覆盖**,反而更频繁地重访已有空间,导致适应性失败。 此外,附录 E 的辅助实验将 CoT 截断至前 5 或前 10 个内容单元,验证了上述启发式分歧在推理**早期阶段**即已出现,排除了晚期漂移的替代解释。 —- ### 实验四:后训练对启发式分布的集中效应诊断 **目的**:检验 RLVR 后训练是扩展了基础模型的启发式 repertoire,还是仅将成功轨迹集中于基础模型已存在的稠密区域。 - **数据集**:MATH-Perturb 测试集,限定为基座模型与后训练模型**均能正确解答**的问题。 - **模型对**: - Qwen3-1.7B-Base arrow Qwen3-1.7B-GRPO - Olmo-3-7B arrow Olmo-3-7B-Think-RL-Zero - Olmo-3-7B arrow Olmo-3-7B-Think-RLVR - 跨模型基线:Olmo-3-7B arrow Qwen3-1.7B-Base(无关分布对照) - **方法**:将每条成功轨迹映射为启发式频率向量 u(h) ,以**基础模型**为参考分布、**后训练模型**为目标分布,计算 Density 与 Coverage(Naeem et al., 2020, k=3 近邻,余弦距离): - **Density > 1**:目标分布集中于参考分布的稠密核心; - **Coverage**:目标分布覆盖参考分布的比例。 - **结果**: | Base | Post-trained | Density | Coverage | |—-|—-|—-|—-| | Qwen3-1.7B-Base | Qwen3-1.7B-GRPO | 1.220 | 0.871 | | Olmo-3-7B | Olmo-3-7B-Think-RL-Zero | 1.250 | 0.707 | | Olmo-3-7B | Olmo-3-7B-Think-RLVR | 1.032 | 0. Q5: 有什么可以进一步探索的点? 基于论文的局限性与所揭示的开放性问题,以下几个方向值得进一步深入探索: ### 1. 跨领域迁移与框架泛化 论文当前对 SHAPE 的验证集中于数学基准。语义空间与启发式的概念在物理推理、形式化逻辑证明、代码合成乃至科学假设生成中同样具有解释力。未来工作可将该框架迁移至这些领域,构建领域特定的启发式分类法与语义空间本体,检验 N(eff)^(space) 与 N(eff)^(trans) 是否仍是跨领域推理质量的稳健预测指标。 ### 2. 标注流水线的精度与粒度提升 现有自动化 pipeline 在稀有启发式类别(如 H7 反证、H10 倒推)上的标注一致性仍较弱。可通过以下方式改进: - 引入更多层次化的少样本示例或链式标注(chain-of-tagging)机制,降低稀有类的混淆; - 将语义空间追踪从当前的基于提示状态机(prompted state machine)升级为带有显式记忆更新的微调模型(fine-tuned transition model),以减少长程依赖中的追踪漂移; - 探索子空间(sub-space)细分,例如在“代数约束”内部进一步区分线性、多项式与模运算解释,以捕捉更微观的框架切换。 ### 3. 从诊断信号到训练目标的深度耦合 论文的 Heuristic-Augmented GRPO 仅在 rollout prompt 中注入启发式信息,属于轻量级干预。更深层的整合路径包括: - 将启发式多样性显式纳入 RLVR 的奖励函数,例如通过辅助奖励(auxiliary reward)鼓励模型在训练过程中覆盖更多启发式类别,从而缓解 mode-seeking; - 构建基于启发式标签的过程奖励模型(Process Reward Model, PRM),在每一步对数学策略的适当性进行评分,而非仅依赖最终答案正确性; - 设计“语义空间正则化”损失,惩罚低效的频繁空间往返(高 rho ),引导模型在单一空间内完成更深度的推理链。 ### 4. 内部表征与外部推理轨迹的联合分析 SHAPE 目前仅分析可观测的 CoT 文本。将其与模型内部计算对接,可验证外部语义空间转换是否对应内部表征空间的结构性跳转: - 利用表示探针(probing)或对比激活分析,检验当 CoT 标注为 NEW 或 RETURN 时,隐藏状态是否确实发生了与代数/几何视角切换相对应的几何变化; - 结合近期关于 CoT 忠实性(faithfulness)的研究,判断模型所声称的启发式行动(如“我尝试倒推”)是否真实驱动了底层的 token 生成概率,抑或仅为事后修饰。 ### 5. 针对适应性失败的“重组”训练策略 实验四揭示,面对 hard perturbation,模型倾向于在已有语义空间内循环而非建立新框架。未来的训练干预可聚焦于: - 显式教授“框架重组”元策略:在训练数据中增加需要根本性视角切换的问题对(如从组合数学切换到数论),并强制模型在 CoT 中显式声明视角转换; - 测试时解码控制:开发基于 SHAPE 指标的动态早期停止或重采样机制——当检测到 rho 超过阈值且 N(eff)^(space) 停滞时,触发模型重新初始化推理框架。 ### 6. 规模化验证与效率权衡 论文的增强训练实验基于 Qwen3-1.7B-Base。需要在更大参数规模(如 30B+)和更复杂任务(如奥林匹克级别几何、形式化定理证明)上验证: - SHAPE 指标是否依然保持对正确率的强预测性; - Heuristic-Augmented RLVR 的增益是否具有规模效应,抑或仅在小型模型上显著; - 语义空间聚焦(低 N_(eff)^(space) )与推理长度之间的最优权衡曲线,以指导高效推理模型的设计。 Q6: 总结一下论文的主要内容 该论文围绕大型语言模型(LLM)数学推理的可解释性与改进,提出了 **SHAPE**(Semantic-space and Heuristic Analysis for Problem-solving Evolution)框架,并基于此展开了一系列诊断分析与训练增强实验。以下是主要内容总结: —- ### 1. 研究动机与核心问题 现有对 Chain-of-Thought(CoT)的分析多停留于表面特征(如轨迹长度、词汇级自我修正标记、宏观推理阶段或图结构),缺乏对模型**具体数学解释框架**与**策略性行动**的追踪。论文指出,理解模型如何在不同数学视角间切换、如何部署启发式策略,是诊断推理失败与指导改进的关键。 —- ### 2. SHAPE 框架与自动化流水线 #### 2.1 理论双轴 SHAPE 建立在数学教育研究的两个核心概念之上: - **语义空间(Semantic Spaces)**:模型对问题的当前数学解释,包括其采纳的对象、目标与约束(如“代数约束”“基于试验的计数”)。 - **启发式(Heuristics)**:模型在特定空间内执行的具体数学行动(如引入符号表示、简化问题、倒推、验证等)。论文整合了 Pólya、Schoenfeld 等经典分类,形成包含 11 类启发式家族与 4 类非启发式标记的统一分类法。 #### 2.2 自动化标注流水线 为规模化分析,论文构建了三级流水线: 1. **内容单元分割**:将 CoT 切分为具有独立策略意图的连贯片段; 2. **启发式标注**:使用经金标数据校准的模型(Grok-4.1-Fast / Qwen3.5-27B)为每个单元分配多标签启发式代码; 3. **语义空间追踪**:将包含表征转换型启发式(H1, H2, H3, H5, H8, H11)的单元视为潜在切换点,通过状态机判断为 **MAINTAIN**(维持)、**NEW**(新建)或 **RETURN**(返回)。 —- ### 3. 过程级诊断度量 基于标注结果,论文定义了量化推理组织方式的指标: - **空间分布** q(i) 与**片段分布** p(k) :分别按语义空间 ID 与连续访问片段聚合启发式使用量; - **有效语义空间数**:

N(eff)^(space) = exp(-∑(i∈I) q(i)log q(i))

  • **有效转移数**:
    N(eff)^(trans) = exp(-∑(k=1)^(K) p(k)log p(k)) - 1
  • **转移比率**: rho = N(eff)^(trans) / N(eff)^(space) ,衡量空间重访强度; - **启发式频率分布**: u(h) ,刻画各类策略的整体使用占比。 —- ### 4. 实验验证与发现 #### 4.1 启发式特征预测正确性(Omni-MATH 子集,15 模型) 以启发式频率 u(h) 为特征训练逻辑回归,预测答案正确性的 AUROC 达到 0.664 ± 0.02 ,显著优于 CoT 长度( 0.504 )、自我修正标记( 0.618 )及宏观阶段标签 ThinkARM( 0.618 )。这表明**过程级数学策略信号强于表面形式特征**。 #### 4.2 语义空间结构揭示推理模式 - **推理模型 vs. 指令模型**:推理模型(如 DeepSeek-R1, QwQ-32B)具有更高的 N(eff)^(space) 与 rho ,说明其扩展推理不仅是“更长”,而是具有质上不同的空间遍历模式。 - **正确 vs. 错误轨迹**:错误轨迹在绝大多数模型中表现出显著更高的 N(eff)^(space) 、 N(eff)^(trans) 与 rho 。这意味着**错误往往源于在多个语义空间间浅层跳跃与无效往返**,而非深度聚焦。 —- ### 5. 后训练(RLVR)的诊断分析 #### 5.1 扰动适应性分析(MATH-Perturb) 对比原始题、简单扰动(解法不变)与困难扰动(需改变解法): - 困难扰动下,模型 Pass@1 显著下降; - 启发式分布的 Jensen–Shannon 散度 D(JS)^(freq) 显著升高,说明模型感知到变化并调整了局部策略; - 然而 Delta N(eff)^(space) 与 Delta rho 亦显著上升,表明模型**未能建立新的语义空间,而是在已有空间内更频繁地循环重访**,最终导致适应性失败。 #### 5.2 RLVR 的 Mode-Seeking 效应 通过 Density 与 Coverage 指标在启发式频率空间中测量: - 所有后训练模型(如 Olmo-3-7B-Think-RLVR, Qwen3-1.7B-GRPO)相对其基座模型均呈现 **Density > 1** 且 **Coverage < 1**; - 这表明 RLVR 使成功轨迹**集中于基座模型启发式分布的稠密核心**,而非拓展策略 repertoire,证实了启发式层面上的**模式坍缩(mode-seeking)**。 —- ### 6. 基于启发式的训练增强 受上述诊断启发,论文提出 **Heuristic-Augmented Plan-GRPO(HA-Plan-GRPO)**:在 GRPO 训练的 rollout 阶段,向模型显式提供 11 类数学启发式列表,要求其先选择启发式、再执行推理。在 MATH-Perturb 测试集上,HA-Plan-GRPO 相比仅添加规划步骤的基线及原始基座模型,在 Original、Simple 与 Hard 拆分上均取得了更高的 Avg@64 与 Pass@64,验证了**将过程级数学结构信号融入 RLVR 可直接提升推理性能**。 —- ### 7. 结论 SHAPE 提供了一个植根于数学教育理论的过程级诊断框架,证明: - 启发式特征是比表面形式更稳健的正确性信号; - 成功推理与语义空间聚焦(低 N(eff)^(space) 、低 rho )密切相关; - RLVR 后训练在拓展策略多样性方面存在局限,易引发 mode-seeking; - 显式引入启发式信息可有效改善后训练效果。 未来工作可将该框架扩展至数学以外的领域,并进一步探索将语义空间与启发式信号深度耦合到训练目标中。 Q7: 想要进一步了解论文 以上只是了解一篇论文的几个基本FAQ。如果你还想与Kimi进一步讨论该论文,请点击 这里 为你跳转Kimi AI网页版,并启动一个与该论文相关的新会话。

Authors: Jonghyun Song, Sangjun Song, Minjae Oh, Haesung Pyun, Sungsik Lee, Yohan Jo

Categories: cs.AI

PDF URL: https://arxiv.org/pdf/2608.28600.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28600

Published: 2026-09-02T01:24:00.958Z


VLM Domain Papers

1. Open-Set Cattle Muzzle Identification: A Leakage-Controlled Benchmark and Evaluation Protocol

Abstract:Reliable individual cattle identification supports disease surveillance, vaccination records, breeding management, and livestock insurance. Although the bovine muzzle provides a stable, non-contact biometric, existing muzzle-recognition systems largely assume a closed set of enrolled animals, limiting their practical deployment. We reformulate cattle muzzle biometrics as an open-set, gallery-based identification problem that can reject previously unseen animals and support incremental enrollment without model retraining. We introduce a leakage-controlled evaluation protocol based on identity-disjoint splits, per-fold retraining, held-out threshold calibration, verified duplicate removal, and bootstrap confidence intervals. We evaluate the framework using two contrasting embedding configurations: a hybrid CNN-ViT metric-learning model and the MegaDescriptor-L foundation model. Under oracle threshold selection, the hybrid model achieves detection-and-identification rates of 98.3%, 96.4%, and 93.6% at target false-acceptance rates of 10^(-1), 10^(-2), and 10^(-3), respectively, while MegaDescriptor-L achieves 99.3%, 98.1%, and 96.1%. However, deployable threshold calibration reveals a substantial difference between oracle and calibrated performance: the hybrid model achieves a false-acceptance rate of 1.03% at a 1% target, whereas MegaDescriptor-L reaches 2.44%. Incremental enrollment further achieves Rank-1 accuracy above 91% with a single reference image and up to 97.3% with eight reference images, without retraining the model or degrading the existing gallery. These results demonstrate that threshold calibration, leakage control, and embedding quality are critical for reliable open-set cattle identification and provide a practical evaluation framework for deployment-oriented animal biometric systems.

中文摘要

摘要:可靠的个体牛只识别支持疾病监测、疫苗接种记录、繁殖管理和畜牧保险。尽管牛嘴提供了一种稳定的非接触式生物特征,但现有的牛嘴识别系统大多假设动物集合是封闭的,这限制了其实际部署。我们将牛嘴生物特征识别重新定义为一个开放集、基于图库的识别问题,该问题可以拒绝先前未见过的动物,并支持增量登记而无需重新训练模型。我们引入了一种基于身份不相交划分、按折数重新训练、保留阈值校准、经过验证的重复移除以及自助法置信区间的泄漏控制评估协议。我们使用两种对比鲜明的嵌入配置对该框架进行了评估:混合 CNN-ViT 度量学习模型和 MegaDescriptor-L 基础模型。在理想阈值选择下,混合模型在目标误接受率为 10^(-1)、10^(-2) 和 10^(-3) 时,检测与识别率分别达到 98.3%、96.4% 和 93.6%,而 MegaDescriptor-L 分别达到 99.3%、98.1% 和 96.1%。然而,可部署的阈值校准显示理想性能与校准性能之间存在显著差异:混合模型在 1% 目标下的误接受率为 1.03%,而 MegaDescriptor-L 为 2.44%。增量登记进一步实现了单参考图像时 Rank-1 准确率超过 91%,使用八张参考图像时可高达 97.3%,且无需重新训练模型或降低现有图库性能。这些结果表明,阈值校准、泄漏控制和嵌入质量对于可靠的开放集牛只识别至关重要,并为面向部署的动物生物特征系统提供了实用的评估框架。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28663 (HTTP 429)

Authors: Lalit BC, Dharmendra Singh Chaudhary, Shovit Nepal

Categories: cs.CV

PDF URL: https://arxiv.org/pdf/2608.28663.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28663

Published: 2026-09-02T01:24:22.326Z


2. Understanding Temporal Semantic Stability in Open-Vocabulary UAV Perception through Metric 3D Fusion

Abstract:Recent open-vocabulary segmentation models have advanced semantic perception for UAVs, but predictions from moving aerial platforms can remain temporally inconsistent across repeated observations of the same physical scene. We investigate temporal semantic stability by associating frame-wise predictions with persistent world-space locations through metric 3D fusion. We introduce a voxel-level evaluation framework that jointly characterises final semantic agreement, Semantic Belief Drift (SBD), Observation Persistence (OP), and semantic uncertainty. Experiments on UAVid-3D reveal substantial frame-wise semantic flicker and show that high aggregate world-space agreement can overstate temporal stability when locations have limited repeated-observation support. Persistence-stratified analysis shows that recurrent voxels expose greater semantic disagreement, while belief drift decreases as additional evidence accumulates. This behaviour is observed across two segmentation backbones and remains consistent under variations in voxel resolution, geometric association, and temporal sampling density. Conditions that reduce world-space recurrence can increase apparent aggregate stability, demonstrating that semantic consistency must be interpreted together with observation support. Our findings highlight observation persistence as an essential conditioning variable for evaluating long-horizon semantic reliability.

中文摘要

摘要:近年来的开放词汇分割模型提升了无人机的语义感知能力,但来自移动空中平台的预测在对同一物理场景的重复观察中可能仍会出现时间不一致性。我们通过通过度量的3D融合将逐帧预测与持久世界空间位置关联,来研究时间语义稳定性。我们引入了一个体素级评估框架,该框架联合刻画最终语义一致性、语义信念漂移(SBD)、观察持久性(OP)和语义不确定性。在UAVid-3D上的实验显示了显著的逐帧语义闪烁,并表明当位置的重复观察支持有限时,高聚合的世界空间一致性可能会夸大时间稳定性。基于持久性的分析显示,重复出现的体素揭示了更大的语义分歧,而随着额外证据的累积,信念漂移会降低。这一行为在两个分割骨干网络中均被观察到,并在体素分辨率、几何关联和时间采样密度变化下保持一致。减少世界空间重复的条件可能会增加表观聚合稳定性,表明语义一致性必须与观察支持一起解读。我们的研究结果强调了观察持久性作为评估长时间语义可靠性的关键条件变量。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28665 (HTTP 429)

Authors: Saurbh Singh Jamwal

Categories: cs.CV

PDF URL: https://arxiv.org/pdf/2608.28665.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28665

Published: 2026-09-02T01:24:22.326Z


3. Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting

Abstract:Video-language models (VLMs) remain brittle on tasks that require tracking events over time and grounding answers in specific spatial regions. We propose that part of this limitation can be addressed through better organization of visual evidence at inference time. We introduce structured video prompting, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding and without altering the question prompt in the main comparison. We evaluate this approach on two complementary video benchmarks and two open video-language models. Across these settings, structured inputs improve performance in several cases, with gains varying by model and task. Our findings suggest that some failures of VLMs arise not only from reasoning capacity, but also from how video evidence is presented at inference time. These results highlight structured video prompting as a simple and practical direction for improving video understanding.

中文摘要

摘要:视频语言模型(VLMs)在需要跟踪时间事件和在特定空间区域中定位答案的任务上仍然表现脆弱。我们提出,这一部分限制可以通过在推理时更好地组织视觉证据来解决。我们引入了结构化视频提示,这是一种无需训练的推理时方法,通过轻量级的空间结构和时间结构增强输入视频,为在空间和时间上组织证据提供明确的锚点,而无需改变模型权重或解码,也不改变主要对比中的问题提示。我们在两个互补的视频基准和两个开放的视频语言模型上评估了该方法。在这些设置中,结构化输入在多个情况下提高了性能,增益因模型和任务而异。我们的发现表明,VLMs 的一些失败不仅源于推理能力,还源于视频证据在推理时的呈现方式。这些结果凸显了结构化视频提示作为提升视频理解的一个简单且实用的方向。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28666 (HTTP 429)

Authors: Sadegh Mohammadian

Categories: cs.CV

PDF URL: https://arxiv.org/pdf/2608.28666.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28666

Published: 2026-09-02T01:24:22.326Z


4. MIRAGE-CAD: Construction-Mediated Multimodal Generation of Executable CAD Programs

Abstract:Recovering an executable parametric CAD program from an observed object is fundamentally ambiguous, because the same final geometry can result from different construction procedures. We study this problem from four types of input: natural-language descriptions, rendered images, point clouds, and STEP/B-Rep geometry. MIRAGE-CAD maps each input to a shared construction representation and mediates program generation through an explicit construction-plan interface. The resulting Python CAD code is executed by an OpenCASCADE kernel to build the solid and export it as STEP. On 2,500 held-out queries per modality, the system achieves 55.4-70.0% build success and 52.3-66.2% STEP export success without retrieval at inference. Controlled comparisons show that strong reconstruction does not depend on expressing the construction representation as text: a decoder conditioned directly on the continuous representation also reconstructs strongly, while an exposure-matched plan-based decoder shows no detected material loss in per-part geometric fidelity. The explicit plan instead provides a readable and separately measurable intermediate representation whose agreement with the reference construction is informative about downstream execution success. Finally, we show that executable validity, geometric fidelity, and parametric responsiveness can diverge substantially and should therefore be evaluated separately.

中文摘要

摘要:从观察到的对象中恢复可执行的参数化 CAD 程序在本质上是模糊的,因为相同的最终几何形状可能源自不同的构建过程。我们从四种输入类型研究该问题:自然语言描述、渲染图像、点云和 STEP/B-Rep 几何。MIRAGE-CAD 将每种输入映射到共享的构建表示,并通过显式的构建计划接口来调节程序生成。生成的 Python CAD 代码由 OpenCASCADE 内核执行,用于构建实物并导出为 STEP 文件。在每种模态下的 2,500 个保留查询上,该系统在不进行推理时检索的情况下实现了 55.4%-70.0% 的构建成功率和 52.3%-66.2% 的 STEP 导出成功率。控制对比表明,强重建不依赖于将构建表示表达为文本:直接以连续表示作为条件的解码器也能实现强重建,而曝光匹配的基于计划的解码器在每个零件的几何精度上未显示出材料损失。显式计划提供了一个可读且可单独测量的中间表示,其与参考构建的一致性能够提供关于后续执行成功的重要信息。最后,我们展示了可执行有效性、几何精度和参数响应性可能存在显著差异,因此应分别进行评估。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28669 (HTTP 429)

Authors: Jizong Zhan

Categories: cs.CV

PDF URL: https://arxiv.org/pdf/2608.28669.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28669

Published: 2026-09-02T01:24:22.326Z


5. Memory-Efficient Training-Free Acceleration of Diffusion Transformers with BaryCache

Abstract:Diffusion Transformers achieve high-fidelity image and video generation, but their iterative sampling remains expensive, for each denoising step requires large matrix operations. Existing cache-based acceleration reduces redundant computation yet increases the VRAM footprint by storing intermediate states, which can directly constrain inference batch size. In this work, we propose a training-free acceleration method that performs stepwise forecasting for DiT sampling using a Barycentric Extrapolator. By leveraging barycentric extrapolation, our predictor is numerically stable and alleviates oscillatory artifacts analogous to the Runge phenomenon during forward forecasting. Across extensive experiments on both image and video generation, our approach provides a favorable trade-off between memory usage and perceptual quality, while delivering up to 3.30x end-to-end sampling speedup compared with baseline DiT inference.

中文摘要

摘要:扩散变换器(Diffusion Transformers)能够实现高保真图像和视频生成,但其迭代采样依旧代价高昂,因为每个去噪步骤都需要进行大规模矩阵运算。现有的基于缓存的加速方法虽然减少了冗余计算,但通过存储中间状态增加了显存占用,这会直接限制推理批量大小。在本工作中,我们提出了一种无需训练的加速方法,使用重心外推器(Barycentric Extrapolator)对 DiT 采样进行逐步预测。通过利用重心外推,我们的预测器在数值上稳定,并缓解了类似龙格现象(Runge phenomenon)在前向预测过程中出现的振荡伪影。在针对图像和视频生成的大量实验中,我们的方法在内存使用和感知质量之间提供了良好的权衡,同时相比基线 DiT 推理实现了高达 3.30 倍的端到端采样加速。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28670 (HTTP 429)

Authors: Chengjie Lu, Tianchi Deng, Zhengqi He, Zhijian Gao, Huisi Wu, Xueliang Li

Categories: cs.CV

PDF URL: https://arxiv.org/pdf/2608.28670.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28670

Published: 2026-09-02T01:24:22.326Z


6. Measuring Similarity between Artistic and AI Generated Images using Siamese Neural Networks

Abstract:AI-generated art has sparked debates around potential plagiarism, as these images may closely resemble existing artworks. This research quantifies the similarity between original pieces and AI-generated counterparts, particularly those produced by the Stable Diffusion XL Refiner 1.0. We use Siamese Networks with frozen CLIP encoders and cosine similarity optimized through triplet loss. A dataset of paired original and generated images was built using image-to-image generation and custom prompts, enriched with semantic descriptors and BLIP-2 captions. Prior studies report up to 81\% style replication and 90\% visual similarity. Our results show high discriminative performance: training accuracy reached 99.9\%, and the best model configuration achieved 99.4\% test accuracy with strong inter-class separation ($\delta \mu$ = 0.677), demonstrating the effectiveness of our semantic-visual embeddings.

中文摘要

摘要:AI生成的艺术作品引发了关于潜在抄袭的讨论,因为这些图像可能与现有艺术作品极为相似。本研究量化了原创作品与AI生成作品之间的相似性,特别是由Stable Diffusion XL Refiner 1.0生成的作品。我们使用Siamese Networks,结合冻结的CLIP编码器和通过三重态损失优化余弦相似度。通过图像对图像生成和自定义提示构建了原创与生成图像配对数据集,并丰富了语义描述符和BLIP-2字幕。先前研究报告显示样式复制率高达81%,视觉相似度高达90%。我们的结果显示判别性能很高:训练准确率达到99.9%,最佳模型配置在强类间分离($\delta \mu$ = 0.677)下达到99.4%测试准确率,展示了我们语义-视觉嵌入的有效性。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28671 (HTTP 429)

Authors: Diego Castro Elvira, Navil Pineda Rugerio, Jesús García-Ramírez, Cecilia Reyes-Peña, Ricardo Ramos-Aguilar

Categories: cs.CV

PDF URL: https://arxiv.org/pdf/2608.28671.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28671

Published: 2026-09-02T01:24:22.326Z


7. FrameScope: Temporal Data Valuation for Stream Active Learning in Autonomous Vehicle Systems

Abstract:Autonomous vehicles operate in dynamic, ever-changing environments where new scenarios and edge cases constantly emerge. As a result, static learning models are inadequate for ensuring safe and reliable operation. Continuous learning is essential for adapting to these evolving conditions and maintaining robust performance across diverse real-world settings. However, autonomous vehicles generate massive streams of visual data during operation, and existing continuous learning approaches typically rely on heuristic sampling methods that fail to capture temporal dynamics, often overlooking critical learning opportunities or selecting redundant frames. In this paper, we introduce FrameScope, a temporal data valuation framework for continuous learning in autonomous vehicles. FrameScope extends neural tangent kernel theory to temporal domains, enabling principled valuation of streaming visual data. Unlike cloud-centric methods that transmit all video data for processing, our approach performs principled, local frame selection on the vehicle and queries a cloud-based oracle model only for labels of those high-value frames. Extensive experiments across multiple domain shifts show that FrameScope consistently outperforms existing methods, achieving higher sample efficiency and significantly reducing catastrophic forgetting in autonomous vehicle perception. By valuing data on the vehicle and querying only labels for selected frames, FrameScope reduces bandwidth requirements, enabling scalable operation with a lightweight cloud labeling service.

中文摘要

摘要:自动驾驶车辆在动态、不断变化的环境中运行,在这些环境中,新场景和边缘情况不断出现。因此,静态学习模型不足以保证其安全和可靠的运行。连续学习对于适应这些不断变化的条件并在各种现实环境中保持稳健的性能至关重要。然而,自动驾驶车辆在运行过程中会生成大量视觉数据流,而现有的连续学习方法通常依赖启发式采样方法,这些方法无法捕捉时间动态,常常忽略关键的学习机会或选择冗余帧。在本文中,我们提出了FrameScope,这是一个用于自动驾驶车辆连续学习的时间数据价值评估框架。FrameScope将神经切线核理论扩展到时间域,从而实现对流式视觉数据的原则性评估。不同于将所有视频数据传输到云端处理的以云为中心的方法,我们的方法在车辆本地进行原则性的帧筛选,仅对这些高价值帧向基于云的模型查询标签。在多个领域迁移实验中,FrameScope始终优于现有方法,实现了更高的样本效率,并显著减少了自动驾驶车辆感知中的灾难性遗忘。通过在车辆端评估数据并仅查询所选帧的标签,FrameScope降低了带宽需求,使其能够通过轻量级云标签服务实现可扩展运行。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28672 (HTTP 429)

Authors: Yuheng Zhu, Man-Ki Yoon

Categories: cs.CV

PDF URL: https://arxiv.org/pdf/2608.28672.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28672

Published: 2026-09-02T01:24:22.326Z


8. AdaptAV: Continuous Adaption of Vision Models for Autonomous Vehicles Using Cloud-based Oracle

Abstract:Deploying vision perception models in autonomous vehicles requires that we prioritize inference speeds, resulting in a model with shallower architectures and lesser model parameters (i.e., more pruned). Such small models do not generalize well, which could result in poor performance when encountered with novel scenarios. We propose a system that overcomes this by continuously retraining the vision models on the cloud with data uploaded by vehicles. We leverage the abundant compute resources, including machine learning accelerators, of the cloud to run a highly-accurate oracle model that will guide the retraining process of the on-vehicle model. This newly trained model is transmitted to the vehicle over the network and is utilized by the vehicle for perceptions, leading to improved inference accuracy over time.

中文摘要

摘要:在自动驾驶车辆中部署视觉感知模型需要我们优先考虑推理速度,从而产生具有较浅架构和较少模型参数(即经过更多剪枝)的模型。这类小型模型泛化能力不强,可能在遇到新场景时表现不佳。我们提出了一种系统,通过持续在云端使用车辆上传的数据重新训练视觉模型来克服这一问题。我们利用云端丰富的计算资源,包括机器学习加速器,运行一个高精度的专家模型,以指导车载模型的再训练过程。新训练的模型通过网络传输至车辆,并被车辆用于感知任务,从而随时间提高推理精度。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28673 (HTTP 429)

Authors: Yuheng Zhu, Dhruva Ungrupulithaya, Boluo Ge, Man-Ki Yoon

Categories: cs.CV

PDF URL: https://arxiv.org/pdf/2608.28673.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28673

Published: 2026-09-02T01:24:22.326Z


9. Multi-exposure HDR Imaging: A Review of Pixel-level and Feature-level Reconstruction Methods

Abstract:Multi-exposure is an efficient way to capture real-world high-dynamic-range (HDR) scenes. However, HDR imaging suffers from severe ghosting artifacts in dynamic scenes due to the temporal gap between sequential exposures. In this article, we categorize the literature on two important topics on HDR imaging: multi-exposure fusion (MEF) and ghost removal. Conventional filter-based and data-driven methods are studied in pixel space and feature space. For popular deep learning-based approaches, we provide a granular taxonomy based on their alignment and fusion domains: pixel-space methods, which typically employ explicit motion compensation such as optical flow or spatial transformers, and feature-space methods, which leverage implicit alignment through deformable convolutions, attention mechanisms, or latent representation merging. Representative works are compared across different supervision settings, and key design principles are summarized. In addition, this survey summarizes commonly used datasets and evaluation metrics, discussing their applicability under diverse output forms. Finally, major bottlenecks and promising directions for future research are outlined.

中文摘要

摘要:多曝光是一种高效捕捉真实世界高动态范围(HDR)场景的方法。然而,由于连续曝光之间的时间间隔,HDR成像在动态场景中容易产生严重的幽灵伪影。本文对HDR成像中两个重要主题——多曝光融合(MEF)与幽灵去除——的文献进行了分类。研究了传统基于滤波器的方法和数据驱动的方法在像素空间和特征空间的应用。对于流行的基于深度学习的方法,我们提供了基于对齐和融合域的详细分类:像素空间方法,通常采用显式运动补偿如光流或空间变换器,以及特征空间方法,通过可变形卷积、注意力机制或潜在表示合并实现隐式对齐。对不同监督设置下的代表性工作进行了比较,并总结了关键设计原则。此外,本综述还总结了常用的数据集和评价指标,讨论了它们在不同输出形式下的适用性。最后,概述了主要瓶颈和未来研究的潜在方向。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28674 (HTTP 429)

Authors: Qian Tao, Wei Wang, Chaobing Zheng, Zhengguo Li

Categories: cs.CV

PDF URL: https://arxiv.org/pdf/2608.28674.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28674

Published: 2026-09-02T01:24:22.326Z


10. Multi-Agent Self-Improving Reinforcement Learning for Video Reasoning

Abstract:Video reasoning tasks such as grounded video question answering and temporal grounding require selecting temporal evidence that supports the query. In many current training setups, temporal supervision is applied through local objectives such as boundary regression or span generation, while verification is used mainly to rerank candidate segments at inference time. We study whether a frozen verifier can also guide training. Our multi-agent framework couples a trainable \emph{Grounder} with a frozen \emph{Verifier}: the Grounder samples candidate trajectories and evidence segments, the Verifier assigns query-conditioned segment scores, a group-relative policy-gradient objective favors trajectories that outperform their within-input peers, and a bootstrapped calibration loss steers temporal predictions toward verifier-preferred spans. Trained on source tasks and evaluated without target-dataset fine-tuning, a two-billion-parameter instantiation transfers zero-shot across grounded question answering, temporal grounding, and long-video question answering, reaching 28.7\% intersection-over-union and 25.4\% answer-grounding accuracy on a grounded-question-answering benchmark, 46.1\% intersection-over-union on a temporal-grounding benchmark, and 54.1\% on a long-video question-answering benchmark. Relative to a strong same-scale baseline, the gains are modest but consistent, with the clearest improvements on relevance-oriented metrics such as intersection-over-union and moderate-overlap recall. Within the tested benchmarks and transfer setting, the results support frozen verification as a training signal for evidence selection, while showing that strict boundary precision remains comparatively weaker. Code and models are available at this https URL

中文摘要

摘要:视频推理任务,如有根视频问答和时间定位,需要选择支持查询的时间证据。在许多当前的训练设置中,时间监督通过局部目标(如边界回归或区间生成)应用,而验证主要用于在推理时重新排序候选段落。我们研究了冻结的验证器是否也可以指导训练。我们的多智能体框架将可训练的\emph{Grounder}与冻结的\emph{Verifier}耦合:Grounder采样候选轨迹和证据片段,Verifier为条件查询的片段分配分数,基于组的策略梯度目标偏向于表现超越输入内部同组对手的轨迹,且引导式校准损失使时间预测趋向Verifier偏好的片段。在源任务上训练,并在未进行目标数据集微调的情况下评估,一个二十亿参数的实现可零样本迁移到有根问答、时间定位和长视频问答,在有根问答基准上达到28.7%的交并比和25.4%的答案定位准确率,在时间定位基准上达到46.1%的交并比,在长视频问答基准上达到54.1%。相较于强大的同规模基线,提升虽适中但稳定,最明显的改进体现在与相关性相关的指标上,如交并比和中度重叠召回。在测试的基准和迁移设置内,结果支持将冻结验证作为证据选择的训练信号,同时显示严格的边界精度仍相对较弱。代码和模型可通过此https URL获取。

LLM Analysis

LLM Analysis Failed: Error: Kimi内容未就绪或抓取失败: 2608.28675 (HTTP 429)

Authors: Mingwen Zhang, Jisheng Dang, Minqiang Yang, Bimei Wang, Bin Hu, Tat-Seng Chua

Categories: cs.CV

PDF URL: https://arxiv.org/pdf/2608.28675.pdf

CoolPaper URL: https://papers.cool/arxiv/2608.28675

Published: 2026-09-02T01:24:22.326Z