Hot News
202608
hot_news 2026-08-07
hot_news 2026-08-06
hot_news 2026-08-05
hot_news 2026-08-03
hot_news 2026-08-02
hot_news 2026-08-01
202607
hot_news 2026-07-31
hot_news 2026-07-30
hot_news 2026-07-29
hot_news 2026-07-27
hot_news 2026-07-26
hot_news 2026-07-21
hot_news 2026-07-20
hot_news 2026-07-19
hot_news 2026-07-18
hot_news 2026-07-17
hot_news 2026-07-16
hot_news 2026-07-15
hot_news 2026-07-14
hot_news 2026-07-13
hot_news 2026-07-12
hot_news 2026-07-11
hot_news 2026-07-09
hot_news 2026-07-08
hot_n ...
ArXiv Domain 2026-06-01
数据来源:ArXiv Domain
LLM Domain Papers1. Lightweight Multimodal LLM-Enabled Cost-Effective Defect Grading of Power Transmission EquipmentAbstract:Defect grading of power transmission equipment (DGPTE) is crucial to the stability of electric energy transmission. Although existing machine learning methods exhibit strong capabilities in defect detection, they are plagued by difficulties in integrating expert experience and facing class imbalance in more refined defect grading field. To address this ...
ArXiv Domain 2026-06-03
数据来源:ArXiv Domain
LLM Domain Papers1. IdiomX A Multilingual Benchmark for Idiom Understanding, Retrieval, and InterpretationAbstract:Idiomatic expressions remain a persistent challenge for natural language processing because their meanings are often non-compositional, context-dependent, and difficult to align across languages. Existing idiom resources are often limited in scale, contextual diversity, or multilingual coverage, restricting their utility for modern language models. We introduce I ...
ArXiv Domain 2026-06-04
数据来源:ArXiv Domain
LLM Domain Papers1. POLARIS: Guiding Small Models to Write Long StoriesAbstract:Small open-weight models struggle at long-form creative writing: their generated stories either fall far short of the requested length, or their quality significantly degrades as length increases, especially when compared to frontier models. We present POLARIS (Policy Optimization with LLM-as-a-judge rewards and Anchored-Reference Injection for Storywriting), a lower-compute GRPO recipe with two k ...
ArXiv Domain 2026-06-05
数据来源:ArXiv Domain
LLM Domain Papers1. POLARIS: Guiding Small Models to Write Long StoriesAbstract:Small open-weight models struggle at long-form creative writing: their generated stories either fall far short of the requested length, or their quality significantly degrades as length increases, especially when compared to frontier models. We present POLARIS (Policy Optimization with LLM-as-a-judge rewards and Anchored-Reference Injection for Storywriting), a lower-compute GRPO recipe with two k ...
ArXiv Domain 2026-06-08
数据来源:ArXiv Domain
LLM Domain Papers1. Improving Cross-Lingual Factual Recall via Consistency-Driven Reinforcement LearningAbstract:Large language models (LLMs) trained predominantly on English data encode substantial world knowledge, yet often fail to express it reliably in other languages, a phenomenon known as cross-lingual factual inconsistency. To study and address this, we introduce PolyFact, a large-scale parallel multilingual factual QA dataset containing 100K Wikidata-grounded facts ac ...
ArXiv Domain 2026-06-10
数据来源:ArXiv Domain
LLM Domain Papers1. Bidirectional Small-Granularity Search between Code and TextAbstract:We introduce the novel task of bidirectional small-granularity search between code and text, where the queries are small snippets of text or code and the results are also small fragments of the opposite modality, i.e., code or text. This task establishes direct links between text in scientific publications and corresponding code segments, in support of better and faster understanding of s ...
ArXiv Domain 2026-06-11
数据来源:ArXiv Domain
LLM Domain Papers1. PoQ-Judge: A Multi-Architecture Evaluation Framework for Cost-Aware Proof-of-Quality in Decentralized LLM InferenceAbstract:Decentralized LLM inference networks need lightweight, reference-free quality evaluation for Proof of Quality (PoQ). We present PoQ-Judge, a framework that trains dedicated judge models to score query-output pairs without ground-truth references. We study three architectures across the quality-cost tradeoff: a TextCNN judge, a MiniLM ...
ArXiv Domain 2026-06-15
数据来源:ArXiv Domain
LLM Domain Papers1. The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge EvaluationAbstract:LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized. We study repeated identical evaluations on 29 tasks spanning 10 categories using two OpenAI judge models (GPT-4o-mini and GPT-4.1-mini), with 50 pairwise trials and 50 pointwise trials per question, supplement ...
ArXiv Domain 2026-06-19
数据来源:ArXiv Domain
LLM Domain Papers1. Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path AggregationAbstract:Large Language Models (LLMs) exhibit representational and syntactic biases that are difficult to evaluate due to the stochastic nature of text generation. Standard auditing methods rely on a single output inspection or static automated metrics. These approaches obscure the underlying probability distributions and fail to capture biases hidden in lower-probability g ...
ArXiv Domain 2026-06-24
数据来源:ArXiv Domain
LLM Domain Papers1. Less is More: Lightweight Prompt Compression for Question Answering Applications on Edge DevicesAbstract:In agent-driven question answering (QA) applications, retrieval-augmented generation (RAG) is commonly introduced to enhance the response accuracy of large language models (LLMs) by providing additional context. Due to the inherent noise in retrieval results and the coarse granularity of document-level retrieval, the retrieved context often contains sub ...
ArXiv Domain 2026-06-25
数据来源:ArXiv Domain
LLM Domain Papers1. EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQLAbstract:Text-to-SQL enables users to query databases using natural language by generating executable SQL queries. Recent methods have increasingly adopted Large Language Models based reinforcement learning (RL) to leverage execution feedback for training. However, existing RL methods assign uniform query-level rewards to all clauses in a SQL query, treating correct and incorrect cla ...
ArXiv Domain 2026-07-08
数据来源:ArXiv Domain
LLM Domain Papers1. Improving LLMs via Validator-to-Generator AlignmentAbstract:Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generator-validator (G-V) gap is one manifestation of this phenomenon, where LLMs generate responses that they then deem as invalid if re-queried to validate them. In this work, we introduce a new formulation of G-V consistency that involves a principled c ...
ArXiv Domain 2026-07-15
数据来源:ArXiv Domain
LLM Domain Papers1. CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time SeriesAbstract:Clinical time series are central to patient monitoring, risk assessment, and clinical decision support. However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Question Answering (QA). Existing benchmarks primarily focus on regularly sampled time-series QA or ...
ArXiv Domain 2026-08-05
数据来源:ArXiv Domain
LLM Domain Papers1. Cost-Effective Automated Judging of Natural-Language Mathematical ProofsAbstract:Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whether cheap open-weight models can serve as reliable judges given a candidate proof, a ground-truth proof, and a human-grading rubric. On a 200-instance validation sample of IMO-GradingBench, three cheap judges (GPT-OSS 120B, De ...