avatar
Articles
273
Tags
23
Categories
15

Home
Content
  • Paper
  • LLMs
  • Jupyter
  • Algorithm
  • PLs
Daily
  • Github
  • HotNews
  • HF
  • Arxiv
Archives
Categories
About
37.2° Blog
Search
Home
Content
  • Paper
  • LLMs
  • Jupyter
  • Algorithm
  • PLs
Daily
  • Github
  • HotNews
  • HF
  • Arxiv
Archives
Categories
About
HuggingFace Papers 2026-09-06
Created2019-06-18|AI
数据来源:HuggingFace Papers Latest Papers1. Compile by Training: Turning Natural-Language Specifications into Local Neural FunctionsAbstract:Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task ...
HuggingFace Papers 2026-09-10
Created2019-06-18|AI
数据来源:HuggingFace Papers Latest Papers1. NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing HarnessAbstract:Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heterogeneous model pool with intelligent ...
HuggingFace Papers 2026-09-14
Created2019-06-18|AI
数据来源:HuggingFace Papers Latest Papers1. NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept PredictionAbstract:We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective w ...
HuggingFace Papers 2026-09-16
Created2019-06-18|AI
数据来源:HuggingFace Papers Latest Papers1. Vidu S2: Real-Time Interactive, Editable, and Spatial Video GenerationAbstract:We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic ...
HuggingFace Papers 2026-09-18
Created2019-06-18|AI
数据来源:HuggingFace Papers Latest Papers1. LimiX-2: A Contextual Mechanism Network Towards General Structured-Data IntelligenceAbstract:We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric predic ...
HuggingFace Papers 2026-09-20
Created2019-06-18|AI
数据来源:HuggingFace Papers Latest Papers1. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache CompressionAbstract:The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the p ...
HuggingFace Papers 2026-09-21
Created2019-06-18|AI
数据来源:HuggingFace Papers Latest Papers1. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache CompressionAbstract:The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the p ...
HuggingFace Papers 2026-09-22
Created2019-06-18|AI
数据来源:HuggingFace Papers Latest Papers1. CodeMidas: Scaling Agentic Coding RL Environments from Code ItselfAbstract:Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipelin ...
ArXiv Domain 2026-08-02
Created2019-06-18|AI
数据来源:ArXiv Domain LLM Domain Papers1. Prompt Chaining in Practice: A Case Study in Automated Scholarly Report GenerationAbstract:The exponential growth of scholarly publications requires automated tools for effective information synthesis. However, simple, single-shot prompting methods often lack the reliability and quality required for complex synthesis tasks. This paper introduces and empirically evaluates a multi-stage prompt chaining methodology as a more reliable architectural pattern for ...
ArXiv Domain 2026-08-03
Created2019-06-18|AI
数据来源:ArXiv Domain LLM Domain Papers1. Prompt Chaining in Practice: A Case Study in Automated Scholarly Report GenerationAbstract:The exponential growth of scholarly publications requires automated tools for effective information synthesis. However, simple, single-shot prompting methods often lack the reliability and quality required for complex synthesis tasks. This paper introduces and empirically evaluates a multi-stage prompt chaining methodology as a more reliable architectural pattern for ...
ArXiv Domain 2026-08-09
Created2019-06-18|AI
数据来源:ArXiv Domain LLM Domain Papers1. Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision SupportAbstract:Wastewater operators need answers grounded in how their plant’s variables interact and how fast effects propagate, not in generic pretraining text, when asking causal questions such as “why is N2O rising?” or “what happens if I cut aeration by 20%?”. We compare three concret ...
ArXiv Domain 2026-08-10
Created2019-06-18|AI
数据来源:ArXiv Domain LLM Domain Papers1. Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision SupportAbstract:Wastewater operators need answers grounded in how their plant’s variables interact and how fast effects propagate, not in generic pretraining text, when asking causal questions such as “why is N2O rising?” or “what happens if I cut aeration by 20%?”. We compare three concret ...
ArXiv Domain 2026-08-16
Created2019-06-18|AI
数据来源:ArXiv Domain LLM Domain Papers1. LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint ReasoningAbstract:When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail — but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional constraint activation: the constraint is internally encoded (Knowledge) symmetrically across constraint-present and -abs ...
ArXiv Domain 2026-08-17
Created2019-06-18|AI
数据来源:ArXiv Domain LLM Domain Papers1. LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint ReasoningAbstract:When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail — but aggregate accuracy conflates genuine constraint inference with conservative defaulting. We formalize the distinction as conditional constraint activation: the constraint is internally encoded (Knowledge) symmetrically across constraint-present and -abs ...
ArXiv Domain 2026-08-23
Created2019-06-18|AI
数据来源:ArXiv Domain LLM Domain Papers1. A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to DeploymentAbstract:We describe the evolution of a virtual assistant, called ATHENA, designed to support the capture, retrieval, and dissemination of knowledge for members of a Community of Practice (CoP) related to the Oil and Gas sector. An evaluation of a first prototype involving 75 professionals from the Society of Petroleum Engineering (SPE) showed th ...
1…171819
avatar
Firefly
A firefly flying freely in the AI domain.
Articles
273
Tags
23
Categories
15
Follow Me
Announcement
Welcome to My Personal Blog!
If Not, Please Visit Gitee Mirror.
Recent Post
检索增强LLM2024-01-13
LLMs公开课 - 6.文本理解和生成大模型2024-01-10
LLMs公开课 - 5.高效训练&模型压缩2024-01-07
Categories
  • AI92
  • Cython1
  • DSA24
  • GitHub45
  • HotNews45
Tags
DSARLTransformerLLMsPaperReadingDeepLearningCVGPTPLdomaingithubhfhot_newsArXivDomainAIGitHubTrendingHuggingFacePapersHotNewsleetcodealgo
Archives
  • January 20245
  • December 202314
  • November 202326
  • October 20231
  • September 20234
Info
Article :
273
Run time :
Total Count :
9791k
UV :
PV :
Last Push :
©2023 - 2026 By Firefly
Search
Loading the Database