ArXiv Domain 2026-07-06
数据来源:ArXiv Domain
LLM Domain Papers1. TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language ModelsAbstract:Understanding how Large Language Models (LLMs) make token-level decisions during code generation remains a major challenge for both researchers and practitioners. While recent tools provide insights into model internals or generation outcomes, they often lack decoding-time signals, fine-grained uncertainty measures, and interactive mechanism ...
ArXiv Domain 2026-07-07
数据来源:ArXiv Domain
LLM Domain Papers1. TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language ModelsAbstract:Understanding how Large Language Models (LLMs) make token-level decisions during code generation remains a major challenge for both researchers and practitioners. While recent tools provide insights into model internals or generation outcomes, they often lack decoding-time signals, fine-grained uncertainty measures, and interactive mechanism ...
ArXiv Domain 2026-07-13
数据来源:ArXiv Domain
LLM Domain Papers1. Unveiling Public Opinion: A Study of Sentiment Analysis Using LSTM and Traditional ModelsAbstract:In this age of social media, sites like Twitter have become meeting places for people to share their views and feelings on a wide range of issues and current events as they unfold in real time. Sentiment analysis, a critical application of NLP, has become indispensable due to the massive influx of user-generated content, enabling the extraction of meaningful i ...
ArXiv Domain 2026-07-19
数据来源:ArXiv Domain
LLM Domain Papers1. Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMsAbstract:Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure. We introduce Just Keep Prompting (JKP), a multi-turn evaluation framework that measures VLM epistemic stability when users repeatedly challenge, question, or contradict a model’s answer. JKP probes models for up to 10 ...
ArXiv Domain 2026-07-20
数据来源:ArXiv Domain
LLM Domain Papers1. Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMsAbstract:Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure. We introduce Just Keep Prompting (JKP), a multi-turn evaluation framework that measures VLM epistemic stability when users repeatedly challenge, question, or contradict a model’s answer. JKP probes models for up to 10 ...
ArXiv Domain 2026-07-26
数据来源:ArXiv Domain
LLM Domain Papers1. What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning TracesAbstract:What makes writing “good” remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality. In Study 1, we construct a benchmark of 30 real texts spanning six quality tiers, from canonical literature to anonymous forum posts, and extract the ...
ArXiv Domain 2026-07-27
数据来源:ArXiv Domain
LLM Domain Papers1. What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning TracesAbstract:What makes writing “good” remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality. In Study 1, we construct a benchmark of 30 real texts spanning six quality tiers, from canonical literature to anonymous forum posts, and extract the ...
ArXiv Domain 2026-08-02
数据来源:ArXiv Domain
LLM Domain Papers1. Prompt Chaining in Practice: A Case Study in Automated Scholarly Report GenerationAbstract:The exponential growth of scholarly publications requires automated tools for effective information synthesis. However, simple, single-shot prompting methods often lack the reliability and quality required for complex synthesis tasks. This paper introduces and empirically evaluates a multi-stage prompt chaining methodology as a more reliable architectural pattern for ...
ArXiv Domain 2026-08-03
数据来源:ArXiv Domain
LLM Domain Papers1. Prompt Chaining in Practice: A Case Study in Automated Scholarly Report GenerationAbstract:The exponential growth of scholarly publications requires automated tools for effective information synthesis. However, simple, single-shot prompting methods often lack the reliability and quality required for complex synthesis tasks. This paper introduces and empirically evaluates a multi-stage prompt chaining methodology as a more reliable architectural pattern for ...