-
Heart failure (HF) care requires repeated decisions across suspected disease, diagnostic confirmation, phenotyping, guideline-directed medical therapy, device consideration, worsening HF, transition care, and advanced HF planning. Large language models (LLMs) …
europepmc
2026
置信度 0.80
-
To address the bottlenecks of missing decision-making closed loop, insufficient experience reuse, and decoupled resource scheduling in industrial LLM deployment, this paper proposes LLM-Conductor, a three-layer collaborative architecture that enables monitorin…
europepmc
2026
置信度 0.80
-
With the rapid growth of unstructured clinical narratives in electronic health records (EHRs), clinical named entity recognition (NER) has become a crucial technique for extracting structured medical information. However, traditional supervised models such as …
pubmed
Tao X, Dong X, Zhu Q, Zhou X
2026
置信度 0.82
-
Anomaly detection is crucial in maintaining the safety, reliability, and optimal performance of complex systems across diverse domains, such as industrial manufacturing, cybersecurity, and autonomous systems. While conventional methods typically handle single …
europepmc
2026
置信度 0.80
-
With the proliferation of Large Language Models (LLMs), the detection of misinformation has become increasingly important and complex. This research proposes an innovative verifiable misinformation detection LLM agent that goes beyond traditional true/false bi…
openalex
Zikun Cui, Tianyi Huang, Chia-En Chiang, C. Du
2025-08-04
置信度 0.72
MisinformationVerifiable secret sharingComputer scienceCredibilityTrustworthiness
-
This paper explores the use of Large Language Models (LLMs) in modeling real-world optimization problems. We concretely define the task of translating natural language descriptions into optimization models (NL2OPT) and provide criteria for classifying optimiza…
openalex
Mahdi Mostajabdaveh, Timothy T. Yu, Rindranirina Ramamonjison, Giuseppe Carenini 等
2024-08-07
置信度 0.72
Computer scienceTask (project management)Modeling languageSolverRepresentation (politics)
-
In general, educational support with Large Language Models (LLMs) faces challenges in knowledge organization, expertise integration, and contextual adaptation. So, we present EduMAS, a novel multi-agent framework that coordinates specialized agents with graph-…
openalex
Qiaomu Li, Ying Xie, Sumit Chakravarty, Dabae Lee
2024-12-15
置信度 0.72
Computer scienceHuman–computer interaction
-
Integrating Large Language Models (LLMs) into autonomous agents marks a significant shift in the research landscape by offering cognitive abilities that are competitive with human planning and reasoning. This article explores the transformative potential of in…
openalex
Junda He, Christoph Treude, David Lo
2025-01-13
置信度 0.72
Computer scienceScalabilitySoftware engineeringSystems development life cycleTransformative learning
-
Large language models (LLMs) have drastically changed the possible ways to design intelligent systems, shifting the focus from massive data acquisition and new model training to human alignment and strategic elicitation of the full potential of existing pre-tr…
openalex
Frank Xing
2024-08-13
置信度 0.72
Leverage (statistics)Summative assessmentContext (archaeology)Computer scienceGenerative grammar
-
Large language models (LLMs) have been increasingly used to interact with external environments (e.g., games, compilers, APIs) as goal-driven agents. However, it remains challenging for these language agents to quickly and efficiently learn from trial-and-erro…
openalex
Noah Shinn, Cassano, Federico, Berman, Edward, Gopinath, Ashwin 等
2023-03-20
置信度 0.72
Computer scienceReinforcement learningCoding (social sciences)Benchmark (surveying)Compiler
-
The automation of resume screening is a crucial aspect of the recruitment process in organizations. Automated resume screening systems often encompass a range of natural language processing (NLP) tasks. This paper introduces a novel Large Language Models (LLMs…
openalex
Chengguang Gan, Qinghao Zhang, Tatsunori Mori
2024-01-16
置信度 0.72
Automatic summarizationComputer scienceArtificial intelligenceGrading (engineering)Engineering
-
AutoGen is an open-source framework that allows developers to build LLM applications via multiple agents that can converse with each other to accomplish tasks. AutoGen agents are customizable, conversable, and can operate in various modes that employ combinati…
openalex
Wu, Qingyun, Gagan Bansal, Jieyu Zhang, Yiran Wu 等
2023-08-16
置信度 0.72
ConverseConversationComputer scienceCoding (social sciences)Entertainment
-
Recent progress with LLM-based agents has shown promising results across various tasks.However, their use in answering questions from knowledge bases remains largely unexplored.Implementing a KBQA system using traditional methods is challenging due to the shor…
openalex
Chang Zong, Yuchen Yan, Weiming Lü, Jian Shao 等
2024-01-01
置信度 0.72
Triad (sociology)Question answeringComputer scienceKnowledge baseOpen Knowledge Base Connectivity
-
openalex
Jack Gallifant, Majid Afshar, Saleem Ameen, Yindalon Aphinyanaphongs 等
2025-01-01
置信度 0.72
Tripod (photography)ChecklistGuidelineStandardizationComputer science
-
In the digital transformation era, the surge of better development technologies and citizen developers disrupted the space of innovation by increasing the number and complexity of applications used in production. This context prompts advanced cybersecurity mea…
openalex
Stanislas G. Bianou, Rodrigue G. Batogna
2024-09-02
置信度 0.72
Computer scienceAutomationPenetration (warfare)Software engineeringEngineering
-
The rapid advancement of Large Language Models (LLMs) has led to substantial investment in enhancing their capabilities and expanding their feature sets. Despite these developments, a critical gap remains between model sophistication and their dependable deplo…
openalex
Ahmed M. Darwish, Essam A. Rashed, Ghada Khoriba
2025-06-21
置信度 0.72
PsychologyPsychotherapistComputer science
-
Large Language Model (LLM) agents are transforming education by automating complex pedagogical tasks and enhancing both teaching and learning processes. In this survey, we present a systematic review of recent advances in applying LLM agents to address key cha…
openalex
Zhendong Chu, Shen Wang, Jianhe Xie, Tinghui Zhu 等
2025-01-01
置信度 0.72
Computer scienceEngineeringMedicineTroubleshootingKey (lock)
-
Advances in large language models (LLMs) have empowered a variety of applications. However, there is still a significant gap in research when it comes to understanding and enhancing the capabilities of LLMs in the field of mental health. In this work, we prese…
openalex
Xuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel 等
2024-03-06
置信度 0.72
Mental healthSet (abstract data type)Computer scienceTask (project management)Variety (cybernetics)
-
In software development, resolving the emergent issues within GitHub repositories is a complex challenge that involves not only the incorporation of new code but also the maintenance of existing code. Large Language Models (LLMs) have shown promise in code gen…
openalex
Wei Tao, Yucheng Zhou, Wang, Yanlin, Wenqiang Zhang 等
2024-03-26
置信度 0.72
Computer scienceResolution (logic)Artificial intelligence
-
In this paper, we present a novel framework for enhancing the capabilities of large language models (LLMs) by leveraging the power of multi-agent systems. Our framework introduces a collaborative environment where multiple intelligent agent components, each wi…
openalex
Yashar Talebirad, Amirhossein Nadiri
2023-06-05
置信度 0.72
Computer scienceScalabilityIntelligent agentMulti-agent systemWork (physics)
-
• AI framework automates MgH 2 catalyst data extraction from literature. • LLM to Agent approach accelerates MgH 2 catalyst discovery and design. • Machine learning predicts MgH 2 dehydrogenation with high accuracy. • Cat-Advisor provides actionable catalyst d…
openalex
Tongao Yao, Yang Yang, Jianghao Cai, Rui Liu 等
2025-10-22
置信度 0.72
DehydrogenationComputer scienceMachine learningArtificial intelligenceHydrogen storage
-
Integrating Large Language Models (LLMs) in healthcare diagnosis demands systematic frameworks that can handle complex medical scenarios while maintaining specialized expertise. We present KG4Diagnosis, a novel hierarchical multi-agent framework that combines …
openalex
Kaiwen Zuo, Yirui Jiang, Fan Mo, Píetro Lió
2024-12-22
置信度 0.72
GraphKnowledge graphComputer scienceArtificial intelligenceData science
-
The recent surge in research interest in applying large language models (LLMs) to decision-making tasks has flourished by leveraging the extensive world knowledge embedded in LLMs. While there is a growing demand to tailor LLMs for custom decision-making tasks…
openalex
Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Gaetan Lin 等
2024-03-24
置信度 0.72
Experiential learningPsychologyExperiential educationMathematics educationMedical education
-
Large language models (LLMs) can be used to serve as agents to simulate human behaviors, given the powerful ability to understand human instructions and provide high-quality generated texts. Such ability stimulates us to wonder whether LLMs can simulate a pers…
openalex
Yunfan Shao, Linyang Li, Junqi Dai, Xipeng Qiu
2023-01-01
置信度 0.72
Character (mathematics)MemorizationWonderCleopatraComputer science
-
Leveraging advanced reasoning capabilities and extensive world knowledge of large language models (LLMs) to construct generative agents for solving complex real-world problems is a major trend. However, LLMs inherently lack embodiment as humans, resulting in s…
openalex
Ye Jin, Yang, Ruoxuan, Yi, Zhijie, Xiaoxi Shen 等
2023-09-22
置信度 0.72
Computer scienceGenerative grammarContext (archaeology)Human–computer interactionDriving simulation
-
Cloud Operations (CloudOps) is a rapidly growing field focused on the automated management and optimization of cloud infrastructure which is essential for organizations nav-igating increasingly complex cloud environments. MontyCloud Inc. is one of the major co…
openalex
Kannan Parthasarathy, Karthik Vaidhyanathan, Rudra Dhar, Venkat Krishnamachari 等
2025-04-27
置信度 0.72
Computer scienceSystems engineeringEngineering
-
Recent developments in large language models (LLMs) have unlocked opportunities for healthcare, from information synthesis to clinical decision support. These LLMs are not just capable of modeling language, but can also act as intelligent “agents” that interac…
openalex
Nikita Mehandru, Brenda Y. Miao, Eduardo Rodriguez Almaraz, Madhumita Sushil 等
2024-04-03
置信度 0.72
WorkflowProcess (computing)Computer scienceFidelityCorporate governance
-
openalex
Junjie Huang, Quanyan Zhu
2024-01-01
置信度 0.72
Environmental remediationComputer scienceEnvironmental scienceBiologyEcology
-
In the era of Industry 4.0, the proliferation of data within manufacturing environments has presented both unprecedented opportunities and challenges. This paper introduces a framework that capitalizes on the capabilities of Large Language Models (LLMs) to rev…
openalex
Cristian I. Garcia, Marcus A. DiBattista, Tomás A. Letelier, Hunter D. Halloran 等
2024-10-01
置信度 0.72
Computer scienceManufacturing engineeringEngineeringEngineering drawing
-
Agentic AI systems are a recently emerged and important approach that goes beyond traditional AI, generative AI, and autonomous systems by focusing on autonomy, adaptability, and goal-driven reasoning. This study provides a clear review of agentic AI systems b…
openalex
Ajay Bandi, Bhavani Kongari, Roshini Naguru, Sahitya Pasnoor 等
2025-09-04
置信度 0.72
Computer scienceArtificial intelligenceData scienceMachine learningSoftware engineering
-
Leveraging advanced reasoning capabilities and extensive world knowledge of large language models (LLMs) to construct generative agents for solving complex real-world problems is a major trend. However, LLMs inherently lack embodiment as humans, resulting in s…
openalex
Jin Ye, Ruoxuan Yang, Zhijie Yi, Xiaoxi Shen 等
2024-10-14
置信度 0.72
Computer scienceGenerative grammarHuman–computer interactionSystems engineeringArtificial intelligence
-
Large Language Models (LLMs) are increasingly being integrated into applications, with versatile functionalities that can be easily modulated via natural language prompts. So far, it was assumed that the user is directly prompting the LLM. But, what if it is n…
openalex
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres 等
2023-11-21
置信度 0.72
Computer scienceExploitComputer securitySoftware deploymentInterface (matter)
-
Automated machine learning (AutoML) accelerates AI development by automating tasks in the development pipeline, such as optimal model search and hyperparameter tuning. Existing AutoML systems often require technical expertise to set up complex tools, which is …
openalex
Patara Trirat, Wonyong Jeong, Sung Ju Hwang
2024-10-03
置信度 0.72
Pipeline (software)Computer scienceProgramming language
-
This study focuses on the utilization of Large Language Models (LLMs) for the rapid development of applications, with a spotlight on LangChain, an open-source software library. LLMs have been rapidly adopted due to their capabilities in a range of tasks, inclu…
openalex
Oğuzhan Topsakal, Tahir Çetin Akıncı
2023-07-22
置信度 0.72
BespokeComputer scienceModular designDebuggingSoftware engineering
-
openalex
Shreya Johri, Jae‐Hwan Jeong, Benjamin A. Tran, Daniel I. Schlessinger 等
2025-01-01
置信度 0.72
CraftSet (abstract data type)Medical historyComputer sciencePsychology
-
Stance detection automatically detects the stance in a text towards a target, vital for content analysis in web and social media research. Despite their promising capabilities, LLMs encounter challenges when directly applied to stance detection. First, stance …
openalex
Xiaochong Lan, Chen Gao, Depeng Jin, Yong Li
2024-05-28
置信度 0.72
Computer sciencePsychology
-
The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area.This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a twodimensiona…
openalex
Mahmoud Mohammadi, Yipeng Li, Jane C. Lo, Wendy Yip
2025-08-03
置信度 0.72
BenchmarkingComputer scienceBusinessMarketing
-
ABSTRACT The discovery and validation of genetic biomarkers across diverse diseases demand intelligent systems capable of integrating complex multi-omics data with clinical relevance. We introduce HEAL-KGGen, an end-to-end framework that enhances Large Languag…
preprints
Kaiwen Zuo, Zixuan Zhong, Peizhou Huang, Shiyan Tang 等
2025
置信度 0.74
BiomarkerComputer scienceGraphComputational biologyArtificial intelligence
-
We introduce a novel reinforcement learning framework of LLM agents named AGILE (AGent that Interacts and Learns from Environments) designed to perform complex conversational tasks with users, leveraging LLMs, memory, tools, and interactions with experts. The …
openalex
Peiyuan Feng, Yichen He, Guanhua Huang, Yuan Lin 等
2024-01-01
置信度 0.72
Computer scienceReinforcement learningArtificial intelligenceAction (physics)Key (lock)
-
(1) Background and objectives: Large language models (LLMs) such as GPT, Mistral, and LLaMA exhibit strong capabilities in text generation, yet assessing the quality of their reasoning—particularly in open-ended and argumentative contexts—remains a persistent …
openalex
Cătălin Anghel, Andreea Alexandra Anghel, Emilia Pecheanu, Ioan Șușnea 等
2025-08-01
置信度 0.72
Dual (grammatical number)DialecticComputer scienceArtificial intelligenceEpistemology
-
This study explores strategies for optimizing the use of large language models (LLMs) in Building Information Modeling (BIM) data retrieval. BIM data retrieval plays a crucial role in enhancing the efficiency and effectiveness of building management and constr…
openalex
Deli Liu, Xiaoping Zhou, Yu Li
2025-01-27
置信度 0.72
Building information modelingBusinessComputer scienceRisk analysis (engineering)Engineering
-
Recent advancements in Large Language Models (LLMs) have exhibited notable efficacy in question-answering (QA) tasks across diverse domains. Their prowess in integrating extensive web knowledge has fueled interest in developing LLM-based autonomous agents. Whi…
openalex
Yangyang Yu, Haohang Li, Cheng Zhi, Yuechen Jiang 等
2024-05-20
置信度 0.72
Computer scienceInterpretabilityTrading strategyFinancial marketArtificial intelligence
-
This paper introduces a Large Language Model (LLM)-based multi-agent framework designed to enhance anomaly detection within financial market data, tackling the longstanding challenge of manually verifying system-generated anomaly alerts. The framework harnesse…
openalex
Taejin Park
2024-03-28
置信度 0.72
Anomaly detectionAnomaly (physics)BusinessFinancial marketFinance
-
ChatMOF is an artificial intelligence (AI) system that is built to predict and generate metal-organic frameworks (MOFs). By leveraging a large-scale language model (GPT-4, GPT-3.5-turbo, and GPT-3.5-turbo-16k), ChatMOF extracts key details from textual inputs …
openalex
Yeonghun Kang, Jihan Kim
2024-06-03
置信度 0.72
Computer scienceVariety (cybernetics)Pipeline (software)Key (lock)Artificial intelligence
-
Modern software systems demand continuous evolution to maintain performance, scalability, and security. Traditional single-agent AI-driven code refactoring approaches are often limited in addressing the multi-faceted constraints (e.g., performance, security, m…
openalex
Vasanth Rajendran, Dinesh Besiahgari, Sachin C. Patil, Manjunath Chandrashekaraiah 等
2025-03-22
置信度 0.72
Code refactoringComputer scienceSoftware engineeringConceptual designConceptual model
-
Artificial Intelligence (AI) has been transformative in the healthcare sector, leading to enhanced precision in medical diagnosis, more effective treatment options, and a significant improvement in patient safety. However, computer-based administrative tasks, …
openalex
Senay A. Gebreab, Khaled Salah, Raja Jayaraman, Muhammad Habib ur Rehman 等
2024-04-29
置信度 0.72
AutomationTask (project management)Health careComputer scienceTask analysis
-
Abstract—Large Language Models (LLMs) suffer from inherent stochasticity, limiting their utility in high-stakes enterprise environments where determinism and auditability are required. This paper introduces the MFOUR Vibe Framework (MVF), a platform-agnostic a…
openalex
OpenAI, Achiam, Josh, Adler, Steven, Agarwal, Sandhini 等
2023-03-15
置信度 0.72
TransformerComputer scienceSecurity tokenProcess (computing)Scale (ratio)
-
openalex
Zixuan Xiao, Jun Ma, Siwei Zhang
2026-01-06
置信度 0.72
Computer scienceModular designUrban planningTask (project management)Quality (philosophy)
-
We introduce HackSynth, a novel Large Language Model (LLM)-based agent capable of autonomous penetration testing. HackSynth's dual-module architecture includes a Planner and a Summarizer, which enable it to generate commands and process feedback iteratively. T…
openalex
Lajos Muzsai, David Imolai, András Lukács
2024-12-02
置信度 0.72
Penetration (warfare)Computer scienceBusinessEngineeringOperations research
-
As Large Language Models (LLMs) are core components in Retrieval-Augmented Generation (RAG) systems for knowledge-intensive tasks, concerns regarding hallucinations, redundancy, and unverifiable outputs have intensified, particularly in high-stakes domains, su…
openalex
George Papageorgiou, Vangelis Sarlis, Manolis Μaragoudakis, Ioannis Magnisalis 等
2025-12-03
置信度 0.72
InterpretabilityComputer scienceRedundancy (engineering)Artificial intelligenceVocabulary
-
Abstract Retrieving relevant medical cases or documents is a critical information retrieval (IR) task in clinical decision support, particularly in cardiology, yet traditional search methods struggle with complex semantic queries in healthcare. Recent advances…
openalex
Lang Deng, Huanhuan Hu, Kongjie Lu, Ping He
2025-10-20
置信度 0.72
Computer scienceInformation retrievalRelevance (law)Task (project management)Question answering
-
Abstract Large Language Models (LLMs) are transforming industrial-organizational psychology and human resource management, with one of their most promising applications being automatic item generation (AIG) for psychological test development. Although recent a…
openalex
Philseok Lee, Mina Son, Zihao Jia
2025-08-26
置信度 0.72
Industrial and organizational psychologyPsychologyApplied psychologyConceptual modelSocial psychology
-
Although large language models (LLMs) have demonstrated remarkable code-generation ability, they still struggle with complex tasks. In real-world software development, humans usually tackle complex tasks through collaborative teamwork, a strategy that signific…
openalex
Yihong Dong, Xue Jiang, Zhi Jin, Ge Li
2024-06-12
置信度 0.72
Computer scienceCode generationCode (set theory)Software engineeringProgramming language
-
Significant progress has been made in automated problem-solving using societies of agents powered by large language models (LLMs). In finance, efforts have largely focused on single-agent systems handling specific tasks or multi-agent frameworks independently …
openalex
Xiao, Yijia, Edward W. Sun, Di Luo, Wei Wang
2024-12-28
置信度 0.72
BusinessFinanceFinancial system
-
Text evaluation has historically posed significant challenges, often demanding substantial labor and time cost. With the emergence of large language models (LLMs), researchers have explored LLMs' potential as alternatives for human evaluation. While these sing…
openalex
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu 等
2023-08-14
置信度 0.72
Computer scienceConstruct (python library)Quality (philosophy)Process (computing)Bridge (graph theory)
-
Digital twins are increasingly used in the Architecture, Engineering, and Construction (AEC) industry, but their adoption is often hindered by the need for specialised knowledge, such as database querying. This paper presents Graph-DT-GPT, a multi-agent framew…
openalex
Yuandong Pan, Mudan Wang, Linjun Lu, Rabindra Lamsal 等
2026-01-27
置信度 0.72
Computer scienceCorrectnessScalabilityModular designGraph
-
BACKGROUND Recent large language models (LLMs) have demonstrated significant advancements, particularly in their ability to serve as agents, thereby surpassing their traditional role as chatbots.These agents can leverage their planning and tool utilization cap…
openalex
Yixing Jiang, Kameron Collin Black, Gloria Geng, Dae-Gyun Park 等
2025-08-14
置信度 0.72
Computer scienceHealth careLeverage (statistics)Software deploymentBenchmark (surveying)
-
Large Language Model (LLM)-based multi-agent systems have demonstrated remarkable capabilities across di- verse applications, yet they face critical security challenges in- cluding backdoor attacks, prompt injection, and privacy leakage. Existing defense mecha…
openalex
Jinyu Chen, Jixiao Yang, Ziyang Zeng, Z. Jennifer Huang 等
2025-12-26
置信度 0.72
Computer scienceAction (physics)Context (archaeology)Key (lock)Perspective (graphical)
-
The justice system has increasingly applied AI techniques for legal judgment to enhance efficiency. However, most AI techniques focus on decision-making outcomes, failing to capture the deliberative nature of the real-world judicial process. To address these c…
openalex
Cong Jiang, Xiaolei Yang
2025-08-01
置信度 0.72
Computer scienceArtificial intelligence
-
openalex
Duzhen Zhang, Yahan Yu, Jiahua Dong, Chenxing Li 等
2024-01-01
置信度 0.72
Computer science
-
Dual-arm coordination is a fundamental problem in humanoid nursing robot. Large language model (LLM)-driven dual-arm collaboration is gradually becoming a research hotspot in this field. However, the single-thread LLM task planner lacks the ability of co-sched…
openalex
Zhendong Zhao, X. Yue, Jiexin Xie, Chuanhong Fang 等
2025-01-24
置信度 0.72
Dual (grammatical number)RobotHuman–computer interactionComputer scienceKnowledge management
-
Rare, yet critical, scenarios pose a significant challenge in testing and evaluating autonomous driving planners. Relying solely on real-world driving scenes requires collecting massive datasets to capture these scenarios. While automatic generation of traffic…
openalex
Yu Yao, Salil Bhatnagar, Markus Mazzola, Vasileios Belagiannis 等
2025-10-19
置信度 0.72
Computer scienceKey (lock)Control (management)Domain (mathematical analysis)Quality (philosophy)
-
Single-cell RNA sequencing (scRNA-seq) data analysis is crucial for biological research, as it enables the precise characterization of cellular heterogeneity. However, manual manipulation of various tools to achieve desired outcomes can be labor-intensive for …
openalex
Yihang Xiao, Jinyi Liu, Yan Zheng, Xiaohan Xie 等
2024-07-13
置信度 0.72
Computer scienceData scienceData mining
-
Scene simulation in autonomous driving has gained significant attention because of its huge potential for generating customized data. However, existing editable scene simulation approaches face limitations in terms of user interaction efficiency, multi-camera …
openalex
Yuxi Wei, Zi Wang, Yifan Lu, Chenxin Xu 等
2024-06-16
置信度 0.72
Computer scienceHuman–computer interaction
-
Integrating large language models (LLMs) into personal assistants, like Xiao Ai and Blue Heart V, effectively enhances their ability to interact with humans, solve complex tasks, and manage IoT devices. Such assistants are also termed LLM-driven agents. Upon r…
arxiv
Guopeng Li, Ruiqi Wu, Haisheng Tan
2025-12-24T18:08:03Z
置信度 0.78
cs.MA
-
Large Language Model (LLM) agents significantly extend the capabilities of standalone LLMs, empowering them to interact with external tools (e.g., APIs, functions) and complete various tasks in a self-directed fashion. The challenge of tool use demands that LL…
arxiv
Weizhou Shen, Chenliang Li, Hongzhan Chen, Ming Yan 等
2024-01-14T16:17:07Z
置信度 0.78
cs.AIcs.CL
-
Large language models (LLMs) drive significant financial innovations, yet their high-concurrency deployment is severely bottlenecked by KV cache memory overhead, which inflates infrastructure costs and throttles scalability. To address this, we propose YouZhi-…
arxiv
PSBC LLM Team, Huawei LLM Team, Ruihan Long, Junjie Wu 等
2026-06-04T08:44:37Z
置信度 0.78
cs.CL
-
Phishing websites remain a major cybersecurity threat, exploiting deceptive structures, brand impersonation, and social engineering to evade detection. Recent advances in large language models (LLMs) have improved phishing detection through contextual understa…
arxiv
Wenhao Li, Selvakumar Manickam, Yung-wey Chong, Shankar Karuppayah
2025-06-18T17:33:18Z
置信度 0.78
cs.CR
-
Large language model (LLM)-based agents are increasingly deployed in applications, such as trip-planning agents and web-use agents, to perform complex planning and execution tasks. Prior work has shown that LLM-based agents are vulnerable to context confusion,…
arxiv
Fengchao Chen, Tingmin Wu, Van Nguyen, Surya. Nepal 等
2026-01-14T03:29:13Z
置信度 0.78
cs.CR
-
We read twelve well-known LLM agent benchmark papers and recorded, dimension by dimension, what each paper actually says about how its evaluation was run. The motivation came from a familiar frustration: two papers will report results on the same benchmark wit…
arxiv
Mahdi Naser Moghadasi, Faezeh Ghaderi
2026-05-20T17:02:36Z
置信度 0.78
cs.LG
-
As Large Language Models (LLMs) transition from static tools to autonomous agents, traditional evaluation benchmarks that measure performance on downstream tasks are becoming insufficient. These methods fail to capture the emergent social and cognitive dynamic…
arxiv
Zarreen Reza
2025-10-01T07:10:28Z
置信度 0.78
cs.AIcs.MA
-
LLM agents increasingly rely on reusable skills (e.g., SKILL markdown files) to execute complex tasks, yet these artifacts lack portability: agent frameworks are highly sensitive to prompt formatting, leading to a large performance variation for the same skill…
arxiv
Yipeng Ouyang, Yi Xiao, Yuhao Gu, Xianwei Zhang
2026-05-05T04:15:48Z
置信度 0.78
cs.CRcs.AI
-
The integration of Large Language Models (LLMs) and knowledge graphs (KGs) has achieved remarkable success in various natural language processing tasks. However, existing methodologies that integrate LLMs and KGs often navigate the task-solving process solely …
arxiv
Lei Sun, Zhengwei Tao, Youdi Li, Hiroshi Arakawa
2024-04-11T12:16:16Z
置信度 0.78
cs.CLcs.AI
-
Large language models (LLMs) have achieved success in acting as agents, which interact with environments through tools such as search engines. However, LLMs are optimized for language generation instead of tool use during training or alignment, limiting their …
arxiv
Renxi Wang, Haonan Li, Xudong Han, Yixuan Zhang 等
2024-02-18T17:10:07Z
置信度 0.78
cs.CL
-
Driven by the rapid development of Large Language Models (LLMs), LLM-based agents have been developed to handle various real-world applications, including finance, healthcare, and shopping, etc. It is crucial to ensure the reliability and security of LLM-based…
arxiv
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen 等
2024-02-17T06:48:45Z
置信度 0.78
cs.CRcs.AIcs.CL
-
Recent position papers argue that the classical aleatoric/epistemic uncertainty framework is insufficient for interactive large language model (LLM) agents and call for underspecification-aware, decomposed, and communicable uncertainty representations that can…
arxiv
Gregory Matsnev
2026-06-17T19:59:32Z
置信度 0.78
cs.AIcs.CL
-
Large-scale learner-task interaction data are crucial for intelligent educational systems but are costly to collect and constrained by privacy and learner engagement. Learner simulators play a critical role in simulating scalable learner behavior without the n…
arxiv
Weibo Gao, Qi Liu, Linan Yue, Zheng Zhang 等
2026-06-13T09:48:25Z
置信度 0.78
cs.LGcs.AIcs.IR
-
Large language models (LLMs) have demonstrated notable potential in conducting complex tasks and are increasingly utilized in various financial applications. However, high-quality sequential financial investment decision-making remains challenging. These tasks…
arxiv
Yangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng 等
2024-07-09T05:52:26Z
置信度 0.78
cs.CL
-
A multi-agent pipeline with N agents typically issues N LLM calls per run. Merging agents into fewer calls (compound execution) promises token savings, but naively merged calls silently degrade quality through tool loss and prompt compression. We present Agent…
arxiv
Aninda Ray
2026-05-01T05:08:14Z
置信度 0.78
cs.CLcs.AIcs.SE
-
While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an efficient conversational framework design…
arxiv
Qi Wu, Yubo Zhao, Yifan Wang, Xinhang Liu 等
2024-05-27T09:57:51Z
置信度 0.78
cs.CV
-
We propose a personal-LLM exchange (LLM-X), a scalable negotiation-oriented environment that enables direct, structured communication across populations of personal agents (LLMs), each representing an individual user. Unlike existing tool-centric protocols tha…
arxiv
Giuliano Lorenzoni, Paulo Alencar, Donald Cowan
2026-05-12T01:04:37Z
置信度 0.78
cs.AI
-
While Large Language Model (LLM) agents are often approached from the angle of action planning/generation to accomplish a goal (e.g., given by language descriptions), their abilities to collaborate with each other to achieve a joint goal are not well explored.…
arxiv
Run Peng, Ziqiao Ma, Amy Pang, Sikai Li 等
2025-10-29T15:03:53Z
置信度 0.78
cs.CLcs.AI
-
LLM based agents have recently demonstrated strong potential in automating complex tasks, yet accurately predicting startup success remains an open challenge with few benchmarks and tailored frameworks. To address these limitations, we propose the Startup Succ…
arxiv
Xisen Wang, Yigit Ihlamur, Fuat Alican
2024-05-29T19:07:42Z
置信度 0.78
cs.AI
-
We present Social Agent, a novel framework for synthesizing realistic and contextually appropriate co-speech nonverbal behaviors in dyadic conversations. In this framework, we develop an agentic system driven by a Large Language Model (LLM) to direct the conve…
arxiv
Zeyi Zhang, Yanju Zhou, Heyuan Yao, Tenglong Ao 等
2025-10-06T09:41:37Z
置信度 0.78
cs.GRcs.CV
-
Large language models (LLMs) have rapidly evolved from single-turn text generators into the foundation of increasingly capable agents. As these agents take on more complex reasoning, decision making, tool use, and long-horizon tasks, reinforcement learning (RL…
arxiv
Mingyue Cheng, Shuo Yu, Daoyu Wang, Qingchuan Li 等
2025-11-18T13:03:15Z
置信度 0.78
cs.CL
-
Do LLM agents act on the reasoning they state? This question of process fidelity is central to LLM-based social simulation, yet hard to measure where no reference for correct behavior exists. We study it in a controlled setting: a Texas Poker simulator with a …
arxiv
Yufeng Wang
2026-05-30T02:02:21Z
置信度 0.78
cs.AI
-
Large language models (LLMs) are evolving into autonomous decision-makers, raising concerns about catastrophic risks in high-stakes scenarios, particularly in Chemical, Biological, Radiological and Nuclear (CBRN) domains. Based on the insight that such risks c…
arxiv
Rongwu Xu, Xiaojian Li, Shuo Chen, Wei Xu
2025-02-17T02:11:17Z
置信度 0.78
cs.CLcs.AIcs.CRcs.CY
-
Large language model (LLM) agents are increasingly built less by changing model weights than by reorganizing the runtime around them. Capabilities that earlier systems expected the model to recover internally are now externalized into memory stores, reusable s…
arxiv
Chenyu Zhou, Huacan Chai, Wenteng Chen, Zihan Guo 等
2026-04-09T13:19:41Z
置信度 0.78
cs.SEcs.MA
-
Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify verification, the ability to determine the correctness of a solution, as a new scaling axis. To unloc…
arxiv
Jacky Kwok, Shulu Li, Pranav Atreya, Yuejiang Liu 等
2026-07-06T17:59:35Z
置信度 0.78
cs.AIcs.CLcs.LGcs.MAcs.RO
-
As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, prior tool-selection studies focus on safety-agnostic metadata preferences, leaving privilege-sensitive choices underexpl…
arxiv
Kaiyue Yang, Yuyan Bu, Jingwei Yi, Yuchi Wang 等
2026-06-18T09:54:48Z
置信度 0.78
cs.SEcs.AIcs.CL
-
This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) evaluation and AI safety: benchmark scores, reward signals, and safety metrics can improve while the capabilities and align…
arxiv
Buğra Alperen Uluırmak, Rifat Kurban
2026-06-29T12:33:06Z
置信度 0.78
cs.AIcs.CLcs.LGcs.SE
-
Graph-based Retrieval-Augmented Generation (GraphRAG) has become the important paradigm for enhancing Large Language Models (LLMs) with external knowledge. However, existing approaches are constrained by their reliance on high-quality knowledge graphs: manuall…
arxiv
Xiaojun Wu, Cehao Yang, Xueyuan Lin, Chengjin Xu 等
2025-09-26T00:13:10Z
置信度 0.78
cs.CL
-
We present MACLA, a framework that decouples reasoning from learning by maintaining a frozen large language model while performing all adaptation in an external hierarchical procedural memory. MACLA extracts reusable procedures from trajectories, tracks reliab…
arxiv
Saman Forouzandeh, Wei Peng, Parham Moradi, Xinghuo Yu 等
2025-12-22T01:56:28Z
置信度 0.78
cs.LGcs.AI
-
LLM recommendation agents increasingly produce structured recommendation reports: sets of items accompanied by natural-language justifications. Yet existing evaluations often reduce this setting to reranking small shortlisted candidate sets or judge reports ma…
arxiv
Imad Aouali, Flavian Vasile, Otmane Sakhi, Alexandre Gilotte 等
2026-05-11T18:55:32Z
置信度 0.78
cs.IRcs.AIcs.LG
-
Multi-agent LLM systems on edge devices face a memory management problem: device RAM is too small to hold every agent's KV cache simultaneously. On Apple M4 Pro with 10.2 GB of cache budget, only 3 agents fit at 8K context in FP16. A 10-agent workflow must con…
arxiv
Yakov Pyotr Shkolnikov
2026-02-17T05:46:20Z
置信度 0.78
cs.LGcs.AI
-
Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely treated as a first-class architectural object. This paper names that boundary the stochastic-deterministic boundary (SDB): a …
arxiv
Vasundra Srinivasan
2026-05-19T17:54:21Z
置信度 0.78
cs.AIcs.SE
-
Traditional approaches -- such as Win Probability Added (WPA)-based ranking or computer vision-driven event detection -- can identify scoring plays but often miss strategic depth, momentum shifts, and storyline progression. Manual curation remains the gold sta…
arxiv
Jeonghun Kang, Soonmok Kwon, Joonseok Lee, Byung-Hak Kim
2025-06-03T01:10:20Z
置信度 0.78
cs.CLcs.AIcs.CV
-
LLM agents increasingly present as conversational collaborators, yet human--agent teamwork remains brittle due to information asymmetry: users lack task-specific reliability cues, and agents rarely surface calibrated uncertainty or rationale. We propose a task…
arxiv
Xingrui Gu
2026-03-11T17:35:44Z
置信度 0.78
cs.HC
-
Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable, the days must be physically traversable, …
arxiv
Jinhu Qi, Wentao Zhang, Siu Man Ng, Feiyang Xu 等
2026-07-29T14:35:29Z
置信度 0.78
cs.CL
-
Agentic intelligence in large language models (LLMs) requires not only model intrinsic capabilities but also interactions with external environments. Equipping LLMs with computers now represents a prevailing trend. However, the computer environment's intrinsic…
arxiv
Daixuan Cheng, Shaohan Huang, Yuxian Gu, Huatong Song 等
2026-01-22T18:57:09Z
置信度 0.78
cs.CLcs.AI