Abstract
This scoping review systematically maps recent peer-reviewed research on Agentic AI, a system-level extension of generative AI in which large language models (LLMs) and other foundation models are embedded within autonomous or semi-autonomous workflows involving planning, tool or API use, persistent memory, external environment interaction, reflection or self-correction, and multi-agent coordination. Despite rapid growth in this area, the literature remains fragmented across domains and uses inconsistent terminology, including LLM agents, language agents, tool-using LLMs, RAG agents, workflow agents, and multi-agent LLM systems. The primary purpose of this review is to provide a structured, evidence-based overview of how Agentic AI is conceptualized, where it is being applied, what architectural and technical patterns are emerging, what risks and limitations are reported, and what future research directions are proposed. The review treats Agentic AI as a system-level paradigm rather than a single algorithm, model architecture, or prompting technique. Systematic electronic searches were conducted across three academic databases, IEEE Xplore, Scopus, and ScienceDirect, for studies published between January 2017 and February 2026. A total of 3,549 records were identified. Following programmatic duplicate removal (n = 241), title and abstract screening (n = 3,308 screened, 2,961 excluded), and full-text eligibility assessment (n = 347 evaluated, 184 excluded), a final corpus of 163 strict-included peer-reviewed studies was retained for data extraction and thematic synthesis. Study selection followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. A system was classified as Agentic AI when it combined a core foundation model for reasoning with autonomous multi-step task execution, enabled by at least one of the following structural mechanisms: automated tool or API use, persistent memory architectures, external environment interaction, active reflection or self-correction, or multi-agent coordination frameworks. This conservative operational definition was used to distinguish Agentic AI systems from passive LLM applications, classical non-foundation-model agents, and purely conceptual frameworks. Expected outcomes include: (1) an operational definition of Agentic AI grounded in empirical evidence from 163 peer-reviewed studies; (2) a cross-domain map of application concentrations, showing that the field is currently strongest in cybersecurity and privacy compliance, networks and telecommunications, robotics and embodied AI, software engineering and EDA, and knowledge and multimodal analytics; (3) a synthesis of six emerging technical trends including governance and safety middleware, collaborative and hierarchical multi-agent architectures, retrieval-augmented and memory-enabled agency, executable reasoning, closed-loop reflection and validation, and multimodal and embodied interaction; (4) a structured risk taxonomy spanning safety-critical reliability, hallucination and grounding failures, privacy and compliance risks, agent coordination failure, security and abuse risks, operational cost and latency, and bias and human oversight concerns; and (5) a set of priority future research directions including real-world longitudinal validation, standardized agentic benchmarks, runtime governance and safety assurance, secure memory management, and human-agent collaboration frameworks. This review is retrospectively registered as required by the target publication venue. All authors are affiliated with the Department of Computer Science and Engineering, The University of Aizu, Aizuwakamatsu, Fukushima, Japan.