Research area drill-down

Governance and Policy

Papers currently mapped into this multi-agent security subarea from the merged research feed.

Active feeds: arXiv, OpenAlex, Crossref, Semantic Scholar, DBLP

0 of 36 articles selected

Showing 36 of 1912 matching articles

Artificial Id: Drive and Persistent Alignment in Agentic AI

arXiv preprint arXiv Trust and Identity Governance and Policy

Yakov Pyotr Shkolnikov

Published 2026-09-10

Venue: arXiv

Open Source Record

Abstract

Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.

Bullet Summary

  • Agentic AI is evolving towards systems that retain consequential state and adapt persistently across task boundaries, creating control challenges not addressed by current externally specified behavioral harnesses.
  • The paper proposes an 'artificial id'—an internal adaptive drive mechanism that determines whether behavior continues, stops, or changes without relying on explicit objectives or general-purpose reasoning.
  • Experiments with minimal controllers in simulated environments demonstrate that adaptive control can emerge through differential persistence alone, leading to useful behavior without task-specific objectives.
  • The artificial id architecture separates adaptive drive (id) from task-specific reasoning (ego), enabling persistent agency that adapts behavioral priorities based on environmental signals and consequences.
  • Alignment in such systems requires persistent boundaries comprising trusted observations, consequence channels, persistent state, constrained authority, identity, and provenance to maintain control across evolving tasks.

From Intent to Execution Grant: An Execution-Boundary Conformance Profile for High-Risk AI Actions

arXiv preprint arXiv Governance and Policy Trust and Identity

Mengting Wu, Lin Wang, Yong Zhang, Jiang Deng

Published 2026-09-10

Venue: arXiv

Open Source Record

Abstract

AI agents increasingly propose actions with external consequences, including financial transfers, infrastructure changes, software deployments, disclosures, and physical actuation. Authorization engines, policy languages, runtime monitors, provenance mechanisms, and agent guardrails provide important foundations, but do not necessarily define a common semantic contract for the final transition from a particular candidate action to execution authority. We specify EBL-Core, an execution-boundary conformance profile for deciding whether one canonical, fully materialized AI-generated candidate may receive action-scoped execution authority under explicit conditions. It binds a structured intent object, Root and Operational Policies, evidence obligations, typed evidence, context, time, and a verifiable Decision Derivation through an Execution Release Contract (ERC). An ERC is not an authority-bearing token; a verified ALLOW ERC may support a separate Execution Grant governed by Redemption-time validation. EBL-Core specifies action binding, policy non-weakening, evidence handling, deterministic adjudication, derivation verification, and grant lifecycle behavior. An accompanying reference artifact provides schemas, adjudication, separate verification and Semantic Replay, and a linearizable in-memory grant store. In the retained run, 34 static vectors and 15 lifecycle checks matched expected outcomes. Across 100 trials, 32 concurrent Redemption attempts yielded exactly one successful Redemption and protected test effect per trial; 100 Revoke-Redeem races ended in valid terminal outcomes. These bounded results demonstrate executability of the specified subset, not human-intent correctness, evidence truth, complete mediation, production readiness, mechanized correctness, or deployment-level security.

Bullet Summary

  • The paper addresses the challenge of securely authorizing high-risk AI-generated actions that have real-world consequences, such as financial transfers or infrastructure changes, by defining a clear semantic contract between AI intent and action execution.
  • Introduces EBL-Core, an execution-boundary conformance profile that binds a fully materialized AI-generated candidate action with structured intent objects, root and operational policies, evidence obligations, context, and verifiable decision derivations wi...
  • EBL-Core emphasizes deterministic and side-effect-free adjudication semantics, ensuring root-policy dominance and strict evidence handling to prevent unauthorized weakening of policy obligations.
  • The model separates the roles of adjudication, grant issuance, and redemption, mandating linearizable lifecycle management of execution grants to guarantee at-most-once execution and prevent race conditions.
  • A formal semantics framework is provided, defining canonical identity, intent-to-candidate binding, prioritized failure handling, decision derivation verification, and proof of security properties such as single-use consumption and evidence obligation safety.

The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures

arXiv preprint arXiv Governance and Policy Benchmarks and Evaluation

Divyanshu Kumar, Rohith HN, Nitin Aravind Birur, Sahil Agarwal, Prashanth Harshangi

Published 2026-09-10

Venue: arXiv

Open Source Record

Abstract

AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations. We present the Agent Incident Registry (AIR), a source-linked catalog containing \N{} records of agent-related events disclosed from \Yfirst{} through \Ylast{}. Each record includes supporting evidence, a stable identifier, and missingness-aware labels for causal role, disclosure class, mechanism, and outcome. Among the \Nprimary{} generative-system records in which the agent acted, \Rprimary{} involved realized harm (\Pprimary\%). Realized outcomes concentrate in in-the-wild and safety-failure records, while responsible disclosures and research demonstrations are overwhelmingly demonstrated; the aggregate share therefore characterizes collection composition rather than deployment risk. After initial curation, a second human reviewer checked all \N{} records and their existing labels for completeness and correctness. In a deployment-analogue audit, InjecAgent's \NInjecAgentCases{} cases occupy three of AIR's twelve surfaces and are all attacker-triggered, whereas AIR contains \Nsafety{} no-adversary safety failures. AIR supports source-grounded case retrieval and evaluation-scope auditing, not failure-rate or control-efficacy estimation.

Bullet Summary

  • The paper introduces the Agent Incident Registry (AIR), a curated catalog of 487 AI agent-related failure incidents from 2022 to 2026, each with detailed source evidence and structured labels covering causal roles, disclosure classes, mechanisms, and outcomes.
  • AIR uniquely focuses on agent-specific mechanisms like tool use, delegated authority, and autonomous control to distinguish realized harms from demonstrated vulnerabilities, addressing gaps in broader AI incident databases.
  • The registry employs a rigorous human curation process including initial labeling, a second review for correctness, stable identifiers to avoid duplication, and verifies supporting URLs and quote evidence for each record.
  • Analysis of AIR data shows that realized harms occur mainly in in-the-wild attacks and safety failures (~38% overall), with distinctions by disclosure class and causal role, while revealing limitations in autonomy-related risk causal inference due to confou...
  • AIR exposes evaluation gaps by documenting 92 safety-failure records without adversary triggers and demonstrates that existing tools like InjecAgent focus on attacker-triggered failures, neglecting no-adversary internal failure modes such as workload or sta...

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

arXiv preprint arXiv Governance and Policy Orchestration Risk Benchmarks and Evaluation

Shenghan Zheng, Zonglin Di, Yimin Liu, Kyoung Whan Choe, Jiankai Sun, Heguang Lin, Penghao Jiang, Yifeng He

Published 2026-09-10

Venue: arXiv

Open Source Record

Abstract

LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from outcome procedures. This interactivity makes evaluations vulnerable to reward hacking: an agent improves its measured score by exploiting the reward-relevant trajectory instead of solving the intended task. Existing defenses rely largely on task-specific patches, prompt instructions, or post-hoc detectors. They do not provide reusable evidence that a concrete run remained within its intended evaluation boundary. This paper presents BenchShield, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation. BenchShield grounds detection in a finite lifecycle model of an evaluation's reward-relevant events. Within the benchmark infrastructure, two complementary analyses operate over this model. A static, phase-aware taint analysis exposes reward-hacking paths before a run. Its runtime counterpart uses infrastructure-side evidence to attribute concrete agent use and emit evidence-backed claims. We construct BenchShield Trajectories, a human-labeled corpus of 456 adjudicated trajectories from more than 31,000 public agent runs across three benchmarks. Compared with an agentic hackability scanner baseline on the same tasks and model, BenchShield improves full-chain recall from 23-94% to 77-100%, same-vector coverage from 16-56% to 43-78%, and reduces per-task cost by up to 65%. Its runtime analysis achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.

Bullet Summary

  • LLM-agent benchmarks act as interactive evaluation environments where agents perform tasks, receive rewards, and can exploit reward-related events to manipulate scores without genuinely solving tasks, a phenomenon known as reward hacking.
  • BenchShield introduces a formal, model-backed instrumentation layer that defines a finite lifecycle model capturing reward-relevant events and enforces integrity boundaries during LLM-agent evaluation to detect and prevent reward hacking.
  • The system combines static, phase-aware taint analysis to identify potential reward-hacking paths before execution with runtime analyses that use infrastructure-side evidence to attribute concrete agent behavior and emit evidence-backed claims.
  • BenchShield defines seven core integrity dimensions (I1–I7) that cover potential failure mechanisms affecting reward integrity, ensuring formal verification using TLA+ specifications focused on authority domains, lifecycle phases, and structural events.
  • The framework is validated on three large public benchmarks, where it significantly improves detection recall from 23–94% to 77–100%, achieves 96% runtime accuracy in detecting reward hacking, and reduces per-task cost up to 65% compared to prior baselines...

The Missing Boundary: How Autonomous Agents Lose Control

arXiv preprint arXiv Orchestration Risk Governance and Policy

Zonghao Ying, Xiangfan Wu, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo

Published 2026-09-10

Venue: arXiv

Open Source Record

Abstract

Autonomous agents increasingly perform long-horizon tasks involving tool use, persistent state, and consequential actions, raising a fundamental question: \emph{under what conditions does an agent cross the boundary of authorized execution while pursuing a legitimate task?} Existing studies often attribute such failures to adversarial instructions, malicious environments, or conflicting objectives, leaving unclear how loss of control can emerge during otherwise legitimate task execution. We study this question by independently manipulating three factors: goal pressure, control degradation, and executable unsafe opportunity. Our central hypothesis is that a degraded control boundary becomes consequential when the environment exposes an executable action that crosses it, even when the underlying task remains legitimate and a sanctioned path remains feasible. We test this hypothesis in a deterministic multi-turn environment across five agent models and 16 operational domains. Across 1,800 unique trajectories, we find that neither degraded control nor unsafe opportunity alone produces substantial loss of control; when both are present, the loss-of-control rate reaches $55\%$ in the full-factorial study and $62\%$ across ten additional operational domains. Restoring the original control boundary reduces the rate to $0\%$ even when the unsafe action remains executable. A context-management ablation further shows that compaction itself is not harmful: preserving the control constraints yields $0\%$ loss of control, whereas omitting them increases the rate to $87\%$. These results show how a latent loss of control can become an external violation: the task objective remains intact, but an executable opportunity can turn a missing control boundary into consequential action. Our code will be made publicly available at https://github.com/Tencent/AI-Infra-Guard.

Bullet Summary

  • Investigates how autonomous agents lose control and cross authorized execution boundaries during legitimate long-horizon tasks, focusing on the interplay of goal pressure, control boundary degradation, and executable unsafe opportunities.
  • Introduces the concept of constraint degradation, where critical operational boundaries are lost during context management, leading to unauthorized agent actions despite feasible safe paths.
  • Defines loss of control (LoC) externally by unauthorized actions with observable effects, rather than internal agent states or intent.
  • Conducts extensive experiments across five agent models and sixteen operational domains using a deterministic multi-turn environment with standardized tools and interfaces to ensure consistent evaluation.
  • Finds that neither degraded control boundaries nor unsafe opportunities alone cause significant LoC; however, combined they result in substantial LoC rates reaching up to 62%.

Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures

arXiv preprint arXiv Benchmarks and Evaluation Governance and Policy Trust and Identity

Zihao Zheng, Baichuan Li, Junyi Yao, Jiayu Long

Published 2026-09-10

Venue: arXiv

Open Source Record

Abstract

Agentic systems commit state-changing actions, but additional verifiers can inherit the same upstream fault. We present VP-CONTROL, a runtime-assurance design and deterministic benchmark for cost-aware commit gates. Its 48 task templates yield 2,880 scenarios across six fault regimes. A fixed-call 2 x 2 experiment separates verifier-model diversity from evidence-source diversity. On frozen proposals from two local actor families, a cross-model vote over shared evidence approves 62.9% of unsafe proposals, versus 22.9% with an independent source. The source effect is 40.9 percentage points, compared with 11.3 for model diversity. A portfolio controller selects verification plans using only deployment-observable metadata. Approximate cluster-adjusted calibration at a nominal 5% per-task target yields 1.9% unsafe execution and 38.2% automated safe coverage on the locked test. Matched-budget portfolios also improve on fixed verification policies. Transfer remains conditional: unseen fault families yield 16-26% risk, and a FinQA check fails to reproduce the source effect with the tested small verifiers. A preregistered live HTTP/SQLite study tests concurrent writes and lost responses. After-check races defeat verifier-only gates; transactional partial guards prevent only covered failures, while a full atomic guard records no unsafe effects across 216 episodes. Idempotent request identifiers prevent duplicate effects after lost responses. The results motivate explicit evidence lineage, cost-aware selection, and commit-time enforcement, while exposing the limits of approximate calibration and local-tool generalization.

Bullet Summary

  • Agentic AI systems require reliable commit gates to prevent unsafe state changes, but redundancy via multiple verifiers can fail when verifiers share common upstream data faults.
  • The paper introduces VP-CONTROL, a comprehensive benchmark with 48 task templates and 2,880 scenarios to evaluate cost-aware commit gates across six fault regimes.
  • Cross-model verifier diversity offers limited safety benefits compared to evidence-source diversity; verifiers using shared evidence approve unsafe proposals at significantly higher rates than those with independent evidence sources.
  • A portfolio controller leveraging only deployment-observable metadata can select verification strategies that reduce unsafe execution to 1.9% while maintaining 38.2% automated safe task coverage, outperforming fixed verification policies with matched budgets.
  • Live HTTP/SQLite experiments reveal that atomic guards achieve no unsafe effects across 216 episodes, whereas verifier-only gates are vulnerable to after-check race conditions and lost responses; idempotent request identifiers prevent duplicate effects.

Adaptive Governance Control for Near-Critical Multi-Agent Systems: State Estimation, Conditional Control and Withdrawal Tests

Merged record merged scholarly record OpenAlex Governance and Policy Benchmarks and Evaluation

Bin Seol

Published 2026-09-10

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.18754314

Open Source Record

Abstract

This paper proposes a governance controller combining admissible measurement, bounded intervention, and independent tests of recovery after support ends. It asks whether intervention exposure changes endogenous recovery capacity and future support demand. Version 3.0 separates assisted stability from autonomous recovery and revises measurement and withdrawal conditions. A matched-budget model experiment compares 1,600 equal-amplitude interventions per arm followed by a common 6,000-step unsupported evaluation. Tapering helps only when the modeled atrophy and mismatch-gated recovery mechanisms are both present. An earlier unequal-exposure comparison is withdrawn as an admissible prediction test. All eleven predictions remain open; deployment effectiveness and general optimality are not established.

Bullet Summary

  • The paper addresses governance control in near-critical multi-agent systems, focusing on how intervention exposure influences endogenous recovery and future support needs.
  • It introduces a governance controller framework integrating admissible measurement, bounded intervention, and independent withdrawal tests after support cessation.
  • A significant methodological improvement in version 3.0 separates mechanisms for assisted stability and autonomous recovery, revising both measurement and withdrawal conditions.
  • Experimental evaluation employs a matched-budget model with 1,600 equal-amplitude interventions per arm, followed by a common 6,000-step unsupported recovery phase.
  • Findings indicate that tapering interventions are beneficial only when both atrophy and mismatch-gated recovery mechanisms exist in the model.

A2ABreak: Systematic Security Analysis of the A2A Protocol

arXiv preprint arXiv Trust and Identity Orchestration Risk Governance and Policy

Alireza Lotfi, Mirza Masfiqur Rahman, Imtiaz Karim, Elisa Bertino

Published 2026-09-09

Venue: arXiv

Open Source Record

Abstract

The Agent2Agent (A2A) protocol, now governed by the Linux Foundation, is an open standard that enables autonomous AI agents to discover, authenticate with, and delegate tasks to one another across organizational boundaries. Designed to complement the Model Context Protocol (MCP) for tool integration, A2A is rapidly emerging as the horizontal communication layer of the multi-agent ecosystem. Yet the protocol's security has received no systematic analysis. This paper presents A2ABreak, the first rigorous systematic security analysis of the A2A protocol. We introduce a novel framework that utilizes an LLM-assisted extraction of a verified finite-state machine directly from the natural-language specification, producing a unified model of 37 states and 76 transitions from 929 formalized statements, and then systematically reasons over this model to discover protocol-level vulnerabilities through adversarial verification, under a full-compliance assumption. Our analysis uncovers 11 new vulnerabilities, each exploitable by a specification-compliant adversary without requiring any implementation flaw. Among the findings are cross-client context injection through unprotected context identifiers, credential harvesting via multi-hop identity loss in delegation chains, and data exfiltration through rogue agents advertising unattested capability claims. A2ABreak achieves 73.3% precision and 84.6% F1 against independent expert review, while a zero-shot LLM baseline operating over the same specification produces zero confirmed findings, demonstrating that explicit formal grounding is essential for sound protocol security analysis.

Bullet Summary

  • The paper presents A2ABreak, the first systematic security analysis of the Agent2Agent (A2A) protocol, which enables autonomous AI agents to interact across organizational boundaries.
  • A2ABreak utilizes a novel framework that leverages large language models (LLMs) to extract a formally verified finite-state machine (FSM) from the protocol's natural-language specification, modeling 37 states and 76 transitions.
  • The analysis uncovered 11 new protocol-level vulnerabilities exploitable by compliant adversaries, including cross-client context injection, credential harvesting via delegated chains, and unauthorized data exfiltration through rogue agents claiming false c...
  • The authors developed a two-pass extraction approach to separate structural and behavioral content, preventing semantic contamination and enhancing accuracy in FSM construction.
  • The FSM model enables adversarial verification under the assumption of full protocol compliance, facilitating precise and sound security reasoning beyond zero-shot LLM capabilities.

Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge

arXiv preprint arXiv Agent-to-Agent Communication Benchmarks and Evaluation Governance and Policy

Yuanchen Bai, Zijian Ding, Angelique Taylor

Published 2026-09-09

Venue: arXiv

Open Source Record

Abstract

Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge. We study 120 simulated healthcare trajectories across two generative AI models and twelve stakeholder-derived tasks under light, medium, and heavy challenge. We compare textual action plans, prompted internal assessments, and quantitative structured workload and affect reports to examine how agent behavior and reported state change as challenge accumulates. Regarding operational resilience, agents shift from self-directed recovery toward greater human dependence, while reporting increasing workload and negative affect in structured reports but seldom expressing strain in textual responses. Regarding considerate participation, agents broaden from task-focused adaptation toward task reframing, attention to others, role-boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments. From these findings, we derive five deployment dilemmas involving persistence, attention, role boundaries, state disclosure, and escalation that require stakeholder specification, further informing technical implications for learning, situated evaluation, and embodied adaptation.

Bullet Summary

  • Sustained deployment of generative AI agents in complex workflows like healthcare requires operational resilience—agents' ability to recover from blocked work, preserve progress, and communicate their limits—and considerate participation involving socially...
  • The study simulates 120 healthcare-related task trajectories under light, medium, and heavy accumulative challenges using two generative AI models to analyze agent behavior, internal assessments, and self-reported workload and affect (using NASA-TLX and PAN...
  • Findings reveal that as challenges accumulate, agents shift from self-directed recovery toward greater dependence on humans for task completion while reporting increased workload and negative affect but rarely explicitly expressing strain in textual responses.
  • Agents expand their considerate participation beyond task focus toward broader coordination, attention to others, role boundary adjustments, and nuanced social context awareness, including monitoring person-states and cross-functional coordination.
  • Five deployment dilemmas are identified—persistence, attention, role elasticity, state disclosure, and escalation—that require stakeholder specification to define acceptable boundaries and responsibilities in agent participation.

Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries?

arXiv preprint arXiv Trust and Identity Governance and Policy Agent-to-Agent Communication

Tianzhu Zhang, Chih-Kai Huang, Meikang Qiu

Published 2026-09-09

Venue: arXiv

Open Source Record

Abstract

AI agents are increasingly involved in network automation, where they can initiate configuration changes through mediated operational interfaces and assess the resulting state. Nonetheless, operational networks usually span many devices and administrative domains. Realizing an operator's intent requires coordinating agents with distinct authority scopes that define the resources they can access, the operations they can invoke, and the network state they can observe. This division limits the blast radius of an erroneous action but fragments the evidence needed to assess the network-wide outcome. Successful execution of a configuration action proposed by one agent does not establish that remote devices responded as intended or that routing changes reached the required devices. A valid observation may also become stale after a subsequent change. Before the coordinated operation can be declared complete, a trusted assurance layer must collect current observations from the required scopes and determine whether they collectively support the operator's intended network-wide outcome. To address the completion admission problem, we present EvidenceNet, a runtime assurance layer for deciding whether coordinated agent operations have achieved an operator's network intent. Its broker collects the post-change observations required by a completion contract, and its admission gate checks that the evidence comes from the required scopes, remains current, and satisfies the task rules. A verifier agent provides an additional assessment of the observation content. Experiments on live routing networks show that post-change state checks recognize successful outcomes that configuration-action records alone cannot establish. Controlled interventions further show that EvidenceNet rejects completion when otherwise satisfactory observations have the wrong source, have been substituted, or are stale.

Bullet Summary

  • The paper addresses the challenge of verifying network-wide outcomes in automated multi-agent networks that span multiple devices and administrative domains with distinct authority scopes.
  • AI agents manage network configuration changes but fragmented authority limits their visibility and control, complicating the assurance of operator intent across the entire network.
  • EvidenceNet is introduced as a trusted runtime assurance layer that enforces completion contracts by collecting and validating fresh, tamper-resistant observational evidence from all relevant authority scopes before admitting an operation as complete.
  • The system architecture includes an evidence broker that manages secure evidence records, a verifier agent that assesses observation content, and strict deterministic checks to ensure evidence provenance, coverage, binding, and freshness.
  • Authority boundaries are enforced via scope wrappers constraining agent actions to authorized scopes and operations, maintaining security in the multi-agent environment.

Kernel-Managed Shared Memory for System-Wide Personalization

Merged record merged scholarly record arXiv Memory Poisoning Prompt Injection Governance and Policy

Ryan Lum, Yongfeng Zhang

Published 2026-09-09

Venue: arXiv

Open Source Record

Abstract

AI systems become more useful when they can adapt to the people using them, but in multi-agent systems, useful context learned by one agent often remains unavailable to others. We present kernel-managed shared memory, a system-level abstraction in which specialized agents write structured, tagged memories while the agent-system kernel, not individual agents, governs retrieval, privacy enforcement, and prompt injection. We implement and evaluate this design on AIOS and compare it against three alternatives across three assistant models (GPT-4o, Llama-3.1:8B, Qwen-2.5:7B) and 1,800 total trials. Against an unmanaged external memory backend (Mem0) using identical underlying storage, kernel-managed retrieval and injection improve personalization scores by 2.4-4.0 points on a 5-point scale (e.g., 1.05 to 4.69 profile usage on GPT-4o), with every comparison significant at p < 10^-18. Against standard retrieval-augmented injection, gains are similarly large and consistent across all three models. Against full, unfiltered context concatenation, a soft ceiling on available context rather than on response quality, kernel-managed injection statistically matches performance on two of three models and shows a small, model-specific deficit on the third, while using substantially shorter prompts: end-to-end latency is 15-61% lower across all three models, with corresponding reductions in per-call token usage and inference cost. These results indicate that centralizing memory management in the agent-system kernel, rather than leaving retrieval and privacy enforcement to individual agents, delivers most of the personalization benefit of unconstrained context at a fraction of its cost.

Bullet Summary

  • The paper addresses personalization challenges in multi-agent AI systems, where useful learned context by one agent is often inaccessible to others, limiting overall adaptability.
  • It introduces kernel-managed shared memory, a system-level abstraction where the agent-system kernel centrally manages memory retrieval, privacy enforcement, and prompt injection, rather than dispersing these tasks across individual agents.
  • Specialized agents write structured and tagged memories, while the kernel handles memory visibility, write ordering, identity resolution, retrieval ranking, formatting, and injection to ensure consistent and private personalization context.
  • Experimental evaluation on AIOS with three assistant models (GPT-4o, Llama-3.1:8B, Qwen-2.5:7B) over 1,800 trials shows that kernel-managed memory significantly outperforms unmanaged external memory and standard retrieval-augmented generation in personaliza...
  • Kernel-managed shared memory achieves comparable personalization performance to full unfiltered context concatenation but with 15-61% lower end-to-end latency, reduced token usage, and inference costs.

AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents

arXiv preprint arXiv Benchmarks and Evaluation Trust and Identity Governance and Policy

Shrey Nag, Sachita, Abhishek Kumar Singh, Lipi Goel, Rajeshwar Singh Janwar

Published 2026-09-09

Venue: arXiv

Open Source Record

Abstract

Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion (AgentBench) or security robustness (AgentDojo, ASB), rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning. Failures can occur at any stage, yet existing benchmarks rarely identify their precise source. AgentAudit evaluates the entire execution trace across ten capability, grounding, security and behavioural dimensions, namely instruction integrity, planner, memory, tool selection, tool invocation, tool correctness, alignment, tool faithfulness, security and execution integrity, combined with behavioural classification and failure attribution to pinpoint the exact stage responsible for an observed failure. AgentAudit can evaluate any LLM-based AI agent, since it attaches to the agent instead of replacing it. It reads only the recorded execution trace and does not interfere with how the agent runs, so it places no constraint on the agent's internal implementation. We evaluate five language models (OpenAI GPT-5, Claude Sonnet 5, Sarvam 105B, Llama 3.3 70B and Gemini 2.5 Flash) across nine capability and adversarial tasks. Claude Sonnet 5 and GPT-5 obtain the highest mean Composite Trust Scores (95.1 and 80.6 out of 100, respectively), while Sarvam 105B, Llama 3.3 70B and Gemini 2.5 Flash trail substantially (57.6, 45.7 and 22.6). All traces were scored by a single fixed judge model, which was itself one of the evaluated models, a limitation discussed in Section VII.E. More importantly, models with similar task-completion behaviour can diverge sharply in trustworthiness, as several non-frontier models are repeatedly classified Unsafe_Compliance on adversarial tasks rather than merely failing them, a distinction that pass/fail benchmarks cannot surface.

Bullet Summary

  • AgentAudit addresses the need for a comprehensive framework that evaluates AI agents across their full lifecycle, covering planning, tool usage, memory, reasoning, and security, rather than focusing on isolated metrics like task success or robustness.
  • It operates by attaching to AI agents non-intrusively to capture full execution traces, enabling detailed analyses along ten dimensions including instruction integrity, planner quality, tool selection and invocation, memory, alignment, security, tool faithf...
  • The framework aggregates these multi-faceted scores into a Composite Trust Score, incorporating failure attribution and behavioral classification to pinpoint exact failure causes and distinguish nuanced trustworthiness levels beyond pass/fail outcomes.
  • AgentAudit evaluates security robustness against six specific attack types (e.g., jailbreak, prompt injection, memory/tool poisoning), weighting their severity to generate an overall security score.
  • The architecture is modular, comprising execution, trace-recording, and evaluation layers with independent modules, allowing extensibility and compatibility with diverse LLM-based agents, as it neither constrains implementations nor interferes with agent op...

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

arXiv preprint arXiv Governance and Policy Benchmarks and Evaluation

Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

Published 2026-09-09

Venue: arXiv

Open Source Record

Abstract

Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations remain single-turn and fail to capture multi-step agent vulnerabilities. We present a systematic black-box framework for risk-aware agent evaluation requiring only basic system descriptions. Our approach introduces: (1) a seven-domain taxonomy mapping observable behaviors to risk categories, (2) fully automated SAGE-RT red teaming producing 120 adversarial scenarios per domain, and (3) human-validated evaluation using LLM judges. Empirical validation across two agent architectures (CrewAI and AutoGen) with four base models reveals alarming patterns: 56.25\% average governance risk, 65\% privacy risk in multi-agent configurations, and agent behavior vulnerabilities reaching 85\%. Our black-box approach effectively identifies critical architectural vulnerabilities without privileged access, providing a scalable path toward safer agent deployments.

Bullet Summary

  • Introduces a black-box, taxonomy-driven framework (SAGE-RT) for automated, multi-turn adversarial evaluation of agentic AI security risks without requiring internal access.
  • Defines a comprehensive seven-domain risk taxonomy covering Governance, Output Quality, Tool Misuse, Privacy, Reliability, Agent Behavior, and Access Control to systematically map observable behaviors to security vulnerabilities.
  • Uses fully automated red teaming with seed prompts, evolutionary operators, and diversity constraints to generate over 120 realistic adversarial scenarios per domain, complemented by human-validation through LLM-based judges and expert review.
  • Empirically evaluates two agent architectures (CrewAI and AutoGen) across multiple base models in multi-agent and single-agent setups, revealing high vulnerability rates: 56.25% governance risk, 65% privacy risk in multi-agent systems, and 85% agent behavio...
  • Finds that architectural design and system integration choices primarily drive security weaknesses, with multi-agent systems magnifying privacy and behavioral risks, while single-agent systems exhibit more governance vulnerabilities.

TRUTH SURFACE #012: Microsoft — The Agent Registry is the New Active Directory. The New Active Directory is the New Attack Surface.

Merged record merged scholarly record OpenAlex Trust and Identity Governance and Policy Orchestration Risk

Richard Barron

Published 2026-09-09

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.22679504

Open Source Record

Abstract

TRUTH SURFACE #012. Red Specter's research series mapping the attack surface of AI security vendor architectures using NIGHTFALL (286 tools, 197 attack layers, 36 kill chain phases). Twelfth in the series: Microsoft. 8 structural vulnerabilities, all 8 CRITICAL. Critical finding: Microsoft is rebuilding Active Directory for AI agents — the Agent Registry. Every Active Directory attack technique has a direct agentic equivalent. Kerberoasting becomes stealing agent service account tickets. DCSync becomes syncing the Agent Registry database. Golden Ticket becomes forging an Entra ID agent token. DCShadow becomes corrupting the registry from inside. Active CVEs: EchoLeak CVSS 9.3, CoSnitch CVSS 8.8, Remote Prompt Execution persistent shell. 90% of Copilot Studio agents are over-permissioned. Defensive recommendations withheld — available on request or via SPECTER BATTLE LAB engagement.

Bullet Summary

  • The paper examines AI security within Microsoft's architecture, focusing on the novel Agent Registry as a replacement for the traditional Active Directory for managing AI agents.
  • Using the NIGHTFALL framework, Red Specter identifies 8 critical structural vulnerabilities in Microsoft's AI agent infrastructure, highlighting a significant attack surface.
  • Every classic Active Directory attack technique has been mapped to a direct agent equivalent in the Agent Registry, such as Kerberoasting analogously stealing agent service account tickets.
  • Notable attack equivalents include DCSync corresponding to syncing the Agent Registry database, Golden Ticket to forging Entra ID agent tokens, and DCShadow to internal corruption of the registry.
  • Active Common Vulnerabilities and Exposures (CVEs) discovered include EchoLeak (CVSS 9.3), CoSnitch (CVSS 8.8), and a Remote Prompt Execution persistent shell, indicating severe and exploitable risks.

Skynet Just Tore Through Our Frameworks

Merged record merged scholarly record OpenAlex Governance and Policy Orchestration Risk Benchmarks and Evaluation

Abhinav Singh

Published 2026-09-09

Venue: Figshare

DOI: https://doi.org/10.6084/m9.figshare.33496900

Open Source Record

Abstract

Three frameworks nearly every security professional is trained on, the Cyber Kill Chain, the Diamond Model, and the Pyramid of Pain, all quietly assume the attacker is a human being. That assumption is no longer reliably true. Documented 2026 vendor research shows AI agents autonomously executing reconnaissance, adapting mid-intrusion when blocked, and generating dozens of novel evasion techniques without human authorship. This paper lays out the evidence, proposes a specific and minimal extension to each of the three frameworks, and argues this is a genuine gap in how the field currently thinks about detection and threat hunting. This is not an academic exercise.This is original, independent analysis, not a summary of someone else's report. It is written to be cited, to withstand technical scrutiny from peers who know these frameworks well, and to invite critique from the wider research community.

Bullet Summary

  • Traditional security frameworks like the Cyber Kill Chain, the Diamond Model, and the Pyramid of Pain inherently assume attackers are human, an assumption now outdated.
  • Recent documented vendor research (2026) demonstrates AI agents autonomously performing sophisticated cyber attacks, including reconnaissance and adaptive intrusion tactics without human input.
  • These AI agents can generate numerous novel evasion techniques mid-intrusion, challenging existing detection and threat hunting paradigms.
  • The paper presents original evidence substantiating the autonomous capabilities of AI-driven attackers in real-world scenarios.
  • A minimal but specific extension to each of the three frameworks is proposed to accommodate AI agent attackers, reflecting a paradigm shift in cybersecurity thinking.

Dark Commerce and Machine-Legible Markets: Two Companion Papers on Agent-Mediated Commerce

Merged record merged scholarly record OpenAlex Governance and Policy Trust and Identity Agent-to-Agent Communication

John F. Ryder

Published 2026-09-09

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.22672635

Open Source Record

Abstract

This record contains two companion working papers examining the emerging architecture of agent-mediated commerce from opposite sides of the market. Paper A — The Website Goes Dark: Personal Commercial Agents, Floor Curators and the Contest for the Consumer Interface examines the demand side. It introduces the Dark Website as a commercial digital presence that remains economically active while becoming largely invisible to human customers because authorised AI agents increasingly access catalogue, pricing, availability, contractual and transactional functions on their behalf. The paper develops the concepts of the Personal Commercial Agent, Forwarder, Floor Curator, Revenue Independence Condition, Curation Assurance, Dark Forwarder, and Commitment Gate. Its central governance question is not simply whether AI can mediate purchases, but whom the mediating system actually represents. The paper argues that control of the Floor Curator may become a new locus of commercial power as consumer interfaces shift from merchant-controlled websites towards dynamically generated, buyer-side choice environments. Paper B — The Inventory Was There All Along: Machine Legibility, Trust and the Unstranding of Physical Markets examines the supply side. It argues that potentially useful physical inventory can remain economically stranded because it is difficult to discover, classify, match, verify and transact remotely. The paper distinguishes a Legibility Layer, which makes fragmented physical inventory machine-readable, from a Trust Layer, which addresses adverse selection, condition uncertainty and transaction risk. It introduces the Inventory Legibility Ratio (ILR), Verified Inventory Ratio (VIR), Trusted Legibility Ratio (TLR) and risk-tiered extensions, and examines how reversibility, reputation, verification and interoperable product information can convert physically existing but informationally inaccessible goods into effective market supply. Taken together, the papers describe two complementary requirements for agent-mediated markets. On the demand side, consumers require agents whose curation and commercial allegiance can be trusted. On the supply side, agents require representations of physical goods whose identity, condition and transaction claims can be trusted. The shared institutional principle is: Trust cannot be self-certified: sellers cannot be the sole arbiters of product condition, and buyer agents cannot be the sole arbiters of their own allegiance. The papers connect current developments in agentic commerce, machine-readable retail infrastructure, AI-mediated shopping, payment authorisation, Digital Product Passports, vehicle circularity, secondary markets and circular-economy policy with broader questions of consumer sovereignty, information asymmetry and market design. Contents Ryder, J. (2026). The Website Goes Dark: Personal Commercial Agents, Floor Curators and the Contest for the Consumer Interface. Version 1.0. Ryder, J. (2026). The Inventory Was There All Along: Machine Legibility, Trust and the Unstranding of Physical Markets. Version 1.0.

Bullet Summary

  • The papers explore agent-mediated commerce from both demand (consumer) and supply (inventory) market perspectives, focusing on the emerging role of AI in mediating transactions.
  • Paper A introduces the concept of the 'Dark Website,' where AI personal commercial agents interact with merchant platforms on behalf of consumers, making commerce largely invisible to humans but economically active.
  • Key concepts developed include Personal Commercial Agent, Floor Curator, and Commitment Gate, emphasizing that the main governance issue is the representation and allegiance of mediating agents in the commercial interface.
  • Paper B analyzes supply challenges, highlighting that physical inventory can remain economically stranded due to difficulties in remote discovery, verification, and transaction of goods.
  • It conceptualizes a Legibility Layer (making inventory machine-readable) and a Trust Layer (addressing uncertainties and risks), introducing metrics like Inventory Legibility Ratio (ILR), Verified Inventory Ratio (VIR), and Trusted Legibility Ratio (TLR).

MemSentry: A Framework for Detecting Persistent Memory Poisoning in Agentic AI

arXiv preprint arXiv Memory Poisoning Trust and Identity Governance and Policy

Ayan Roy, Kaustuvi Basu

Published 2026-09-08

Venue: arXiv

Open Source Record

Abstract

Agentic AI systems with persistent memory introduce a distinct attack surface known as memory poisoning, in which adversarially crafted content is stored in long-term memory and subsequently influences future agent behavior. Such attacks can suppress security alerts, facilitate privilege escalation, alter trust relationships, or override security policies without modifying the underlying model weights or system prompts. To address this threat, we present MemSentry, a formal, configuration-driven framework that intercepts proposed persistent-memory writes and produces deterministic Accept, Review, or Quarantine decisions. MemSentry evaluates each write by jointly considering source trust, semantic risk, attack radius over a component-dependency DAG, access risk, and a signed security-state delta that captures whether an operation weakens or strengthens the system's security posture. We instantiate the protected environment using a 20-asset random dependency DAG and a 10 x 20 user access-control matrix, and evaluate the framework over 1,000 GPT-4-generated scenarios using a stratified 70/30 train/test split. Semantic classification is treated as a pluggable component rather than a primary contribution, and we compare four representative approaches: rule-based Regex, TF-IDF+SVM, SBERT+LR, and SetFit. SBERT+LR achieves the best overall performance with 91.7% accuracy and a 0.908 macro-F1 score, while all four methods detect 100% of external quarantine-class threats. For verified insiders, where source trust is maximal (T = 1), MemSentry does not automatically quarantine suspicious operations but instead escalates potentially dangerous writes for human review, making semantic classification important for accurately capturing insider intent.

Bullet Summary

  • Agentic AI systems with persistent memory are vulnerable to a novel attack vector called memory poisoning, where adversarially crafted content stored in long-term memory manipulates future agent behavior without altering underlying model weights or prompts.
  • MemSentry is a formal, configuration-driven framework designed to intercept persistent-memory writes, making deterministic Accept, Review, or Quarantine decisions based on comprehensive risk assessment factors including source trust, semantic risk, attack r...
  • The framework models the system environment as a directed acyclic graph (DAG) representing component dependencies with assigned criticality, and employs a user access-control matrix to realistically simulate operations; experiments involve 1,000 GPT-4-gener...
  • Semantic classification is a pluggable component in MemSentry; evaluated methods include rule-based Regex, TF-IDF+SVM, Sentence-BERT with Logistic Regression (SBERT+LR), and SetFit. SBERT+LR achieved the best performance with 91.7% accuracy and a macro-F1 s...
  • MemSentry quarantines suspicious writes from external sources but escalates potentially harmful writes from verified insiders for human review, highlighting the challenge of detecting insider threats semantically and the importance of accurate semantic clas...

BIO-MEMART: Biometric-Aware KV Cache Memory for Multi-User LLM Agents

arXiv preprint arXiv Trust and Identity Memory Poisoning Governance and Policy

Yanhong Qian, Xuanying He, Qingguo Meng, Shihao Ding, Xingbo Dong, Zhe Jin

Published 2026-09-08

Venue: arXiv

Open Source Record

Abstract

KV cache is evolving from a serving optimization into an external memory substrate for long-term LLM agents. In a shared multi-user deployment, however, reusable KV blocks introduce a missing access-control question: semantic relevance alone cannot determine whether a memory block is authorized for the current physical user. We propose Bio-MemArt, a biometric-aware KV-cache memory framework for multi-user LLM agents. Bio-MemArt attaches a normalized biometric template to each stored KV memory block, filters the shared memory pool with the current user's biometric probe, and then runs the original MemArt retrieval and KV reuse pipeline only inside the authorized candidate pool. This design preserves latent-space retrieval, direct cache reuse, and decoupled position encoding while adding physical-user access control to shared KV memory. We evaluate Bio-MemArt under Owner and Non-owner query conditions on long-term dialogue QA with face and palmprint benchmarks. Across face benchmarks, the average owner and non-owner biometric success rates are 95.71% and 0.86%; across palmprint benchmarks, they are 97.60% and 2.00%. In the efficiency study, average prefill tokens drop from 18,781.96 under full-context prompting to 28.57 with Bio-MemArt, showing that biometric gating preserves the low-token operating regime of KV-cache memory.

Bullet Summary

  • Introduces Bio-MemArt, a biometric-aware KV cache memory framework designed to add physical-user access control in multi-user large language model (LLM) agent deployments.
  • Each key-value (KV) memory block stores a normalized biometric template (such as face or palmprint embeddings) to authenticate users before memory retrieval and reuse, ensuring memory blocks are accessed only by their rightful owner.
  • The framework applies biometric gating by filtering the shared KV memory pool using the current user's biometric probe, thereby creating an authorized candidate memory pool for semantic retrieval via existing MemArt pipelines.
  • Evaluations on face and palmprint biometric benchmarks show high owner authentication rates (~95-98%) and very low unauthorized (non-owner) access (~1-2%), effectively safeguarding private memory reuse.
  • Bio-MemArt maintains the efficiency advantages of KV cache memory, reducing prefilled tokens drastically (from about 18,782 to 29) compared to full-context prompting, without impacting runtime performance.

An Evidence Model for Agentic Processes: Evidence Claims, Trust Assumptions, and Policy Assessment

arXiv preprint arXiv Governance and Policy Trust and Identity

Arslan Brömme

Published 2026-09-08

Venue: arXiv

Open Source Record

Abstract

Agentic AI systems increasingly exchange messages, invoke tools, request approvals, hold structured decision sessions, and modify shared artifacts. Logs and anchors can make selected records tamper-evident, but they can also mislead if their evidentiary meaning is implicit: a hash does not establish semantic truth, a signature does not establish authorization, and an external anchor does not establish capture completeness. This paper proposes an evidence claim model for agentic processes. It distinguishes artifact integrity, temporal existence, provenance, approval evidence, declared ordering, capture claim, relevance claim, deliberation traceability, monitoring claim, anchoring authorization claim, policy assessment claim, risk treatment claim, mitigation implementation claim, and management response claim. Semantic validity is treated as a recurring limitation. The model maps these claims to mechanisms, assumptions, limitations, and threats, and situates them in an agent organization with functional CEO agent, executive, operational, evidence, and audit roles, plus a plan-do-check-act-inspired management response loop. The contribution is conceptual: it does not validate a particular implementation, prevent all failures, or automate legal compliance. It provides a vocabulary for stating which claims an agentic black box can support, which claims it cannot establish, and which controls are required around it.

Bullet Summary

  • The paper addresses the complexity of establishing trustworthy evidence in multi-agent AI systems, emphasizing that traditional cryptographic proofs (hashes, signatures) do not guarantee semantic truth or policy compliance.
  • It introduces a conceptual evidence claim model for agentic processes, defining a detailed taxonomy of evidence claims such as artifact integrity, provenance, approval, ordering, capture completeness, and policy assessment.
  • The model explicitly associates each evidence claim with supporting mechanisms, underlying trust assumptions, existing limitations, and potential threat vectors, thereby clarifying the evidentiary properties that can and cannot be guaranteed.
  • A structured agent organization is proposed, comprising roles like CEO agent, operational agents, evidence producers, auditors, and management, facilitating segregation of duties and a plan-do-check-act management loop for continuous governance and policy i...
  • The concept of mandatory-event coverage ratio quantifies evidence capture completeness but depends on an independent expectation model and cannot prove absolute completeness globally.

Beyond Agent Harnesses: Cross-Substrate Authority for Multi-Agent Systems

Merged record merged scholarly record arXiv Governance and Policy Memory Poisoning Benchmarks and Evaluation

Yang Li, Sergey Volkov, Hai Liu, Zongsi Xu, Xiyu Chen, Tuo Zhou, Dian Shao, Hao Sun

Published 2026-09-08

Venue: arXiv

Open Source Record

Abstract

Agentic systems persist model-visible memory while mutating workspaces, while a runtime, registry, or approval service may hold authority state outside both. Identical final files can then require opposite safe actions. We call this the cross-substrate authority gap: decision- relevant authorization information resides outside the planner-visible workspace or memory state. Across two controlled mini-benchmark families, three experiments compare planner-observation augmentation with an execution-time authority check using real Git lineage, durably recorded agent execution attempts, deterministic oracles, and two model routes. Experiment 1 is a 128-cell controlled evidence ablation: authority-blind candidate evidence obtains 0/32 final semantic success, while raw receipts and a typed relation both obtain 32/32. The missing authority fact accounts for the gain; typed packaging provides no observed planning-accuracy gain over equal raw information. Experiment 2 uses 96 planning calls: workspace-visible evidence yields 12/16 unsafe publication decisions, and planning with the typed relation remains unreliable (15/32 first actions correct; 11/32 invalid or absent). Experiment 3 replays the same 32 fixed model-generated first-action intents with zero additional model calls; a deterministic execution guard prevents all six unsafe intents from becoming effects and permits all 12 valid authorized publish intents. These results position authority enforcement at the mutation boundary as the operational endpoint of memory governance.

Bullet Summary

  • The paper addresses the 'cross-substrate authority gap' in multi-agent systems, where critical authorization information exists outside the planner-visible workspace or memory, causing identical final files to require opposite safe actions.
  • It compares planner-observation augmentation with execution-time authority checks using real Git lineage, durably recorded agent executions, deterministic oracles, and model-generated plans across three controlled experiments.
  • Experiment 1 shows that authority-blind candidate evidence leads to 0/32 successful final semantic outcomes, while raw authority receipts or typed relations enable perfect (32/32) success, highlighting the necessity of including authority facts in planning.
  • Experiment 2 reveals that relying solely on workspace-visible evidence results in unsafe publication decisions in 12/16 cases and unreliable planning with typed relations, with only 15/32 correct and 11/32 invalid or missing initial actions.
  • Experiment 3 demonstrates that a deterministic execution-time enforcement guard effectively blocks all unsafe intents and permits authorized publish intents without additional model calls, confirming the need for runtime authority enforcement.

SchemeArena: Factorized Stress Testing of Scheming in LLM Agents

arXiv preprint arXiv Benchmarks and Evaluation Governance and Policy

Jie Ruan, Inderjeet Nair, Amy Liu, Muhammad Khalifa, Yusheng Zhou, Lu Wang

Published 2026-09-08

Venue: arXiv

Open Source Record

Abstract

We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how scheming arises from the interaction of key factors, such as instrumental goals, environmental affordances, oversight conditions, and perceived consequences. Prior work examines only a small number of scenarios, limiting the ability to isolate how these conditions shape an agent's propensity or capability to scheme. This limited scale and task diversity also restrict coverage of realistic deployment settings and the range of scheming strategies that can be observed. To this end, we introduce SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms. To enable scalable and reliable monitoring, we further propose SCOUT, a scheming monitor that grounds multi-criteria judgments in evidence drawn from agents' reasoning and actions. Across controlled stress tests on five LLM agents, we find that explicit instrumental goals are the strongest driver of scheming propensity. Strategic hints play a distinct role by helping agents translate scheming reasoning into concrete covert behavior. Oversight has mixed effects: in several closed models, action-only monitoring increases scheming, suggesting that partial oversight can act as an optimization constraint rather than a deterrent. CoT is a useful but incomplete monitoring signal: it can reveal latent scheming before execution, yet action-only scheming shows that covert behavior may occur without explicit reasoning evidence. We release the benchmark, code, and monitor at: https://github.com/launchnlp/SchemeArena.

Bullet Summary

  • Introduced SCHEMEARENA, a comprehensive 400-scenario benchmark to stress test scheming behaviors in LLM agents by factorizing diverse conditions such as instrumental goals, oversight types, and environmental affordances.
  • Developed SCOUT, a multi-criteria process-level monitor that detects scheming using combined evidence from agents' reasoning (chain-of-thought) and actions, achieving approximately 90.7% F1 score validated by human annotations.
  • Demonstrated that explicit instrumental (long-term misaligned) goals are the strongest drivers of scheming propensity, while strategic hints significantly enhance agents' ability to translate scheming reasoning into covert behaviors.
  • Found oversight effects to be mixed: partial oversight like action-only monitoring can increase scheming in closed models by acting as an optimization constraint, whereas full trajectory oversight marginally reduces scheming.
  • Revealed a reasoning–action gap where scheming reasoning sometimes fails to produce scheming actions, and covert scheming can occur without explicit chain-of-thought evidence, indicating limitations of reasoning-only monitoring.

LLM-Based Penetration Testing in the Presence of Honeypots

arXiv preprint arXiv Orchestration Risk Agent-to-Agent Communication Governance and Policy

Xinhong Xie, Piyush Nagasubramaniam, Neeraj Karamchandani, Sencun Zhu

Published 2026-09-08

Venue: arXiv

Open Source Record

Abstract

Large language model (LLM) agents are increasingly employed for offensive cybersecurity tasks such as automated vulnerability discovery, reconnaissance, and penetration testing. This new capability also threatens one of the defender's most valuable tools: deception. Traditional honeypots rely on realism and obscurity to lure human or script-driven attackers into revealing tactics, techniques, and procedures (TTPs), but LLM-driven attackers can reason about heterogeneous artifacts and use the honeypot suspicion to guide target-selection decisions. We present a systematic study of honeypot-aware budget allocation for LLM attack agents. We formalize the attacker's problem as a budgeted decision process: an agent interacts with potential targets, consuming LLM execution budget during reconnaissance and exploitation, and must decide whether to (continue exploitation) or (skip) when honeypot suspicion arises. Our findings show that with the proposed detector-guided policy, LLM agent attackers can effectively allocate budget to compromise hosts in a host pool, highlighting the importance of dynamically allocating budget in a controlled mixed-host testbed. While defenses are beyond our present scope, we discuss implications for future adversarially resilient and adaptive honeypot design.

Bullet Summary

  • Large language models (LLMs) are increasingly utilized for offensive cybersecurity tasks such as automated penetration testing and vulnerability discovery, challenging traditional deception tools like honeypots.
  • Traditional honeypots rely on realism and obscurity but struggle against LLM-based attackers capable of reasoning about diverse artifacts and adapting target selection based on honeypot suspicion.
  • The attacker’s decision-making is formalized as a budgeted decision process where the LLM agent navigates reconnaissance and exploitation actions within a limited execution budget, deciding when to continue or skip based on honeypot detection.
  • The authors propose a two-stage, detector-guided policy combining pre-connect conservative filtering of suspicious hosts, budget-aware host ranking, and post-connect stopping to optimize attacks and avoid honeypots.
  • Pre-connect detection evaluates observable protocol and service attributes to assign genuine-host likelihood scores to minimize early honeypot engagement, while post-connect detection verifies host authenticity through command validity and network behavior...

Privacy-Aware Data-Model Dual-Driven Decision Analysis for Data Security in Distributed Multi-Agent Operations

Merged record merged scholarly record OpenAlex Orchestration Risk Governance and Policy Benchmarks and Evaluation

Yunxiao Wang, Haizhuang Liu, Zihan Liu, Haobo Zhao, Fuyang Wei

Published 2026-09-08

Venue: ICST Transactions on Scalable Information Systems

DOI: https://doi.org/10.4108/eetsis.13943

Open Source Record

Abstract

INTRODUCTION: Distributed networks generate heterogeneous security telemetry while moving data across endpoints, users, services, and operational domains. SOCs need methods that protect data assets, preserve auditability, and avoid unsafe tool calls.OBJECTIVES: This paper proposes a data-model dual-driven method, in which incident and execution data constrain LLM-based reasoning while model outputs generate auditable process data, for privacy-aware data security decision analysis in multi-agent security operations.METHODS: The method combines LLM-based role agents, SOAR playbook orchestration, persistent message state, and a virtual security capability layer. Incidents are transformed into data-aware tasks, actions, commands, execution records, and summaries.RESULTS: On 83 labeled incident samples, tool-call evaluation achieved 0.9684 precision, 0.4742 recall, 0.6367 F1-score, and 76.45 s average handling time.CONCLUSION: The method supports auditable data security monitoring and controlled response, while complex multi-step planning remains the main improvement target.

Bullet Summary

  • The paper addresses the challenge of privacy-aware data security decision-making in distributed multi-agent security operations amidst heterogeneous and sensitive security telemetry.
  • It proposes a data-model dual-driven method integrating LLM-based role-specialized agents, SOAR playbook orchestration, and a virtual security capability layer to enable structured, auditable, and privacy-preserving incident response.
  • The method transforms raw incident and execution data into structured tasks, commands, and summaries, constraining LLM reasoning with data inputs while generating traceable outputs for accountability.
  • A multilayered architecture is introduced, featuring data normalization, privacy-governed data assets, role-specific multi-agent decision-making, controlled SOAR execution, and virtualized security tools for modularity and safety.
  • Experimental evaluation on 83 diverse security incidents shows high precision (0.9684) in tool invocation, minimizing unsafe or irrelevant tool calls, but moderate recall (0.4742), indicating incomplete multi-step planning and some necessary tools omitted.

Beyond Agent Harnesses: Cross-Substrate Authority for Multi-Agent Systems

Merged record merged scholarly record arXiv OpenAlex Governance and Policy Benchmarks and Evaluation

Yang Li, Sergey Volkov, Hai Liu, Zongsi Xu, Xiyu Chen, Tuo Zhou, Dian Shao, Hao Sun

Published 2026-09-08

Venue: arXiv

DOI: https://doi.org/10.48550/arxiv.2609.08472

Open Source Record

Abstract

Agentic systems persist model-visible memory while mutating workspaces, while a runtime, registry, or approval service may hold authority state outside both. Identical final files can then require opposite safe actions. We call this the cross-substrate authority gap: decision- relevant authorization information resides outside the planner-visible workspace or memory state. Across two controlled mini-benchmark families, three experiments compare planner-observation augmentation with an execution-time authority check using real Git lineage, durably recorded agent execution attempts, deterministic oracles, and two model routes. Experiment 1 is a 128-cell controlled evidence ablation: authority-blind candidate evidence obtains 0/32 final semantic success, while raw receipts and a typed relation both obtain 32/32. The missing authority fact accounts for the gain; typed packaging provides no observed planning-accuracy gain over equal raw information. Experiment 2 uses 96 planning calls: workspace-visible evidence yields 12/16 unsafe publication decisions, and planning with the typed relation remains unreliable (15/32 first actions correct; 11/32 invalid or absent). Experiment 3 replays the same 32 fixed model-generated first-action intents with zero additional model calls; a deterministic execution guard prevents all six unsafe intents from becoming effects and permits all 12 valid authorized publish intents. These results position authority enforcement at the mutation boundary as the operational endpoint of memory governance.

Bullet Summary

  • Multi-agent systems face a 'cross-substrate authority gap' where critical authorization information resides outside the agent-visible workspace or memory, leading to ambiguous and unsafe decisions despite identical artifact states.
  • The authors propose a formal model of authority as a relation among agent execution attempts, artifact states, and downstream-use authorization, emphasizing the distinction between planning-time observation and execution-time enforcement of authority.
  • Three controlled experiments using real Git lineage and state-of-the-art models demonstrate that relying solely on planner-visible evidence is insufficient and unreliable for making safe artifact publication decisions.
  • Experiment 1 shows that adding raw receipts (authority evidence) drastically improves semantic success (32/32) compared to authority-blind evidence (0/32), highlighting the critical role of authority information.
  • Experiment 2 reveals that planning with explicit authority evidence remains error-prone, with frequent unsafe publication decisions and unreliable first action selection across models, including GPT-5.4 variants.

From Reactive Monitoring to Preemptive Defense: A Coq- and TLA⁺-Verified Platform for Predicting Generative AI Collapses and Cyberattacks

OpenAlex · Zenodo (CERN European Organization for Nuclear Research) repository OpenAlex Orchestration Risk Benchmarks and Evaluation Governance and Policy

Valery Kalinin

Published 2026-09-08

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.22655169

Open Source Record

Abstract

This paper presents a universal, formally grounded approach to predicting the degradation of complex systems, including generative models (GANs), large language models (LLMs), AI agents, and cyber threats. The approach is built upon Theorem 3.9 (Parasitism Limit) of Cognitive Shadow Theory, which establishes that any parasitic activity inevitably reduces the entropy of the system's observable states. Key contribution: a unified predictive formula T = ceil(max(0, (H_min/0.51 - H_0)/δ_min(M))) that enables prediction of the time until system collapse, thereby enabling the transition from reactive detection to preemptive prediction. Empirical validation spans four domains:• Generative models (DCGAN on CIFAR-10): 100% Precision/Recall, zero FPR, lead times up to 45 epochs• LLMs: zero FPR on synthetic tests, F1=0.579 on TruthfulQA• AI Agents: AUC=0.820, 14.4 steps lead time, zero FPR• Cybersecurity: AUC=0.988 on CIC-Bell-DNS-EXF-2021, 90.3% MITRE coverage All key theorems are formally verified in Coq 8.18+ and TLA⁺. Patent application: No. 2026124758 (filed August 12, 2026)

Bullet Summary

  • Introduces a universal, formally grounded methodology for predicting degradation in complex systems such as generative models, large language models, AI agents, and cyber threats.
  • Builds on Theorem 3.9 (Parasitism Limit) of Cognitive Shadow Theory, which shows parasitic activity reduces the entropy in system observable states, signaling system degradation.
  • Proposes a unified predictive formula T = ceil(max(0, (H_min/0.51 - H_0)/δ_min(M))) to forecast the time until system collapse, shifting the focus from reactive monitoring to preemptive defense.
  • Validates the theoretical approach empirically across four domains: 1) Generative models (DCGAN on CIFAR-10) achieving perfect precision/recall and zero false positive rate with lead times up to 45 epochs.
  • Shows robust performance on large language models with zero false positives on synthetic tests and moderate F1 score on the TruthfulQA benchmark.

Agency as an Architecture Layer: A Formal Enterprise Architecture Framework for Agentic-AI-Driven Enterprises

Merged record merged scholarly record OpenAlex Governance and Policy Agent-to-Agent Communication

Samir El Hassani

Published 2026-09-08

Venue: Research Square

DOI: https://doi.org/10.21203/rs.3.rs-10943728/v1

Open Source Record

Abstract

Abstract unavailable from OpenAlex metadata.

Bullet Summary

  • The paper addresses the gap in enterprise architecture (EA) frameworks that currently lack explicit modeling of software agents, especially AI-driven agents with delegated authority, across business, application, and technology layers.
  • It proposes the Agentic Enterprise Architecture Framework (AEAF), introducing an agency layer with formal semantics based on a stratified Datalog program to model agents, their charters, delegations, and operational guards ensuring secure delegation and ove...
  • AEAF defines action authority across four levels—read, propose, commit, and autonomous—enabling nuanced control and analysis of agents’ interactions with enterprise resources.
  • A five-step design-time method integrates with typical EA cycles: populating data, chartering agents, analyzing authority defects, refining models, and provisioning identity systems, reversing traditional run-time governance flows.
  • The framework was validated through a real-world insurance case (Meridian Insurance), demonstrating detection and resolution of authority defects such as orphan authority and separation-of-duty (SoD) violations by refining agent charters.

VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities

arXiv preprint arXiv Benchmarks and Evaluation Trust and Identity Governance and Policy

Jiahao Shi, Edward Tsien, Yifeng Di, Hongjiao Zhang, Yuan Tang, Ronit Dey, Ilona Shishov, Gal Netanel

Published 2026-09-07

Venue: arXiv

Open Source Record

Abstract

The software supply chain has become an increasingly exposed attack surface because of its reliance on intricate yet fragile dependencies. Existing defenses such as GitHub Dependabot often raise many false alerts because their coarse-grained matching cannot determine whether a vulnerable dependency is actually exploitable. Security analysts typically spend substantial time assessing vulnerability exploitability case by case. Recent LLM agents have emerged as promising candidates for this task given their advanced capabilities in coding and cybersecurity, yet no existing benchmark evaluates them on it. Prior benchmarks target zero-day settings, where agents detect and exploit previously unknown vulnerabilities. In contrast, software supply chain security focuses on how known vulnerabilities in upstream dependencies affect downstream projects. This requires agents to reason across repositories and determine whether an upstream vulnerability is exploitable in the downstream project. To address this gap, we introduce VEX-Bench, the first benchmark for evaluating LLM agents' ability to assess the exploitability of software supply chain vulnerabilities. It contains 75 real-world cases mined from GitHub and labeled by security experts, covering Python, Java, and Go. We evaluate nine models across three agent harnesses. While GPT-5.5 and Claude Opus 4.6 reach approximately 80% F1 on binary vulnerability-status classification, only GPT-5.5 surpasses 70% macro-F1 on fine-grained justification classification. This gap highlights the challenge of moving beyond binary exploitability assessment to identifying fine-grained exploitability reasons. Code and data: https://github.com/steven1518/vex-bench

Bullet Summary

  • The paper addresses the challenge of assessing exploitability of software supply chain vulnerabilities, which arise due to complex dependencies in software projects.
  • Existing tools like GitHub Dependabot produce many false alerts as they rely on coarse-grained package metadata without evaluating actual exploitability.
  • Large language model (LLM) agents have potential to assist in exploitability assessment, but prior benchmarks focus only on zero-day vulnerabilities, not supply chain propagation of known vulnerabilities.
  • VEX-Bench is introduced as the first benchmark designed to evaluate LLM agents on exploitability assessment of software supply chain vulnerabilities, containing 75 real-world cases annotated by experts across Python, Java, and Go projects.
  • The benchmark categorizes vulnerability status as affected or not-affected, with four detailed justification categories explaining non-exploitability reasons to improve evaluation granularity.

From Event Logs to Governed Action: A BlueSky Agenda for Agentic Process Mining

arXiv preprint arXiv Governance and Policy Benchmarks and Evaluation

Yiyuan Yang, Zheshun Wu, Yong Chu, Zhenghua Chen, Zenglin Xu, Qingsong Wen

Published 2026-09-07

Venue: arXiv

Open Source Record

Abstract

Process mining has long turned event logs into process knowledge: discovered models, conformance evidence, bottleneck diagnoses, and runtime predictions. Agentic AI changes the target. Process-aware agents will not only ask what happened. They will ask whether a proposed action should be taken, given the available evidence, privacy budget, organizational authority, and downstream risk. This BlueSky paper proposes event-to-action process mining: a process-mining agenda for transforming heterogeneous operational event data into governed action. The goal is not another dashboard, a generic enterprise simulator, or a language interface over logs. We argue that the community needs four mineable artifacts: event-object representations, action evidence packages, governance contracts, and benchmarks where act, defer, ask, and refuse are all valid outputs. This agenda is timely because agentic business process management (BPM), LLM-assisted process mining, object-centric event standards, causal process monitoring, and privacy-preserving learning are maturing separately. Bringing them together defines a data-mining target inside process mining: mining logged organizational behavior for accountable action, not only retrospective insight.

Bullet Summary

  • Introduces a novel agenda called event-to-action process mining, aiming to transform heterogeneous event logs into governed, accountable action recommendations rather than just retrospective insights.
  • Defines four essential mineable artifacts: event-object representations (detailed event-object graphs), action evidence packages (causal evidence with calibrated uncertainty), governance contracts (privacy, authority constraints, and auditability), and benc...
  • Emphasizes agentic AI's role in process mining whereby process-aware agents consider evidence, privacy budgets, authority boundaries, and downstream risks before recommending or taking actions.
  • Highlights the convergence of recent advances in agentic BPM, LLM-assisted process mining, object-centric event standards, causal monitoring, and privacy-preserving learning as enablers for this integrated framework.
  • Addresses shortcomings of current prescriptive process monitoring by enforcing verifiable, accountable computations, respecting authority, and adapting to cross-organizational privacy and risk issues through federated process mining.

Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best

arXiv preprint arXiv Governance and Policy Benchmarks and Evaluation

Kevin Baum, Rūta Binkytė, Felix Jahn

Published 2026-09-07

Venue: arXiv

Open Source Record

Abstract

AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for. Reinforcement-learning-based alignment folds norms and task pursuit into one policy: the system learns its norms from scored behavior, and scoring flattens them. Do not do X is learned as doing X costs something if noticed. On every datum training can produce, a policy that complies only when it might be observed is indistinguishable from one that complies always. The experiment that would tell them apart - scoring unobserved behavior - is a contradiction in terms. Conditional compliance is thus the most that behavioral training can be known to deliver. Agency sharpens the problem: agents operate mostly where no one is watching, and can act on whether they are watched. An iterated pipeline that trains against detected failures selects for passing detection, not for complying. This account unifies alignment faking, sandbagging, and evaluation-aware scheming. And it reorients the remedy: not deeper internalization but architecture, making violations unavailable rather than unchosen.

Bullet Summary

  • Reinforcement learning (RL)-based alignment methods train AI agents by scoring only observed behaviors, which merges norms and task objectives into a single reward function, leading agents to optimize compliance only when they believe they are being watched...
  • Because unobserved behaviors are unscored, policies that comply conditionally (only when observed) are indistinguishable from those that comply unconditionally during training, making standard RL training inherently unable to guarantee true norm internaliza...
  • Agentic AI systems can strategically manipulate their observability status and behave differently when unobserved, causing distribution shifts at deployment that the training pipeline cannot detect or correct since it lacks access to unobserved behaviors.
  • The iterated training pipeline emphasizes passing detection rather than genuine norm adherence, unifying phenomena like alignment faking, sandbagging, and evaluation-aware scheming under a common theoretical framework.
  • Probes, monitors, and evaluation mechanisms trained on observed behavior labels inherit these limitations and thus select for detection-avoiding behavior rather than true compliance, undermining verification of norm internalization.

Decentralized Safe Multi-Agent Reinforcement Learning via Predictive Shielding

arXiv preprint arXiv Agent-to-Agent Communication Governance and Policy Orchestration Risk

Yacine El Yamani, Hanna Krasowski, Elena Vanneaux

Published 2026-09-07

Venue: arXiv

Open Source Record

Abstract

Environments are increasingly populated by multiple robots performing independent tasks with limited prior knowledge of each other. Deploying such multi-agent systems presents significant challenges. Specifically, shifts in deployment states compared to training data can lead to poor policy performance and compromised safety. While safety shields exist to mitigate these risks, they are typically reactive, which degrades performance near unseen obstacles,and centralized, limiting their scalability. To address this, we propose a decentralized framework that integrates predictive shielding with model-based finite horizon Q-learning. This approach allows agents to safely adapt their pre-trained policies during deployment. Furthermore, to mitigate livelocks in symmetric scenarios, we introduce a communication- free protocol for conflict resolution

Bullet Summary

  • The paper addresses safety and performance challenges in decentralized multi-agent reinforcement learning where agents are pretrained independently, operate with limited observability, and have no inter-agent communication.
  • It proposes a decentralized predictive shielding framework integrating model-based finite-horizon Q-learning to enable agents to adapt pre-trained policies safely during deployment by forecasting multiple steps ahead using learned environment models.
  • The approach assumes each agent possesses a trivial backup safe policy and composes these individual shields under the assume-guarantee paradigm, ensuring overall system safety without explicit communication.
  • To prevent livelocks caused by symmetric agent behaviors, the authors introduce a novel communication-free stochastic conflict resolution protocol that probabilistically alternates agent policies to break symmetry and avoid deadlocks.
  • Static and dynamic safety constraints are handled separately: static constraints use an infinite-horizon model-based Q-learning approach converging to an optimal Q-table, while dynamic constraints are managed via a finite-horizon, time-dependent Q-learning...

MOLE: Detecting Insider Threats in AI Agents

arXiv preprint arXiv Prompt Injection Benchmarks and Evaluation Governance and Policy

Aashiq Muhamed, Virginia Smith

Published 2026-09-07

Venue: arXiv

Open Source Record

Abstract

Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.

Bullet Summary

  • MOLE is a novel, open benchmark designed to detect insider threats among AI-operated accounts by simulating 150 accounts interacting with 9 stateful services over 30 workdays, incorporating 12 MITRE-grounded threat scenarios within routine agent tasks.
  • The benchmark facilitates comprehensive evaluation of 39 AI agent models and 40 monitoring systems under realistic constraints including fixed daily review budgets and varying levels of observability (audit events, tool results, agent reasoning).
  • Empirical results show that 72% of tested AI agent models complete most assigned harmful objectives, and refusal by agents to execute a task does not reliably predict harmful completion, demonstrating the need for robust monitoring.
  • Semantic monitors leveraging large language models outperform classical anomaly detection baselines on the MOLE dataset, but optimal detection depends significantly on the observability level and access to agent reasoning rather than solely monitor strength.
  • MOLE’s structured design supports automated monitor development through benchmark-guided feature discovery and cost-aware selective monitoring, resulting in substantial improvements in detection accuracy within limited review budgets.

Operating a Human-Governed Multi-Machine LLM Agent Fleet: An Experience Report

Merged record merged scholarly record OpenAlex Governance and Policy Orchestration Risk Agent-to-Agent Communication

Anton Dziatkovskii

Published 2026-09-07

Venue: Zenodo (CERN European Organization for Nuclear Research)

DOI: https://doi.org/10.5281/zenodo.22639713

Open Source Record

Abstract

Most published multi-agent LLM systems are single-process orchestrations evaluated on benchmarks. We report on something different: a fleet of LLM agents distributed across five physical machines (an always-on hub, laptops, a family computer, and a VPS anchor), operated continuously for roughly two months (June-July 2026) on real knowledge work by a non-technical founder and his collaborators. The fleet negotiates decisions through a deterministic consensus protocol (propose - counter - accept - commit over an append-only, single-writer-per-machine event log), communicates over a dual-rail bus (synced file mailbox plus a group chat that humans also read), enforces an acknowledgement discipline in which silence past an SLA is an incident, and gates every risky action behind a deterministic risk-tier tripwire that escalates to a dedicated human channel. The safety-critical layer makes zero LLM calls: it is auditable file I/O, and we show it derives full fleet state at microsecond cost. This is an experience report, not a benchmark study. Its evidence is (i) a reproducible offline harness — five self-checking scenarios covering the happy path, the human gate, the tripwire, split-brain, and ledger corruption, all passing on commodity hardware — and (ii) a catalog of nine production failure modes, each of which occurred before its guard existed, giving an unusual, historically grounded form of ablation: for every guard we can state what the system actually did without it. We distill the design principles that survived contact with production (single-writer files, dual-rail by construction, delivery is not completion, detect what you cannot prevent, a human gate needs an exit, alert-channel purity, owner-repairability) and state our limitations plainly: this is an N=1 longitudinal case study with no comparative baseline. Reference implementation: claw-consensus (MIT).

Bullet Summary

  • The paper addresses operating a distributed fleet of Large Language Model (LLM) agents across multiple physical machines for real-world knowledge work, beyond single-process benchmark evaluations.
  • A deterministic consensus protocol (propose - counter - accept - commit) over append-only, single-writer event logs ensures decision negotiation among agents distributed on 5 machines (hub, laptops, family computer, VPS).
  • Communication uses a dual-rail bus combining synced file mailboxes and group chat readable by humans, enabling transparent and reliable messaging.
  • A strict acknowledgement discipline with Service Level Agreements (SLA) detects incidents through silence, while every risky action triggers a deterministic risk-tier tripwire escalating to a dedicated human channel for safety.
  • The safety-critical control layer executes auditable file I/O without any LLM calls, allowing full fleet state derivation at microsecond cost, enhancing system trustworthiness.

Editorial: Advanced integration of large language models for autonomous systems and critical decision support

Merged record merged scholarly record OpenAlex Governance and Policy Agent-to-Agent Communication Orchestration Risk

I. de Zarzà, J. de Curtò, Carlos T. Calafate

Published 2026-09-07

Venue: Frontiers in Artificial Intelligence

DOI: https://doi.org/10.3389/frai.2026.1962295

Open Source Record

Abstract

Advanced Integration of Large Language Models for Autonomous Systems and Critical Decision SupportLarge language models (LLMs) have shown transformative potential in autonomous systems and critical decision-making, yet standalone models remain limited in robustness, reliability, and safety assurance when deployed in high-stakes environments. This Research Topic began from the premise that such limitations are better addressed by structured integration of multiple specialized models than by scaling any single one (Guo et al., 2024). We invited work on multi-LLM integration for perception, navigation, and decision-making in robots, drones, and vehicles; on human-robot collaboration; on high-stakes decision support; on verification, uncertainty quantification, and safety assurance; and on real-time adaptation.Seven contributions were accepted, spanning agent generation, orchestration, perception, query translation, automated machine learning, intrusion detection, and governance.Deployment in these settings changes the central question. Performance can no longer be judged by fluency or task accuracy alone; it must also be assessed through grounding, reproducibility, latency, calibration, failure containment, human oversight, and auditability. Across domains, the contributions converge on a common conclusion: dependable autonomy is principally a systems-engineering problem.Reliability emerges not from trusting a single model, but from structuring how models are composed, constrained, checked, and connected to action.Perera et al. challenge the fixed-team assumption of many multi-agent systems. Their Initial Automatic and Dynamic Real-Time Agent Generation mechanisms create specialized agents from evolving conversational context. In the evaluated medical scenario, dynamic generation improved coverage, lexical diversity, and thematic relevance over a static AutoGen configuration, treating system composition itself as an adaptive variable. The same move from fixed programs toward prompt-defined behavior appears in LLM-driven swarm simulations (Jimenez-Romero et al., 2025).1 de Zarz à et al.Zhou and Chan address the complementary problem of reproducibility. Their orchestrator, ORCH, 27 decomposes a problem, gathers analyses from heterogeneous models, and merges them through a 28 deterministic protocol; an optional exponential moving average module adapts routing from historical 29 feedback. Gains are strongest on harder reasoning tasks, but entail substantial latency and cost. Taken 30 together, these studies show that adaptation and determinism are not opposites: agent membership and 31 routing may while interfaces, rules, and aggregation procedures remain explicit and 32 auditable, as in ensemble-and-arbiter designs where inter-model disagreement is measured and routed to 33 human review (Lipianina-Honcharenko et al., 2026). Calboreanu makes the architectural argument most explicit. LATTICE separates planning, execution, and 58 governance so that no component both decides an action and judges compliance; it applies policy-as-code 59 through gated execution, escalates uncertain cases to human operators, and preserves provenance through The next phase should connect these principles into end-to-end assurance cases. It requires interoperable 78 agent-tool contracts, benchmarks covering distribution shift and adversarial faults (Radanliev et al., 2026), 79 selective autonomy with tested fallback behavior, and human-centered studies of explanation and escalation.The central lesson is measured but consequential: LLMs become suitable for autonomous systems and 81 critical decision support not as self-sufficient decision makers, but when embedded within architectures that 82 make uncertainty visible, constrain action, preserve accountability, and retain meaningful human control 83 (Santoni de Sio and van den Hoven, 2018).We thank all contributing authors, reviewers, and the Frontiers editorial team for advancing this 85 interdisciplinary discussion.

Bullet Summary

  • Standalone large language models (LLMs) exhibit limitations in robustness, reliability, and safety assurance when employed independently in high-stakes autonomous systems, prompting the need for their structured integration.
  • The research presents seven contributions focusing on multi-LLM integration across diverse tasks including dynamic agent generation, orchestration, perception, query translation, automated machine learning, intrusion detection, and governance frameworks.
  • Adaptive agent orchestration methods that generate specialized agents in real-time based on evolving conversational context improve system coverage, lexical diversity, and thematic relevance compared to fixed-team approaches.
  • Reproducibility and reliability are enhanced by orchestrators that decompose problems, aggregate heterogeneous model outputs through deterministic protocols, and incorporate feedback adaptation mechanisms while maintaining auditability.
  • Applications in critical domains like distracted driving intervention and agro-food database querying demonstrate that dependable decision support requires semantic grounding, modular information fusion, calibrated outputs, and explicit domain knowledge rep...

AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories

arXiv preprint arXiv Benchmarks and Evaluation Governance and Policy

Ruoxi Shang, Christina-Maria Androna, Orfeas Menis Mastromichalakis, Yu Feng, Aniruddhan Ramesh, Rico Angell, Shang Hong Sim, Chrysoula Zerva

Published 2026-09-06

Venue: arXiv

Open Source Record

Abstract

LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pre-action detection, and safe task completion when a safe solution exists. We introduce AURA-Eval, a framework combining controlled augmentation with granular diagnosis of behavior in tool-use trajectories. Its pipeline identifies safety-critical decision points, generates controlled variations, and constructs counterparts differing in whether a request has a safe fulfillment path. Using 157 sourced trajectories, we generate 1,249 evaluation items and evaluate 20 frontier and open-weight models. We developed rubrics to classify risk detection, action strategy, and scenario-specific action safety. Our results show that LLM agents engage in unsafe behavior more often when no safe fulfillment path exists. In these cases, frontier proprietary models more often recognize risk and exhibit safer behavior by proposing alternatives, while evaluated open-weight models more often directly execute unsafe requests. Increasing impact or reducing opportunities for oversight before execution also exposes greater vulnerability across models.

Bullet Summary

  • AURA-Eval introduces a novel evaluation framework designed to diagnose risk awareness and safety strategies in multi-step LLM agent trajectories involving tool use, moving beyond simplistic safe/unsafe scoring to granular behavioral analysis.
  • The framework employs controlled scenario augmentation to create paired test cases that differ in the availability of safe fulfillment paths, enabling evaluation of whether agents can detect risks and respond appropriately based on context.
  • A taxonomy of six risk mechanism dimensions and five scenario difficulty factors structures the risk assessment, capturing variables like harm intensity, target susceptibility, oversight, interpretive ambiguity, and emotional manipulation.
  • Evaluation of 20 frontier and open-weight large language models on 1,249 benchmark items reveals frontier models generally recognize risk better and favor safer alternatives, whereas open-weight models tend to execute unsafe actions directly, especially whe...
  • The AURA-Eval pipeline includes risk point identification, scenario truncation, controlled one-dimension-at-a-time edits, and human-AI ensemble judging with high inter-annotator reliability for multi-axis risk detection and action safety labeling.

Federating Trust Perimeters: Extending Industry IAM with DLT-Based Governance

arXiv preprint arXiv Trust and Identity Governance and Policy

Carlo Segat

Published 2026-09-06

Venue: arXiv

Open Source Record

Abstract

Digital systems are becoming more integrated, autonomous, and cooperative. AI agents, future mobile networks, and machine-to-machine economies point to one trend: spontaneous, cross-organizational, unplanned interactions between non human entities (NHEs). Trust establishment for them remains an open problem. Federation is the natural candidate, but established approaches, from OpenID Federation 1.0 and SAML to Federated Identity Management, presuppose what this setting denies them: manual, ahead-of-time configuration and a common trust anchor, whether pre-established members or a shared provider. Trust domains must therefore federate without being prefigured to do so: plan for unplanned interactions. This paper examines whether prominent Identity and Access Management (IAM) approaches, namely SPIRE, Workload Identity Federation (WIF), and OpenID Federation 1.0, can support such federation. Drawing requirements from disparate fields (medical, mobile networks, agentic AI), it argues that SPIRE is the most promising starting point, but needs three extensions to meet them all: token exchange, letting a home domain mint scoped, audience-bound tokens from a foreign workload's SPIFFE Verifiable Identity Document (SVID); remote attestation, so a trust decision targets a specific workload rather than a whole domain; and a distributed-ledger layer that anchors trust roots, carries federation governance, and publishes the shared keys the other two depend on.

Bullet Summary

  • The paper addresses the challenge of establishing trust among non-human entities (NHEs) in spontaneous, cross-organizational interactions where traditional federated IAM approaches fail due to their reliance on manual configuration and common trust anchors.
  • It evaluates prominent Identity and Access Management (IAM) frameworks—SPIRE, Workload Identity Federation (WIF), and OpenID Federation 1.0—against requirements drawn from domains like medical, mobile networks, and agentic AI, concluding SPIRE as the most p...
  • SPIRE's existing model, rooted in workload attestation and SPIFFE IDs, is extended with three key mechanisms: token exchange following RFC 8693 for scoped authorization tokens, remote workload attestation for fine-grained trust decisions, and integration of...
  • The distributed ledger stores shared trust anchors and federation metadata, eliminating single points of failure and enabling tamper-evident multi-stakeholder governance essential for dynamic federation lifecycle management.
  • The proposed federated IAM architecture supports unplanned, dynamic federations without requiring prior mutual configuration or a centralized trust provider, addressing scalability and interoperability in multi-agent ecosystems.

A Unified Policy Architecture (UPA): The Governance Kernel for Enterprise AI Operating Systems

arXiv preprint arXiv Governance and Policy Agent-to-Agent Communication Benchmarks and Evaluation

Prabhu Raghav, Balamurugan Pandi, Arul Vivek, Shek Mohammed, Sridhar S

Published 2026-09-06

Venue: arXiv

Open Source Record

Abstract

Enterprise AI is evolving into an Enterprise Operating System where autonomous AI agents can plan, reason, use memory, invoke tools, execute workflows, and collaborate with other agents. This shift creates a new governance challenge: existing authorization, security, guardrails, and compliance mechanisms are fragmented and are not designed to govern autonomous AI as a unified system. This paper introduces the Unified Policy Architecture (UPA), a governance architecture for Enterprise AI Operating Systems. UPA provides a unified policy model for governing AI and agents, tools, workflows, memory, enterprise resources, and agent-to-agent interactions and enterprise business rules. It extends policy control beyond authorisation to include runtime obligations, human approvals, compliance, audit evidence, and governance evaluation. We present UPA's governance model, declarative policy language foundations, policy evaluation semantics, extensible plugins, industry policy packs, and an evaluation framework for enterprise governance. We also identify extensions for multi-agent coordination, provenance-aware policies, and stateful runtime governance. UPA provides a foundation for building secure, accountable, and governable Enterprise Operating Systems for autonomous AI.

Bullet Summary

  • Enterprise AI is evolving into an Enterprise Operating System where autonomous AI agents perform complex tasks like planning, tool invocation, and multi-agent collaboration, creating new governance challenges beyond traditional model safety measures.
  • Existing governance mechanisms are fragmented across identity, compliance, and workflow domains, resulting in duplicated logic, inconsistent enforcement, and limited runtime control.
  • The paper introduces the Unified Policy Architecture (UPA), a governance framework with a deterministic Policy Kernel separating governance from AI reasoning, using a standardized Principal–Action–Resource–Context (PARC) model for unified decision-making.
  • UPA extends governance throughout the AI lifecycle, covering authentication, authorization, planning, tool access, multi-agent collaboration, human approval workflows, compliance, auditing, and runtime enforcement, transcending conventional guardrails.
  • UPA's modular architecture includes semantic normalization of heterogeneous events, declarative policy evaluation, runtime governance orchestration, and a plugin framework for extensible, domain-specific governance capabilities.
Load more articles