SANS Internet Storm Center — Reconstructing AI Agent Activity: New Forensic Review Scripts

Publication date: October 08, 2026
Category: AI attacks (LLM/LocalAI)

Introduction

The widespread adoption of Large Language Model (LLM) coding assistants and autonomous agents has introduced novel risk vectors and complex challenges for incident response and digital forensics teams. To address the lack of visibility into the internal activity of these tools, two new Python scripts have been released to examine telemetry, chat histories, and local artifacts generated by popular assistants. This methodological development, reported through the analysis channels of the SANS Internet Storm Center, enables forensic investigators to accurately reconstruct interactions, tool invocations, and decisions made by autonomous agents on compromised or audited systems.

What is AI Agent Forensic Activity? (General Analysis)

Unlike traditional applications, AI agents and local development assistants (such as OpenCode, Hermes, Claude Code, Cursor, Copilot, and others) generate complex data structures containing conversation histories, internal model reasoning, local tool executions, error logs, and token/cost telemetry. When a development environment is compromised or misused—for example, via indirect prompt injection or unauthorized code execution through agent tools—precise timeline reconstruction becomes critical.

Since this methodological case corresponds to a defensive analysis utility rather than a software vulnerability, no CVE identifier, CVSS score, or direct CWE classification is assigned, framing it strictly as an enhancement to forensic readiness and investigative capabilities.

How Does It Work? (Technical Analysis)

The two developed scripts (opencode-chat-replay.py and hermes_forensic_extract.py) require Python 3.10 or later and operate exclusively using the language’s standard library, minimizing external dependencies. Their internal operation is divided into two distinct architectural approaches for evidence collection:

  • Temporary Snapshot and Extraction Flow (opencode-chat-replay.py):

    • Before querying OpenCode’s SQLite database (typically located in user configuration paths), the script creates a temporary copy including associated -wal (Write-Ahead Log) and -shm (Shared Memory) sidecar files. This prevents tampering with or corrupting the original evidence, allowing safe read-only opening via SQLite URI modes.
    • Supports both separate message/part tables and newer consolidated session message storage schemas (session_message).
    • Processes complex turns containing text, assistant reasoning, tool calls, completion info, errors, and cost/token accounting metrics.
  • Multi-Artifact Collection Flow (hermes_forensic_extract.py):

    • Focuses on a broad extraction of multiple artifact types within the Hermes configuration directory (~/.hermes/).
    • Gathers state databases (state.db), full LLM API request and response payload dumps (request_dump_*.json), and agent, gateway, and GUI activity logs (logs/*.log).
    • Generates NDJSON/JSONL (Newline-Delimited JSON) output, where each record identifies its source category through the _extraction_type field (session, message, model_usage, request_dump, log_entry), preserving source metadata and original timestamps.
  • Temporal Filtering and Structuring:

    • Both scripts implement time range selectors (--start and --end) and file/directory switches (-f, -d) to narrow down analysis to specific investigative windows, simplifying work with mounted disk images or collected tarballs.

Affected Systems / Environments

The forensic analysis tools are designed to examine development environments and workstations utilizing local agent-based coding assistants. Specifically, the analyzed artifacts correspond to:

  • OpenCode Environment: Local SQLite databases of chat sessions and agent execution records.
  • Hermes Environment: User configuration directories (~/.hermes/), specific profiles, state databases (state.db), API network request dumps, and application logs.
  • Execution Requirements: Forensic workstations or systems under investigation with Python 3.10 or later installed.

Mitigation and Detection

Remediation

  • Artifact Access Control: Chat histories and API request dumps stored by AI assistants often contain leaked credentials, access tokens, or sensitive data inadvertently entered by developers. It is vital to enforce strict file permission controls on user configuration directories (~/.opencode/, ~/.hermes/).
  • Periodic Cleanup: Establish retention policies to purge historical chat logs and request dumps on high-risk development workstations.

Detection

  • Local Database Access Monitoring: Security teams should monitor for unauthorized access or bulk read operations targeting SQLite databases in hidden user paths associated with AI utilities.
  • Agent Activity Auditing: Proactively utilize forensic extraction scripts during internal audits to detect anomalous agent behaviors or unauthorized tool calls during software development.

“The useful forensic question is not just what appears in the structured export, but which source supports it—and what critical context still remains outside the tool’s scope.”

Wrapping Up

The introduction of specialized scripts for reconstructing AI agent activity marks a significant step forward in enhancing defensive teams’ capabilities to audit modern development environments. As LLM-based assistants gain greater autonomy and local execution power, having secure, standardized mechanisms to extract chat histories, reasoning traces, and API dumps without altering original evidence becomes a cornerstone for cyber incident investigations.

References