How to use AgentSight

Updated at:

AgentSight is an eBPF-based AI Agent observability tool that collects fine-grained data and correlates events across the Agent execution lifecycle without modifying business logic.

How to use AgentSight

Product introduction

AgentSight is an eBPF-based AI Agent observability tool that collects fine-grained data and correlates events across the Agent execution lifecycle without modifying business logic. You do not need to modify Agent code or configure an agent. After installation, AgentSight automatically discovers AI Agents running on the system and collects their LLM calls, token usage, and process behavior.

Core capabilities

AgentSight mainly includes the following capabilities:

  • Token consumption analysis: Comprehensive measurement and attribution of Token consumption during Agent operation. Supports flexible query by time period or the last N hours, and can automatically perform comparisons. It supports splitting consumption sources by multiple dimensions such as agents, tasks, roles, etc. The analysis granularity can be accurate to a single LLM call, and supports separate accounting of cached tokens.

  • Behavioral audit: Full-link recording and tracking of Agent's LLM calls and process execution behaviors. During data collection, key metadata such as provider and model version of each LLM call are completely retained, and the command line parameters of the process are synchronously captured. In addition, the system supports multi-dimensional flexible filtering by time dimension, process identification and event type, and provides visual summary statistical analysis capabilities.

  • Session interruption diagnostics: Automatically detects and attributes session exceptions across 18 categories, including rate limiting, authentication failure, truncated streaming responses, context overflow, process crashes, retry storms, and infinite loops. Each event includes a severity and can be handled in the Dashboard or queried through the CLI.

  • Optimization analysis: The built-in workspace analyzes a specified session for accuracy, performance, and cost. Some dimensions require an LLM to be configured. It identifies issues, recommends optimizations, and retains the analysis history for later review.

  • Trajectory collection and export: Converts Agent session traces to standard ATIF v1.7 for storage, export, offline analysis, replay, and evaluation. The Dashboard provides a trace list and sub-agent topology view.

  • Dashboard visualization: Web visual interface that provides intuitive display of Token consumption trends, Agent status monitoring, session interruption processing, Session details, optimization analysis, and Token savings. Token authentication is enabled by default and can be securely deployed on remote servers and accessed directly through local browsers. The interface supports bilingual display (Chinese and English), automatically matching the browser's preferred language.

Usage Scope

AgentSight automatically discovers and tracks the following Agents through built-in rules: OpenClaw, Copilot Shell (cosh and cosh-ng, non-AK/SK authentication scenarios), Claude Code, Codex CLI, QwenCode, Hermes, AgentScope, etc. If you need to track other Agents, you can customize process matching rules through the configuration file.

For Claude Code built with Bun, version 2.1.113 or later is required to capture its LLM traffic.

Installation method

For details, please refer toQuick start. After the system service is installed, eBPF tracing (trace) and API service (serve) are started with the system by default, and there is no need to manually execute them.

Conversational interaction method

AgentSight provides conversational interaction Skill, supporting installation and use in various AI Agents. Users do not need to memorize CLI commands and can complete operations directly through natural language:

  • View Token consumption: Such as "How many tokens were used today?"

  • Query audit logs: Such as "Help me check today's LLM call record"

  • Troubleshoot session exceptions: Such as "Have any sessions been interrupted recently? Please help me find out why."

If you are using Copilot Shell (cosh), the Skill is built-in, and you can directly use the above natural language instructions. The system will automatically call AgentSight to complete the query and return the analysis conclusion.

CLI command details

agentsight trace — Start eBPF Tracing

Note: This service has been started by default in the system and does not need to be executed manually.

Start eBPF-based AI Agent activity tracking.

sudo agentsight trace

agentsight serve — Start API and Dashboard

Note: This service has been started in the system by default and is bound to 0.0.0.0:7396 by default. There is no need to execute it manually.

Start an HTTP API server to provide an embedded Dashboard UI.

sudo agentsight serve --host 0.0.0.0 --port 7396 

This command will bind all network interfaces and can be accessed through the server's public IP: http://<server public IP>:7396

Make sure that the server firewall or security group allows inbound traffic on port 7396.

agentsight summary — One-stop status overview

This command summarizes session and token usage for the latest time window, interruption events by severity, and token savings. Use it for routine checks or to quickly understand the overall status of Agents on the host. If a data source is unavailable, only the corresponding result is omitted; the remaining results are unaffected.

agentsight summary            # Overview of the last 24 hours
agentsight summary --last 48  # Last 48 hours
agentsight summary --json     # JSON output

agentsight token — Query token usage

Query Token usage data.

# Check today's usage
agentsight token

agentsight audit — Query audit events

Query audit events (LLM calls, process operations).

# View recent events
agentsight audit
# Filter by PID and type
agentsight audit --pid 12345 --type llm
# Summary statistics
agentsight audit --summary

agentsight discover — Scan Agent

Discover the AI Agents running on the system.

# Scan Agent
agentsight discover
# List known types
agentsight discover --list-known

agentsight interruption — Query and manage session interruptions

Query session interruption events, filter them by severity, collect statistics by type, view event details, and mark events as resolved.

agentsight interruption list                    # List outage events in the last 24 hours
agentsight interruption list --severity high    # Filter by severity
agentsight interruption stats                   # Statistics by type
agentsight interruption get <ID>                # View individual outage details
agentsight interruption resolve <ID>            # Mark as solved

agentsight dashboard — Get dashboard access information

Display Dashboard access address, authentication status, and login token; in the ECS environment, security group release guidance is also output.

agentsight dashboard

Dashboard visual interface

Dashboard is AgentSight's Web visual interface for viewing conversation history, Trace details, and Token statistics. Dashboard enables token authentication by default, and the first access requires obtaining the login token through the agentsight dashboard command;

Dashboard functionality

Dashboard provides the following core functionality:

  • Token consumption overview: View the current machine's token consumption within the selected time period. A time range selection control is provided at the top of the Dashboard to switch between different time periods; below, statistical cards display input Token, output Token, and total Token usage.

  • Agent status: The right-side status bar can view the current Agent process status and provides Agent process hang and restart functionality. Health cards simultaneously display the Agent's LLM response latency metrics and retain historical activity records, allowing review of its operational interval even after process exit.

  • Session interruption diagnostics: Independent interruption handling page, where session lists identify exceptions with tags, and interruption detail panels display the type, severity level, and error message, and support Resolve or Hide actions.

  • Session details: Click "Details" to view detailed token usage for each session and trace.

  • Optimization analysis: Select session running accuracy, performance, and cost dimensions for analysis, with cost dimension supporting "detour"-based waste analysis, and analysis history being reviewable.

  • Token savings: Integrated with the Tokenless component (automatically effective when both are installed), displaying total consumption, reduced tokens, and reduction rate, supporting baseline comparison, breakdown by optimization strategy, and row-level comparison before and after optimization.

Data management

Database management

Automatic capacity limit and cleanup: To prevent database infinite growth and excessive disk space occupation, the system defaults to setting the database maximum capacity to 200 MB. When the database size reaches the upper limit, automatic cleanup is triggered.

Users can customize the maximum capacity (unit: MB) via the environment variable AGENTSIGHT_GENAI_DB_MAX_SIZE_MB, for example, setting it to 500 MB.

export AGENTSIGHT_GENAI_DB_MAX_SIZE_MB=500 

Capacity exceeding triggers row-by-row trimming of oldest data, rather than full database clearance, ensuring historical data isn't lost entirely. The interruption event database retains data for 30 days by default with a 100 MB upper limit, adjustable in the configuration file.

Clean historical data

To clean historical data, execute the following:

rm -rf /var/log/sysak/.agentsight

Then restart AgentSight.

Configuration Management

Default path: /etc/agentsight/config.json (alternative paths can be specified via runtime --config).

Configuration Block

Description

cmdline.allow / cmdline.deny

Agent process command-line matching rules that determine which processes to track

https

Domain filtering rules that capture only LLM traffic for matching domains

features

Feature switches that enable or disable interruption detection, trajectory collection, auditing, and other capabilities

runtime_limits

Memory and buffer limits that prevent unbounded resource usage

The default https rules cover Bailian (DashScope, including *.maas.aliyuncs.com) and OpenAI domains. Add domains manually for other LLM services.

Note: User configuration files completely replace built-in default rules, rather than merging or appending. Missing rules for an Agent in custom configurations means it won't be discovered; ensure retention of all required monitoring Agent rules before modification. Configuration format upgrades are automatically managed by the schema_version field. When a version becomes outdated, the program performs an automatic backup and migration.

FAQ

Q1: Why can't I obtain OpenClaw's Token consumption data?
A: AgentSight monitors the openclaw-gateway daemon process. Please check whether the connection status between the client and Gateway is normal. If the following exception log appears, it indicates pairing failure:
Gateway agent failed; falling back to embedded: Error: gateway closed (1008): pairing required.
It is recommended to execute the command openclaw devices approve to complete device pairing.


Q2: Why does the Token savings page not display the current Session ID, or show zero Token savings?
A: This may be caused by the following two reasons:

1. The current version does not support Cosh's AK/SK authentication method;

2. The Session ID format is non-standard UUID, causing system matching failure.

Q3: Why is the "Optimization item savings" displayed on the Token savings page greater than the difference between "Tokens before optimization" minus "Tokens after optimization"?

A: This is because the Agent incorporates historical messages into context for each conversation. Therefore, the current conversation's statistical results include optimization benefits from historical messages, causing cumulative savings to exceed the immediate difference of a single conversation.

Q4: Why does the browser prompt login when opening the Dashboard?

A: New versions enable token authentication by default to protect data security. Execute agentsight dashboard on the server to obtain the login token and enter it on the login page.

Q5: The Agent is running, but no LLM call data is being collected?

A: Troubleshoot in this order: confirm that the Agent command line matches a cmdline.allow rule; confirm that the domain used for LLM calls is included in the https rules; if you use a Bun-based build of Claude Code, upgrade to version 2.1.113 or later.