---
name: arthas-diagnosis
description: "Diagnose Java applications, JVM metrics (CPU, memory, threads), trace methods, and inspect classes via Arthas MCP server."
homepage: https://github.com/alibaba/arthas
license: Apache-2.0
---

# Arthas Java Runtime Diagnosis Skill

This skill guides the AI assistant in performing real-time Java application diagnostics and troubleshooting using the **Arthas MCP (Model Context Protocol) Server**.

## When to Use

Activate this skill when:
- Investigating Java application performance issues (CPU spikes, memory leaks, slow latency).
- Finding deadlocks, blocked threads, or high-CPU threads.
- Inspecting loaded classes, classloaders, bytecode decompilation (`jad`), or method signatures (`sm`).
- Observing live method parameters, return values, and exceptions (`watch`, `trace`, `stack`).
- Interacting with live Java objects or triggering JVM-level tools (`vmtool`, `ognl`, `dashboard`).

---

## ⚠️ Safe Production Diagnostic Rules (Safety Boundary)

1. **Limit Observation Depth & Counts**:
   - For `watch`, `trace`, `stack`, ALWAYS specify `-n <count>` (e.g., `-n 5` or `-n 10`) to prevent unbounded logging and high overhead.
   - Avoid tracing heavily invoked framework methods (e.g., `String.equals`, `HashMap.get`, high-QPS filter chains) without specific conditional filters.
2. **Caution with Heapdump & Memory Allocation**:
   - Check available disk space before triggering `heapdump`.
   - Use `--live` flag where appropriate to dump only live reachable objects and reduce file size.
3. **Session Lifecycle Management**:
   - Remember to stop/reset active trace sessions to ensure bytecode instrumentation overhead is removed once the diagnosis is complete.

---

## Standard Diagnostic Playbooks

### 1. High CPU Troubleshooting Playbook
1. **Locate Hot Threads**: Run `thread -n 3` to find top 3 CPU consuming threads and their stacks.
2. **Check Deadlocks & Contention**: Run `thread -b` to identify threads holding blocking locks.
3. **CPU Sample Window**: Run `thread -i 1000 -n 3` to calculate real-time CPU percentage over 1 second.

### 2. Slow Response / Latency Tracing Playbook
1. **Locate Target Class & Method**: Run `sc -d *OrderService*` and `sm *OrderService* createOrder`.
2. **Trace Execution Path**: Run `trace com.example.service.OrderService createOrder -n 5 '#cost > 50'` to filter sub-calls exceeding 50ms.
3. **Inspect Callers**: Run `stack com.example.dao.OrderDao queryOrder -n 3` to find upstream callers.

### 3. Exception & Method Data Inspection Playbook
1. **Observe Parameters & Return Value**: Run `watch com.example.service.UserService getUser '{params, returnObj, throwExp}' -x 2 -n 5`.
2. **Capture Exceptions Only**: Run `watch com.example.service.PaymentService pay '{params, throwExp}' -e -x 2 -n 5`.
3. **Time-Tunnel Historical Record**: Run `tt -t com.example.service.OrderService calculateDiscount -n 5`, then inspect with `tt -l` and `tt -i <INDEX>`.

### 4. Class & Environment Verification Playbook
1. **Decompile Bytecode**: Run `jad com.example.config.AppProperties --source-only` to verify running bytecode matches source.
2. **Classloader Hierarchy**: Run `classloader -t` to inspect classloader delegation trees.
3. **Static Fields & Context**: Run `getstatic com.example.Constant HOLDER` or `vmtool --action getInstances --className com.example.service.UserService --limit 1`.
