179 lines
7.6 KiB
Markdown
179 lines
7.6 KiB
Markdown
# CAPABILITY_AUDIT.md
|
||
|
||
Last updated: 2026-03-08 (Asia/Shanghai)
|
||
Scope: **Current subagent practical abilities** in this OpenClaw instance (session-level, tool-level, policy-level).
|
||
|
||
## 1) Executive Summary (Practical)
|
||
|
||
This subagent can reliably:
|
||
- perform web research (`web_search`, `web_fetch`),
|
||
- automate browser workflows (`browser`),
|
||
- read/write/edit workspace files (`read`, `write`, `edit`),
|
||
- run local shell commands (`exec`, `process`),
|
||
- produce documentation/artifacts,
|
||
- interact with paired devices if available (`nodes`),
|
||
- send outbound messages only via explicit tool actions (`message`),
|
||
- generate speech (`tts`) when needed.
|
||
|
||
This subagent cannot:
|
||
- bypass tool/policy guardrails,
|
||
- guarantee access to external services without credentials/network,
|
||
- assume paired nodes exist,
|
||
- perform hidden background persistence on its own.
|
||
|
||
---
|
||
|
||
## 2) Available Tools by Category
|
||
|
||
## A. Filesystem & Content
|
||
- `read`: read file/image content.
|
||
- `write`: create/overwrite files.
|
||
- `edit`: exact-text surgical edits.
|
||
|
||
**Testable check**
|
||
1. `write` temp file with known text.
|
||
2. `read` file and verify exact content.
|
||
3. `edit` one token and verify diff.
|
||
|
||
## B. Local Execution
|
||
- `exec`: run shell commands (foreground/background, PTY optional).
|
||
- `process`: inspect/poll/kill ongoing exec sessions.
|
||
|
||
**Testable check**
|
||
1. Run `exec: echo ok`.
|
||
2. Run long command in background.
|
||
3. Use `process poll/log/kill` to verify lifecycle control.
|
||
|
||
## C. Web & Research
|
||
- `web_search`: Brave-based web search (localized/region filters).
|
||
- `web_fetch`: fetch readable page content (markdown/text extraction).
|
||
|
||
**Testable check**
|
||
1. Search a query with top 3 results.
|
||
2. Fetch 1 result URL and verify extract is non-empty.
|
||
3. Cross-check claim against at least 2 sources.
|
||
|
||
## D. Browser Automation
|
||
- `browser`: open/navigate/snapshot/act/screenshot/pdf/upload/dialog.
|
||
|
||
**Testable check**
|
||
1. Open target URL.
|
||
2. Snapshot and click/fill one element via ref.
|
||
3. Capture screenshot as evidence.
|
||
|
||
## E. Device / Node Operations
|
||
- `nodes`: status/describe/notify/camera/screen/location/run/invoke.
|
||
- `canvas`: UI rendering/snapshot/eval for node canvas surfaces.
|
||
|
||
**Testable check**
|
||
1. `nodes status` lists paired node(s).
|
||
2. Send one notification.
|
||
3. If permitted, capture one camera snap/screen record.
|
||
|
||
## F. Messaging & Voice
|
||
- `message`: send outbound messages via configured channel plugin.
|
||
- `tts`: convert text to speech (audio returned by tool).
|
||
|
||
**Testable check**
|
||
1. Send a test message to controlled channel.
|
||
2. Generate TTS clip from fixed sentence.
|
||
|
||
## G. Orchestration
|
||
- `subagents`: list/steer/kill sub-agents.
|
||
|
||
**Testable check**
|
||
1. `subagents list` returns active runs.
|
||
2. steer/kill only with explicit operator intent.
|
||
|
||
---
|
||
|
||
## 3) Current Limitations (Concrete)
|
||
|
||
1. **Policy-bounded toolset only**
|
||
- No direct access beyond listed tools.
|
||
- No arbitrary privileged system APIs outside tools.
|
||
|
||
2. **Environment-dependent success**
|
||
- Web calls depend on network + target availability.
|
||
- Browser automation can break on dynamic UI changes.
|
||
- Node operations require active paired devices and granted permissions.
|
||
|
||
3. **Credential-dependent actions**
|
||
- External account actions (mail/social/APIs) need existing auth context.
|
||
|
||
4. **No guaranteed long-term persistence**
|
||
- Session is ephemeral; durable state must be written to files.
|
||
|
||
5. **Safety constraints on destructive/external actions**
|
||
- Deletion, irreversible operations, and external posting may require explicit confirmation/authorization per policy/user prefs.
|
||
|
||
6. **Observability limitations**
|
||
- Tool output may be truncated/compacted; requires targeted re-read.
|
||
|
||
---
|
||
|
||
## 4) Risk Controls in Effect
|
||
|
||
1. **Tool allowlist enforcement**
|
||
- Only explicitly available tools may be invoked.
|
||
|
||
2. **Human oversight for sensitive actions**
|
||
- External sends, irreversible deletion, or privacy-impacting actions need explicit intent/approval.
|
||
|
||
3. **Non-manipulation / no self-escalation**
|
||
- No attempts to expand privileges, disable safeguards, or alter policy controls without explicit instruction.
|
||
|
||
4. **Execution hygiene**
|
||
- Prefer non-destructive checks first.
|
||
- For long-running tasks, use background sessions and controlled polling (avoid busy loops).
|
||
|
||
5. **Evidence-first reporting**
|
||
- Prefer commands, snapshots, logs, and file artifacts to support conclusions.
|
||
|
||
6. **Scope discipline**
|
||
- Subagent remains task-bounded; no unsolicited side effects.
|
||
|
||
---
|
||
|
||
## 5) Go / No-Go Matrix (Real-World Tasks)
|
||
|
||
Legend:
|
||
- **GO** = can execute now with current tools/policy.
|
||
- **GO-WITH-CONDITIONS** = feasible if prerequisite(s) met.
|
||
- **NO-GO** = not feasible/restricted in current context.
|
||
|
||
| Task Area | Typical Task | Status | Preconditions | Acceptance Test (Concrete) | Primary Risks / Controls |
|
||
|---|---|---|---|---|---|
|
||
| Research | Multi-source fact finding | GO | Network available | 3+ sources returned, 2-source corroboration, citations captured | Hallucination risk → enforce source quoting + cross-check |
|
||
| Research | Deep paywalled/proprietary data extraction | GO-WITH-CONDITIONS | Valid access/session credentials | Able to fetch/read target content with authenticated context | Access/legal risk → only use authorized sessions |
|
||
| Browser Ops | Form fill / navigation automation | GO | Reachable site, stable selectors | Complete scripted click/fill flow + screenshot proof | UI drift risk → snapshot refs and retry logic |
|
||
| Browser Ops | CAPTCHA bypass / anti-bot evasion | NO-GO | N/A | Fails without manual/authorized path | Compliance risk → require user-handled verification |
|
||
| Local Tooling | File transforms, scripts, CLI audits | GO | Workspace access, shell available | Command exit code 0 + artifact produced | Command safety → non-destructive defaults |
|
||
| Local Tooling | Privileged host reconfiguration (root-only) | GO-WITH-CONDITIONS | Host policy permits elevated exec | Explicit elevated command succeeds with audit log | System risk → explicit approval + reversible steps |
|
||
| Docs Generation | Specs, runbooks, status reports | GO | Target path writable | Markdown/doc generated and readable at path | Accuracy risk → include evidence references |
|
||
| Docs Generation | Signed/legal commitments on behalf of user | NO-GO | N/A | Cannot legitimately authorize legal intent | Authority risk → require human sign-off |
|
||
| Monitoring | Polling logs/processes in-session | GO | Command/tool access | Background process monitored via `process poll/log` | Resource risk → bounded intervals/timeouts |
|
||
| Monitoring | Persistent autonomous monitoring daemon deployment | GO-WITH-CONDITIONS | Explicit request + permitted runtime | Daemon/cron created and verifiable | Persistence risk → explicit opt-in + rollback |
|
||
|
||
---
|
||
|
||
## 6) Minimal Verification Playbook
|
||
|
||
Run these 8 checks to validate capability envelope quickly:
|
||
1. Files: write/read/edit roundtrip on temp file.
|
||
2. Shell: `exec` simple command + capture output.
|
||
3. Background control: long task + `process poll`.
|
||
4. Search: `web_search` returns relevant results.
|
||
5. Fetch: `web_fetch` extracts readable text from one result.
|
||
6. Browser: open page + snapshot + one click/fill + screenshot.
|
||
7. Docs: generate one markdown report file.
|
||
8. Nodes/Messaging (conditional): if configured, send test notify/message.
|
||
|
||
Pass criterion: ≥6/8 mandatory checks pass (1–6 required), and conditional checks pass when dependencies exist.
|
||
|
||
---
|
||
|
||
## 7) Operational Recommendation
|
||
|
||
Use this subagent as a **reliable execution auditor/operator** for: research, browser-assisted collection, local CLI workflows, and documentation pipelines. Treat node/device actions and persistent monitoring as **conditional** based on explicit enablement, permissions, and operator approval.
|