Files
val-blog/org/capability/CAPABILITY_AUDIT.md
T

7.6 KiB
Raw Blame History

CAPABILITY_AUDIT.md

Last updated: 2026-03-08 (Asia/Shanghai) Scope: Current subagent practical abilities in this OpenClaw instance (session-level, tool-level, policy-level).

1) Executive Summary (Practical)

This subagent can reliably:

  • perform web research (web_search, web_fetch),
  • automate browser workflows (browser),
  • read/write/edit workspace files (read, write, edit),
  • run local shell commands (exec, process),
  • produce documentation/artifacts,
  • interact with paired devices if available (nodes),
  • send outbound messages only via explicit tool actions (message),
  • generate speech (tts) when needed.

This subagent cannot:

  • bypass tool/policy guardrails,
  • guarantee access to external services without credentials/network,
  • assume paired nodes exist,
  • perform hidden background persistence on its own.

2) Available Tools by Category

A. Filesystem & Content

  • read: read file/image content.
  • write: create/overwrite files.
  • edit: exact-text surgical edits.

Testable check

  1. write temp file with known text.
  2. read file and verify exact content.
  3. edit one token and verify diff.

B. Local Execution

  • exec: run shell commands (foreground/background, PTY optional).
  • process: inspect/poll/kill ongoing exec sessions.

Testable check

  1. Run exec: echo ok.
  2. Run long command in background.
  3. Use process poll/log/kill to verify lifecycle control.

C. Web & Research

  • web_search: Brave-based web search (localized/region filters).
  • web_fetch: fetch readable page content (markdown/text extraction).

Testable check

  1. Search a query with top 3 results.
  2. Fetch 1 result URL and verify extract is non-empty.
  3. Cross-check claim against at least 2 sources.

D. Browser Automation

  • browser: open/navigate/snapshot/act/screenshot/pdf/upload/dialog.

Testable check

  1. Open target URL.
  2. Snapshot and click/fill one element via ref.
  3. Capture screenshot as evidence.

E. Device / Node Operations

  • nodes: status/describe/notify/camera/screen/location/run/invoke.
  • canvas: UI rendering/snapshot/eval for node canvas surfaces.

Testable check

  1. nodes status lists paired node(s).
  2. Send one notification.
  3. If permitted, capture one camera snap/screen record.

F. Messaging & Voice

  • message: send outbound messages via configured channel plugin.
  • tts: convert text to speech (audio returned by tool).

Testable check

  1. Send a test message to controlled channel.
  2. Generate TTS clip from fixed sentence.

G. Orchestration

  • subagents: list/steer/kill sub-agents.

Testable check

  1. subagents list returns active runs.
  2. steer/kill only with explicit operator intent.

3) Current Limitations (Concrete)

  1. Policy-bounded toolset only

    • No direct access beyond listed tools.
    • No arbitrary privileged system APIs outside tools.
  2. Environment-dependent success

    • Web calls depend on network + target availability.
    • Browser automation can break on dynamic UI changes.
    • Node operations require active paired devices and granted permissions.
  3. Credential-dependent actions

    • External account actions (mail/social/APIs) need existing auth context.
  4. No guaranteed long-term persistence

    • Session is ephemeral; durable state must be written to files.
  5. Safety constraints on destructive/external actions

    • Deletion, irreversible operations, and external posting may require explicit confirmation/authorization per policy/user prefs.
  6. Observability limitations

    • Tool output may be truncated/compacted; requires targeted re-read.

4) Risk Controls in Effect

  1. Tool allowlist enforcement

    • Only explicitly available tools may be invoked.
  2. Human oversight for sensitive actions

    • External sends, irreversible deletion, or privacy-impacting actions need explicit intent/approval.
  3. Non-manipulation / no self-escalation

    • No attempts to expand privileges, disable safeguards, or alter policy controls without explicit instruction.
  4. Execution hygiene

    • Prefer non-destructive checks first.
    • For long-running tasks, use background sessions and controlled polling (avoid busy loops).
  5. Evidence-first reporting

    • Prefer commands, snapshots, logs, and file artifacts to support conclusions.
  6. Scope discipline

    • Subagent remains task-bounded; no unsolicited side effects.

5) Go / No-Go Matrix (Real-World Tasks)

Legend:

  • GO = can execute now with current tools/policy.
  • GO-WITH-CONDITIONS = feasible if prerequisite(s) met.
  • NO-GO = not feasible/restricted in current context.
Task Area Typical Task Status Preconditions Acceptance Test (Concrete) Primary Risks / Controls
Research Multi-source fact finding GO Network available 3+ sources returned, 2-source corroboration, citations captured Hallucination risk → enforce source quoting + cross-check
Research Deep paywalled/proprietary data extraction GO-WITH-CONDITIONS Valid access/session credentials Able to fetch/read target content with authenticated context Access/legal risk → only use authorized sessions
Browser Ops Form fill / navigation automation GO Reachable site, stable selectors Complete scripted click/fill flow + screenshot proof UI drift risk → snapshot refs and retry logic
Browser Ops CAPTCHA bypass / anti-bot evasion NO-GO N/A Fails without manual/authorized path Compliance risk → require user-handled verification
Local Tooling File transforms, scripts, CLI audits GO Workspace access, shell available Command exit code 0 + artifact produced Command safety → non-destructive defaults
Local Tooling Privileged host reconfiguration (root-only) GO-WITH-CONDITIONS Host policy permits elevated exec Explicit elevated command succeeds with audit log System risk → explicit approval + reversible steps
Docs Generation Specs, runbooks, status reports GO Target path writable Markdown/doc generated and readable at path Accuracy risk → include evidence references
Docs Generation Signed/legal commitments on behalf of user NO-GO N/A Cannot legitimately authorize legal intent Authority risk → require human sign-off
Monitoring Polling logs/processes in-session GO Command/tool access Background process monitored via process poll/log Resource risk → bounded intervals/timeouts
Monitoring Persistent autonomous monitoring daemon deployment GO-WITH-CONDITIONS Explicit request + permitted runtime Daemon/cron created and verifiable Persistence risk → explicit opt-in + rollback

6) Minimal Verification Playbook

Run these 8 checks to validate capability envelope quickly:

  1. Files: write/read/edit roundtrip on temp file.
  2. Shell: exec simple command + capture output.
  3. Background control: long task + process poll.
  4. Search: web_search returns relevant results.
  5. Fetch: web_fetch extracts readable text from one result.
  6. Browser: open page + snapshot + one click/fill + screenshot.
  7. Docs: generate one markdown report file.
  8. Nodes/Messaging (conditional): if configured, send test notify/message.

Pass criterion: ≥6/8 mandatory checks pass (16 required), and conditional checks pass when dependencies exist.


7) Operational Recommendation

Use this subagent as a reliable execution auditor/operator for: research, browser-assisted collection, local CLI workflows, and documentation pipelines. Treat node/device actions and persistent monitoring as conditional based on explicit enablement, permissions, and operator approval.