Repository navigation
fix(tool): record locally executed tools in OpenCode V2 - #134
Conversation
In OpenCode V2 tool events, `executed` means the model provider ran the tool. Locally executed tools always report `executed: false`, so the plugin skipped their spans, duration histogram, and commit detection, and reported `duration_ms: 0` in `tool_result`. Treat a tool as executed when the plugin observes `session.tool.called`. Malformed-input rejections and calls cancelled before publication never emit that event, so they are still not recorded as executions. Record provider execution as the `tool.provider_executed` span attribute, and accept both the top-level `executed` field and the nested `provider.executed` shape.
|
Non-blocking test coverage suggestion: strengthen the disabled tool-tracing test in Add a local tool sequence (
This would protect the behavior changed here: execution metadata and command capture must happen before the tracing-disabled early return, so disabling tool spans does not also disable metrics or commit detection. |
…bled Send a complete local shell call with tool tracing disabled, and assert that the plugin creates no tool span but still records duration, the tool_result log, and the git commit counter and log.
|
Thanks @dialupdisaster 🏓 Good catch: the old test only sent I replaced it in 23a6c22 and 7ae54b2 with a full local
I confirmed the test fails if the tracing-disabled early return moves above the metadata capture in |
## Description Add deterministic integration tests that run the real OpenCode SDK and telemetry plugin in isolated Bun subprocesses. A scripted OpenAI-compatible Chat Completions endpoint drives actual tool execution, while a temporary loopback OTLP/JSON receiver captures exported spans, metrics, and logs. No model credentials, Docker, or external collector are required. - Pin the OpenCode SDK and plugin development dependencies to matching `2.0.20` versions. - Add seven integration cases covering session/LLM telemetry, token attributes, trace parenting, real local `read` execution, and disabled tool tracing. - Include three explicit expected-failure regression cases for #134. These assert local tool spans and duration metrics; scenario setup and actual execution remain ordinary passing tests. Change both `test.failing` registrations to `test` when #134 lands. - Add `bun run test:integration` and documentation in `tests/integration/README.md`. The existing `bun test` CI step also runs this suite. - Isolate home, XDG directories, project config, database, and inherited environment; enforce a process deadline and clean up servers and temporary files. No production plugin code changes. ## Type of change - [ ] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to not work as expected) - [x] Documentation update - [ ] Refactoring (no functional changes) - [x] Chore (dependency updates, etc.) ## Checklist - [x] I have read the [CONTRIBUTING.md](../CONTRIBUTING.md) document - [x] My code follows the style guidelines of this project - [x] `bun run lint` passes with no errors - [x] `bun run check:jsdoc-coverage` passes with no errors - [x] `bun run typecheck` passes with no errors - [x] `bun test` passes with no errors - [x] I have added tests that prove my fix is effective or that my feature works - [x] I have updated the documentation accordingly - [x] My commits follow the [Conventional Commits](https://www.conventionalcommits.org/) specification ## Related issues Relates to #134. ## Additional context Validation on this branch: 198 tests pass, including the three expected-failure cases; typecheck, lint, and JSDoc coverage pass. Separately ran all seven integration cases against an isolated copy of the #134 commit with expected-failure markers removed, and all seven passed normally. The temporary receiver covers OTLP HTTP/JSON only. gRPC/protobuf transport coverage and additional scenarios (permissions, failures, commits, subagents) remain follow-up work.
|
Thanks for the contribution! |
🤖 I have created a release *beep* *boop* --- ## [2.0.1](v2.0.0...v2.0.1) (2026-10-06) ### Bug Fixes * **tool:** record locally executed tools in OpenCode V2 ([#134](#134)) ([129feb9](129feb9)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Description
With OpenCode V2, the plugin drops telemetry for every tool that OpenCode runs locally, such as
shell,read,edit, MCP tools, Code Mode, and subagents:opencode.tool.durationis never recorded.opencode.tool.*spans are created.opencode.commit.countand thecommitlog event never fire.tool_resultlog event reportsduration_ms: 0.Root cause: In V2 tool events,
executedmeans the model provider ran the tool, not that OpenCode ran it. OpenCode publishessession.tool.calledwithexecuted: tool.providerExecuted(publish-llm-event.ts#L458-L476). It then publishes the terminal event for a locally executed tool with the samefalsevalue (#L573-L591). Only provider-hosted tools, such as OpenAIweb_searchor Anthropiccode_execution, reportexecuted: true(#L477-L513).handleToolCalledandfinishToolgated spans, durations, and commit detection onexecuted, so the plugin measured only hosted tools.Fix: A tool counts as executed once the plugin observes
session.tool.called. OpenCode starts a local execution only after that event is published (step.ts#L100-L128). Calls that never ran don't publish it:#L315-L338).#L360-L375).Neither case is recorded as an execution.
session.tool.called, including a permission rejection or user decline (step.ts#L199-L207), is recorded as a failed execution (success=false). Its duration includes any permission wait.tool.provider_executedspan attribute. It accepts both the current top-levelexecutedand the nestedprovider.executedshape, which OpenCodedevuses (session-event.ts#L312-L372). That covers only the payload shape:devalso renames the events tosession.next.tool.*, which this PR doesn't address.session.tool.calledinstead of at the first progress event that carries a child session ID. Child-session linking is unchanged.Tests now cover local success and failure, local
git commitdetection, permission rejection, malformed input without a call, provider-hosted success and failure, the nestedprovider.executedshape, and a subagent that fails before creating a child. Two existing tests encoded the old assumption thatexecuted: falsemeant the tool didn't run, so I updated them.I also checked it end to end with OpenCode
2.0.20and a local OTLP/HTTP JSON sink. Areadcall produced a 33 ms span, histogram sample, andduration_ms. Ashellcall runningsleep 2recorded 2,834 ms. Agit committhroughshellincrementedopencode.commit.countand emitted thecommitlog event.Type of change
Checklist
bun run lintpasses with no errorsbun run check:jsdoc-coveragepasses with no errorsbun run typecheckpasses with no errorsbun testpasses with no errorsRelated issues
Independent of #112, which covers V1
message.part.updatedtimestamp restamping and orphan spans. This PR fixes V2 tool events that were never measured.Additional context
The
executedgating was introduced in #132, which treatedexecuted: falsecalls as invalid. That holds for malformed calls but also matches every successful local call.