Target Workflow: smoke-otel-tracing
Source report: #9065
Estimated cost per run: $47.06 (avg across 3 runs, $141.19 total)
Total tokens per run: ~13K avg (10.3K–16.3K range)
Cache hit rate: N/A (not present in pre-downloaded run data)
LLM turns/invocations: 15–26 per run (from working_set.invocations), well above the ~4.8K cumulative_input_tokens this implies heavy re-sending of context each turn (rebuild_factor ~1.0058–1.0104, so context reuse itself is fine, but invocation count is high for a smoke test)
Current Configuration
| Setting |
Value |
| Tools loaded |
bash: ["*"] (unrestricted), github: {toolsets: [repos, pull_requests]} |
| Tools actually used |
Only local bash (grep/node/npm/jest) for validation; github toolset used solely to post add_comment via safe-outputs, not via direct GitHub tool calls |
| Network groups |
defaults, node, github, *.ingest.us.sentry.io |
| Pre-agent steps |
Yes — 4 steps: already run OTEL tests/validation before the agent starts |
| Post-agent steps |
Yes — 2 post-steps: collect diagnostics and validate safe-outputs |
| Prompt size |
~2.35K chars (body) + ~5.9K chars (frontmatter) |
| Model |
auto (via GH_AW_MODEL_AGENT_COPILOT/GH_AW_DEFAULT_MODEL_COPILOT, defaults to auto → large-tier model) |
Recommendations
1. Restrict bash tool to specific commands instead of "*"
Estimated savings: ~1–2K tokens/run (~10-15%)
The workflow already precomputes all deterministic work in steps: (OTEL tests, module validation, env var grep checks) and the prompt instructs the agent to only read the output of those steps ("Check the output from... Report whether..."). The agent needs zero live bash execution — it's a read-and-summarize task. Restricting bash to an empty/minimal set (or removing the tool entirely and instead surfacing step outputs via env vars / step summary) removes an unbounded tool surface and reduces the risk of the agent re-running expensive commands (e.g., npm ci, npx jest) that were already run in pre-steps.
tools:
bash:
- "echo *" # only for optional debug echoing, or omit `bash:` entirely
github:
toolsets: [repos, pull_requests]
If the agent truly never needs bash (it's just reporting pre-computed results), remove bash: from tools: altogether — this is the highest-confidence, lowest-risk cut since all 5 scenarios in the prompt reference "the output from" a named pre-step, not live investigation.
2. Move result reporting into post-steps: to skip the agent turn entirely for non-PR runs
Estimated savings: ~8–13K tokens/run (~70-100%) for schedule/workflow_dispatch runs
The 5 "Scenario" prompt sections ask the agent to re-read step logs and produce a formatted summary — this is deterministic aggregation, not reasoning. For schedule and workflow_dispatch triggers (2 of 3 recent runs were pull_request, but this workflow also runs weekly on schedule), a post-steps: script can parse the step outputs (module load status, jest pass/fail counts, grep results, span counts) directly and call noop/write a step summary without invoking the LLM agent at all. For pull_request runs, keep the agent (needed to format an add_comment with human-readable ✅/❌ summary), but shrink the prompt (see #3).
post-steps:
- name: Emit deterministic scenario summary (schedule/dispatch only)
if: github.event_name != 'pull_request'
run: |
echo "## OTel Smoke Summary" >> "$GITHUB_STEP_SUMMARY"
# aggregate pass/fail from prior step outputs/log files here
This is the single biggest lever since schedule-triggered runs don't need a natural-language PR comment.
3. Trim the prompt body — collapse 5 verbose "Scenario" sections into a single instruction
Estimated savings: ~300-500 tokens/run input (~3-5%)
The prompt repeats "Check the output from the '(step name)' step. Report..." five times with step names duplicated from the steps: section. Since step names are already visible to the agent via the run context, a single consolidated instruction reduces prompt tokens without losing information:
## Reporting
Summarize the 5 validation steps executed above (module load, test suite, env var
forwarding, token-tracker integration, OTEL diagnostics) as ✅/❌ one-liners.
- pull_request trigger: add a comment with the summary
- otherwise: use noop
File an issue only for genuinely unexpected failures (not expected-pending states).
4. Drop the github network group if github: MCP toolset traffic is already covered by defaults
Estimated savings: ~0 tokens directly, but reduces attack surface / squid ACL size
Verify whether defaults already includes api.github.com/github.com (used elsewhere in the repo's smoke workflows without an explicit github group). If so, the explicit github entry in network.allowed is redundant. This doesn't save agent tokens but simplifies the generated Squid ACL and is worth a quick check during the same PR.
Expected Impact
| Metric |
Current |
Projected |
Savings |
| Total tokens/run (PR trigger) |
~13K |
~10-11K |
~20% |
| Total tokens/run (schedule/dispatch trigger) |
~13K |
~0.5K (no agent turn) |
~95% |
| Cost/run (avg) |
$47.06 |
~$35-38 (PR) / ~$2-5 (schedule) |
20-90% |
| LLM invocations/run |
15-26 |
5-10 (PR) / 0 (schedule) |
60-100% |
Implementation Checklist
Generated by Daily Copilot Token Optimization Advisor · copilot · auto · 44.2 AIC · ⊞ 10.6K · ◷
Target Workflow:
smoke-otel-tracingSource report: #9065
Estimated cost per run: $47.06 (avg across 3 runs, $141.19 total)
Total tokens per run: ~13K avg (10.3K–16.3K range)
Cache hit rate: N/A (not present in pre-downloaded run data)
LLM turns/invocations: 15–26 per run (from
working_set.invocations), well above the ~4.8Kcumulative_input_tokensthis implies heavy re-sending of context each turn (rebuild_factor~1.0058–1.0104, so context reuse itself is fine, but invocation count is high for a smoke test)Current Configuration
bash: ["*"](unrestricted),github: {toolsets: [repos, pull_requests]}bash(grep/node/npm/jest) for validation;githubtoolset used solely to postadd_commentvia safe-outputs, not via direct GitHub tool callsdefaults,node,github,*.ingest.us.sentry.iosteps:already run OTEL tests/validation before the agent startspost-steps:collect diagnostics and validate safe-outputsauto(viaGH_AW_MODEL_AGENT_COPILOT/GH_AW_DEFAULT_MODEL_COPILOT, defaults toauto→ large-tier model)Recommendations
1. Restrict
bashtool to specific commands instead of"*"Estimated savings: ~1–2K tokens/run (~10-15%)
The workflow already precomputes all deterministic work in
steps:(OTEL tests, module validation, env var grep checks) and the prompt instructs the agent to only read the output of those steps ("Check the output from... Report whether..."). The agent needs zero live bash execution — it's a read-and-summarize task. Restrictingbashto an empty/minimal set (or removing the tool entirely and instead surfacing step outputs via env vars / step summary) removes an unbounded tool surface and reduces the risk of the agent re-running expensive commands (e.g.,npm ci,npx jest) that were already run in pre-steps.If the agent truly never needs bash (it's just reporting pre-computed results), remove
bash:fromtools:altogether — this is the highest-confidence, lowest-risk cut since all 5 scenarios in the prompt reference "the output from" a named pre-step, not live investigation.2. Move result reporting into
post-steps:to skip the agent turn entirely for non-PR runsEstimated savings: ~8–13K tokens/run (~70-100%) for
schedule/workflow_dispatchrunsThe 5 "Scenario" prompt sections ask the agent to re-read step logs and produce a formatted summary — this is deterministic aggregation, not reasoning. For
scheduleandworkflow_dispatchtriggers (2 of 3 recent runs werepull_request, but this workflow also runsweeklyon schedule), apost-steps:script can parse the step outputs (module load status, jest pass/fail counts, grep results, span counts) directly and callnoop/write a step summary without invoking the LLM agent at all. Forpull_requestruns, keep the agent (needed to format anadd_commentwith human-readable ✅/❌ summary), but shrink the prompt (see #3).This is the single biggest lever since
schedule-triggered runs don't need a natural-language PR comment.3. Trim the prompt body — collapse 5 verbose "Scenario" sections into a single instruction
Estimated savings: ~300-500 tokens/run input (~3-5%)
The prompt repeats "Check the output from the '(step name)' step. Report..." five times with step names duplicated from the
steps:section. Since step names are already visible to the agent via the run context, a single consolidated instruction reduces prompt tokens without losing information:4. Drop the
githubnetwork group ifgithub:MCP toolset traffic is already covered bydefaultsEstimated savings: ~0 tokens directly, but reduces attack surface / squid ACL size
Verify whether
defaultsalready includesapi.github.com/github.com(used elsewhere in the repo's smoke workflows without an explicitgithubgroup). If so, the explicitgithubentry innetwork.allowedis redundant. This doesn't save agent tokens but simplifies the generated Squid ACL and is worth a quick check during the same PR.Expected Impact
Implementation Checklist
bash: ["*"]intools:— the agent only needs to read step outputs, not execute commandspost-steps:deterministic summary path for non-pull_requesttriggers, gated byif: github.event_name != 'pull_request'githubnetwork group is redundant withdefaultsand remove if sogh aw compile .github/workflows/smoke-otel-tracing.mdnpx tsx scripts/ci/postprocess-smoke-workflows.ts