Cost Management
If a task can be solved using deterministic tools, use deterministic tools. Only use agents when necessary, as they incur higher cost. Agentic workflows allow you to run deterministic tools first and gate agent execution using conditions, enabling workflows that avoid triggering agents most of the time and only use them when needed.
The cost of running an agentic workflow is the sum of two components: GitHub Actions minutes consumed by the workflow jobs, and inference costs charged by the AI provider for each agent run. For an overview of how billing works across providers and Copilot licensing, see Billing.
AI Credits (AIC)
Section titled “AI Credits (AIC)”AI Credits (AIC) are the primary metric for monitoring and budgeting inference costs in gh-aw. One AIC equals $0.01 USD. AIC values are computed from pricing data sourced from the models.dev catalog and appear in gh aw logs, gh aw audit, and run footer messages.
AIC is shown in the gh aw logs output table under the AIC column, in audit reports alongside raw token counts, and as {ai_credits_suffix} in workflow footer templates. For structured output, each run under .runs[] includes an aic field and each episode under .episodes[] includes total_aic.
Cost Components
Section titled “Cost Components”A typical run includes a short pre-activation job (10–30 seconds) and an agent job (1–15 minutes), each with ~1.5 minutes of runner setup overhead billed at standard GitHub Actions pricing. Inference is charged by the AI provider based on the engine and API token type used.
Monitoring Costs with gh aw logs
Section titled “Monitoring Costs with gh aw logs”The gh aw logs command surfaces per-run metrics — elapsed duration, token usage, AIC (AI Credits), and turn count — before you decide what to optimize. Use gh aw audit <run-id> to deep-dive into a single run’s token usage, tool calls, and inference spend; its Metrics and Performance Metrics sections cover token counts, AIC, turn counts, and estimated cost in one place. For cost trends across multiple runs, use gh aw logs --format markdown [workflow] to generate a cross-run report with anomaly detection.
View recent run durations
Section titled “View recent run durations”# Overview table for all agentic workflows (last 10 runs)gh aw logs
# Narrow to a single workflowgh aw logs issue-triage-agent
# Last 30 days for Copilot workflowsgh aw logs --engine copilot --start-date -30dThe overview table includes a Duration column showing elapsed wall-clock time per run. Because GitHub Actions bills compute time by the minute (rounded up per job), duration is the primary indicator of Actions spend.
Export metrics as JSON
Section titled “Export metrics as JSON”Use --json to get structured output suitable for scripting or trend analysis:
# Write JSON to a file for further processinggh aw logs --start-date -1w --json > /tmp/logs.json
# List per-run duration, tokens, and AIC across all workflowsgh aw logs --start-date -30d --json | \ jq '.runs[] | {workflow: .workflow_name, duration: .duration, tokens: .token_usage, aic: .aic}'
# AIC spend grouped by workflow over the past 30 daysgh aw logs --start-date -30d --json | \ jq '[.runs[]] | group_by(.workflow_name) | map({workflow: .[0].workflow_name, runs: length, total_aic: (map(.aic // 0) | add)})'Each run under .runs[] includes duration, token_usage, aic, workflow_name, and agent. For orchestrated workflows, the same JSON includes deterministic lineage under .episodes[] and .edges[] — see the next section.
Interpret Episode-Level Usage
Section titled “Interpret Episode-Level Usage”gh aw logs --json emits three views of the same data: .runs[] for individual workflow runs, .episodes[] for related runs grouped into one logical execution, and .edges[] for inferred parent-child lineage. Use .runs[] to find which run was resource-heavy and .episodes[] to answer “what did this job use end-to-end?”. For non-orchestrated workflows, an episode collapses to a single run.
For usage analysis, focus on total_runs, total_tokens, total_aic, total_duration, primary_workflow, resource_heavy_node_count, and blocked_request_count. For Claude, Codex, and Copilot runs, total_aic is the preferred cost metric because it maps to provider billing in AI Credits (1 AIC = $0.01 USD).
# Top 10 costliest logical executions over the past 30 days by AICgh aw logs --start-date -30d --json | \ jq '[.episodes[] | {episode: .episode_id, workflow: .primary_workflow, runs: .total_runs, aic: (.total_aic // 0)}] | sort_by(.aic) | reverse | .[:10]'
# Top 10 heaviest Copilot executions by AICgh aw logs --start-date -30d --engine copilot --json | \ jq '[.episodes[] | {episode: .episode_id, workflow: .primary_workflow, runs: .total_runs, aic: (.total_aic // 0)}] | sort_by(.aic) | reverse | .[:10]'Track Costs at Scale with OpenTelemetry
Section titled “Track Costs at Scale with OpenTelemetry”Use observability.otlp to stream run telemetry into a central OpenTelemetry backend when one repository or one gh aw logs report is no longer enough. This is best for organization-wide dashboards, alerting, and cross-repository cost analysis.
observability: otlp: endpoint: ${{ secrets.OTLP_ENDPOINT }} headers: Authorization: ${{ secrets.OTLP_TOKEN }}The exported spans include workflow and model metadata such as gh-aw.engine.id, gen_ai.request.model, gen_ai.usage.input_tokens, and gen_ai.usage.output_tokens. Use these attributes to group usage by workflow, engine, model, repository, or team in the backend of your choice. For inference cost, derive AIC from raw token counts using provider pricing.
OpenTelemetry helps answer questions like “Which repositories are driving the most token usage?”, “Which model change caused a cost spike?”, and “Which workflows should move to a smaller model or stricter trigger policy?” See the OpenTelemetry guide for collector configuration and the OpenTelemetry attribute reference for emitted fields.
Trigger Frequency and Cost Risk
Section titled “Trigger Frequency and Cost Risk”The primary cost lever for most workflows is how often they run.
| Risk level | Triggers | Why |
|---|---|---|
| High | push, check_run, check_suite | Fires on nearly every code change; busy repositories can generate many runs per day. |
| Medium-high | pull_request, issues | Fires on many lifecycle events such as open, sync, label, edit, and close. |
| Medium | issue_comment, pull_request_review_comment | Scales with discussion activity. |
| Low / predictable | schedule, workflow_dispatch | Fixed cadence or human-initiated, so easier to budget. |
Reducing Cost
Section titled “Reducing Cost”Use Deterministic Checks to Skip the Agent
Section titled “Use Deterministic Checks to Skip the Agent”The most effective cost reduction is skipping the agent job entirely when it is not needed. The skip-if-match and skip-if-no-match conditions run during the low-cost pre-activation job and cancel the workflow before the agent starts:
on: issues: types: [opened] skip-if-match: 'label:duplicate OR label:wont-fix'on: issues: types: [labeled] skip-if-no-match: 'label:needs-triage'Use these to filter out noise before incurring inference costs. See Triggers for the full syntax.
Skip the Agent from Steps Using noop
Section titled “Skip the Agent from Steps Using noop”When a condition is too complex for a GitHub search query — for example, when you need to call an API, inspect a file, or apply custom business logic — write a noop entry to $GH_AW_SAFE_OUTPUTS from a steps: block. The harness checks for this entry before starting the AI engine and exits cleanly without incurring any AI Credits.
steps: - name: Skip if no open issues exist run: | count=$(gh issue list --state open --json number --jq length) if [ "$count" -eq 0 ]; then echo '{"type":"noop","message":"No open issues to process"}' >> "$GH_AW_SAFE_OUTPUTS" fi env: GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}When a noop entry is present at harness startup, the agent is never started and no AI Credits are charged. The noop message appears in the workflow conclusion comment or step summary. The same check also suppresses retries: if a noop is written during a failed run, the harness exits 0 instead of retrying.
Use pre-agent-steps: instead of steps: when the check must run right before the engine starts (for example, after MCP configuration is complete).
Compared to skip-if-match and skip-if-no-match:
skip-if-match / skip-if-no-match | noop in steps: | |
|---|---|---|
| Evaluated in | Pre-activation job (earliest, cheapest) | Agent job (after checkout and steps) |
| Condition type | GitHub search query | Arbitrary shell or script logic |
| Actions minutes saved | Yes — agent job is never scheduled | No — agent job still runs through setup |
| AI Credits saved | Yes | Yes |
| Best for | Simple label/status/title filters | Complex API calls or file-based conditions |
For maximum savings, prefer skip-if-match / skip-if-no-match when possible. Reserve noop in steps: for conditions that require full scripting access or the agent job environment.
Choose a Cheaper Model
Section titled “Choose a Cheaper Model”The engine.model field selects the AI model. Smaller or faster models cost significantly less per token while still handling many routine tasks:
engine: id: copilot model: gpt-4.1-miniengine: id: claude model: claude-haiku-4-5Reserve frontier models (GPT-5, Claude Sonnet, etc.) for complex tasks. Use lighter models for triage, labeling, summarization, and other structured outputs.
Limit Context Size
Section titled “Limit Context Size”Inference cost scales with prompt size. Write focused prompts, avoid whole-file reads when only a few lines matter, cap result counts in tool calls, and use imports to compose a smaller subset of prompt sections at runtime.
Prevent Runaway Costs from Agents
Section titled “Prevent Runaway Costs from Agents”GitHub Agentic Workflows includes default guardrails: a 20-minute timeout on the agentic step, 1000 AI Credits per workflow run, and 5000 AI Credits per workflow per day (24-hour window). Override them with frontmatter (timeout-minutes, max-ai-credits, max-daily-ai-credits) or enterprise environment variables (GH_AW_DEFAULT_TIMEOUT_MINUTES, GH_AW_DEFAULT_MAX_AI_CREDITS, GH_AW_DEFAULT_MAX_DAILY_AI_CREDITS).
Cap AI Credits per Run
Section titled “Cap AI Credits per Run”Use the top-level max-ai-credits frontmatter field to cap the AI Credits (AIC) budget for a single workflow run. It provides a hard stop for unusually expensive runs and a consistent guardrail across all supported engines. The field accepts plain integers or K/M suffixes such as 100M.
max-ai-credits: 500When the budget is approached, gh-aw emits steering warnings before the run reaches the limit. Set a negative value only to disable budget enforcement explicitly.
Cap Turns per Run
Section titled “Cap Turns per Run”Use the top-level max-turns frontmatter field to cap the number of chat iterations (model responses and tool calls) for a single workflow run. Each additional turn consumes more tokens and Actions compute time, so a turn limit bounds both runaway loops and cost.
max-turns: 20max-turns works across Claude, Codex, Copilot, Gemini, and Pi. When set, gh-aw exports the compiled value as GH_AW_MAX_TURNS for the engine runtime, so you do not need to set CLAUDE_CODE_MAX_TURNS or an equivalent variable separately.
The field accepts integer literals or GitHub Actions expressions, making it composable with workflow_call inputs:
max-turns: ${{ inputs.max-turns || 15 }}An enterprise-wide default can be set with the compiler process environment variable GH_AW_DEFAULT_MAX_TURNS. Individual workflows override it by setting max-turns in frontmatter.
Cap Daily AI Credits per Workflow
Section titled “Cap Daily AI Credits per Workflow”Use max-daily-ai-credits to set a 24-hour AI Credits cap for one workflow. The guardrail sums runs from the past 24 hours of the same workflow across the repository, regardless of who triggered them.
max-daily-ai-credits: 15MWhen the total from the past 24 hours already meets or exceeds this threshold, the activation job warns, creates an issue, skips the agent job, and lets the conclusion job report the failure context.
The guardrail is disabled by default when omitted. Positive values accept plain integers or K/M suffixes such as 100M.
To disable the guardrail explicitly, set -1:
max-daily-ai-credits: -1You can also disable the guardrail at runtime by setting the env variable to -1:
env: GH_AW_MAX_DAILY_AI_CREDITS: "-1"For org-wide or enterprise-wide control, prefer setting the default once at the
org or enterprise level instead of per-workflow. Set
GH_AW_DEFAULT_MAX_DAILY_AI_CREDITS as an Actions variable
at the organization or enterprise scope so every workflow in the org
inherits it without any frontmatter change:
# Disable the daily guardrail org-widegh aw env update - --scope org --org MY_ORG <<'EOF'default_max_daily_ai_credits: "-1"EOFRoll out org/repo defaults with enterprise controls
Section titled “Roll out org/repo defaults with enterprise controls”For large installations, set baseline model and token guardrails once and let individual workflows override them only when needed:
- Export current defaults:
gh aw env get defaults.yml --scope org --org MY_ORG- Update and apply shared defaults in batch:
default_max_ai_credits: "5M"default_max_daily_ai_credits: "15M"default_model_copilot: "gpt-5-mini"default_model_claude: "claude-haiku-4-5"default_model_codex: "gpt-5.4-mini"gh aw env update defaults.yml --scope org --org MY_ORGgh aw env update shows a confirmation preview before applying changes. Pass --yes to skip the prompt in automation or --dry-run to preview without changing variables. Set a field to null to delete the corresponding variable from the target scope. Unknown YAML keys are rejected, default_max_turns and default_timeout_minutes must be positive integers, and default_max_ai_credits and default_max_daily_ai_credits must be non-zero integers; negative values disable the corresponding guardrail.
- If you compile workflows in CI, pass compiler-read defaults into the compiler process environment, for example via
${{ vars.* }}:GH_AW_DEFAULT_MAX_AI_CREDITS,GH_AW_DEFAULT_MAX_DAILY_AI_CREDITS,GH_AW_DEFAULT_MAX_TURNS,GH_AW_DEFAULT_TIMEOUT_MINUTES, andGH_AW_DEFAULT_DETECTION_MODEL.
Rate Limiting and Concurrency
Section titled “Rate Limiting and Concurrency”Use user-rate-limit to cap how many times a user can trigger the workflow in a given window, and rely on concurrency controls to serialize runs rather than letting them pile up:
user-rate-limit: max-runs-per-window: 3 window: 60 # 3 runs per hour per userSee Rate Limiting Controls and Concurrency for details.
Use Schedules for Predictable Budgets
Section titled “Use Schedules for Predictable Budgets”Scheduled workflows fire at a fixed cadence, making cost easy to estimate and cap. The less often a workflow runs, the lower the cost:
# Once per day on weekdays — 5 runs/weekschedule: daily on weekdays# Every two days — roughly 15 runs/monthschedule: every 2 days# Weekly on Monday mornings — 4–5 runs/monthschedule: weeklyWhen an event-based trigger fires far more often than the agent actually needs to act, a schedule is almost always cheaper. Replace push or issues triggers with a daily or weekly schedule and let the agent work through a backlog of items in one run.
See Schedule Syntax for the full fuzzy schedule syntax.
Batch Instead of Reacting to Events
Section titled “Batch Instead of Reacting to Events”Reactive triggers like issues or pull_request launch one agent run per event. When many events arrive in a short window, that adds up quickly. A scheduled batch run groups all pending items into a single invocation — and because the shared system prompt and instructions are sent once for the whole batch, AI providers can cache that context across items, further reducing AI Credits consumption.
description: Nightly issue triage (replaces reactive issues trigger)on: schedule: daily workflow_dispatch:
permissions: issues: readengine: id: copilot model: gpt-4.1-minitools: github: toolsets: [issues]---
Fetch all issues opened in the past 24 hours with no labels.For each issue, apply the most appropriate label. Process them in a single pass.Use Inline Sub-Agents with Smaller Models
Section titled “Use Inline Sub-Agents with Smaller Models”When a workflow delegates specialized tasks to sub-agents, each sub-agent can use a different model. Assign cheap, fast models to high-frequency sub-tasks (summarization, labeling, classification) and reserve frontier models only for the orchestrator.
engine: id: copilot model: smallpermissions: pull-requests: read
---
Use the `summarizer` sub-agent to summarize the diff, then post the result as a review comment.
## agent: `summarizer`---model: smalldescription: Summarizes a pull request diff in one paragraph---Read the diff and return a single paragraph describing what changed and why.See Inline Sub-Agents for the full syntax.
Use Inline Skills to Reduce Context
Section titled “Use Inline Skills to Reduce Context”Move large instruction blocks out of the main prompt body using inline skills. At runtime, each ## skill: block is extracted and written to engine-specific skill locations — the agent can invoke the skill on demand instead of receiving the guidance upfront, keeping the ambient context slim.
Each block ends at a matching ## end skill: \name`marker if present, or otherwise at the next##` heading or EOF. Add the explicit end marker when the skill is imported into the middle of a document, so content following it is not swallowed into the skill block.
Treat the main prompt as an execution plan and sub-skills as deferred detail:
- Main prompt: concise plan, sequencing, and decision points.
- Sub-skills: verbose checklists, report templates/layout rules, and domain rubrics.
- Invoke sub-skills only when needed (for example, at final report generation), not at startup.
This progressive-disclosure pattern keeps early turns focused and reduces per-run token overhead:
engine: id: copilot model: smallpermissions: issues: readtools: github: toolsets: [issues]
---
Triage the issue using the `triage-rules` skill.
## skill: `triage-rules`---description: Classify issues and suggest next actions.---Classify by bug / feature / question, identify missing information, and suggestthe smallest actionable next step.Agentic Cost Optimization
Section titled “Agentic Cost Optimization”The agentic-workflows MCP tool exposes the same operations as the CLI (logs, audit, status) to any workflow agent, so a scheduled meta-agent can inspect and optimize other agentic workflows automatically — fetching aggregate cost data, deep-diving into individual runs, and proposing frontmatter changes (cheaper model, tighter skip-if-match, lower user-rate-limit) via a pull request.
description: Weekly Actions minutes cost reporton: weeklypermissions: actions: readengine: copilottools: agentic-workflows:What to Optimize Automatically
Section titled “What to Optimize Automatically”Target workflows with high AIC per run by moving to a smaller model or reducing context size, cap high turn counts with max-turns, tighten skip-if-match when runs often produce no safe output, reduce queue pressure with user-rate-limit.max-runs-per-window or concurrency, and switch overly chatty workflows to schedule or workflow_dispatch.
Optimize at Scale with github/agentic-ops
Section titled “Optimize at Scale with github/agentic-ops”The githubnext/agentic-ops repository is the reference implementation for organization-wide agentic workflow monitoring and optimization. It applies the MonitorOps pattern to summarize spend, escalate failures, and propose workflow improvements on a schedule.
Common Scenario Estimates
Section titled “Common Scenario Estimates”These are rough budgeting estimates; actual costs vary by prompt size, tool usage, model, and provider pricing.
| Scenario | Frequency | Actions minutes/month | Inference/month |
|---|---|---|---|
| Weekly digest (schedule, 1 repo) | 4×/month | ~1 min | Varies by model and prompt size |
| Issue triage (issues opened, 20/month) | 20×/month | ~10 min | Varies by model and prompt size |
| PR review on every push (busy repo, 100 pushes/month) | 100×/month | ~100 min | Varies by model and prompt size |
| On-demand via slash command | User-controlled | Varies | Varies |
Related Documentation
Section titled “Related Documentation”- Audit Commands - Single-run analysis, diff, and cross-run reporting
- Artifacts - Artifact names, directory structures, and token usage file locations
- OpenTelemetry - Exporting workflow telemetry to centralized observability backends
- Triggers - Configuring workflow triggers and skip conditions
- Rate Limiting Controls - Preventing runaway workflows
- Concurrency - Serializing workflow execution
- AI Engines - Engine and model configuration
- Inline Sub-Agents - Defining sub-agents with per-task model selection
- Imports - Sharing workflow components across multiple workflows
- BatchOps - Grouping work items into scheduled batch runs
- MonitorOps - Scheduled monitoring and escalation for agentic workflows
- Compiler Enterprise Environment Controls - Default model and guardrail precedence
- Environment Variables - Variable scopes and compiler-managed defaults
- Schedule Syntax - Schedule format
- GH-AW as an MCP Server -
agentic-workflowstool for self-inspection - FAQ - Common questions about cost and billing