<!--
Sitemap:
- [Installation](/installation)
- [Upgrading](/upgrading): Version-specific steps for upgrading an existing Bento install.
- [Concepts](/concepts)
- [Build your first pipeline](/tutorials/pipeline-args)
- [Target a specific issue or PR from a URL](/tutorials/url-targeting)
- [Keep state across runs](/tutorials/pipeline-state)
- [Fire a pipeline on a schedule or on demand](/tutorials/schedule-and-fire)
- [Deploy a box to Railway](/tutorials/deploy-to-railway)
- [Operate a hosted daemon](/tutorials/operate-a-hosted-daemon)
- [Configuration](/configuration)
- [Members](/members)
- [Knowledge base](/knowledge-base/)
- [Method and delivery](/knowledge-base/modes)
- [Config](/knowledge-base/config)
- [MCP](/knowledge-base/mcp)
- [Pipeline configuration reference](/pipelines/config)
- [Filters](/pipelines/filters)
- [Triggers](/triggers/)
- [GitHub trigger](/triggers/github)
- [Linear trigger](/triggers/linear)
- [Webhook trigger](/triggers/webhook)
- [Schedule trigger](/triggers/schedule)
- [Manual trigger](/triggers/manual)
- [Traces](/pipelines/traces)
- [Slack](/integrations/slack)
- [Public access](/public-access)
- [Context engineering](/context-engineering)
- [Best practices](/best-practices)
- [Troubleshooting](/troubleshooting)
- [Architecture](/architecture/vision)
- [Workspaces](/workspaces)
- [Authentication](/authentication)
- [Identity](/identity)
- [Security](/security)
- [References](/references)
- [Changelog](/changelog): Bento release history.
- [CLI reference](/cli/)
- [Setup](/cli/setup)
- [Secrets](/cli/secrets)
- [Lifecycle](/cli/lifecycle)
- [Sandbox image](/cli/image)
- [Sandboxes](/cli/sandbox)
- [Observability](/cli/observability)
- [Diagnostics](/cli/diagnostics)
- [Triggers](/cli/triggers)
- [Workbench](/cli/workbench)
- [Auth](/cli/auth)
- [Knowledge](/cli/knowledge)
- [Evals](/cli/evals)
- [Bento](/index)
- [Runtime wrapper](/architecture/runtime-wrapper)
- [Skill evolve](/architecture/skill-evolve)
-->

# Pipeline configuration reference

Each file under `.bento/pipelines/` defines one pipeline: a configuration that connects a trigger to an agent, an optional skill, an `instructions` directive, and optional structured result projections.

## Pipeline structure

Minimal pipeline — a scheduled notification with no event payload:

```yaml
# .bento/pipelines/morning-standup.yaml
trigger:
  schedule: "0 9 * * 1-5"
agent: notifier
instructions: |
  Post a short summary of yesterday's merged PRs and open review requests.
result:
  schema:
    type: object
    properties:
      message:
        $ref: bento://schemas/slack/message/v1
    required: [message]
outputs:
  message:
    handler: slack.message
    config:
      bot: teamBot
      channel: "#standup"
```

Full-featured pipeline — a webhook trigger with filtering, skill, guardrails, sandbox mounts, and orchestrator override:

```yaml
# .bento/pipelines/pr-review.yaml
trigger:
  github:
    - pull_request.opened
    - pull_request.synchronize
  filter:
    repos: [my-org/backend]
    mentioned: false
    not:
      comment.user.login: my-bot
agent: reviewer
skill: agentic-review
targets: [codex-high, claude-high]
instructions: |
  Review PR #{{event.pull_request.number}} ("{{event.pull_request.title}}")
  on {{event.repository.full_name}}, opened by
  @{{event.pull_request.user.login}}, for merge readiness.
result:
  schema:
    $ref: bento://schemas/pipelines/review-replies-result/v1
outputs:
  review:
    handler: github.review-replies
guardrails:
  timeout: 900
sandbox:
  backend: podman
  mounts:
    - host: ~/.config/gh
      container: /home/node/.config/gh
    - host: ~/.gitconfig
      container: /home/node/.gitconfig
orchestrator:
  strategy:
    type: auto
context:
  diff: true
```

## trigger

What causes the pipeline to fire automatically. Declare at least one of `github`, `linear`, `webhook` (see [Webhook trigger](/triggers/webhook)), or `schedule`. You declare `schedule` alongside an event trigger. The other three are event arms, and they combine too. Omit `trigger` entirely for a pipeline that only ever runs on demand, because you fire every pipeline with `bento trigger fire` whatever it declares here. Add `patterns` (a list of URL prefixes) to route a pasted URL to this pipeline — see [URL routing](/triggers/manual#url-routing).

### github

A list of GitHub webhook event entries. Each entry is either a bare event type string or a per-event object with its own predicate. The pipeline fires when any entry matches the arriving event and passes the `filter`.

```yaml
trigger:
  github:
    - pull_request.opened
    - pull_request.synchronize
    - pull_request_review_comment.created
    - pull_request_review.submitted
    - issue_comment.created
    - issues.assigned
    - deployment_status
```

Event strings follow the pattern `<event_type>.<action>`. The event type comes from the `X-GitHub-Event` request header of GitHub. The action comes from the `action` field in the payload body.

The object form adds a per-event `when`/`not` predicate and inline `options` overrides evaluated only when that entry fires:

```yaml
trigger:
  github:
    - pull_request.opened
    - event: pull_request.review_requested      # object form
      when: { requested_reviewer.login: ava-lh }
      skip_reviewed_head: false                 # overrides pipeline-level options
  filter:
    when:
      pull_request.draft: false                 # ANDs with per-event when
```

The per-event `when` and `not` combine with the shared `filter.when` under AND. The daemon evaluates the entries in declared order. The first whose event type matches and whose predicate passes wins, and its inline options override the pipeline-level `options` for that dispatch only. When the predicate of one matching entry fails, evaluation continues to the later entries for the same event type. The pipeline fails to fire only when no matching entry passes its predicate.

### linear

A list of Linear webhook event strings.

```yaml
trigger:
  linear:
    - Issue.create
```

`AgentSessionEvent.created` and `AgentSessionEvent.prompted` fire when someone delegates an issue to the installed Linear application, or replies in the session thread. A Linear agent session names an issue and carries no repository, so a pipeline on either event must also set `repo` (and usually `branch`) — the same two fields a `schedule` trigger uses. Config load rejects a session pipeline that names no repo, or one whose repo is not the `url` of a `repos:` entry.

```yaml
trigger:
  linear:
    - AgentSessionEvent.created
    - AgentSessionEvent.prompted
  repo: my-org/backend
  branch: main
```

Every run on one session shares one workspace, keyed `linear:<organization>:session/<session>`, so a `prompted` follow-up sees the notes and prior runs of the delegation that opened the thread.

A `linear` entry takes the same object form as a `github` entry, with its own `when`, `not`, and the subject allowlists `labels`, `not_labels`, and `states`. Two `created` entries split a delegation from a mention. Linear records the comment a mention opened the session from under `agentSession.sourceMetadata`, and leaves it null on a delegation. Every session carries a `commentId`, so that field does not tell the two apart.

```yaml
trigger:
  linear:
    - event: AgentSessionEvent.created          # a mention, ungated
      when:
        agentSession.sourceMetadata.type: { exists: true }
    - event: AgentSessionEvent.created          # a delegation, gated on the issue
      when:
        agentSession.sourceMetadata.type: { exists: false }
      states: [Candidate]
      labels: [agent-ready]
      not_labels: [deferred]
    - AgentSessionEvent.prompted
  repo: my-org/backend
  branch: main
```

The daemon reads the issue's state and labels from Linear at match time, because the session payload carries neither. See [filters](./filters#filterlabels).

Other Linear events are data webhooks. They name no session, so the built-in extractor resolves no workspace for them; declare a [`workspace`](#workspace) template when such a pipeline needs a checkout.

### schedule

A 5-field cron pattern: `minute hour day month weekday`.

```yaml
trigger:
  schedule: "0 9 * * 1-5"   # 09:00 Monday–Friday
```

```yaml
trigger:
  schedule: "0 3 * * *"     # 03:00 daily
```

```yaml
trigger:
  schedule: "*/5 * * * *"   # every 5 minutes
```

The string form uses a five-minute maximum tick age. Use the object form to change it:

```yaml
trigger:
  schedule:
    cron: "0 3 * * *"
    max_age: 15m
```

`max_age` accepts a positive duration in seconds, minutes, hours, or days, such as `30s`, `5m`, `2h`, or `1d`. At dequeue, the daemon records a cron tick older than this limit as `discarded`, and it creates no workload. The age includes queue wait. Use a larger value for sparse schedules that share workers with long-running pipelines. Manual `bento trigger schedule <name>` requests do not expire.

Schedule triggers have no incoming event payload. `{{event.*}}` substitutions in `instructions` are unavailable, and the daemon uses the directive unchanged.

Two fields apply to `schedule` triggers and to `linear` session triggers, and to no others — a `github` trigger reads the repository from its payload:

| Field | Description |
|-------|-------------|
| `repo` | Repository to clone before the run, as `owner/name`. Must match the `url` field (not the map key) of a `repos:` entry in `daemon.yaml`. |
| `branch` | Branch to check out. Defaults to the repo's configured branch. |

```yaml
trigger:
  schedule: "0 3 * * *"
  repo: my-org/backend
  branch: main
```

Omitting `repo` means no checkout — the agent runs without a codebase.

### debounce

Collapses a burst of events on one subject into a single run. A pipeline on `pull_request.synchronize` starts one run for each push, so three pushes in one minute start three runs on the same pull request.

```yaml
trigger:
  github:
    - event: pull_request.synchronize
  debounce:
    wait: 120s
    max_wait: 10m
```

| Field | Description |
|-------|-------------|
| `wait` | **Required.** The quiet period the subject must go without a new event before the run starts. A positive duration in seconds, minutes, hours, or days, such as `30s`, `5m`, `2h`, or `1d`. There is no default. |
| `max_wait` | **Optional.** The limit on how long new events defer the run, measured from the first event in the burst. Omit it for no limit. The format is the same as `wait`, and the value must not be shorter than `wait`. |

A pipeline with `debounce` holds its run. The daemon queues the job with a start time `wait` after the event. Each later event on the same subject replaces that job and moves the start time out again. The burst becomes one run, and that run carries the last event. The subject is the pipeline and its workspace. For a pull request event, the workspace is the head branch. The daemon keeps the hold in the queue table, so a restart does not release the run early and does not lose it. Each event the daemon collapses gets the `superseded` status, which names the event that replaced it. A pipeline without `debounce` starts its run immediately.

`max_wait` limits that hold. Without it, a subject that never goes quiet never runs, because each event moves the start time out again. The daemon records the time of the first event in the burst and keeps that time on the held job. A later event that would move the start time past `max_wait` after the first event does not move it. The daemon takes only that event's payload. The run then starts no later than `max_wait` after the first event, and it still carries the last event.

A missing `wait`, a duration the daemon cannot parse, an unknown field under `debounce`, and a `max_wait` shorter than `wait` are startup errors. A typo stops the config load. A `max_wait` shorter than `wait` is an error because the subject could never go quiet for the full `wait`. Every run would start at the limit, and `wait` would decide nothing.

### filter

Narrows which events fire the pipeline. Every field you specify matches.

| Field | Type | Description |
|-------|------|-------------|
| `when` | map | Payload predicate. Every entry matches. Keys are dot-path expressions into the payload object. Values use the leaf vocabulary: scalar = strict equality, array = OR (matches any member), `{ exists: bool }` = presence test. A missing payload path fails a scalar/array match and satisfies `{ exists: false }`. |
| `not` | map | Inverse predicate — rejects the event if ANY entry matches. Same leaf vocabulary as `when`. |
| `mentioned` | boolean | `true` — bot was @-mentioned in the triggering text. `false` — bot was not mentioned. Omit to match either. |
| `repos` | list | The event comes from one of these repositories (`owner/repo` format). |
| `labels` | list | The pull request carries at least one of these label names. Reads `pull_request.labels[]` plus the triggering label on a `pull_request.labeled` event. A label sits in an array of objects, which a `when` path does not reach, so this is a named field. |
| `assignee` | string | The assignee field of the event equals this value. |

```yaml
filter:
  when:
    review.state: approved
    pull_request.user.login: my-bot
  not:
    review.user.login: my-bot
    comment.user.login: my-bot
  mentioned: true
  repos: [my-org/backend, my-org/frontend]
  assignee: bot
```

## agent

**Required.** The name of the agent persona to run. Must match a directory under `agents/<name>/` containing a `PERSONA.md` file.

```yaml
agent: reviewer
```

## skill

**Optional.** The name of a skill to give the agent. Must match a directory under `skills/<name>/` containing a `SKILL.md` file. Several pipelines reference the same skill with different agents.

```yaml
skill: agentic-review
```

## companions

**Optional.** Additional skills mounted alongside the governing [`skill`](#skill) as read-only reference. Each one matches a directory under `skills/<name>/`. Requires a governing `skill`. The governing skill stays the authoritative procedure. Companions carry no pointer, and the runtime surfaces them on relevance through its own skill discovery.

Delivery puts the skill trees where the runtime discovers them: Docker and Podman bind-mount them, and a remote backend (Cloudflare, Daytona) uploads them. Bento refuses a run that declares companions where neither is available, such as a non-`claude` runtime or the host backend. It never runs the pipeline without them.

```yaml
skill: solve-linear-issue
companions: [linear]
```

## version

**Optional.** SemVer string identifying this pipeline's procedure revision, surfaced in the attribution footer of everything it posts.

When the pipeline governs exactly one skill (a `skill` with no `companions` and no orchestrator capabilities), you omit `version`, and the daemon inherits the skill's own declared version. For every other pipeline — multi-skill, companion-bearing, orchestrator-enabled, or skill-free — declare `version` explicitly. Validation fails a pipeline that is ambiguous and undeclared.

```yaml
skill: agentic-review
version: 1.4.2
```

## principals and principal

**Optional.** `principals` binds pipeline-owned role names to canonical provider principals. A value is one `<provider>.principal.<id>` reference or a list containing at most one reference per provider. `principal` selects the binding used by a direct actor run.

```yaml
skill: agentic-review
principals:
  pull-request-author: github.principal.your-org-author-agent
  pull-request-reviewer: github.principal.your-org-review-agent
principal: pull-request-reviewer
```

A skill declares several required binding names in its `SKILL.md` frontmatter:

```yaml
principals:
  - pull-request-author
  - pull-request-reviewer
```

Every required name exists in the pipeline. A direct `principal` selects a role that the pipeline skill or a companion declares. A capability `principal` selects a role that the skill of that capability declares. Skills and companions name roles only. The pipeline supplies concrete provider references. Generic orchestrator tasks inherit the pipeline's selected role. Model-created task input never replaces it. Observer and dry-run execution neither select nor receive a principal. The legacy `github.username` configuration supplies one implicit GitHub principal only for pipelines without explicit bindings.

## env

**Optional.** The environment variables injected into this pipeline's run sandbox (BIP-20). This manifest is the single, operator-authored declaration of what a run receives. Each entry is one of three things: a variable name, forwarded from the daemon's environment; a computed `{name, command, ttl?, required?}` object, whose value a host command mints at spawn time; or a stored-secret `{name, secret, required?}` object, whose value the daemon reads from its [secret source](/cli/secrets) at spawn time.

```yaml
skill: solve-linear-issue
companions: [linear]
env:
  - SENTRY_AUTH_TOKEN # forwarded from the daemon environment
  - name: NPM_TOKEN   # minted at spawn time
    command: vault kv get -field=token secret/ci/npm
    ttl: 3600
    required: true
  - name: CLOUDFLARE_API_TOKEN            # read from the secret source at spawn time
    secret: sandbox/cloudflare-api-key
    required: true
```

For a computed entry, `command` runs on the daemon host — the same trust domain as `setup:` steps — and its trimmed stdout becomes the value. Commands have a 60-second timeout; a timeout follows the same required-or-optional failure behavior as a nonzero exit. Errors omit command output and secret-source exception text to avoid exposing credentials. Empty output is a failure, not an empty value. `ttl` caches the minted value in memory, keyed by pipeline, variable name, command, and TTL, for that many seconds. Omit it (or set `0`) to mint every run. With `required: true`, a failed mint fails the invocation before the agent spawns. Without it, a failed mint withholds the variable and the run proceeds.

For a stored-secret entry, `secret` is a full secret name, `<kind>/<leaf>`. The daemon reads it through its secret source when it prepares the run. The value travels on the injection channel, and the daemon does not log it. An entry carries `secret` or `command`, not both, and a `secret` entry carries no `ttl`. An entry MUST NOT name a secret of kind `oauth`, because a principal binding provisions that credential, and validation rejects the entry. With `required: true`, the daemon checks the name at startup and refuses to boot when the source does not hold it or is not reachable. An absent value at spawn time then fails the invocation before the agent spawns. Without `required`, the daemon does not check the name at startup, and an absent value withholds the variable.

The daemon validates the manifest at startup, where every name entry is set in its environment. It mints a computed entry at runtime, so that entry skips the check but carries a non-empty `command`. It reads a stored-secret entry at runtime, so that entry skips the environment check too. The manifest covers every variable that the skill, the agent, or a companion of the run declares in its own `env:`, and a declared need the manifest omits is a config error that the daemon names at boot. Injection is actor-mode only: the daemon rejects a non-empty manifest on a `read_only` pipeline, because an observer run receives no secrets.

A manifest entry may not name a variable in the `BENTO_` namespace. The daemon sets those variables itself, about the run — which pipeline matched, which invocation carries it, where the checkout sits — and injects them after the manifest. A manifest entry naming one fails startup validation.

Agent processes default to `TZ=UTC`. To select another time zone, declare `TZ` in the pipeline environment manifest. For example, a computed entry with `name: TZ` and `command: printf Europe/London` selects that zone.

Bento interprets Codex reset times without a zone in the time zone of the agent process. Claude reset messages supply their own zone.

## targets

**Optional.** An ordered list of names from the daemon's [`targets`](/configuration#targets) registry. The first target is primary. A target's `credential_pool` entries are tried in order before later targets when the agent CLI reports a quota or authentication failure. Omit the field to use the agent's frontmatter runtime and model with the runtime's implicit host login.

```yaml
agent: reviewer
targets: [codex-high, claude-high]
```

Bento preflights authentication and tries the primary first. It advances within the same run only when the agent CLI reports a quota or authentication failure. A subscription cap — session, daily, weekly or monthly — cools that credential down until the reset time it reports, so concurrent runs skip it. Cooling the last live credential of a runtime shortens that window to 15 minutes instead of honouring the reported reset, so a runtime that loses its whole pool recovers without an operator. Bento also retests a cooling credential about every 45 minutes, so a reset read longer than the real cap clears early rather than standing until the reported instant. A transient server-side failure (Anthropic 529 overloaded, dropped API socket) retries the same candidate in-process up to two times before stopping — it does not advance to the next candidate. Infrastructure and unclassified failures stop the candidate loop immediately. Authentication failure removes later candidates that use the same runtime and credential. The pipeline's `guardrails.timeout` is one deadline shared by all attempts. A fallback receives only the remaining time.

## target\_routes

**Optional.** A map of GitHub label name to an ordered target list, used to pick the specialist route by a label on the triggering PR. When the PR carries a label that is a key here, that route's targets are used in place of [`targets`](#targets) — resolved identically (first primary, rest fallbacks). The first key in declaration order the PR carries wins. A PR matching no key falls back to `targets`.

```yaml
agent: reviewer
targets: [codex-high, claude-high] # default: no matching label
target_routes:
  tier:low: [claude-sonnet-high]
  tier:high: [claude-opus-high, codex-high]
```

Keys are arbitrary label names, so a pipeline routes on any label family (`tier:*`, `priority:*`, …) — the daemon has no built-in label vocabulary. The daemon rejects an integer-like key, such as a label named `1`, at startup, because JavaScript enumerates those numerically rather than in declaration order. At startup the daemon also validates each route target for its target-name, credential, and runtime-auth references, the same as `targets`. The label set comes from the webhook payload's `pull_request.labels` plus the triggering `label`.

## instructions

**Required.** The user message sent to the agent. Supports `{{event.*}}` template substitutions, where `event` is the full webhook payload object, and `{{args.*}}` for declared [inputs](#args). `{{lists.<name>}}` expands a named list from the daemon's [`lists`](/configuration#lists) config (items joined by newlines).

A schedule trigger carries no event data. Use a static directive or declared `args:` defaults.

Examples:

```yaml
# PR review
instructions: |
  Review PR #{{event.pull_request.number}} ("{{event.pull_request.title}}")
  on {{event.repository.full_name}}, opened by
  @{{event.pull_request.user.login}}, for merge readiness.
```

````yaml
# Inline review comment
instructions: |
  A reviewer left a comment on a pull request.

  Repo:     {{event.repository.full_name}}
  PR:       #{{event.pull_request.number}} — {{event.pull_request.title}}
  Reviewer: @{{event.comment.user.login}}
  File:     {{event.comment.path}}:{{event.comment.line}}

  Diff hunk:
  ```
  {{event.comment.diff_hunk}}
  ```

  Comment:
  {{event.comment.body}}
````

```yaml
# Issue assignment
instructions: |
  You were assigned to an issue.

  Repo:   {{event.repository.full_name}}
  Issue:  #{{event.issue.number}} — {{event.issue.title}}
  Author: @{{event.issue.user.login}}
  URL:    {{event.issue.html_url}}

  Body:
  {{event.issue.body}}
```

## args

**Optional.** Typed inputs the pipeline accepts from synthesized surfaces (manual / url / schedule). Each input is keyed by name with a `type` and optional `required` / `default`.

| Field | Type | Description |
|-------|------|-------------|
| `type` | `string` | `number` | `boolean` | The input's value type. `number`/`boolean` coerce from string inputs. Anything unrecognized is treated as `string`. |
| `required` | boolean | When `true`, supply the input or bento rejects the fire. |
| `default` | string | Applied when the input is omitted (an empty string `""` is a valid sentinel). |
| `from` | string (template) | Derive the input from the trigger itself — a template rendered against `{{event.*}}` (parsed payload) and, on the url surface, `{{url.*}}` (`host`, `pathname`, `segments.N`, `query.X`). An explicit `--arg` overrides it. An empty render falls back to `default`. |

```yaml
args:
  persona:
    type: string
    required: true
  tone:
    type: string
    default: terse
  issue:
    type: string
    default: ""                      # blank → self-select
    from: "{{event.issue.number}}"   # a pasted issue URL targets this issue
instructions: |
  Draft a persona ({{args.tone}} voice): {{args.persona}}
```

Validated values substitute into `instructions` as `{{args.<name>}}`, distinct from the `{{event.*}}` payload. Validation is **strict**. Bento rejects an undeclared key, a missing required input, and a wrong type, and a `default` fills an omitted input. A pipeline with no `args:` accepts no inputs. Pass them with `bento trigger fire --arg key=value`, or map them off the trigger with `from:` — `{{url.*}}` lets a pipeline target off any URL shape with no daemon-side code (see [manual triggers](/triggers/manual#typed-inputs)). Targeting flags (`--branch`/`--sha`/`--pr`) stay separate and never collide with a declared input.

Trust follows the source. Bento trusts an operator-supplied `--arg`. An arg that `from:` fills on a webhook or URL surface carries untrusted payload text, because validation checks the type and not the origin, so bento wraps it in `<untrusted>`, as it does for `{{event.*}}`.

## Terminal projections

Use `result:` and `outputs:` for daemon-managed terminal delivery. The daemon rejects the legacy `output:` field at startup. A pipeline without `outputs:` keeps its stdout opaque. Actor workflows and local run artifacts stay outside terminal projections.

## result and outputs

Use `result:` when a pipeline emits structured terminal data. The daemon validates the complete JSON object once after the agent exits. A result is one whole JSON object or one fenced `json` object. Other stdout remains opaque text.

`outputs.<key>` sends only `result.<key>` to a daemon-owned handler. A pipeline does not run handler code. It only references a handler schema and name.

```yaml
result:
  schema:
    $ref: bento://schemas/pipelines/review-replies-result/v1
outputs:
  review:
    handler: github.review-replies
```

The `github.review-replies` handler posts replies, resolves the replied threads, posts a summary, and then re-requests review from each standing `CHANGES_REQUESTED` reviewer whose comments the run answered, when the run moved the pull request head. The summary carries one daemon-written line above the agent's text, naming the commits pushed since the run was queued or saying none were: the head comparison is read from GitHub, not from the agent's account of its own run. The daemon records one projection claim and one effect delivery for each provider request. A pipeline with no `result:` keeps its existing opaque-text behavior.

The `github.review` handler publishes one pull request review from a verdict. The agent returns the verdict, the inline findings, and the labels; the daemon posts the review, dismisses the previous Bento reviews, and applies the labels after the review lands. Its claim key contains the pull request and its head commit, thus a second agent that reviews the same commit publishes nothing.

```yaml
result:
  schema:
    type: object
    properties:
      review:
        $ref: bento://schemas/github/review/v1
    required: [review]
outputs:
  review:
    handler: github.review
```

## guardrails

Limits applied to the agent run. Merged per-key over `defaults.guardrails` from `daemon.yaml` — the pipeline wins where both set the same field.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `timeout` | integer (seconds), or a map of label name to seconds | `600` | Kill the run after this many seconds. Primary and fallback attempts share this one deadline. A map routes the budget by a label on the triggering event — see below. |
| `max_prompt_tokens` | integer (estimated tokens) | `50000` | Refuse to dispatch when the assembled prompt (system + user) exceeds this budget. Tokens are estimated as chars / 4 — an estimate, not an exact count. Prior runs shrink to fit the budget first (see `options.history.max_bytes`), so the error fires only when the remaining blocks are themselves over budget. The run fails before any sandbox is provisioned. The error itemizes per-block sizes and lands in the trigger's terminal error (`bento trace`). `bento trigger inspect` warns when a matched pipeline's prompt is over budget. |
| `allowed_tools` | list of strings | — | Restrict the agent to these tools (passed to the runtime as `--allowedTools`). Omit for no restriction. |
| `read_only` | boolean | `false` | Sets `BENTO_READ_ONLY=1` in the agent's environment. Advisory only — agent prose honors it, but the daemon does not mount the sandbox read-only and nothing blocks a write, a commit, or a push. |

```yaml
guardrails:
  timeout: 900
  allowed_tools: [Read, Grep, Glob]
  read_only: true
```

### Routing the timeout by label

`timeout` also takes a map, so one pipeline can give a large change a longer budget than a small one. `default` is required and is the budget for an event that carries none of the label keys. The first label key the event carries wins, in declaration order — the same match [`target_routes`](#target_routes) makes, against the same labels, so a run's model and its budget always agree on which label won.

```yaml
targets: [codex-high, claude-high]
target_routes:
  tier:high: [claude-opus-high, codex-high]
guardrails:
  timeout:
    default: 1800
    tier:high: 2700
```

A PR labelled `tier:high` runs for 2700 seconds; any other PR runs for 1800. The daemon resolves the map to one number at dispatch, so the run's `timeout=` log line and `bento trace` show the resolved seconds, not the map. The daemon rejects a map with no `default`, a non-integer value, and an integer-like label key, at startup and in `bento config validate`.

## sandbox

Overrides the sandbox configuration for this pipeline.

| Field | Type | Description |
|-------|------|-------------|
| `backend` | `docker` | `podman` | `daytona` | Use this sandbox backend instead of the daemon default. The named backend's connection config still comes from the top-level [`sandboxes:`](/configuration#sandbox) block — `backend: daytona` here needs `sandboxes.daytona` declared. |
| `class` | string | map | Name of the [compute class](/configuration#compute-classes) this pipeline's sandboxes are sized from, declared in the backend's `sandboxes.<backend>.classes` registry. A map routes the class by label on the triggering event and requires a `default` key. Falls back to the registry's own `default`. |
| `image` | string | Override the sandbox image for this pipeline's runs. Rendered as a template against `{ payload, event }`. Falls back to the matched repo's [`sandbox.image`](/configuration#repos), then the daemon-wide image (or `agent:default`). Accepts the same three forms as the repo-level setting: `devcontainer`, a `./`-prefixed repo-relative path, or an image reference. Used to pin a pipeline to an instance-specific image. |
| `snapshot` | string | Select a ready repository snapshot by its `owner/repository:tag` alias. Bento uses the matching backend and compute class. `none` selects no snapshot. |
| `mounts` | list | Additional host paths delivered into the sandbox. Each entry has `host` (path on the host), `container` (path inside the sandbox), and optional `readOnly` (boolean — mount the path read-only). Docker and Podman bind-mount the path. A remote backend (Cloudflare, Daytona) uploads a copy instead, so the agent sees no later host changes and writes nothing back; an entry with `readOnly: false` asks for write-back and fails the run. |
| `cache` | list | Dependency paths persisted across runs, keyed by a fingerprint. Each entry has `path` (the in-sandbox path to persist), `key` (the fingerprint template), and optional `restoreKeys` (prefix fallbacks tried in order on an exact-key miss). |
| `prewarm` | string | Command run once inside the sandbox, before the agent, to populate a cold or key-missed `cache`. |

```yaml
sandbox:
  backend: docker
  mounts:
    - host: ~/.config/gh
      container: /home/node/.config/gh
    - host: ~/.gitconfig
      container: /home/node/.gitconfig
```

Mount host credentials the agent needs at runtime.

When `snapshot` is omitted and the pipeline uses `agent:default`, Bento tries the alias that matches the checked-out repository and branch. If no ready snapshot matches, Bento starts from the base image.

Set `snapshot: none` to start from the base image with no snapshot. The run skips the repository's [declared snapshot](/configuration#repos) and the branch alias. On Cloudflare, the run opens in the bridge's default `Sandbox` container. The run takes no slot from the declared snapshot's `max_concurrent`, and it still counts against the daemon-wide [concurrency limit](/configuration#queue) of its pool. [`options.skip_checkout`](#options) does not imply `none`.

### Compute class

`class` names the [compute class](/configuration#compute-classes) the pipeline's sandboxes are sized from. A string names it outright.

```yaml
sandbox:
  class: large
```

A map selects the class from event labels, like [`guardrails.timeout`](#guardrails). It requires a `default` key for events with no matching label. The first matching key in declaration order wins. The class, specialist route, and run budget use the same event labels. Each map selects its first matching key independently.

```yaml
sandbox:
  class:
    default: small
    tier:high: large
```

A PR labelled `tier:high` runs on `large`; any other PR runs on `small`. The daemon resolves the map to one name at dispatch, so `bento trigger inspect` shows the resolved class, not the map. The daemon rejects a map with no `default`, an integer-like label key, and a name the backend's `classes` registry does not hold, at startup and in `bento config validate`.

### Dependency cache

`cache` persists a dependency path across runs so a cold install is paid once per key instead of once per run ([BIP-26](https://github.com/1a35e1/bento/blob/develop/docs/bip/0026-sandbox-dependency-caching.md)).

```yaml
sandbox:
  cache:
    - path: /pnpm-store                          # in-sandbox path to persist
      key: "pnpm-{{ hashFiles('pnpm-lock.yaml') }}"
      restoreKeys: ["pnpm-"]                     # nearest-prefix fallback
  prewarm: pnpm fetch                            # optional, populate on a miss
```

The following on-demand pipeline keeps the pnpm content store between Daytona runs:

```yaml
# .bento/pipelines/daytona-tests.yaml
agent: solver
instructions: |
  Install dependencies from the cached pnpm store:

    pnpm install --offline --frozen-lockfile --store-dir /home/agent/.cache/pnpm-store

  Run the repository test suite and report the result.
sandbox:
  backend: daytona
  cache:
    - path: /home/agent/.cache/pnpm-store
      key: "pnpm-{{ hashFiles('pnpm-lock.yaml') }}"
  prewarm: pnpm fetch --store-dir /home/agent/.cache/pnpm-store
guardrails:
  timeout: 1200
options:
  stateless: true
```

Run `bento trigger fire daytona-tests` to start it. The first run downloads the packages. Later runs with the same lockfile restore them before the install.

`key` is a template. `hashFiles('<glob>')` is the one function it may call, and it hashes the content of the files the globs match. Those files are usually the lockfile. A key with no `{{ }}` is a valid constant that only a config edit invalidates. The daemon scopes each store by repository, in-sandbox path, and key, so a key names content, not the triggering event: dot-path interpolation such as `{{ payload.repository.name }}` is rejected.

Validation runs in two phases. An unreadable template or an unknown function fails startup, naming the pipeline and the entry index. An absolute glob or one that traverses outside the checkout (`../`) fails startup too. Globs resolve per run against the cloned workspace, so a glob matching no file fails that run rather than the boot.

`path` must be absolute, and it may not overlap a path the daemon mounts itself: `/workspace`, `/home/agent/.claude`, `/home/agent/.claude.json`, `/home/agent/.config/gh`, `/home/agent/.gitconfig`, or `/home/agent/.codex`. A cache at one of those would take the mount the run's own credentials need, and a store persists across runs and pipelines. Startup rejects it, naming the entry.

`restoreKeys` never copies a store another run has mounted — a live run writes into its store, and copying one mid-install yields a torn cache. The prefix falls through to the next candidate, or to an empty store, which costs one cold install.

Cache the package manager's content-addressable store (the pnpm, cargo, or go store), not `node_modules`. The store is keyed by content hash, so a changed lockfile adds new content and everything already present is reused. `node_modules` is specific to one lockfile and goes stale when the two drift apart.

`prewarm` runs inside the sandbox before the agent. Docker and Podman run it when a store is cold or its exact key missed. Daytona runs it for each cache attachment because a volume can survive an interrupted populate. A warm package store makes the repeated command inexpensive. The command receives the pipeline's [`env:`](#env) manifest, so a private-registry populate reads its token from there. Its stdin is closed and its output goes to the run's stderr log. Bento stops the populate after 60 seconds, keeps at most 64 KiB of its output, and reports a non-zero exit as `bento: sandbox prewarm exited <code>` without failing the run. A custom image without the `timeout` utility, or with a `timeout` implementation that lacks the required `--signal` and `--kill-after` controls, skips the populate. The container backend limits the complete stderr log to 1 MiB. Bento never infers an install command. Without `prewarm`, only the agent command fills the cache. A `setup:` step cannot fill it. Setup steps run on the daemon host and never receive the cache path.

Size and retention bounds for Docker and Podman host stores are daemon-wide, under [`sandboxes.cacheStore`](/configuration#cachestore). Daytona volume retention is not yet automatic.

On Docker and Podman, a run resolves each key and bind-mounts the host store read-write at `cache.path`. Prefixes in `restoreKeys` select the nearest host store after an exact-key miss.

On Daytona, Bento resolves the key before it creates the sandbox. Bento attaches a repository-scoped managed volume and restores the newest complete archive to `cache.path` on the sandbox disk. This local path avoids the small-file cost of the remote volume. After the agent exits, Bento publishes a new archive and retains the two newest complete archives for that exact key. Daytona ignores `restoreKeys`. An exact-key miss starts with an empty cache.

An observer-mode run mounts no cache because a cache is writable. A run with no repository also mounts no cache because the repository is part of the cache identity. A run that mounts no cache does not run `prewarm`. The `just-bash` backend uses the package-manager store on the host and does not use a keyed Bento store.

Merged per-key over `defaults.sandbox` from `daemon.yaml` — the pipeline wins where both set the same field (`backend`, `image`, `mounts`, `cache`, `prewarm`), so a daemon-wide default `backend` combines with a pipeline-specific `image`. Per-key means a pipeline's own `cache` replaces the default list rather than adding to it.

## retain\_checkout

**Optional.** When `true`, the thread's `checkout/` directory — the per-thread worktree and any deps installed by `setup:` steps — is kept on disk between runs. Default: the daemon discards it when the thread's last active run ends (the checkout is regenerable and, at review volume, is the dominant disk cost). Set this for pipelines whose repeated runs amortize a warm dependency install.

Workbench threads (`workspace.checkout`) always retain regardless of this setting.

```yaml
retain_checkout: true
```

## knowledge

How knowledge base content reaches this pipeline's runs — two settings, `method` (how the daemon retrieves) and `delivery` (how results reach the agent). See [Method and delivery](/knowledge-base/modes) for what each value does and how they compose. Merged per-key over `defaults.knowledge` from `daemon.yaml` (`filter` replaces wholesale). Omit both and runs get no knowledge injection.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `method` | `bm25` | `vector` | `hybrid` | `none` | `hybrid` | How the daemon retrieves — see the table below. `none` means no daemon-side retrieval (the agent fetches knowledge itself). |
| `delivery` | `prompt` | `file` | `browse` | `search` | `none` | `prompt` | How results reach the agent. A primary (`prompt`/`file`/`browse`) composes with `search` as `prompt+search` etc. (content for grounding, the MCP tool as escape hatch). |
| `instructions` | string | per-delivery default | Operator guidance on how to use the knowledge. When set, it renders in the user message and replaces the default system instruction. |
| `filter.categories` | list of strings | — | Only docs whose frontmatter `category` is listed. ORs within, ANDs with `tags`. |
| `filter.tags` | list of strings | — | Only docs carrying at least one listed tag. |
| `topK` | integer | `5` | Results per source (retrieving methods). |
| `rerank` | boolean | `false` | Enable the retrieval engine's optional reranker. This adds query latency. |
| `maxTokens` | integer | `8000` | Injection budget (chars/4 estimate). |

| `method` | How the daemon retrieves |
|------|---------------------|
| `bm25` | Lexical BM25 retrieval over the indexed sources. |
| `vector` | Vector search. Requires embeddings (`qmd embed`). Errors loudly when missing. |
| `none` | No daemon-side retrieval. |

| `delivery` | What the agent gets |
|------|---------------------|
| `prompt` | Retrieved excerpts as `<context>` blocks in the user message. Each block carries the doc's path, category, and tags as attributes. |
| `file` | With a retrieving method: the excerpts written to `.bento/retrieved-knowledge.md` in the checkout. With `method: none`: knowledge dirs delivered read-only at `/knowledge/<type>` plus a generated `index.tsv` and `README.md` — bind-mounted on docker/podman, uploaded to `/tmp/bento-knowledge/<type>` on cloudflare/daytona (`filter` is a config error). |
| `browse` | The filtered doc listing. The agent reads docs on demand via MCP resources. |
| `search` | No content — a system nudge toward the `ask_knowledge` MCP tool. |
| `none` | Nothing. |

A delivery whose prerequisites are missing fails the run at dispatch rather than spawning an agent without them: `search` requires the `ask_knowledge` tool (qmd installed, knowledge sources configured), `browse` requires discovered docs, both require a local sandbox backend (the daemon's MCP endpoint is loopback-only). Whole-corpus `file` (`method: none`) requires the knowledge directories on disk and a backend that bind-mounts or uploads them — docker, podman, cloudflare, or daytona, but not the host. The error names the missing capability and the fix.

```yaml
knowledge:
  method: none
  delivery: browse
  filter:
    categories: [conventions]
  instructions: |-
    Prefer recently-edited docs; cite the doc URI for anything you rely on.
```

Each retrieving run records method, delivery, query, and injected counts on the retrieve task — visible in [`bento trace`](/pipelines/traces).

## options

Per-pipeline behavior switches.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `review_on_push` | boolean | `true` | When `false`, `pull_request.synchronize` events do not fire this pipeline. |
| `skip_reviewed_head` | boolean | `false` | When `true`, discard a dequeued PR job when the PR's live head already carries a review marker (`<!-- bento-review sha=... -->`) posted by the daemon's own login. Rapid pushes enqueue one job per push, but every run reviews the live head — once one run covers it, the remaining jobs are duplicates. Dismissed reviews do not count, so dismissing bento's review lets the next job run. Best-effort: if the GitHub API call fails, the run proceeds. |
| `skip_unchanged_head` | boolean | `false` | Scheduled pipelines only. When `true`, skip a fire when the branch head equals the head recorded by this pipeline's last completed run. The trigger is discarded before any agent spawns. Best-effort: if the fetch of the head fails, the run proceeds. |
| `skip_checkout` | boolean | `false` | When `true`, the run gets no repository checkout even though the event names a repo. For a pipeline that reasons over the event payload alone — triage, labelling, notification — the clone is latency and disk spent on a working tree the agent never reads. The thread is unaffected: the workspace key still derives from the repo and subject, so notes and prior runs behave exactly as they do for a checkout run. The prompt states that no repository is checked out, and the pre-clone GitHub auth check is skipped. A scheduled pipeline says the same thing by omitting `trigger.repo`. Config load rejects this field paired with a `devcontainer` or `./`-prefixed [`sandbox.image`](#sandbox): both forms resolve from the checkout the pipeline declines. Use a plain image reference instead. |
| `stateless` | boolean | `false` | When `true`, the pipeline carries no cross-run memory: prior runs and notes are neither injected into the prompt nor persisted from agent output. For recurring pipelines whose runs are unrelated (e.g. a liveness probe). |
| `state` | `true` | `{ schema }` | — | Persistent per-pipeline JSON state (BIP-22). `true` enables untyped state. `{ schema }` adds a JSON Schema surfaced to the agent in the tool description. Exposes `read_pipeline_state` / `write_pipeline_state` tools to the orchestrator — never prompt-injected automatically. Orchestrator pipelines only. Local backends only (state lives on the daemon host). |
| `graph` | `true` | `{ outputs }` | — | Durable work graph (BIP-32). `true` gives the orchestrator the graph tools — `ingest_graph`, `read_graph`, `claim_node`, `settle_node`, `add_nodes` — bound to graphs whose `pipeline` field is this pipeline's name. The graph persists in Postgres, so a later run resumes it. `{ outputs }` names handlers to fire on `graph_completed`, `graph_failed`, or `graph_cancelled` — each declaration is one graph-anchored claim per graph and event. Orchestrator pipelines only. Config load rejects a `graph` value with no `orchestrator` block, or with no `trigger.schedule` arm — the driver resumes the owner through the schedule surface, so a graph pipeline with no arm stops after its first run. |
| `history.max_runs` | integer | `10` | Maximum number of prior runs injected into the prompt, newest first. The on-disk `runs/` history stays complete — only prompt injection is capped. |
| `history.max_bytes` | integer (bytes) | derived | Byte budget across the rendered prior-runs block. Older runs are dropped whole (never truncated mid-entry) until the kept runs fit. When unset, the budget is whatever remains under `guardrails.max_prompt_tokens` after every other prompt block — prior runs shrink to fit rather than tripping the prompt-budget error. Set a value to pin the budget explicitly. |

When the budget drops runs, the rendered block states how many older runs were omitted. `stateless: true` disables history entirely. `history` applies only to stateful pipelines.

```yaml
options:
  stateless: true
```

```yaml
options:
  history:
    max_runs: 20
    max_bytes: 100000
```

## setup

Commands to run after workspace setup succeeds and before the agent spawns — for anything that mutates the worktree from outside the agent (fixture loading, symbol indexing, …).

| Field | Type | Description |
|-------|------|-------------|
| `run` | string | Shell command. Runs from the daemon's cwd with workspace coordinates in the environment: `WORKTREE`, `REPO`, `SHA`, `BRANCH`, `EVENT_NAME`, `EVENT_SOURCE`, `RUN_DIR`, `SURFACE_DIR`, `PAYLOAD_JSON`. Must not contain `{{`. Config load rejects a templated `run`. Read a workspace coordinate from the environment instead. Soft-fails on non-zero exit — the agent still runs — unless `on_failure: abort`. |
| `timeout` | integer (seconds) | Hard timeout per step. Default `360`. |
| `on_failure` | `abort` | What a non-zero exit does. Unset (the default) soft-fails. `abort` makes the step's success a precondition for the run: the remaining steps are skipped, no agent spawns, and the invocation fails with the diagnostic `setup_precondition_failed`. The trigger is not retried — a failed precondition is not transient. |

```yaml
setup:
  - run: ./scripts/load-fixtures.sh
    timeout: 120
```

Use `on_failure: abort` for a preflight check, a precondition the run needs before it proceeds, such as an authorization check on the subject that triggered it. A step that times out or fails to spawn aborts too: an unproven precondition is not a passed one.

```yaml
setup:
  - run: ./scripts/check-authorization.sh
    on_failure: abort
```

A pipeline's `setup:` replaces `defaults.setup` from `daemon.yaml` wholesale — the array is a single field, so a pipeline that declares its own steps does not inherit the defaults' steps. Omit `setup` here to run the daemon-wide default steps.

The daemon records each step like any other phase, including the step that aborts a run. It gets a task row named `setup[<index>]` in run order, carrying the command, the exit code, and the duration, and its stdout and stderr go to `runs/<run-id>/steps/setup_<index>_/` in the thread directory. [`bento trace`](/pipelines/traces) shows the row and the paths, so you read a failed step from its own files rather than from the daemon log.

## workspace

Templated workspace extraction for generic [webhook](/triggers/webhook) pipelines. Each template renders against `{ payload, event }` to produce the workspace the daemon clones. Omit it for a `github` or `linear` pipeline, which uses a built-in source-aware extractor.

| Field | Description |
|-------|-------------|
| `key` | Template for the workspace key — the dedup / on-disk identifier, e.g. `"swe-bench:{{payload.instance_id}}"`. Must render to a stable non-empty string. |
| `repo` | Template for `owner/name`. |
| `branch` | Template for the branch ref. Required iff `sha` is absent. |
| `sha` | Template for the commit SHA. Required iff `branch` is absent. |

```yaml
workspace:
  key: "swe-bench:{{payload.instance_id}}"
  repo: "{{payload.repo}}"
  sha: "{{payload.base_commit}}"
```

## Queue placement

The sandbox backend selects the worker pool. Daytona and Cloudflare runs use `remote`. All other backends use `host`. Each pool has a daemon-wide [concurrency limit](/configuration#queue), with a default of 2.

Jobs have equal queue priority. Each pool starts eligible jobs in `run_at` order. Debounce, jitter, and retry delays determine when a job becomes eligible. Jobs for the same workspace remain serial.

Remove `lane` and `priority` from pipeline files. Both fields cause a configuration error. The daemon no longer has a fast pool.

## presence

What the daemon posts on the GitHub subject of an event as soon as it queues a run. The agent can wait minutes behind other work and a sandbox start. Presence shows in that time that bento accepted the event.

```yaml
presence:
  github:
    reaction: eyes
    status: bento/review
```

| Field | Type | Description |
|-------|------|-------------|
| `github.reaction` | string | The reaction to add when the run is queued. One of `+1`, `-1`, `laugh`, `confused`, `heart`, `hooray`, `rocket`, `eyes`. The daemon adds it to the triggering comment. An event with no comment gets it on its pull request or issue. |
| `github.status` | string | The commit status context to set on the head of the pull request. A non-empty string of at most 255 characters. An issue event has no head, so it gets no status. |

The daemon sets the status on each stage of the run:

| Stage | State | Description |
|-------|-------|-------------|
| The run is queued | `pending` | `queued` |
| The run starts | `pending` | `running` |
| The run finishes | `success` | `finished` |
| The run stands down, for example on a head that it already reviewed | `success` | `stood down: <reason>` |
| The run fails on its last attempt | `error` | `failed: <diagnostic code>` |
| The daemon discards the run | `error` | `discarded: <reason>` |
| A daemon restart stops the run | `error` | `interrupted by a daemon restart` |

The description holds a diagnostic code, not an error message, because anyone who can read the pull request can read the status. A run that the daemon retries keeps its `pending` status until an attempt ends it. An agent can set the same context while it works. The daemon posts `finished` only while the status is still `pending`, so the verdict of the agent stays.

A pipeline with `presence.github` needs a [`trigger.github`](#github) event. `bento config validate` reports an unknown field, an unknown reaction, a context that is empty or too long, and a missing `trigger.github`.

## orchestrator

Explicitly enables the orchestrator agent during the spawn phase. The orchestrator is a daemon-internal LLM. It decides how to dispatch the work: whether to call the named agent directly, fan out in parallel, or assemble a more complex pattern. A pipeline with no `orchestrator` field runs its named agent directly.

| Value | Description |
|-------|-------------|
| Omitted | Spawn the named agent directly. |
| `{ strategy: { type: auto }, targets?, guidance?, capabilities?, max_capability_calls? }` | Enable parent-model orchestration. Every referenced target uses the `claude` runtime. Listing `capabilities` enables typed capability mode. `max_capability_calls` defaults to `5` and takes a positive integer. |

```yaml
# Direct spawn: omit the orchestrator field.

# Enable auto orchestration and use daemon target defaults.
orchestrator:
  strategy:
    type: auto

# Override the ordered target list and provide static guidance.
orchestrator:
  targets: [claude-low]
  strategy:
    type: auto
  guidance: |
    This is a single-shot notification. No fan-out needed.
    Delegate directly to the named agent and stop.
```

The daemon-level `orchestrator` block supplies target defaults only. It does not enable orchestration for pipelines that omit the field. `guidance` is a static, non-empty string and is not rendered against `event` or `args` templates. `auto` is the only supported strategy type. Boolean values, `assess`, and unknown keys fail configuration validation.

The orchestrator runs as a direct API call in the daemon. Tasks it dispatches run inside sandboxes.

### Task compute classes

In generic orchestration mode, `task` and each entry in `parallel` accept an optional `class` from the active backend registry. A task without `class` inherits the invocation class. Docker and Podman apply the selected class to that task container. Daytona keeps the invocation size because all tasks share one sandbox. It logs a warning when a task requests a different class.

An unknown or empty class returns `dispatch_refused` for that task. An unknown target also fails only its task. Other tasks in the parallel call still execute. Typed capability tools do not expose the generic `class` field.

### Capability mode

Capability definitions live in `.bento/capabilities/*.yaml` or `.yml`. Each file defines one operation. The daemon validates and freezes the registry at startup. Restart the daemon after you change a capability definition. Pipeline changes also require validation and a restart.

```yaml
# .bento/capabilities/canon-review.yaml
name: canon_review
description: Review code against the engineering canon.
input:
  type: object
  properties:
    scope: { type: string, enum: [files, pull_request] }
    question: { type: string }
  required: [scope]
  additionalProperties: false
output:
  type: object
  properties:
    summary: { type: string }
    findings: { type: array, items: { type: object } }
  required: [summary, findings]
  additionalProperties: false
agent: reviewer
skill: agentic-review
targets: [claude-high, claude-backup]
access: observer
timeout: 900
```

An actor capability that performs provider actions selects one pipeline binding:

```yaml
access: actor
principal: pull-request-author
```

The selected binding exists on every pipeline that registers the capability. An observer capability never sets `principal`.

Register capabilities on the pipeline:

```yaml
orchestrator:
  strategy:
    type: auto
  guidance: Select the smallest capability that satisfies the request.
  capabilities:
    - canon_review
    - edit_pull_request
    - run_verification
  max_capability_calls: 5
```

Capability mode exposes one tool per registered capability and a capability-aware `parallel` tool. It does not expose the generic `task` tool. The orchestrator chooses gate, route, sequential, parallel aggregation, orchestrator-worker, generator/evaluator, and adaptive fan-out patterns at runtime. Bento schema-validates each call and each result, and records them in the existing Task trace.

Capability `targets` are ordered: the first is primary and later entries are pre-dispatch fallbacks. Omit `targets` to inherit the parent invocation route.

`observer` capabilities do not receive forge credentials, git configuration, pipeline credential mounts, or actor-only environment variables. An `actor` capability requires actor access plus an exact named grant from outside the model. For a GitHub mention pipeline, include `/allow capability_name` in the mention-bearing comment. Each grant names a capability that the pipeline registers. A dry run never grants an actor capability.

Every capability declares `access: observer` or `access: actor`. Bento does not infer an access tier.

Do not place credentials in capability inputs or outputs. Inject them through the pipeline `env:` manifest or credential mounts. A capability implementation does not echo a credential into a result or into free-form output. Bento redacts a structured field marked `writeOnly: true` or `x-bento-sensitive: true` from the Task attributes, but it does not scan arbitrary Task output or error text. Treat Task traces and agent logs as trusted operator data.

### Orchestrated output

A capability returns structured data to the orchestrator. Bento does not send a capability result to GitHub.

An orchestrated run delivers through `result:` and `outputs:`, the same contract a direct spawn uses. The orchestrator's final response takes the pipeline's declared `result:` schema. A pipeline with no `result:` keeps its stdout opaque. See [result and outputs](#result-and-outputs).

## context

Injects additional context into the agent's prompt at spawn time. For a webhook trigger, bento always includes the trigger event as an `Event` section. The fields below add more.

| Field | Type | Description |
|-------|------|-------------|
| `diff` | boolean | For GitHub pull-request triggers, fetch the PR diff via `gh api` and append it as a context section. Truncated at 100 KB. |

The schema also accepts `git_history` and `related_prs` booleans. They gather nothing.

```yaml
context:
  diff: true
```

## eval

**Optional.** Opts the pipeline into evals (BIP-29): when a run completes, bento asks one nominated person a closed question over Slack and stores their answer against the invocation. Omit `eval` for a pipeline that is never evaluated.

| Field | Type | Description |
|-------|------|--------------|
| `ask` | string | **Required.** The question posted to the respondent, verbatim. |
| `options` | list | **Required.** Closed set of answers, at least two, no duplicate values. An entry is a string, or `{ value, label }` where `label` is the button text. The label, or the value when there is no label, is 75 characters or fewer (Slack's button-text limit). The value is 1980 characters or fewer, so that it fits in a Slack button value with the eval id. Rendered as one button per option. |
| `of` | string | list | `{ contributors }` | **Required.** Who to ask. A `{{event.*}}` template, a literal provider handle, a list of either — one eval is raised per person — or `{ contributors: <result field> }` naming the field of the agent's structured result that lists who worked on the thing being judged. |
| `expires` | duration | How long the request stays open before it moves to `expired`. Default `7d`. |
| `when` | map | Gates whether a completed run is worth asking about, evaluated against the agent's structured result. Same leaf-predicate vocabulary as [`filter.when`](#filter). Omit to ask about every completed run. |
| `labels` | string | Name of the field in the run's structured result holding what the run applied — an array of strings, surfaced as the `labels` artifact. |
| `pull_request` | string | Name of the field in the run's structured result holding the URL of the pull request the run produced. Use it when the trigger payload carries no pull request, such as a Linear session or a schedule. The field must be a string in `result.schema`. |
| `correction` | map | Forced follow-up that captures the correct output for the run, opened when the respondent picks one of the answers listed in `correction.on`. `correction.title` names the modal and defaults to `What was right?`. |

```yaml
name: grade-pull-request
trigger:
  github: [pull_request.closed]
agent: reviewer
instructions: Grade the risk of this merged PR.
eval:
  ask: "Was this risk grade right?"
  options: [correct, too_high, too_low, right_call_wrong_reason]
  of: "{{event.pull_request.user.login}}"
  expires: 7d
  labels: labels
  correction:
    on: [too_high, too_low]
    title: "What was the grade?"
    fields:
      - key: correct_grade
        label: What should the grade have been?
        choices: [low, medium, high]
```

`correction` fires the modal only when the respondent picks an option listed in `on`. Every field in `fields` is required at submit, and a partial correction is not recorded. `fields[].choices` are opaque strings — a pipeline correcting GitHub labels lists them exactly as GitHub spells them.

`of` accepts a literal handle for pipelines with no human in the trigger payload — a schedule-fired pipeline has nothing to template against:

```yaml
of: arnoldporter
```

```yaml
of: ["{{event.pull_request.user.login}}", cosmicallycooked]
```

Each entry in `of` resolves through the [member directory](/members). An entry is a member key, any handle that member carries, or a group name, which asks each of its members once. A literal entry naming none of those fails config load. A `{{event.…}}` template does not, since it resolves to whatever the payload carries. At runtime an unresolved handle, a member who is not an active person, or someone already at their open-eval cap records a skipped eval rather than failing the run. Only an invocation that reaches `completed` raises an eval — a failed or cancelled run produced nothing to judge. Answers are curated with the [`bento evals`](/cli/evals) CLI commands.

`{ contributors: <field> }` reads the named field from the run's structured result instead of a fixed recipient. Use it when the right person to ask depends on who did the work, not on the trigger payload:

```yaml
eval:
  ask: "Was this risk grade right?"
  options: [correct, too_high, too_low]
  of:
    contributors: contributors
```

Entries match every name the directory knows a member by — key, handle, and provider account id. Bento records a field that the agent's result never emits, or one that is not a list of strings, as a skipped eval with the reason `lookup_failed`. It records a field that resolves to no valid recipient — every entry is a bot, or none resolve — as `no_recipients`.

Where the pipeline declares [`result.schema`](#result-and-outputs), `bento config validate` resolves the paths named by `of.contributors` and `labels` against that schema, and fails the load when a path names a property the schema does not declare, when a segment drills into a node declared as anything but an object, or when the property it names is declared as anything but an array of strings. A `$ref` to a bento schema id is resolved first. An `anyOf` fails the load only when every branch is declared and at least one of them permits something other than an array of strings.

The check only refutes a path. Where the schema leaves the shape open — a node with no `type`, an object with no `properties`, an array with open or untyped `items`, an `anyOf` with an undeclared branch, or a `$ref` the catalogue does not hold — the path passes the load, and a value the eval cannot read is still a runtime `lookup_failed`. Where the pipeline declares no `result.schema`, both paths stay unchecked.
