<!--
Sitemap:
- [Installation](/installation)
- [Upgrading](/upgrading): Version-specific steps for upgrading an existing Bento install.
- [Concepts](/concepts)
- [Build your first pipeline](/tutorials/pipeline-args)
- [Target a specific issue or PR from a URL](/tutorials/url-targeting)
- [Keep state across runs](/tutorials/pipeline-state)
- [Fire a pipeline on a schedule or on demand](/tutorials/schedule-and-fire)
- [Deploy a box to Railway](/tutorials/deploy-to-railway)
- [Operate a hosted daemon](/tutorials/operate-a-hosted-daemon)
- [Configuration](/configuration)
- [Members](/members)
- [Knowledge base](/knowledge-base/)
- [Method and delivery](/knowledge-base/modes)
- [Config](/knowledge-base/config)
- [MCP](/knowledge-base/mcp)
- [Pipeline configuration reference](/pipelines/config)
- [Filters](/pipelines/filters)
- [Triggers](/triggers/)
- [GitHub trigger](/triggers/github)
- [Linear trigger](/triggers/linear)
- [Webhook trigger](/triggers/webhook)
- [Schedule trigger](/triggers/schedule)
- [Manual trigger](/triggers/manual)
- [Traces](/pipelines/traces)
- [Slack](/integrations/slack)
- [Public access](/public-access)
- [Context engineering](/context-engineering)
- [Best practices](/best-practices)
- [Troubleshooting](/troubleshooting)
- [Architecture](/architecture/vision)
- [Workspaces](/workspaces)
- [Authentication](/authentication)
- [Identity](/identity)
- [Security](/security)
- [References](/references)
- [Changelog](/changelog): Bento release history.
- [CLI reference](/cli/)
- [Setup](/cli/setup)
- [Secrets](/cli/secrets)
- [Lifecycle](/cli/lifecycle)
- [Sandbox image](/cli/image)
- [Sandboxes](/cli/sandbox)
- [Observability](/cli/observability)
- [Diagnostics](/cli/diagnostics)
- [Triggers](/cli/triggers)
- [Workbench](/cli/workbench)
- [Auth](/cli/auth)
- [Knowledge](/cli/knowledge)
- [Evals](/cli/evals)
- [Bento](/index)
- [Runtime wrapper](/architecture/runtime-wrapper)
- [Skill evolve](/architecture/skill-evolve)
-->

# Method and delivery

Knowledge injection has two independent settings:

* **`method`** — how the daemon retrieves: `bm25` (lexical), `vector` (semantic), `hybrid` (lexical and semantic), or `none` (no daemon-side retrieval).
* **`delivery`** — how retrieved knowledge reaches the agent: `prompt`, `file`, `browse`, `search`, or `none`.

They appear on a pipeline's [`knowledge:` block](/pipelines/config#knowledge) (or `defaults.knowledge` in `daemon.yaml`). The one-shot `/ask` endpoint and `bento ask` use the configured retrieval engine and take a `granularity` (they always return results to the caller, so they have no delivery axis — see [On `bento ask`](#on-bento-ask)).

```yaml
knowledge:
  method: hybrid
  delivery: prompt   # composable: prompt+search
  filter:
    categories: [conventions]
```

Defaults are `method: hybrid`, `delivery: prompt`.

## Where content lands

Content uses two channels:

* **System prompt** — delivery manifests (the browse listing, mount paths) and the delivery's default instruction.
* **User message** — retrieved `<context>` excerpts, and custom `instructions` when set. Setting `instructions` replaces the default system instruction.

## Method

### `bm25`

The daemon rewrites the trigger prompt into a short search query (see [rewrite](/knowledge-base/config#rewrite)), the selected retrieval engine retrieves matching chunks, and the results feed the delivery. `topK` sets results per source, `maxTokens` caps the injection budget, and `filter` drops chunks from non-matching docs after retrieval. QMD runs `qmd search`. Chroma filters on document content and ranks the matches by BM25 (see [Chroma](/knowledge-base/config#chroma)). Requires a configured retrieval engine and discovered knowledge sources.

### `vector`

The selected retrieval engine retrieves by semantic similarity. QMD requires embeddings (`qmd embed`). Chroma uses its configured embedding function. Errors surface when the selected retrieval engine cannot query.

### `hybrid` (default)

The selected retrieval engine combines lexical and semantic matches. QMD runs a structured `qmd query` with `lex:` and `vec:` clauses. The daemon disables QMD's optional language-model reranker by default. Set `knowledge.retrieval.rerank: true`, `knowledge.rerank: true`, or pass `--rerank` to `bento ask` when you accept the added latency. Chroma runs its lexical and its dense search, then combines the two by [reciprocal rank fusion](/knowledge-base/config#chroma).

### `none`

The daemon retrieves nothing. Pair with `file` to mount the whole corpus, then use `grep` to search the files. Pair with `browse` or `search` when the agent uses MCP instead.

## Delivery

### `prompt` — retrieved excerpts inlined (default)

The retrieved chunks land in the user message as `<context>` blocks:

```xml
<context path="wiki/git-workflow.md" category="conventions" tags="git, process">
Feature branches PR into develop; develop merges to main for release…
</context>
```

The run's retrieve task records method, query, and injected counts — visible in [`bento trace`](/pipelines/traces). Requires a retrieving method (`bm25`/`vector`/`hybrid`).

### `file` — results on the filesystem

`method` selects one of two forms for `file`:

* **Retrieved subset** (`method: bm25`/`vector`/`hybrid`) — the daemon writes the retrieved `<context>` blocks to `.bento/retrieved-knowledge.md` in the run checkout instead of inlining them. The agent reads that one file. Keeps large result sets out of the prompt. `filter` applies. Works on every backend; a remote backend uploads the file into its checkout.
* **Whole corpus** (`method: none`) — the daemon delivers the knowledge directories read-only at `/knowledge/<type>` (`/knowledge/wiki`, …), with `/knowledge/index.tsv` (one tab-separated row per doc: path, title, category, tags, summary) and `/knowledge/README.md` (the layout). The agent greps the index, then reads files with its normal tools. Docker and Podman bind-mount the directories. A remote backend (Cloudflare, Daytona) has no host bind mount, so the daemon uploads the same directories to `/tmp/bento-knowledge/<type>` and the system prompt names that root instead. `filter` is a config error, because whole directories are delivered.

### `browse` — listing, agent reads on demand

The system prompt carries the filtered doc listing, grouped by URI scheme:

```
wiki://
- wiki://git-workflow — Git Workflow [git, process]

recipe://
- recipe://setup-knowledge-base — Setting Up the Knowledge Base [knowledge, setup]
```

The agent reads whole docs via MCP resources when relevant. Use when docs are few enough to pick from a list and whole-doc context beats excerpts. Requires discovered docs and a local sandbox backend. The MCP endpoint of the daemon is loopback-only, so a remote backend such as Daytona never reaches it. Pairs with `method: none`.

### `search` — agent's MCP tool

The daemon injects no content. The system prompt points the agent at the `ask_knowledge` MCP tool:

> A team knowledge base is available through your MCP tools. Search it with the `ask_knowledge` tool before deciding questions of convention, architecture, or process.

The agent searches when it decides it needs to. Requires the `ask_knowledge` tool (the selected retrieval engine ready, with sources configured) and a local sandbox backend (same loopback constraint as `browse`). Pairs with `method: none`.

The tool uses `knowledge.retrieval.method`, or `hybrid` when that key is unset. A tool call overrides this setting with its own `method`. The pipeline's daemon-side `method` remains `none` for this delivery. See [MCP](/knowledge-base/mcp).

### `none`

The daemon injects nothing. Use it on a pipeline to opt out of a delivery that `defaults.knowledge` sets. With no `knowledge:` block and no defaults, runs get no injection anyway.

## Composing with search

`prompt`, `file`, and `browse` compose with `search` as `prompt+search`, `file+search`, and `browse+search`: the primary grounds the agent up front, and `ask_knowledge` stays available as the escape hatch when that content does not cover the question.

```yaml
knowledge:
  method: hybrid
  delivery: prompt+search
  topK: 3
```

A composite run satisfies the prerequisites of both deliveries. `none` does not compose, and a delivery has at most one primary.

## Failure semantics

A delivery whose prerequisites are missing fails the run at dispatch rather than spawning an agent without them. The error names the delivery, the missing capability, and the fix:

```
knowledge delivery "file" with method "none" delivers the whole corpus and
requires bind-mounts or remote uploads — sandbox backend is "none (host)"; set
sandbox.backend to docker, podman, cloudflare, or daytona, or set a retrieval
method
```

| Delivery | Checked at dispatch |
|------|---------------------|
| `search` | `ask_knowledge` registered, local sandbox backend, sandbox MCP bundle issued |
| `browse` | docs discovered, local sandbox backend, sandbox MCP bundle issued |
| `file` (`method: none`) | Docker, Podman, Cloudflare, or Daytona backend, no `filter` set |
| `file` (subset) | no backend check — the results file goes into the run checkout, and a remote backend uploads it there |
| `prompt` | checked at the retrieve step, not dispatch — unavailable retrieval errors the task and the run continues without context |

## On `bento ask`

`bento ask` and `/ask` use the configured retrieval engine and take `--granularity` (`chunks` default, or `docs` to inline whole filtered docs — there is no MCP session to read them through). `--tag`, `--category`, `--top-k`, and `--instructions` mirror the pipeline fields.

```bash
bento ask "how do we test the daemon" --granularity docs --category conventions
bento ask "branch rules" --tag git --top-k 3
bento ask "branch rules" --rerank
```
