<!--
Sitemap:
- [Installation](/installation)
- [Upgrading](/upgrading): Version-specific steps for upgrading an existing Bento install.
- [Concepts](/concepts)
- [Build your first pipeline](/tutorials/pipeline-args)
- [Target a specific issue or PR from a URL](/tutorials/url-targeting)
- [Keep state across runs](/tutorials/pipeline-state)
- [Fire a pipeline on a schedule or on demand](/tutorials/schedule-and-fire)
- [Deploy a box to Railway](/tutorials/deploy-to-railway)
- [Operate a hosted daemon](/tutorials/operate-a-hosted-daemon)
- [Configuration](/configuration)
- [Members](/members)
- [Knowledge base](/knowledge-base/)
- [Method and delivery](/knowledge-base/modes)
- [Config](/knowledge-base/config)
- [MCP](/knowledge-base/mcp)
- [Pipeline configuration reference](/pipelines/config)
- [Filters](/pipelines/filters)
- [Triggers](/triggers/)
- [GitHub trigger](/triggers/github)
- [Linear trigger](/triggers/linear)
- [Webhook trigger](/triggers/webhook)
- [Schedule trigger](/triggers/schedule)
- [Manual trigger](/triggers/manual)
- [Traces](/pipelines/traces)
- [Slack](/integrations/slack)
- [Public access](/public-access)
- [Context engineering](/context-engineering)
- [Best practices](/best-practices)
- [Troubleshooting](/troubleshooting)
- [Architecture](/architecture/vision)
- [Workspaces](/workspaces)
- [Authentication](/authentication)
- [Identity](/identity)
- [Security](/security)
- [References](/references)
- [Changelog](/changelog): Bento release history.
- [CLI reference](/cli/)
- [Setup](/cli/setup)
- [Secrets](/cli/secrets)
- [Lifecycle](/cli/lifecycle)
- [Sandbox image](/cli/image)
- [Sandboxes](/cli/sandbox)
- [Observability](/cli/observability)
- [Diagnostics](/cli/diagnostics)
- [Triggers](/cli/triggers)
- [Workbench](/cli/workbench)
- [Auth](/cli/auth)
- [Knowledge](/cli/knowledge)
- [Evals](/cli/evals)
- [Bento](/index)
- [Runtime wrapper](/architecture/runtime-wrapper)
- [Skill evolve](/architecture/skill-evolve)
-->

# Config

`knowledge.retrieval` in `daemon.yaml` configures the retrieval engine and the query rewriter. The [method and delivery](/knowledge-base/modes) knobs decide how results reach an agent. This block decides how results are produced — its settings apply to retrieving pipelines (`method: bm25`/`vector`/`hybrid`), `bento ask`, and the `ask_knowledge` MCP tool alike.

The default retrieval engine is QMD 2.8.3 or later. The daemon initializes it when it discovers knowledge sources and finds a supported `qmd` on PATH. QMD uses GPU acceleration for local embeddings on supported systems. QMD 2.8.3 disables an unsafe macOS Metal residency mode for short-lived commands, so hybrid queries keep GPU acceleration without the llama.cpp exit crash. Set `engine: chroma` to use a Chroma server instead. The daemon fails startup when the selected retrieval engine cannot connect or build its index. There is no enable switch — to keep retrieval out of pipeline runs, omit `knowledge:` blocks (and `defaults.knowledge`), or set `delivery: none`.

Indexing runs before the daemon accepts work. It reports progress to the daemon log. When `qmd doctor` finds that the QMD embedding model is missing, the daemon names the model that `qmd embed` downloads. A cold first boot can spend minutes on this download. After the index update, one line gives the number of documents that the embed step must cover. The embed step reports elapsed time every 30 seconds until `knowledge: index ready`. On a CPU-only host, an embed can take minutes. With `engine: chroma`, one line gives the number of chunks and documents to embed, and each collection's upsert reports elapsed time every 30 seconds in the same way. These reports separate a slow index from a stuck one.

```yaml
knowledge:
  retrieval:
    engine: qmd # qmd or chroma
    topK: 20 # default 5
    contextLines: 60 # default 60
    method: hybrid # bm25, vector, or hybrid — agent search
    rewrite:
      targets: [claude-fast, codex-low]
```

| Field | Default | Description |
|-------|---------|-------------|
| `engine` | `qmd` | Retrieval engine. Use `qmd` for the local QMD index or `chroma` for Chroma Cloud or a Chroma server. |
| `topK` | `5` | Default results per source. Raise it to `20` for a small knowledge base of fewer than about 20 documents. Per-run `knowledge.topK` overrides it, and a `limit` on an `ask_knowledge` call overrides it. |
| `minScore` | `0` | Drop results scoring below this. Read the note below first: it applies to each retrieval component on its own native scale, before fusion, and does not gate a fused `hybrid` score, which is rank-derived on both engines. |
| `rerank` | `false` | Enable QMD's optional language-model reranker. This adds query latency. |
| `contextLines` | `60` | Document lines returned around each match. QMD's own snippet is about 300 characters, which locates a document without answering from it. Raise this for long documents; lower it to spend fewer tokens per result. QMD only — Chroma chunks the corpus itself. |
| `method` | `hybrid` | How an `ask_knowledge` search runs: `bm25`, `vector`, or `hybrid`. Set it to `bm25` or to `vector` on a host where a hybrid pass is too slow. A `method` on the call overrides it. |
| `chroma.cloud.tenant` | — | Chroma Cloud tenant ID. The `cloud` block selects [Chroma Cloud](#chroma-cloud). |
| `chroma.cloud.database` | — | Chroma Cloud database. |
| `chroma.cloud.apiKey` | — | Chroma Cloud API key, usually `${CHROMA_API_KEY}`. Bento reads the key from this field only. |
| `chroma.embedding` | `qwen` with `cloud`, else `minilm` | Embedding model for each chunk and each query. `qwen` needs the `cloud` block. |
| `chroma.host` | `localhost` | Chroma server host. Do not set it with `cloud`. |
| `chroma.port` | `8000` | Chroma server port. Do not set it with `cloud`. |
| `chroma.ssl` | `false` | Use HTTPS when `true`. Do not set it with `cloud`. |
| `chroma.tenant` | client default | Chroma tenant. |
| `chroma.database` | client default | Chroma database. |
| `chroma.collection` | project-specific | Base collection name. Bento adds the embedding model to it. On a self-hosted server, Bento adds one collection for each source type. |
| `rewrite.targets` | — | Ordered execution targets. On quota or authentication failure, Bento tries each target and its credential pool in order. Omit to search with the raw prompt. |
| `rewrite.timeoutMs` | `15000` | Hard cap on the rewriter spawn. On timeout the daemon searches the raw prompt unchanged. |

## What `minScore` compares against

On QMD, a hybrid score is a reciprocal-rank score, not a similarity score. It shows where a result placed, not how well it matched the query, so the first result scores the same whether or not it answers the query. Measured over one corpus with `bge-base-en-v1.5`, ranks one to five scored `0.76`, `0.38`, `0.25`, `0.15`, and `0.12` — close to `1/rank`. A `minScore` on that scale works like `topK`: `0.6` admits only the first result, and `0.2` admits about three.

QMD's `hybrid` method is the default retrieval path. `SourceRetriever.query()` resolves the method to `hybrid` when a caller passes none, and the `ask_knowledge` tool passes `knowledge.retrieval.method`, which defaults to `hybrid` — so a QMD-backed agent search reaches hybrid unless that key or a `method` on the call names `bm25` or `vector`. [Chroma](#chroma)'s `hybrid` method is rank-derived too, by reciprocal rank fusion rather than QMD's fusion, so the same threshold-by-position limit applies there. Chroma's `bm25` method scores by BM25 instead, and `vector` by dense similarity — see [Chroma](#chroma) for how `minScore` reads on each.

A cosine similarity score does separate matches. Over one corpus, model, and query set, `vector` scored relevant matches around `0.67` and irrelevant ones around `0.62`. A threshold of `0.64` kept 87% of the relevant results and admitted 22% of the irrelevant ones. QMD's `hybrid` method separated the same two sets by only `0.08`, against `0.64` for `vector`.

So set `minScore` when retrieval runs `vector` or Chroma's `bm25`, and calibrate the number against the embedding model in use — the number belongs to the model as much as to the corpus, and it does not survive a change of either. Leave it at `0` on any `hybrid` method, and use `topK` to bound the result count instead.

## Chroma

Bento connects to Chroma Cloud or to a self-hosted Chroma server. Chroma Cloud embeds the text on its servers, so the daemon host needs no GPU and no Python.

Bento adds the embedding model to each collection name. A change of `embedding` therefore builds a new collection and does not mix vectors from two models. At startup, the daemon embeds only the chunks that changed since the last index.

### Chroma Cloud

Create a database in Chroma Cloud. Then set the `cloud` block in `daemon.yaml`:

```yaml
knowledge:
  retrieval:
    engine: chroma
    chroma:
      cloud:
        tenant: ${CHROMA_CLOUD_TENANT_ID}
        database: bento-kb
        apiKey: ${CHROMA_API_KEY}
```

The `cloud` block sets the host, the port, and TLS. Do not set `host`, `port`, or `ssl` with it. An empty `tenant`, `database`, or `apiKey` stops the daemon at startup, and an unset `${VAR}` expands to an empty value.

With the `cloud` block:

* Qwen3-Embedding-0.6B embeds each chunk and each query through the Chroma Cloud embedding API.
* One collection holds every source type, and each chunk records its type.
* `hybrid` sends one request to the Chroma Cloud Search API. The request fuses a dense search and a BM25 search by reciprocal rank, and Chroma computes the BM25 statistics over the whole collection.
* `bm25` sends one BM25 search, and `vector` sends one dense search.

Chroma Cloud bills each write, each GiB of storage, and each query. After the first index, a restart writes only the changed chunks.

### Self-hosted Chroma

The TypeScript client connects to a Chroma server. `services.yml` defines a `chroma` service on port 8000. Start it before you start the daemon:

```bash
curl -fsSL https://install.getbento.sh/services.yml -o docker-compose.yml
docker compose up -d
```

Chroma data lives in `./.resources/chroma`. `docker compose down -v` does not delete it.

Chroma embeddings run on the host, not inside the daemon process. The embedder is `uvx --from chromadb==1.5.9`, so `uv` must be on the PATH of the daemon. Install it with `curl -LsSf https://astral.sh/uv/install.sh | sh`.

Then select Chroma in `daemon.yaml`:

```yaml
knowledge:
  retrieval:
    engine: chroma
    chroma:
      host: localhost
      port: 8000
      collection: bento-knowledge
```

Without the `cloud` block, Bento uses Chroma's default `all-MiniLM-L6-v2` embedding function (`embedding: minilm`) and stores source paths as metadata. The first index build downloads the embedding model. `vector` retrieves by dense similarity.

Chroma has no BM25 endpoint, so `bm25` generates its candidates with Chroma's `where_document` substring filter — one filter for each term of the rewritten query — and ranks that candidate set by BM25, against the corpus size and mean document length the daemon captured while indexing. Each filter returns at most 200 rows, so ranking runs over a bounded candidate set rather than over the whole corpus: for a term that matches more rows than the cap, the best-scoring document can stay out of the candidates. Terms under three characters and closed-class words such as `the` and `with` are dropped, because a short substring matches inside unrelated words and a closed-class word matches nearly every document. A query left with no terms retrieves nothing under `bm25`.

`$contains` is case-sensitive, so each term is filtered in its lowercase, capitalized, and uppercase forms, plus the case you typed it in. That last form is what matches an internal capital such as `GitHub` or `OAuth`, provided the query spells it the way the document does. A substring filter cannot enumerate every capitalization, so a document that spells a term differently from both your query and the three prose forms is not retrieved lexically. Write proper nouns, file names, and error strings the way the documents spell them — the [query rewriter](#rewrite) preserves them.

`hybrid` runs the lexical and the dense search, then combines them by reciprocal rank fusion. A Chroma distance and a BM25 score share no scale, so the fused score comes from the rank a chunk holds in each list, never from the values. Each list contributes `1 / (60 + rank)`. Appearing in both lists therefore raises a chunk's fused score, but it does not by itself outrank a chunk that only one search returns: at a high `topK`, a chunk ranked first in one list scores above a chunk ranked low in both.

On Chroma Cloud, `bm25` scores are Chroma's BM25 values, and `hybrid` scores are rank-fusion values near `0`.

Chroma distances are converted with `1 / (1 + distance)`, so `vector` scores are between `0` and `1`. `bm25` scores are BM25 values with no upper bound, and `hybrid` scores are rank-fusion values near `0`. QMD returns its native lexical or hybrid score. No two of these are comparable. `minScore` applies to each retrieval component on its own native scale, before fusion — a dense result by its converted distance, a lexical result by its BM25 score. It does not gate the fused `hybrid` score, which is a rank score. Retune it when you switch engine or method.

A `where_document` failure under `hybrid` is logged and the query answers from the dense results alone. Under `bm25` there is no dense list to fall back to, so the failure surfaces instead of quietly returning semantic matches for a lexical method.

## rewrite

When `rewrite.targets` is set, the daemon spawns one-shot CLI calls in route order. Each call condenses the prompt into at most three queries of 3–6 words each, and preserves proper nouns, file names, and error strings. On quota or authentication failure, the daemon tries the next credential or target. On another failure, or when every candidate fails, the daemon searches the raw prompt. The rewrite route records no cooldown between calls.

The rewriter serves retrieving pipelines and `ask_knowledge`. The daemon-backed `bento ask` uses the query that you provide.

## Per-run settings

Injection behavior — mode, filter, instructions, `topK`, `maxTokens` (budget, default `8000`) — lives in the per-pipeline [`knowledge:` block](/pipelines/config#knowledge) and `defaults.knowledge`. Engine settings here apply underneath: `rewrite.*` and `minScore` always, `topK` unless the run overrides it.
