Index Settings
The [index] section controls which files are indexed.
Configuration
[index]
include = [
"**/*.rs",
"**/*.ts",
"**/*.tsx",
"**/*.js",
"**/*.jsx",
"**/*.py",
"**/*.go",
"**/*.java",
"**/*.cpp",
"**/*.cc",
"**/*.hpp",
"**/*.md",
]
exclude = [
"**/node_modules/**",
"**/target/**",
"**/dist/**",
"**/.git/**",
"**/build/**",
"**/__pycache__/**",
]
use_gitignore = true
# Multimodal ingest (opt-in). When true, the indexer also walks PDFs,
# extracts their text, and chunks it like a plain-text document.
multimodal = false
# Documents source (opt-in). When true, the indexer also walks HTML files,
# converts them to markdown-ish text, and chunks them by heading structure.
documents = false
Options
| Key | Type | Default | Description |
|---|---|---|---|
include | string[] | See above | Glob patterns for files to include |
exclude | string[] | See above | Additional exclusion patterns (on top of .gitignore) |
use_gitignore | bool | true | Whether to respect .gitignore files |
multimodal | bool | false | Enable multimodal ingest (PDF text extraction). See below. |
documents | bool | false | Enable the documents source (HTML text extraction). See below. |
Notes
- Include patterns determine which file extensions are parsed and indexed. Add patterns to index additional file types.
- Exclude patterns are applied in addition to
.gitignore. Use them to skip generated code, vendor directories, or other non-useful content. - When
use_gitignoreistrue, files matched by.gitignoreare automatically excluded even if they match an include pattern.
Multimodal ingest
By default bobbin indexes code, markdown, and beads. Set multimodal = true to
also ingest PDFs (runbooks, design docs, specs):
- The indexer automatically walks
**/*.pdf— you do not need to add it toinclude. Toggling the flag is the only knob. - Text is extracted with a pure-Rust extractor (no Python, no native toolchain)
and chunked like a plain-text document. Chunks are tagged with
language = "pdf", so you can filter on them in search. - Image-only or encrypted PDFs may yield little or no text; those files are skipped the same way an empty file is.
- Image captioning (vision LLM) is not yet supported and is tracked as a follow-up.
Documents source
Set documents = true to also ingest HTML files (.html, .htm) —
exported wikis, generated API docs, saved pages:
- The indexer automatically walks
**/*.htmland**/*.htm— you do not need to add them toinclude. Toggling the flag is the only knob. - Conversion is a deterministic, dependency-free tag stripper (no model in the
loop):
<script>/<style>/comments and<head>are dropped, headings become markdown#headings, lists become bullets,<pre>becomes a fenced code block, and entities are decoded. The result runs through the markdown chunker, so headings become section chunks with breadcrumb names. - Chunks are tagged with
language = "html", so you can filter on them in search. - Unparseable or text-free HTML degrades to a skipped file (same as an empty file), never a failed index run.
- Incremental indexing works as for any other file: unchanged files are skipped by content hash.