Ask AI
unlistedby Dumky de Wilde
Ask a coding-agent CLI — Claude Code, Codex, Gemini, or your own — a question about the note you are reading, from the right-click menu.
Ask AI for Obsidian
An Obsidian plugin. Right-click inside a note, ask a coding-agent CLI a question about it, and read the answer in a sidebar or a modal without leaving the app. Conversations are kept as notes in your vault, so they are searchable, linkable and yours.
You bring the agent. The plugin spawns a CLI you already have installed and
signed in: adapters for Claude Code, Codex and Gemini CLI are built in, and
anything else with a non-interactive mode — opencode, crush, a shell script —
goes in as a custom command. There is no API key to paste and nothing is sent
anywhere the CLI would not send it; questions cost whatever your agent
subscription or API usage costs. Whichever you pick runs with the vault as its
working directory, so its own context file applies and [[wikilinks]] resolve.
Desktop only — it spawns a process, which Obsidian mobile cannot do.
What you get
- Ask AI about this note in the editor right-click menu, the file explorer right-click menu, and the command palette.
- Ask AI about the selection when text is selected. The passage stays with the question in the conversation and in the saved note, so "what does this do?" still makes sense a week later.
- Ask AI about this image with the caret on an embedded image — a diagram, a
screenshot, a photo of a whiteboard. Same menu item, same command palette; what
the caret is on decides what the question is about, and a selection wins over an
image because you chose it. The agent opens the file and answers from what is in
it, and the conversation note keeps the image as an
![[embed]]above the question, so it shows the thing that was asked about rather than a path to it. Works with![[wikilink]]andembeds alike — right-click the picture itself, or put the caret on the line it is written on. - Asking opens a conversation of its own, so a question about a passage does not land in the middle of the thread you had going about something else.
- Follow up in the same conversation, from the box under the answer or from the right-click menu later. Resumed by the agent that started it. Switch agents mid-conversation and the questions and answers so far are handed to the new one, since only the agent that opened a thread can resume it.
- A cold thread is not resumed. Resuming replays the whole thread to the agent — every note it read, every search it ran — which is nearly free while the provider still has it cached and full price once it does not. So a conversation you come back to hours later shows a snowflake on the Ask button, hover it for why, and the next question opens a new thread with the questions and answers replayed into it instead. Every agent caches, each for its own length of time, so each has its own window.
- Leave a thread whenever you like, with the branch icon next to the question
box. Same thing on demand: the conversation carries on, the agent's thread does
not, so the next question does not drag along everything it read to get here.
Either way that answer's footer reads
new thread, which is why it may not remember something the one above it did. - Conversations are notes in your vault. Each one is a file with the
questions as
## headings, linked from the note it is about and carryingtype: ask-ai-conversationand a one-linedescriptionin its frontmatter. The description indexes the conversation's question headings and updates with each answer. That file is the record, not an export of one: the sidebar reads it back, so a conversation survives a restart, is searchable, shows up in the graph, can be listed by a Base, and moves or goes away when you move or delete it. Edit an answer and the sidebar shows what you wrote. - As many conversations per note as you want. The sidebar lists a note's
conversations one collapsed line each — title, when it was last asked in, how
many questions — and the open one below them. Click a line to continue that
one; the
+in the header starts another, so two lines of enquiry about one note stay two threads instead of becoming one long one. - Named by the agent. It ends its first answer with a short title for the conversation, and that title is the filename. A first question makes a poor name for a thread ("In one sentence, what is this note about?").
- Where they go is a setting, in the same terms as Obsidian's attachments: a
folder you name (
askai-conversationsby default) or beside the note itself, and a folder per note or all of them side by side. Later answers are appended to the same file, and once there are two questions it grows a## Contentslist of heading links. - Keep conversations in the vault is on by default. Turned off, a conversation is only in the pane until you press save, and gone when the window closes.
- The sidebar follows whichever note is open: switch notes and it switches with you, back to that note's conversations and the scroll position you left them at, still streaming if one was streaming.
- Suggested follow-ups. The agent ends an answer with up to three next
questions when there are useful ones, and they turn into buttons under the
answer. The buttons are shortcuts in the open conversation. The
agentandsessionfrontmatter fields resume a conversation, not the suggested questions. Answers are short by default because of it: the detail is a click away instead of pre-emptive. Cut the paragraph about it from the system prompt in settings and both the block and the buttons stop appearing. - Agent, model, thinking effort, and whether to search the web picked per question: in the question box, and behind the cog in the conversation footer. Each choice sticks as the default for the next one. The controls redraw for the agent you pick, because one CLI's model names mean nothing to another.
- A sidebar or a modal, whichever you set. The sidebar stays open beside the note; the modal covers it and closes on Escape.
- Answers render as markdown while they stream, with a copy button in the
turn's corner and a second one that follows the pointer from block to block —
that one copies the block's source, so a paragraph lifted into a note keeps
its
[[wikilinks]]and its code. The line under each question shows what the agent is reading while it works, then the model, how long it took, and tokens in and out. - Citations are links you can click. A claim from the vault is cited as
[[note name#heading]], and clicking it opens that note at that heading — Cmd or Ctrl for a new tab, as everywhere else in Obsidian.
Install
Not in the community plugin list yet, so either of these:
With BRAT. Add
dumkydewilde/obsidian-askai as a beta plugin. BRAT reads the latest release
and keeps it updated.
By hand. Download main.js, manifest.json and styles.css from the
latest release
into <your vault>/.obsidian/plugins/ask-ai/.
Then enable Ask AI under Settings → Community plugins, and set the command
for your agent if it is not on the PATH Obsidian inherits — which, launched from
Finder, is almost none of your shell's. That one setting is the usual reason a
fresh install cannot find claude or codex.
From source, which is also how you develop against it:
npm install
npm run build
npm run install-local -- "/path/to/your/vault"
After a later npm run build, rerun install-local and reload the vault window
(Cmd+R). The reload is not optional: a new main.js under a running Obsidian is
only picked up on Cmd+R.
Agents
| Agent | Command | What it can do |
|---|---|---|
| Claude Code | claude | Read, Grep, Glob only. No writes, no shell, no MCP, your own settings files ignored |
| Codex | codex | Read-only sandbox: no writes, network off unless web is on. Shell commands do run |
| Gemini CLI | gemini | Reads run; writes and shell are denied unprompted in headless mode. Web is asked for, not enforced |
| Custom command | yours | Whatever your command allows. The plugin cannot confine it |
Each agent keeps its own command path and its own model in settings, so you can switch between them without retyping either.
Claude Code
claude --print --output-format stream-json --include-partial-messages --verbose \
--restricted --strict-mcp-config --permission-prompts none \
--tools Read,Grep,Glob --allowedTools Read,Grep,Glob ...
--tools Read,Grep,Globis an exact allowlist from the built-in set. This is not the same as--allowedTools, which pre-approves those tools but leaves everything else available. An earlier version of this plugin used--allowedToolsalone and Claude reached forBashandcatanyway.- Both flags are needed, for different jobs.
--toolsdecides which tools exist at all;--allowedToolspre-approves those same ones so none of them trips a permission prompt that--permission-prompts nonewould auto-deny. Without the second flag,WebSearchis present but silently refused. - With Search the web on,
WebSearchandWebFetchjoin both lists. Nothing else changes: still no writes, still no shell. --restricteddrops Bash and the other code-running tools, ignores your user and project settings files, and confines the file tools to the working directory. Your own permissive~/.claude/settings.jsondoes not widen what a note question can reach.- Reads
CLAUDE.mdfrom the vault root, and takes the plugin's system prompt through--append-system-prompt. - A question about an image is answered by
Read, which opens images as images. The prompt names the file by absolute path, because that is whatReadtakes.
Codex
codex exec --json --ignore-user-config --skip-git-repo-check \
--sandbox read-only --config tools.web_search=false ...
Codex is sandboxed rather than tool-restricted. --sandbox read-only blocks
every write and, by default, the network — but it does not stop Codex running
shell commands to read, so in practice it will sed a note rather than call a
read tool. Reads only, but a wider door than Claude's.
--ignore-user-config is the closest thing to Claude's --restricted: it drops
your ~/.codex/config.toml and with it the plugins, hooks and MCP servers a note
question has no business loading. It costs a little — on a stock setup the same
question went from 52k input tokens to 18k.
A question about an image is the one place that sandbox costs something: no shell
command shows a model a PNG, so the file is attached with --image instead of
named in the prompt. It goes in as --image=<file> and not --image <file> — on
a fresh run the flag takes many values, and given a space it swallows the prompt
positional after it as a second image, leaving Codex waiting on stdin for a
question that was never asked.
Codex has no flag for a system prompt, so the plugin's rides in ahead of the
question. Prepended alone it ignored the follow-up block on every question tried,
which is what the trailing reminder is for. It reads AGENTS.md from the vault
root. Follow-ups resume through
codex exec resume <thread id>. It never names the model it ran, so the footer
falls back to "Codex".
Gemini CLI
gemini --output-format stream-json --approval-mode default --skip-trust ...
What keeps Gemini read-only is that in headless mode every tool needing
confirmation — write_file, replace, run_shell_command — is treated as
denied, while the read tools never ask. There is no tool allowlist on the command
line to make that explicit.
Its web tools never ask either, and there is no flag to withhold them, so
Search the web off is an instruction appended to the question rather than a
restriction. --skip-trust trusts the vault folder for the run, without which
the folder-trust check can disable tools in a vault you have not opened in Gemini
before. Like Codex, it takes the system prompt ahead of the question, and it
reads GEMINI.md from the vault root.
This adapter is written to Gemini's documented headless contract and its
stream-json event schema. It has not been run against a live gemini binary —
npm run smoke -- gemini will tell you.
Custom command
Set the command and its arguments. {prompt} is replaced by the question and
{model} by the model, quoted runs stay together, and a template with no
{prompt} gets the question appended. Stdout is read as the answer with terminal
escape codes stripped, so there are no tool calls to show, no session to follow up
on, and no confinement the plugin can promise.
Command opencode
Arguments run {prompt}
Settings
| Setting | Default | Why you would change it |
|---|---|---|
| Agent | Claude Code | Ask a different CLI. Also selectable per question |
| Agent command | the agent's own name | Your binary is somewhere unusual |
| Arguments | run {prompt} | Custom command only |
| Extra PATH entries | ~/.local/bin:/opt/homebrew/bin:/usr/local/bin | Obsidian launched from Finder has almost no PATH, so the binary is not found |
| Model | empty | Pin a model instead of using the agent's default. Kept per agent |
| Thinking effort | empty | Pin an effort level. Hidden for agents that have none |
| Search the web | off | Let answers cite sources outside the vault |
| Open answers in | Modal | Keep the conversation beside the note instead of over it |
| Keep conversations in the vault | on | Off, a conversation is not written unless you press save |
| Conversation location | In the folder specified below | Keep conversations beside the note they are about |
| Conversation folder | askai-conversations | Somewhere else, or your existing research folder |
| A folder per note | on | Off, they sit side by side as Note — Title.md |
| Research heading | ## Research | Match your own note conventions |
| Timeout | 180s | Long questions over a large vault |
| Resume a thread for | each agent's own window | You know better than the published cache TTLs. 0 always resumes |
| System prompt | see below | Change how answers are written |
Model names
Claude Code has no command that lists models, so its dropdown offers aliases
rather than exact ids: fable, opus, sonnet, haiku. An alias keeps
pointing at the newest model in its family as Claude Code updates, which a pinned
id does not. Gemini's aliases work the same way (pro, flash, flash-lite).
Type an exact id into the Model setting if you want one; it stays selectable in
the dropdown. Whichever ran is printed in the footer under each answer.
The system prompt
The prompt tells the agent to answer as a researcher: lead with the answer, use
its own knowledge of the subject rather than treating the vault as the limit of
what is knowable, cite vault notes and URLs as links you can click, close with a
Sources section that is a bare list of those links, keep it short, offer up
to three follow-up questions in a fenced follow-ups block when there are useful
ones, and name the conversation in a title block on its first answer. The
plugin lifts both blocks out of the answer, and turns them into buttons and into
the conversation's filename.
It is a setting, so it is saved in your vault. The settings file also records the default it was given, so an unedited prompt is replaced when the default improves and an edited one is left alone — without the plugin having to carry a copy of every prompt it has ever shipped. Claude Code takes it as a real system prompt; the others have no flag for one, so it goes in ahead of the question — and for those, the lines about the two trailing blocks are repeated after the question, because by the time they reach the end of the answer the prompt is a page behind them, and on a resumed turn it is not sent at all. Delete a paragraph from the prompt and its reminder stops too.
Vault context file
Each agent reads its own file from the vault root: CLAUDE.md, AGENTS.md,
GEMINI.md. A short one tells it how your vault is organised:
# Vault conventions
- Notes link with [[wikilinks]]. Resolve a link by searching for a file named
`<link>.md` anywhere in the vault.
- Frontmatter `status:` is one of seed, growing, evergreen.
- Daily notes live in `journal/` and are not sources of truth.
- When asked about a note, cite sections by heading, not by line number.
Development
npm run dev # rebuild on change; reload the vault window to pick it up
npm run build # typecheck, then bundle main.js
npm run check # the trailing-block parser and the note format; free and instant
npm run smoke # end-to-end against every agent on PATH, costs a few cents
npm run smoke -- codex # just one
npm run smoke covers everything outside Obsidian's UI: spawning the CLI,
parsing its output, resuming a session, the read-only confinement, whether the
agent actually emits the follow-up and title blocks the prompt asks for, a
missing binary, and cancelling a run. It runs every agent whose binary it can find, so it is also
how you check an agent this repository has not been able to test.
For the UI, harness/ loads the real built main.js against a stubbed Obsidian
API and mounts the real sidebar view, inside Obsidian's own app.css and your
vault's theme — both extracted from the installed app. Its vault is a map of
paths to strings and its agent is a script emitting Claude's stream-json, so a
question can be asked and the note it writes read back, in a browser:
npm run build && node harness/prepare.mjs
python3 -m http.server 8901 # from the plugin root, not from harness/
open http://127.0.0.1:8901/harness/index.html
A throw in onOpen shows up in the page with a stack. document.title holds the
measurements that are hard to eyeball — the padding that actually won, the
sidebar font size against the note's, whether the footer clears the status bar.
await window.ask("What happens above the cutoff?") // typed into the box
await window.askAbout("What is this?", { selection: "…" }) // the right-click path
await window.askAbout("What is this?", { image: window.image }) // and about an image
window.focused() // where the caret is
window.spawned // what each CLI was asked
window.thread(); window.branch() // the cold Ask button, and leaving the thread
window.dump() // every file, as written
window.conversations() // the list, as rendered
window.newConversation(); window.openNote(window.notes[1])
That is how the round trip is checked: ask, read the file the plugin wrote, ask
again and see it appended, start a second conversation and watch the list grow.
askAbout goes through the real question modal, so which conversation a question
lands in and where the caret ends up are checked the way the menu drives them.
A conversation already on disk is seeded before the plugin loads, so restoring
one is exercised as well as writing one — as is the one-time move of
conversations out of an old data.json.
It is a stub, not a simulator: anything the plugin reaches for that
harness/obsidian-stub.js does not define throws with its own name. That also
means it cannot vouch for the real API's behaviour — only that the plugin's own
code runs. harness/turns.html is the same CSS with hand-written markup for a
conversation that already has answers in it.
To drive the real thing instead, launch Obsidian with a debug port:
osascript -e 'tell application "Obsidian" to quit'
open -a Obsidian --args --remote-debugging-port=9222
node scripts/cdp.mjs 'app.commands.executeCommandById("ask-ai:ask-about-note"); return "opened";'
node scripts/cdp.mjs --screenshot shot.png --file probe.js
scripts/cdp.mjs evaluates an expression inside the running window over the
DevTools protocol and can grab a screenshot, which is how the modal layout and
the streaming render were checked. Restart Obsidian without the flag when you are
done; the debug port is unauthenticated.
Releasing
npm version minor && git push --follow-tags
The tag is what ships: .github/workflows/release.yml builds it and attaches
main.js, manifest.json and styles.css, which is what BRAT and the community
plugin list read. Nothing about merging to main reaches an installed vault.
Every pull request must bump the version with npm version patch, npm version minor, or npm version major. The Version metadata check compares the pull
request with its base and requires the package, manifest, lockfile, and Obsidian
version map to agree. Run it locally with npm run check:version -- --base <base>.
npm version only knows package.json, so scripts/version-bump.mjs runs as its
version lifecycle script and carries the number into the two files Obsidian reads
— manifest.json, which it installs against, and versions.json, which tells an
older Obsidian the newest build it can still run. All three land in the one commit
npm tags. The workflow refuses to build a tag whose manifest disagrees with it, so
a bump that skipped this would tag a release that never ships.
Notes
- Every CLI's output is reduced to the same handful of events in
providers.ts— append text, replace the message, clear the turn, a tool label, usage, an answer, a failure. Claude streams tokens, Codex delivers whole messages, Gemini streams tokens with no marker between turns;runner.tsdoes not know which. - Which image a question is about is read twice over, because neither way covers
the other.
editor-menusays which editor was right-clicked and not what in it was, and live preview draws an embed as a widget the caret does not move into — so a right-click on the picture is caught by a capturingcontextmenulistener and read off the.internal-embedwrapper'ssrc. With the caret on the line instead,embeds.tsfinds the embed by position. Both resolve throughgetFirstLinkpathDest, so a bare name, a vault path and a note-relative path all land on the same file, and a remote URL lands on none. - What a question is about beyond the note — a passage, an image — is one
AskContextthreaded from the right-click menu to the prompt, the pane and the conversation note. It reaches a resumed session too: the agent already has the note, but not the thing you just picked, and sending the bare question dropped it. - A session belongs to the agent that opened it. Switching agent mid-conversation starts a new one rather than handing a Codex thread id to Claude, and the questions and answers so far go into that first prompt so the new agent is not answering a follow-up cold. Questions and answers only: the notes behind them are in the vault, and it reads what it needs itself.
- Age is the second reason to start a new thread instead, and the branch icon is
the third; all three take the same path. No CLI reports a cache hit, so
cache.tsguesses from the clock: theupdatedstamp in the frontmatter against that agent's window — an hour for Claude Code, which asks Anthropic for the hour-long cache TTL, a quarter of one for Codex and Gemini CLI, whose providers cache for minutes. That makes the conversation note the record of the agent's memory as well as of the conversation, which it already was: the session id lives there too. Both errors are cheap. Too short a window replays two pages of text that were still cached; too long re-sends the whole transcript once. Neither loses an answer, and only the second costs real money, which is why the windows lean short. ModalandItemViewboth carry undocumented fields the type definitions do not declare. A subclass field of the same name wins, becausetarget: ES2022means class fields are defined rather than assigned, so a baretitleEl;redefines it asundefinedaftersuper()set the real one.Modalhasselectionandtitle, which is why the question modal calls themselectedTextandheading.ItemViewhastitleEl, andItemView.load()callsthis.titleEl.setText(...)— beforeonOpen()and outside the promiseonOpenreturns. Naming the sidebar's own headingtitleEltherefore threw inView.open(), skippedonOpen()entirely, and left a pane so blank it had no header either, with the error only in the developer console. It isnoteTitleEl. The harness reproduces this one on purpose.- The editor buffer is flushed to disk before each question. The agent reads the note off disk, so without that a question asked seconds after typing would be answered against the previous text.
- The sidebar keeps a host per note and a pane per conversation, hidden rather than unmounted, so a note you come back to still has its answers and its scroll position. Notes you never asked anything about are dropped past eight open, since rebuilding an empty one costs nothing.
- A note's conversations are found by frontmatter, not by folder: the walk is
over Obsidian's in-memory metadata cache, filtering on
type: ask-ai-conversationand resolving each one'ssourcelink. So moving a conversation, renaming it, or renaming the note it is about does not lose it, and nothing has to be kept in sync. The list of them is drawn from that cache too — the##headings Obsidian already parsed are the question count — so a file is only read when you open it. - New answers are appended to the file rather than rewriting it, and only the
frontmatter keys the plugin owns are rewritten. An answer you have edited
stays as you edited it, and a
tags:you added stays where you put it. - Reading a conversation back is deliberately forgiving. A heading with no answer
under it is still a question, prose above the first question is kept and shown,
a
##inside a code fence is not mistaken for a question, and an answer's own headings drop a level going in and come back up coming out.npm run checkcovers all of it, including a file edited by hand. - The conversation note is the only record. There is no second copy in
data.json, which is what the 20-note and 60k-character caps in the previous version were working around. - The agent's own thread is separate from the file. It resumes from a session id in the frontmatter, so editing an answer changes what the sidebar shows but not what the agent remembers. The conversation is a note in the vault, though, so asking it to read that note is enough when you want it to see your edits.
- Conversations left in an older
data.jsonare written out as notes once, on the first load of a new build, and linked from the notes they are about. That is a write into your vault on load; the alternative was dropping transcripts the previous version had promised to keep. - Navigating an internal link in an answer is the plugin's own job. Obsidian's
markdown renderer produces the anchors, but following them belongs to the
markdown view, so in a pane of its own every
[[note#heading]]was inert until one delegated click handler on the turn list calledopenLinkText. Delegated, so it survives an answer being re-rendered as it streams, and resolved against the note the conversation is about, so a bare[[#heading]]means that note's. - Copying one block hands back that block's markdown, matched to the rendered element by walking both in order and resyncing when they disagree — a list split by blank lines is several source blocks and one element. When they stop agreeing it hands back the rendered text rather than the wrong block.
- Obsidian's status bar is fixed to the bottom-right of the window, over the sidebar's footer, and its own panes clear it with a flat 32px of padding. This one measures the bar instead, so the footer sits right on top of it — and against the very bottom when the status bar is hidden.
.view-content.ask-ai-viewis deliberately two class names..workspace-leaf-content .view-contentin Obsidian's own stylesheet sets the padding, and a single class loses to it.- The trailing blocks are stripped from the answer as it streams, not only at the
end, so a half-written fence never flashes up as an empty code block. Only
prefixes of
follow-upsandtitleare stripped, so a half-typed ```sql fence is left where it is.npm run checkcovers that parser without spending agent tokens. - Claude's token counts add
input_tokens,cache_read_input_tokensandcache_creation_input_tokenstogether, because on a large note almost everything arrives as a cache read. Codex'sinput_tokensis already the whole prompt, so it is used as it comes. - Desktop only. It spawns a process, which Obsidian mobile cannot do.
Contributing
Issues and pull requests are welcome. Two things worth knowing before you open one:
npm run checkis free and instant and covers the parsers.npm run smokespends a few cents of agent usage per agent and is how an adapter gets verified — if you add one, run it and paste the output.- The Gemini CLI adapter is written to its documented headless contract and has
never been run against a live
geminibinary.npm run smoke -- geminiwill tell you, and that report is the single most useful thing anyone could send.
Licence
MIT. See LICENSE.
For plugin developers
Search results and similarity scores are powered by semantic analysis of your plugin's README. If your plugin isn't appearing for searches you'd expect, try updating your README to clearly describe your plugin's purpose, features, and use cases.