Skip to content

Tab completion and next-edit prediction

Tab completion is a fill-in-the-middle request to a local model — Ollama or llama.cpp — with the text before and after the cursor. It never goes to a hosted provider, so no line of your code leaves the machine. It is off by default.

sirius.ai.enable is a per-language map. "*" is the default for every language; a language id overrides it:

"sirius.ai.enable": { "*": true, "markdown": false }

Then pick the model with sirius.ai.completions.model — auto finds a running Ollama’s first code model by name; ollama/<model> names one; llamacpp uses the server at sirius.ai.llamacpp.baseUrl. Local models lists the models that answer FIM well. Changing the model setting takes effect after a window reload.

The editor’s inline-suggestion controls — the status-bar entry’s Inline Suggestions checkboxes for all files and the current language — are wired to the same setting, so either place works; settings.json is the one that is always there.

Suggestions appear as ghost text after a 180 ms pause in typing; Tab accepts, Esc dismisses. Sirius sends up to 2,000 characters before and 600 after the cursor, asks for at most 256 tokens at a low temperature, stops at a blank-line gap, and trims the result to sirius.ai.completions.maxLines (default 12). A newer keystroke cancels the request in flight; recent contexts are cached; nothing is requested in an empty file; a suggestion that merely repeats the text after the cursor is dropped. It works in files on disk and in untitled buffers.

If no backend is reachable, nothing is shown — no error. Sirius re-checks for one every 30 seconds.

sirius.ai.nextEditSuggestions.enabled (default false, experimental) adds a second kind of suggestion: after you make a small edit — a single-line change of up to 120 characters — Sirius asks the same local model what the next edit should be, looking at about 40 lines either side of the cursor, and offers it as an inline diff you take with Tab. The prediction is only requested within 8 seconds of the edit, it never targets the line you just changed, and the model must answer in a strict format or say NONE, so a vague answer produces nothing rather than noise. It needs Tab completion to be on for the language.

Setting Default Meaning
sirius.ai.enable {"*": false} Tab completion per language
sirius.ai.completions.model auto Which local FIM model answers
sirius.ai.completions.maxLines 12 Longest suggestion offered
sirius.ai.nextEditSuggestions.enabled false Predict the next edit after each change
sirius.ai.inlineCompletions false Deprecated: true still turns completion on everywhere