Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
43 changes: 43 additions & 0 deletions .changeset/local-llm-codes-in-process.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
---
"@openspec-ui/core": minor
"@openspec-ui/webui": minor
"@openspec-ui/server": minor
"@openspec-ui/cli": minor
"openspec-ui-vscode": minor
---

The local LLM agent is built in. `local-llm-acp` no longer starts an
external `coding-agent` that you had to find and install: it is a coding
agent inside the product, against your OpenAI-compatible server, with
tools to read, write, replace in a file, list, search and run a command,
all confined to the change's working directory. Each tool call and its
result shows in the run. It reads a tool call the model wrote as text
when the server's parser did not recognise it, as SGLang's `hermes`
parser does with Qwen3.6. Turn on
`openspec-ui.localLlm.agent.askBeforeCommands`
(`OPENSPEC_UI_LOCAL_LLM_ASK_BEFORE_COMMANDS=1`) to allow each command
yourself.

The model is optional for `local-llm` and `local-llm-acp`: a stage may
name one, the settings may, and otherwise the server is asked which model
it serves. A model id may now contain `/`, as Hugging Face names are
written.

Agents can ignore the system proxy: `openspec-ui.agents.ignoreSystemProxy`
(`OPENSPEC_UI_IGNORE_SYSTEM_PROXY=1`). The local LLM agents then connect
directly, and CLI agents start without the proxy variables and with
`NO_PROXY=*`.

A run that failed inside an ACP agent (any of them, not only
`local-llm-acp`) used to say only "Internal error" — the Agent Client
Protocol's own fixed text for an unhandled exception, with the actual
cause (an HTTP status, a bad key, a timeout) discarded. It now says that
cause.

With `openspec-ui.localLlm.agent.askBeforeCommands` on, clicking Allow on
a direct (non-chain) run used to remove the prompt and then go nowhere —
the answer reached a fresh, unrelated agent instance instead of the one
actually waiting on it, so the run sat there until cancelled by hand. It
now reaches the right one.


39 changes: 34 additions & 5 deletions HARNESS.md
Original file line number Diff line number Diff line change
Expand Up @@ -739,6 +739,7 @@ carried here so a reader sees the value before choosing):
| `gemini-cli` | — | — | — | — |
| `gemini-cli-acp` | — | — | — | — |
| `local-llm` | — | — | — | — |
| `local-llm-acp` | — | — | — | — |
| `vscode-chat` | — | — | — | — |

Even thirds — 1, 2/3, 1/3, 0 — rather than the tidier-looking 1, 0.75,
Expand Down Expand Up @@ -959,8 +960,8 @@ binary here" column repeats `README.md`'s own agent table.
| `copilot-cli` | `copilot` | Yes (`--model`) | `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max` | `maxAiCredits` (`--max-ai-credits`, minimum 30) | Yes |
| `codex-cli` | `codex` | No | `minimal`, `low`, `medium`, `high` (from OpenAI's documented config, not live-verified here) | No | **No — never** |
| `gemini-cli` | `gemini` | No | No mechanism | No | **No — never** |
| `local-llm` | HTTP to an OpenAI-compatible `/v1/chat/completions`, with a bearer key where one is set; where, which model and which key are set outside the harness file, see "The local LLM" below | No: the model is set with the address | No mechanism | No | Yes, 2026-09-25: a Qwen server on the LAN answered a prompt with its key, and refused it without one. Chat only: it answers in text and edits no file |
| `local-llm-acp` | `coding-agent --base-url <url> --model <name> [limit flags] acp` (the Python coding agent, 0.3.0 or later: the first version that speaks ACP), with the address, model and key of "The local LLM" below; the key travels in the environment, never the command line | No: the model comes from the local LLM settings, not harness model selection | No mechanism | No mechanism | Yes, 2026-10-01, through this runner and its ACP driver against SGLang serving Qwen3.6 on the LAN: an implement run edited code, added tests, ran them, ticked its tasks and reported its tokens. No permission request is ever sent |
| `local-llm` | HTTP to an OpenAI-compatible `/v1/chat/completions`, with a bearer key where one is set; where, which model and which key are set outside the harness file, see "The local LLM" below | Yes, optional (in the request); see "The local LLM" for the order | No mechanism | No | Yes, 2026-09-25: a Qwen server on the LAN answered a prompt with its key, and refused it without one. Chat only: it answers in text and edits no file |
| `local-llm-acp` | Nothing: a coding agent built into the product (ADR 0038), run in process against the local LLM of "The local LLM" below, with tools to read, write, replace in a file, list, search and run a command, all confined to the run's working directory | Yes, optional (in the request); see "The local LLM" for the order | No mechanism | No mechanism | Yes, see the change `local-llm-codes-in-process` for the run that verified it. Asks before a command only when told to (below) |
| `claude-cli-acp` | `claude --input-format stream-json --output-format stream-json` | Yes (`--model`) | Same as `claude-cli` | Same as `claude-cli` (`maxCostUsd`) | Progress only — no permission gate, see below |
| `copilot-cli-acp` | `copilot --acp` | Yes (`--model`) | Same as `copilot-cli` | Same as `copilot-cli` (`maxAiCredits`) | Yes |
| `codex-cli-acp` | externally installed `codex-acp` | No | No mechanism (deliberately empty — see below) | No mechanism (deliberately empty) | **No — never** |
Expand All @@ -970,15 +971,43 @@ binary here" column repeats `README.md`'s own agent table.

### The local LLM

`local-llm` is set outside `agent-harness.json`: that file is committed, an address on the LAN is one machine's, and a key committed is a key published.
`local-llm` and `local-llm-acp` are set outside `agent-harness.json`: that file is committed, an address on the LAN is one machine's, and a key committed is a key published.

| Setting | In VS Code | In the standalone and the CLI | When unset |
| --- | --- | --- | --- |
| Base URL, with its `/v1` or without | `openspec-ui.localLlm.baseUrl` | `OPENSPEC_UI_LOCAL_LLM_BASE_URL` | `http://localhost:30000` |
| Model, as its server names it | `openspec-ui.localLlm.model` | `OPENSPEC_UI_LOCAL_LLM_MODEL` | `default` |
| Model, as its server names it | `openspec-ui.localLlm.model` | `OPENSPEC_UI_LOCAL_LLM_MODEL` | the server is asked (below) |
| API key | **OpenSpec Workbench: Set Local LLM API Key...**, kept in the editor's secret storage | `OPENSPEC_UI_LOCAL_LLM_API_KEY` | no `Authorization` header |
| Ask before each command (`local-llm-acp`) | `openspec-ui.localLlm.agent.askBeforeCommands` | `OPENSPEC_UI_LOCAL_LLM_ASK_BEFORE_COMMANDS=1` | commands run without asking |

In VS Code the editor's value wins and the environment variable is the fallback. All three are read when the window opens, or when the server or the CLI starts. The key goes to the request's `Authorization: Bearer` header and nowhere else: not to the audit log, not to a run log. The agent answers in text: it can review, and it cannot write a proposal or tick a task.
In VS Code the editor's value wins and the environment variable is the fallback. All of them are read when the window opens, or when the server or the CLI starts. The key goes to the request's `Authorization: Bearer` header and nowhere else: not to the audit log, not to a run log.

**The model is optional.** A run takes the first of: the stage's own `model` (`stepAgents.<stage>.model`, which both agents accept); the setting above; the model the server lists at `/v1/models` (the first, where it lists several; asked once per address while the window or the server is open); `default`. The run's first output names the model and where it came from, for example `Model QuantTrio/Qwen3.6-35B-A3B-AWQ (the model the server serves).`

**`local-llm`** answers in text: it can review, and it cannot write a proposal or tick a task.

**`local-llm-acp`** is a coding agent built into the product (ADR 0038): nothing to install. It runs the model in a loop with six tools, `read_file`, `write_file`, `replace_text`, `list_dir`, `search_text` and `run_command`, and streams each call and its result as the run's updates.

- Every path is resolved by its real location, links included, and refused when that lies outside the run's working directory.
- A command runs in the working directory, is ended at `OPENSPEC_UI_LOCAL_LLM_ACP_COMMAND_TIMEOUT_SECONDS` (60 s by default), and its output is cut at `OPENSPEC_UI_LOCAL_LLM_ACP_MAX_COMMAND_OUTPUT_CHARS` (12000 by default).
- With **ask before each command** on, a command waits for Allow in the run's permission prompt; leave it off for a chain nobody watches. Writes inside the working directory never ask.
- A tool call the model wrote as text, which a server whose tool-call parser does not match the model passes through in `content` (Qwen3-Coder's `<function=...><parameter=...>` form, or Hermes' JSON, inside `<tool_call>`), is read as a call.
- Its loop is bounded by the `OPENSPEC_UI_LOCAL_LLM_ACP_*` limits in `LIMITS.md`.

### Ignoring the system proxy

**`openspec-ui.agents.ignoreSystemProxy`** in VS Code, or **`OPENSPEC_UI_IGNORE_SYSTEM_PROXY=1`** for the standalone server and the CLI, tells agents to ignore the system proxy. Off by default. Turn it on when the proxy cannot reach your model, such as a server on your LAN behind a proxy that resets such requests.

- `local-llm`, `local-llm-acp` and the check that says whether the local LLM is there connect directly, through a connection pool of their own, whatever proxy the editor applies to its own networking.
- A CLI agent is started with `HTTP_PROXY`, `HTTPS_PROXY` and `ALL_PROXY` removed from its environment, in either case, and `NO_PROXY=*`. Only a CLI that reads those variables honours that.

| Agent | Honours the switch |
| --- | --- |
| `local-llm`, `local-llm-acp` | Yes: they run in the product |
| `claude-cli`, `claude-cli-acp`, `copilot-cli`, `copilot-cli-acp` | Through the environment: their vendors document reading the proxy variables |
| `codex-cli`, `codex-cli-acp`, `gemini-cli`, `gemini-cli-acp`, `deepseek-cli-acp`, `vscode-chat` | Unknown: not checked here; such a CLI may keep proxy settings of its own |

The column is `HARNESS_AGENT_CAPABILITIES[*].systemProxy`.
**On the "run against the real binary here" column, plainly: `codex` and
`gemini` have never been run by this project at all**, raw or
ACP-flavored — see `README.md`'s "Agent Selection" section for the full
Expand Down
25 changes: 24 additions & 1 deletion LIMITS.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,7 +147,28 @@ conversion happens at all, so nothing to silently drift.
| --- | --- | --- | --- |
| `claude-cli`, `claude-cli-acp` | `maxCostUsd` | `--max-budget-usd` | Requires Claude Code v2.1.217 or later. |
| `copilot-cli`, `copilot-cli-acp` | `maxAiCredits` | `--max-ai-credits` | Minimum 30 — a configured value below this is rejected before any run starts. |
| `codex-cli`, `gemini-cli`, `local-llm`, `codex-cli-acp`, `gemini-cli-acp`, `deepseek-cli-acp`, `vscode-chat` | Neither | — | No spending-cap mechanism at all; a `stepAgents` entry setting either field for one of these is rejected. |
| `codex-cli`, `gemini-cli`, `local-llm`, `local-llm-acp`, `codex-cli-acp`, `gemini-cli-acp`, `deepseek-cli-acp`, `vscode-chat` | Neither | — | No spending-cap mechanism at all; a `stepAgents` entry setting either field for one of these is rejected. |

**`local-llm-acp`'s own loop is bounded instead** (ADR 0038). It runs in
the product, so these limits act on the loop itself rather than on a
process. Each is read from the environment when the window, the server or
the CLI starts; an absent or invalid value means the default.

| Variable | Bounds | Default |
| --- | --- | --- |
| `OPENSPEC_UI_LOCAL_LLM_ACP_MAX_ITERATIONS` | Model turns in one run | 40 |
| `OPENSPEC_UI_LOCAL_LLM_ACP_MAX_TOOL_CALLS` | Tool calls in one run | 120 |
| `OPENSPEC_UI_LOCAL_LLM_ACP_MAX_SECONDS` | Seconds the loop may run | 1800 |
| `OPENSPEC_UI_LOCAL_LLM_ACP_COMMAND_TIMEOUT_SECONDS` | Seconds one command may run | 60 |
| `OPENSPEC_UI_LOCAL_LLM_ACP_MAX_COMMAND_OUTPUT_CHARS` | Characters of a command's output the model is given | 12000 |
| `OPENSPEC_UI_LOCAL_LLM_ACP_MAX_PROMPT_TOKENS`, `..._MAX_COMPLETION_TOKENS`, `..._MAX_TOTAL_TOKENS` | Tokens over the run, as the server reports them; `..._MAX_COMPLETION_TOKENS` is also sent as each request's `max_tokens` | none |

A run that reaches one stops, says which, and ends with the protocol's
`max_turn_requests` or `max_tokens`. The harness's own `timeout` and
`maxStageAttempts` still bound the stage around it. The context limits
the external agent took (`..._MAX_CONTEXT_*`, `..._MIN_FREE_CONTEXT_TOKENS`)
are read but act on nothing: no OpenAI-compatible answer says how full the
context is.

**A mismatched field is refused when the configuration resolves, not
minutes into a run.** Setting `stepAgents.apply.budget.maxAiCredits` while
Expand Down Expand Up @@ -332,6 +353,7 @@ spend, as `deepseek-cli-acp` does; the two are not the same column.
| --- | --- | --- |
| `deepseek-cli-acp` | Yes, on every turn | **Measured** 2026-09-23: `{"used":8202,"size":1000000,"sessionUpdate":"usage_update"}`, and `dsh-acp`'s own `usageUpdate` builds it from a context meter |
| `claude-cli`, `copilot-cli`, `codex-cli`, `gemini-cli`, `local-llm`, `vscode-chat` | No | Certain — they speak no ACP at all, so no update of any kind arrives |
| `local-llm-acp` | No | Certain — it is the product's own agent and sends none: no OpenAI-compatible answer says how full the context is |
| `copilot-cli-acp`, `claude-cli-acp`, `gemini-cli-acp`, `codex-cli-acp` | Not known | *Unobserved here.* Nothing is claimed either way, and no warning is raised: an ACP agent nobody has watched may well send one |

| Agent | Reports usage | Evidence | Source |
Expand All @@ -340,6 +362,7 @@ spend, as `deepseek-cli-acp` does; the two are not the same column.
| `claude-cli-acp` | Cost (USD), input/output/cache tokens, per-model split | **Measured** — see below | `claude`'s own terminal `"result"` line (`total_cost_usd`, `usage`, `modelUsage`) |
| `gemini-cli-acp`, `codex-cli-acp` | Whatever that CLI sends over ACP — token totals, a cost, or nothing | *Unobserved* | ACP's `PromptResponse.usage` and `usage_update` notifications |
| `deepseek-cli-acp` | Nothing countable: no cost, no credits, no token split | **Measured** 2026-09-23, correcting 2026-09-22: a run through `dsh` 0.1.5-rc.2 does send `usage_update`, and what it carries is `used` and `size` - the tokens now in the session's context, against the model's 1,000,000-token window. One number, growing as the conversation grows: no input/output/cache/thought split, no per-model figures, no currency. The prompt's own answer carries `stopReason` and nothing else. `dsh`'s ACP layer builds that notification from its context meter alone, so there is nothing further to read. A context gauge is not a spend, and nothing here converts one into the other | ACP's `PromptResponse.usage` and `usage_update` notifications |
| `local-llm-acp` | Input and output **tokens**, summed over the run's model calls. **No cost.** | Certain — the product's own agent adds up what the server reports in each answer's `usage` | ACP's `PromptResponse.usage` |
| `claude-cli`, `copilot-cli`, `codex-cli`, `gemini-cli`, `local-llm` | Nothing | Certain — plain text carries no figure to record | Plain text output |
| `vscode-chat` | Nothing | Certain | The run is handed to VS Code chat; this project never sees its cost |

Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -456,7 +456,7 @@ streams events over the same protocol already used for
| Codex CLI | `codex` | **No — never** |
| Gemini CLI | `gemini` | **No — never** |
| Local LLM (OpenAI-compatible) | HTTP to `http://localhost:30000` by default | Not exercised live |
| Local LLM (ACP, OpenAI-compatible) | `coding-agent --base-url <url> --model <name> acp` (the Python coding agent, 0.3.0+) | Yes, against SGLang serving Qwen3.6; edits files, never asks permission |
| Local LLM agent (OpenAI-compatible, built in) | Nothing to install: a coding agent built into the product, against the same server (ADR 0038) | Yes, against SGLang serving Qwen3.6; edits files in the change's directory, asks before commands when told to |
| Claude CLI (ACP) | `claude --input-format stream-json --output-format stream-json` | Progress only — no permission gate, see below |
| GitHub Copilot CLI (ACP) | `copilot --acp` | Yes |
| Codex CLI (ACP) | externally installed `codex-acp` | **No — never** |
Expand Down
Loading
Loading