Chat Assistant — design
A tabbed chat panel in the code-editor screen, backed by a LangGraph
agent that can answer EdgeWeave questions, generate Python, and
eventually author full nodes / .weave files. Provider-abstracted
so users can plug in OpenAI, Anthropic, Gemini, or local models with
their own API keys.
This doc is the design proposal — no code yet — to lock in the shape before we wire it up.
Goals
- Chat panel co-located with the Variables panel on the right
of
code_editor.html, with tabs to switch between "Variables" and "Chat" (and room for more tabs later). - LLM-assisted Python authoring — replies stream into the chat, and code blocks expose an "Insert into editor" button.
- EdgeWeave-aware — the agent has enough SDK context to write a
correct
@register_block(...)definition or a syntactically valid.weavefile when asked. - Skill-like context routing — small prompt packs are loaded only when the router decides they're relevant, so simple chitchat doesn't pay for the whole SDK reference in tokens.
- Provider abstraction — OpenAI first (
gpt-4o-mini), but the integration sits behind an interface so Anthropic / Gemini / Ollama etc. can be added without touching agent logic. - Secrets handled safely — a
.env-reader node that exposes keys but never round-trips values through saved graph state. Threat model anticipates a future cloud deployment of EdgeWeave where DevTools is no longer trusted.
UI: tabbed right panel
frontend/src/code_editor.html currently has:
<div id="variables">
<div id="variables-header">Variables</div>
<div id="variables-list">…</div>
</div>
The proposal is to split that into a tab strip plus two stacked panes:
<div id="right-panel">
<div id="right-tabs" role="tablist">
<button class="right-tab active" data-tab="variables">Variables</button>
<button class="right-tab" data-tab="chat">Chat</button>
</div>
<div id="right-tab-panes">
<div id="variables" class="right-pane active">…</div>
<div id="chat" class="right-pane">
<div id="chat-messages"></div>
<div id="chat-composer">
<textarea id="chat-input" rows="3"
placeholder="Ask about your code, EdgeWeave, or paste an error…"></textarea>
<div class="chat-controls">
<select id="chat-model">…</select>
<button id="chat-send">Send</button>
</div>
</div>
</div>
</div>
</div>
Tab switching is just display: none / flex — keep it framework-free
to match the rest of the codebase. Persist the active tab to
localStorage('rightTab') so it survives reloads.
CSS hooks live in a new frontend/src/css/chat.css; the tab strip
uses the existing --panel-bg-alt / --border-color theme variables
so it themes correctly out of the box.
The chat pane is wired in frontend/src/js/chat_ui.js (new):
renderMessage({role, content, code_blocks})— appends a chat bubble; code blocks render through the same CodeMirror that the editor uses (read-only, syntax-highlighted) and get an "Insert into editor" button next to a "Copy" button.streamFromBackend(prompt)— opens a websocket to/ws/chatand pushes incremental tokens into the latest assistant bubble.
Tab system, not just two tabs
The tab strip should be generic so we can add e.g. an "AI Logs"
tab (for showing token usage, the active prompt pack list, retrieval
hits) without rewriting it every time. A tiny tabsRegister(tabId, {label, paneEl, onActivate, onDeactivate}) helper in chat_ui.js is
plenty — no router lib needed.
Backend: /api/chat
Two new routes alongside the existing cmd_router /
explorer_router:
POST /api/chat/message # one-shot, JSON in/out (testing fallback)
WS /ws/chat # streaming tokens + tool-call events
backend/api/chat_router.py is thin — it only marshals the request
into a ChatSession and pipes events back. The actual graph lives
under backend/ai/:
backend/ai/
├── __init__.py
├── graph.py # LangGraph StateGraph factory
├── providers/
│ ├── base.py # LLMProvider protocol, Message/StreamEvent/ModelInfo
│ ├── registry.py # name → provider, picker metadata, model cache, .env refresh
│ ├── openai_compat.py # OpenAI, Gemini, Ollama, any OpenAI-compatible server
│ └── anthropic_provider.py # Claude (official SDK)
├── prompt_packs/
│ ├── __init__.py
│ ├── registry.py # discovery + relevance scoring
│ ├── core.md # always-loaded — tool name, response style
│ ├── sdk.md # @register_block reference, field types
│ ├── nodes.md # cheat-sheet of existing nodes (auto-built)
│ ├── weave_format.md # the .weave JSON schema
│ └── codemirror.md # editor-insert formatting rules
├── tools/
│ ├── __init__.py
│ ├── read_file.py
│ ├── list_nodes.py # → live snapshot of NODE_REGISTRY
│ └── insert_code.py # asks the frontend to insert into editor
├── secrets.py # the in-memory keystore (see below)
└── session.py # per-tab conversation state
Why this shape:
providers/is justclass LLMProvider(Protocol): async def stream(messages, tools) -> AsyncIterator[Event]. Everything above it (graph, packs, tools) consumes the protocol — swapping OpenAI for Anthropic is one factory line.prompt_packs/are markdown files, loaded on demand. This matches the way Claude Code skills work and lets you edit the packs without restarting Python (hot reload vialoader.py'sos.walkwatcher pattern, same as user_blocks).tools/are LangGraph tool nodes the agent can call;list_nodesin particular is justNODE_REGISTRY.values()filtered to the serializable keys, so the agent always sees the currently registered blocks, not a stale doc.
LangGraph topology
┌─────────────┐
│ user_input │
└──────┬──────┘
▼
┌─────────────┐
│ router │ ← classifies: chitchat | python_help |
└──────┬──────┘ sdk_authoring | weave_authoring | tool_call
▼
┌─────────────┐
│ pack_loader │ ← reads prompt_packs/*.md per route, caches
└──────┬──────┘
▼
┌─────────────┐
│ generator │ ← provider.stream(...) — emits tokens
└──────┬──────┘
▼
┌─────────────┐
│ tool_runner │ ← optional; loops back to generator
└──────┬──────┘
▼
[stream out]
The router is itself a small LLM call (cheap model, gpt-4o-mini
or haiku) returning {"route": "...", "packs": ["sdk", "nodes"]}
JSON. For low-effort optimisations, cache routing decisions on a
hash of the last-3-message window so back-to-back questions on the
same topic don't pay routing cost twice.
pack_loader deduplicates packs across the conversation so we
don't re-send the SDK reference every turn after the first time it's
been pulled in.
Provider abstraction
# backend/ai/providers/base.py
class LLMProvider(Protocol):
name: str
default_model: str
async def stream(
self,
messages: list[Message],
*,
model: str | None = None,
tools: list[ToolSpec] | None = None,
temperature: float = 0.2,
) -> AsyncIterator[StreamEvent]: ...
StreamEvent is a tagged union: TextDelta, ToolCall, ToolResult,
Done, Error. This is just OpenAI's streaming shape generalised —
the Anthropic and Gemini SDKs all map cleanly onto it.
A single ProviderRegistry keyed on lowercase name ("openai",
"anthropic", …) is the only place agent code touches concrete
providers. Switching is a settings dropdown that writes
localStorage('chatProvider') and a request header — no graph
rebuild required.
Per-user API keys live alongside the secrets store described below;
they are never persisted to a .weave or any other graph file.
Prompt packs — skill-like routing
Each pack is a markdown file with optional YAML frontmatter:
---
id: sdk
description: Reference for writing EdgeWeave nodes with @register_block.
triggers: ["register_block", "create a node", "node decorator", "fields"]
priority: 10
---
# EdgeWeave SDK Reference
…content…
The router looks at:
- Explicit triggers in the user's message (cheap, deterministic).
- A short LLM classify call if no explicit trigger fires.
- Recent tool calls (e.g. if the agent just called
list_nodes, thenodespack is implicitly relevant).
Packs not selected are not sent to the model, so a chitchat turn costs ~1k tokens of system prompt instead of ~10k.
Build the initial pack set by extracting from the existing docs we have:
core.md— voice, response style, when to ask vs. answer.sdk.md— distilleddecorators.py+ the@nodewrapper, plus the canonical "good" example (e.g.numpy_io.py).weave_format.md— the.weaveJSON shape, field conventions, how loops are represented.codemirror.md— when to wrap code in fenced blocks vs. emit whole-file replacements.nodes.md— auto-generated at startup fromNODE_REGISTRYso it reflects whatever is actually registered, including user blocks.
Code generation surface
Three increasingly intrusive levels — ship them in this order:
- Chat-only — code lives in fenced blocks in the assistant bubble. Two buttons: "Copy", "Insert into editor".
- Targeted insertion — agent emits a structured tool call
(
insert_code) carrying{path, range, replacement}. Frontend shows a diff modal before applying. - Inline ghost completions — CodeMirror ghost-text completions driven by a debounced "complete this" call. Only worth doing after (1) and (2) feel reliable.
For (2) and (3) we'll need to extend preload.cjs with a
window.chatBridge.insert(path, range, replacement) IPC method so
the agent can drive editor mutations through the same secure boundary
as the file ops you already have.
Secrets handling
The .env-reader node and the chat assistant's API keys share one
problem: a value typed in by the user must never end up serialised
into a .weave, an exported Python script, or a log line.
Design
- New module
backend/ai/secrets.pyprovides a process-localKeystore:set(name, value)— store in memory only.get(name)— return the value, intended only for handler runtime.redacted(name)— return"***"for previews.list_keys()— return key names only.
- New node
read_env(incore/nodes/env_io.py):- Field:
path(file picker, defaults to.envin the project root). - Parses the file with
python-dotenv(dotenv.dotenv_values(path)) — preferred over hand-rolled parsing soexport FOO=bar, comments, quoted values, and multi-line keys all behave as developers expect. - Sets each key/value into the keystore and emits a list of key names. The values never enter the graph's data dict.
- Preview shows
KEY = ***rows. - Canonical use: a
.envcontainingOPENAI_API_KEY=sk-...,ANTHROPIC_API_KEY=...,GEMINI_API_KEY=...is read once; downstreamsecret_get("OPENAI_API_KEY")flows the value into whichever provider node consumes it.
- Field:
- The chat assistant's own provider clients read keys via
keystore.get("<PROVIDER>_API_KEY")first and fall back toos.environonly if the key is unset — so a developer running from source with their shell env exported still works without authoring aread_envnode, but a packaged install reads from the user's.envthrough the keystore path. - Companion node
secret_get(name):- One field:
name. - At runtime, returns
keystore.get(name). Output flows into downstream nodes as a normal string but is filtered in preview HTML and node logs.
- One field:
parse_drawflowand the save path already only touchdata— since the values aren't indata, save is automatically safe.export.py(Python codegen) emitsos.environ.get("KEY")rather than baking the literal — even when the value is in memory at export time.
Threat model now (Electron)
- DevTools is the renderer process; the keystore lives in Python so
values aren't reachable from
window.*. - The
read_envnode uses Electron's existing file picker (already scoped) to read the.env, then hands the bytes to Python over the existing IPC. - Worker subprocess sees secrets only via
keystore.get; it should not be able to enumerate the keystore wholesale. Add aKEYSTORE_DENY_LIST_KEYSflag for paranoid users.
Threat model later (cloud)
If EdgeWeave goes multi-tenant in the browser:
- The whole
Keystoremust move to a per-session server-side store that the browser holds only an opaque session ID for. No values ever transit to the renderer. - The
secret_getnode returns a handle ("secret://uuid") on the client, the worker resolves it server-side at execution time. - All
console.log/WebSocketpayloads pass through a redactor that scans for any in-memory secret values and replaces them with***. Cheap to implement, catches accidental log leaks. .envupload is opt-in and one-shot — files never persist on disk.
The first two bullets above (handle indirection + redactor) are the
ones worth prototyping early, even on Electron, because they're the
ones that bite the hardest if retrofitted. The handle pattern also
naturally lets the read_env node display only key names in
previews without any special-casing.
Phasing
- M1 — Tabbed UI shell. (shipped) Tab system, empty chat pane, no backend. Zero risk; lets us iterate on layout / theming.
- M2 — One-shot OpenAI chat. (shipped)
POST /api/chat/message,gpt-4o-mini, no streaming, no tools. Provider abstraction (backend/ai/providers/{base,registry,openai_provider}.py) landed alongside the OpenAI implementation so M4 can register Anthropic / Gemini without touching the router. LangGraph deliberately deferred to M3 — for one-shot chat with no tools or routing, wrapping the single provider call in aStateGraphadds an import dep and a factory file but doesn't constrain the M3 design. The.envis loaded at backend startup viadotenv.load_dotenv(), soOPENAI_API_KEYfrom the project-root.envworks without a UI step. - M2.5 — Markdown rendering + multi-language code blocks.
(shipped) Assistant replies now flow through
marked→DOMPurify(strict allow-list) into the bubble, and every fenced```langblock is post-processed into a read-only CodeMirror view with a hover "Copy" button. Language dispatch is centralised infrontend/src/js/code_languages.js, which also feedscodemirror_main.jsso opening a.json/.yaml/.md/.js/.tsfile in the editor pane gets the correct language extension instead of the Python parser. EdgeWeave is still Python-first; an unknown extension falls back topython()to preserve existing behaviour. An "Insert into editor" button is deferred to M6 (it needs theinsert_codetool/IPC plumbing anyway). - M3 — Streaming. (shipped — streaming only; LangGraph still
deferred until there are tools/routing for it to orchestrate.)
POST /api/chat/streamreturns newline-delimited JSON (delta…done|error) instead of a WebSocket: the frontend is afetch()+ stream reader (frontend/src/js/chat_providers.js), and failures before the first byte are still ordinary HTTP errors. The provider's blocking stream runs in a worker thread with a stop flag; the async relay pollsrequest.is_disconnected(), so the panel's Stop button closes the upstream request. (Handing Starlette the sync generator directly doesn't do this — uvicorn silently drops writes after a disconnect, so the generator ran to completion and a single-slot server like Ollama kept generating the abandoned reply while the next request queued behind it.) - M4 — Providers + model picker. (shipped; key entry still via
.env— a Settings screen is a later step.) Five providers inbackend/ai/providers/:openai,gemini,ollamaandopenai_compatibleshareopenai_compat.py(Gemini, Ollama and any Groq/Mistral/OpenRouter/LM Studio-style server speak the Chat Completions format, so no extra SDKs), andanthropicuses the official SDK (anthropic_provider.py). Each provider has a class-levelstatus()(key present / Ollama reachable — no network to the vendor) andlist_models()(live from the vendor's models endpoint, filtered to chat models, with a built-in fallback list); the registry caches model lists for 10 minutes. The panel shows a provider dropdown and a model dropdown (plus "Custom model id…"), remembered per provider inlocalStorage; unconfigured providers show their setup hint and disable Send.refresh_env()re-reads.envbefore each request, so adding a key needs no restart. Provider quirks handled: OpenAI reasoning models rejecttemperature(retried once without it, then remembered); current Claude models reject sampling parameters (never sent), needmax_tokens(taken from the Models API's published output cap, capped at 64k) and opt into server-side refusal fallbacks (fallbacks: "default") where supported. - M5 — Knowledge base (graph-RAG). (shipped instead of the
prompt-pack router — retrieval picks context per question, which is
what the packs were for, without hand-maintained packs.)
backend/ai/knowledge.pyindexes three kinds of card: one per registered node, built live fromNODE_REGISTRY(label, registered name, category, doc, typed ports, fields, export support — accurate by construction thanks to the metadata audit, and including hot-reloaded user blocks); one per example.weavein the demo project and the active project (sticky-note text, nodes used, wiring); and the markdown docs split by section (config.DOCS_DIR;docs/docs/features-overview.mdwas written as the user-facing "how do I…" page). Links form a small graph: a demo links to the nodes it uses, a doc section to the nodes it names. Search scores every card, boosts nodes the question names explicitly (registered name, multi-word label, or "the Retrieve node" — entity linking, added after TF-IDF ranked a node's own card below docs sharing the phrase "inputs and fields"), then follows links: a node hit brings a demo that uses it, a demo/doc hit brings the nodes it mentions. Scoring reuses the AI nodes' toolkit (backend/ai/toolkit.py): TF-IDF always, plus a hybrid with Ollama'snomic-embed-textwhen it is pulled (built in a background thread). The index rebuilds when the registry or demo/doc files change and is warmed when the chat panel loads./api/chat/streaminserts the hits as a system message before the latest user turn (short follow-ups also use the previous question) and sends asourcesevent first; the panel shows them under the reply and has an "EdgeWeave" toggle (on by default) to turn grounding off.GET /api/chat/knowledge?q=shows what would be retrieved. Measured on gpt-4.1-mini: asked which nodes build a US-states choropleth from a CSV, it namedcsv_read+plotly_choropleth(Location mode "USA-states") andmaps/plotly_choropleth_demo.weave; with grounding off it invented "GeoJSON Reader", "Join" and "Color Scale" nodes. Cost: ~1.5–2k extra input tokens per question. - M6 — Tools.
list_nodes,read_file,insert_code. - M7 — Secrets node (
read_env,secret_get) and the keystore. Land the redactor and handle pattern from day one. - M8 — Inline ghost completions. Optional — gauge after M6.
Each milestone is independently shippable to dev and useful on its
own, so we can pause anywhere.
Decisions logged
- Conversation persistence — localStorage + summarise (2026-10-04).
(Supersedes "none for now", 2026-04-27.) Chats are saved in the
renderer's
localStorage(frontend/src/js/chat_history.js): up to 30 conversations, capped at ~2 MB total (oldest dropped first), never written to a.weave, the backend or disk files we manage. The chat toolbar has a saved-chat dropdown, New, Summarise and Delete; each conversation remembers its provider/model. This is deliberately per-machine and unsynced — moving it server-side is a later decision (it brings back the redaction question). Summarise / compact: once a chat reaches 20 user+assistant turns or ~40k characters (~10k tokens), a banner suggests continuing from a summary, because every turn resends the whole history (cost grows with length) and models lose track of early details. Summarise asks the selected model for a structured markdown summary (Goal / Key facts & decisions / Code / Open questions) and starts a new chat seeded with it as asystemturn (kind: "summary", shown as a collapsible bubble); the original chat is kept. This is the "markdown summary seeds the next session" idea from the earlier decision, done client-side and provider-agnostic. (Anthropic's server-side compaction isn't used: it only applies to Claude and to contexts far larger than a chat panel reaches.)
Open questions
- Cost telemetry. Show running token/cost counter in the chat header? Useful for users with their own keys; mandatory for any future hosted plan.
- Tool approval UX. Auto-approve
read_file, prompt forinsert_code, never auto-approve anything that touches.env. This is the same model Claude Code uses and it works. - Local model story. Ollama is easy via the abstraction; do we ship a "local-only" mode that disables network egress for the paranoid?
- Chat in the graph editor too? The right panel currently
exists only on
code_editor.html. Worth mirroring onindex.html(graph editor) so users can ask "what does this node do?" while viewing the graph. The samechat_ui.jsshould drop into both.
Appendix — small unrelated heads-up
While reading sdk/decorators.py I noticed that register_block
explicitly whitelists meta keys when populating NODE_REGISTRY, and
the round-2 resizable flag isn't in the whitelist — so it never
reaches the frontend via `/api