Skip to main content

Chat Assistant — design

A tabbed chat panel in the code-editor screen, backed by a LangGraph agent that can answer EdgeWeave questions, generate Python, and eventually author full nodes / .weave files. Provider-abstracted so users can plug in OpenAI, Anthropic, Gemini, or local models with their own API keys.

This doc is the design proposal — no code yet — to lock in the shape before we wire it up.

Goals​

  1. Chat panel co-located with the Variables panel on the right of code_editor.html, with tabs to switch between "Variables" and "Chat" (and room for more tabs later).
  2. LLM-assisted Python authoring — replies stream into the chat, and code blocks expose an "Insert into editor" button.
  3. EdgeWeave-aware — the agent has enough SDK context to write a correct @register_block(...) definition or a syntactically valid .weave file when asked.
  4. Skill-like context routing — small prompt packs are loaded only when the router decides they're relevant, so simple chitchat doesn't pay for the whole SDK reference in tokens.
  5. Provider abstraction — OpenAI first (gpt-4o-mini), but the integration sits behind an interface so Anthropic / Gemini / Ollama etc. can be added without touching agent logic.
  6. Secrets handled safely — a .env-reader node that exposes keys but never round-trips values through saved graph state. Threat model anticipates a future cloud deployment of EdgeWeave where DevTools is no longer trusted.

UI: tabbed right panel​

frontend/src/code_editor.html currently has:

<div id="variables">
<div id="variables-header">Variables</div>
<div id="variables-list">…</div>
</div>

The proposal is to split that into a tab strip plus two stacked panes:

<div id="right-panel">
<div id="right-tabs" role="tablist">
<button class="right-tab active" data-tab="variables">Variables</button>
<button class="right-tab" data-tab="chat">Chat</button>
</div>

<div id="right-tab-panes">
<div id="variables" class="right-pane active">…</div>
<div id="chat" class="right-pane">
<div id="chat-messages"></div>
<div id="chat-composer">
<textarea id="chat-input" rows="3"
placeholder="Ask about your code, EdgeWeave, or paste an error…"></textarea>
<div class="chat-controls">
<select id="chat-model">…</select>
<button id="chat-send">Send</button>
</div>
</div>
</div>
</div>
</div>

Tab switching is just display: none / flex — keep it framework-free to match the rest of the codebase. Persist the active tab to localStorage('rightTab') so it survives reloads.

CSS hooks live in a new frontend/src/css/chat.css; the tab strip uses the existing --panel-bg-alt / --border-color theme variables so it themes correctly out of the box.

The chat pane is wired in frontend/src/js/chat_ui.js (new):

  • renderMessage({role, content, code_blocks}) — appends a chat bubble; code blocks render through the same CodeMirror that the editor uses (read-only, syntax-highlighted) and get an "Insert into editor" button next to a "Copy" button.
  • streamFromBackend(prompt) — opens a websocket to /ws/chat and pushes incremental tokens into the latest assistant bubble.

Tab system, not just two tabs​

The tab strip should be generic so we can add e.g. an "AI Logs" tab (for showing token usage, the active prompt pack list, retrieval hits) without rewriting it every time. A tiny tabsRegister(tabId, {label, paneEl, onActivate, onDeactivate}) helper in chat_ui.js is plenty — no router lib needed.

Backend: /api/chat​

Two new routes alongside the existing cmd_router / explorer_router:

POST /api/chat/message # one-shot, JSON in/out (testing fallback)
WS /ws/chat # streaming tokens + tool-call events

backend/api/chat_router.py is thin — it only marshals the request into a ChatSession and pipes events back. The actual graph lives under backend/ai/:

backend/ai/
├── __init__.py
├── graph.py # LangGraph StateGraph factory
├── providers/
│ ├── base.py # LLMProvider protocol, Message/StreamEvent/ModelInfo
│ ├── registry.py # name → provider, picker metadata, model cache, .env refresh
│ ├── openai_compat.py # OpenAI, Gemini, Ollama, any OpenAI-compatible server
│ └── anthropic_provider.py # Claude (official SDK)
├── prompt_packs/
│ ├── __init__.py
│ ├── registry.py # discovery + relevance scoring
│ ├── core.md # always-loaded — tool name, response style
│ ├── sdk.md # @register_block reference, field types
│ ├── nodes.md # cheat-sheet of existing nodes (auto-built)
│ ├── weave_format.md # the .weave JSON schema
│ └── codemirror.md # editor-insert formatting rules
├── tools/
│ ├── __init__.py
│ ├── read_file.py
│ ├── list_nodes.py # → live snapshot of NODE_REGISTRY
│ └── insert_code.py # asks the frontend to insert into editor
├── secrets.py # the in-memory keystore (see below)
└── session.py # per-tab conversation state

Why this shape:

  • providers/ is just class LLMProvider(Protocol): async def stream(messages, tools) -> AsyncIterator[Event]. Everything above it (graph, packs, tools) consumes the protocol — swapping OpenAI for Anthropic is one factory line.
  • prompt_packs/ are markdown files, loaded on demand. This matches the way Claude Code skills work and lets you edit the packs without restarting Python (hot reload via loader.py's os.walk watcher pattern, same as user_blocks).
  • tools/ are LangGraph tool nodes the agent can call; list_nodes in particular is just NODE_REGISTRY.values() filtered to the serializable keys, so the agent always sees the currently registered blocks, not a stale doc.

LangGraph topology​

┌─────────────┐
│ user_input │
└──────┬──────┘
▼
┌─────────────┐
│ router │ ← classifies: chitchat | python_help |
└──────┬──────┘ sdk_authoring | weave_authoring | tool_call
▼
┌─────────────┐
│ pack_loader │ ← reads prompt_packs/*.md per route, caches
└──────┬──────┘
▼
┌─────────────┐
│ generator │ ← provider.stream(...) — emits tokens
└──────┬──────┘
▼
┌─────────────┐
│ tool_runner │ ← optional; loops back to generator
└──────┬──────┘
▼
[stream out]

The router is itself a small LLM call (cheap model, gpt-4o-mini or haiku) returning {"route": "...", "packs": ["sdk", "nodes"]} JSON. For low-effort optimisations, cache routing decisions on a hash of the last-3-message window so back-to-back questions on the same topic don't pay routing cost twice.

pack_loader deduplicates packs across the conversation so we don't re-send the SDK reference every turn after the first time it's been pulled in.

Provider abstraction​

# backend/ai/providers/base.py
class LLMProvider(Protocol):
name: str
default_model: str

async def stream(
self,
messages: list[Message],
*,
model: str | None = None,
tools: list[ToolSpec] | None = None,
temperature: float = 0.2,
) -> AsyncIterator[StreamEvent]: ...

StreamEvent is a tagged union: TextDelta, ToolCall, ToolResult, Done, Error. This is just OpenAI's streaming shape generalised — the Anthropic and Gemini SDKs all map cleanly onto it.

A single ProviderRegistry keyed on lowercase name ("openai", "anthropic", …) is the only place agent code touches concrete providers. Switching is a settings dropdown that writes localStorage('chatProvider') and a request header — no graph rebuild required.

Per-user API keys live alongside the secrets store described below; they are never persisted to a .weave or any other graph file.

Prompt packs — skill-like routing​

Each pack is a markdown file with optional YAML frontmatter:

---
id: sdk
description: Reference for writing EdgeWeave nodes with @register_block.
triggers: ["register_block", "create a node", "node decorator", "fields"]
priority: 10
---
# EdgeWeave SDK Reference

…content…

The router looks at:

  1. Explicit triggers in the user's message (cheap, deterministic).
  2. A short LLM classify call if no explicit trigger fires.
  3. Recent tool calls (e.g. if the agent just called list_nodes, the nodes pack is implicitly relevant).

Packs not selected are not sent to the model, so a chitchat turn costs ~1k tokens of system prompt instead of ~10k.

Build the initial pack set by extracting from the existing docs we have:

  • core.md — voice, response style, when to ask vs. answer.
  • sdk.md — distilled decorators.py + the @node wrapper, plus the canonical "good" example (e.g. numpy_io.py).
  • weave_format.md — the .weave JSON shape, field conventions, how loops are represented.
  • codemirror.md — when to wrap code in fenced blocks vs. emit whole-file replacements.
  • nodes.md — auto-generated at startup from NODE_REGISTRY so it reflects whatever is actually registered, including user blocks.

Code generation surface​

Three increasingly intrusive levels — ship them in this order:

  1. Chat-only — code lives in fenced blocks in the assistant bubble. Two buttons: "Copy", "Insert into editor".
  2. Targeted insertion — agent emits a structured tool call (insert_code) carrying {path, range, replacement}. Frontend shows a diff modal before applying.
  3. Inline ghost completions — CodeMirror ghost-text completions driven by a debounced "complete this" call. Only worth doing after (1) and (2) feel reliable.

For (2) and (3) we'll need to extend preload.cjs with a window.chatBridge.insert(path, range, replacement) IPC method so the agent can drive editor mutations through the same secure boundary as the file ops you already have.

Secrets handling​

The .env-reader node and the chat assistant's API keys share one problem: a value typed in by the user must never end up serialised into a .weave, an exported Python script, or a log line.

Design​

  • New module backend/ai/secrets.py provides a process-local Keystore:
    • set(name, value) — store in memory only.
    • get(name) — return the value, intended only for handler runtime.
    • redacted(name) — return "***" for previews.
    • list_keys() — return key names only.
  • New node read_env (in core/nodes/env_io.py):
    • Field: path (file picker, defaults to .env in the project root).
    • Parses the file with python-dotenv (dotenv.dotenv_values(path)) — preferred over hand-rolled parsing so export FOO=bar, comments, quoted values, and multi-line keys all behave as developers expect.
    • Sets each key/value into the keystore and emits a list of key names. The values never enter the graph's data dict.
    • Preview shows KEY = *** rows.
    • Canonical use: a .env containing OPENAI_API_KEY=sk-..., ANTHROPIC_API_KEY=..., GEMINI_API_KEY=... is read once; downstream secret_get("OPENAI_API_KEY") flows the value into whichever provider node consumes it.
  • The chat assistant's own provider clients read keys via keystore.get("<PROVIDER>_API_KEY") first and fall back to os.environ only if the key is unset — so a developer running from source with their shell env exported still works without authoring a read_env node, but a packaged install reads from the user's .env through the keystore path.
  • Companion node secret_get(name):
    • One field: name.
    • At runtime, returns keystore.get(name). Output flows into downstream nodes as a normal string but is filtered in preview HTML and node logs.
  • parse_drawflow and the save path already only touch data — since the values aren't in data, save is automatically safe.
  • export.py (Python codegen) emits os.environ.get("KEY") rather than baking the literal — even when the value is in memory at export time.

Threat model now (Electron)​

  • DevTools is the renderer process; the keystore lives in Python so values aren't reachable from window.*.
  • The read_env node uses Electron's existing file picker (already scoped) to read the .env, then hands the bytes to Python over the existing IPC.
  • Worker subprocess sees secrets only via keystore.get; it should not be able to enumerate the keystore wholesale. Add a KEYSTORE_DENY_LIST_KEYS flag for paranoid users.

Threat model later (cloud)​

If EdgeWeave goes multi-tenant in the browser:

  • The whole Keystore must move to a per-session server-side store that the browser holds only an opaque session ID for. No values ever transit to the renderer.
  • The secret_get node returns a handle ("secret://uuid") on the client, the worker resolves it server-side at execution time.
  • All console.log / WebSocket payloads pass through a redactor that scans for any in-memory secret values and replaces them with ***. Cheap to implement, catches accidental log leaks.
  • .env upload is opt-in and one-shot — files never persist on disk.

The first two bullets above (handle indirection + redactor) are the ones worth prototyping early, even on Electron, because they're the ones that bite the hardest if retrofitted. The handle pattern also naturally lets the read_env node display only key names in previews without any special-casing.

Phasing​

  1. M1 — Tabbed UI shell. (shipped) Tab system, empty chat pane, no backend. Zero risk; lets us iterate on layout / theming.
  2. M2 — One-shot OpenAI chat. (shipped) POST /api/chat/message, gpt-4o-mini, no streaming, no tools. Provider abstraction (backend/ai/providers/{base,registry,openai_provider}.py) landed alongside the OpenAI implementation so M4 can register Anthropic / Gemini without touching the router. LangGraph deliberately deferred to M3 — for one-shot chat with no tools or routing, wrapping the single provider call in a StateGraph adds an import dep and a factory file but doesn't constrain the M3 design. The .env is loaded at backend startup via dotenv.load_dotenv(), so OPENAI_API_KEY from the project-root .env works without a UI step.
  3. M2.5 — Markdown rendering + multi-language code blocks. (shipped) Assistant replies now flow through marked → DOMPurify (strict allow-list) into the bubble, and every fenced ```lang block is post-processed into a read-only CodeMirror view with a hover "Copy" button. Language dispatch is centralised in frontend/src/js/code_languages.js, which also feeds codemirror_main.js so opening a .json / .yaml / .md / .js/.ts file in the editor pane gets the correct language extension instead of the Python parser. EdgeWeave is still Python-first; an unknown extension falls back to python() to preserve existing behaviour. An "Insert into editor" button is deferred to M6 (it needs the insert_code tool/IPC plumbing anyway).
  4. M3 — Streaming. (shipped — streaming only; LangGraph still deferred until there are tools/routing for it to orchestrate.) POST /api/chat/stream returns newline-delimited JSON (delta … done | error) instead of a WebSocket: the frontend is a fetch() + stream reader (frontend/src/js/chat_providers.js), and failures before the first byte are still ordinary HTTP errors. The provider's blocking stream runs in a worker thread with a stop flag; the async relay polls request.is_disconnected(), so the panel's Stop button closes the upstream request. (Handing Starlette the sync generator directly doesn't do this — uvicorn silently drops writes after a disconnect, so the generator ran to completion and a single-slot server like Ollama kept generating the abandoned reply while the next request queued behind it.)
  5. M4 — Providers + model picker. (shipped; key entry still via .env — a Settings screen is a later step.) Five providers in backend/ai/providers/: openai, gemini, ollama and openai_compatible share openai_compat.py (Gemini, Ollama and any Groq/Mistral/OpenRouter/LM Studio-style server speak the Chat Completions format, so no extra SDKs), and anthropic uses the official SDK (anthropic_provider.py). Each provider has a class-level status() (key present / Ollama reachable — no network to the vendor) and list_models() (live from the vendor's models endpoint, filtered to chat models, with a built-in fallback list); the registry caches model lists for 10 minutes. The panel shows a provider dropdown and a model dropdown (plus "Custom model id…"), remembered per provider in localStorage; unconfigured providers show their setup hint and disable Send. refresh_env() re-reads .env before each request, so adding a key needs no restart. Provider quirks handled: OpenAI reasoning models reject temperature (retried once without it, then remembered); current Claude models reject sampling parameters (never sent), need max_tokens (taken from the Models API's published output cap, capped at 64k) and opt into server-side refusal fallbacks (fallbacks: "default") where supported.
  6. M5 — Knowledge base (graph-RAG). (shipped instead of the prompt-pack router — retrieval picks context per question, which is what the packs were for, without hand-maintained packs.) backend/ai/knowledge.py indexes three kinds of card: one per registered node, built live from NODE_REGISTRY (label, registered name, category, doc, typed ports, fields, export support — accurate by construction thanks to the metadata audit, and including hot-reloaded user blocks); one per example .weave in the demo project and the active project (sticky-note text, nodes used, wiring); and the markdown docs split by section (config.DOCS_DIR; docs/docs/features-overview.md was written as the user-facing "how do I…" page). Links form a small graph: a demo links to the nodes it uses, a doc section to the nodes it names. Search scores every card, boosts nodes the question names explicitly (registered name, multi-word label, or "the Retrieve node" — entity linking, added after TF-IDF ranked a node's own card below docs sharing the phrase "inputs and fields"), then follows links: a node hit brings a demo that uses it, a demo/doc hit brings the nodes it mentions. Scoring reuses the AI nodes' toolkit (backend/ai/toolkit.py): TF-IDF always, plus a hybrid with Ollama's nomic-embed-text when it is pulled (built in a background thread). The index rebuilds when the registry or demo/doc files change and is warmed when the chat panel loads. /api/chat/stream inserts the hits as a system message before the latest user turn (short follow-ups also use the previous question) and sends a sources event first; the panel shows them under the reply and has an "EdgeWeave" toggle (on by default) to turn grounding off. GET /api/chat/knowledge?q= shows what would be retrieved. Measured on gpt-4.1-mini: asked which nodes build a US-states choropleth from a CSV, it named csv_read + plotly_choropleth (Location mode "USA-states") and maps/plotly_choropleth_demo.weave; with grounding off it invented "GeoJSON Reader", "Join" and "Color Scale" nodes. Cost: ~1.5–2k extra input tokens per question.
  7. M6 — Tools. list_nodes, read_file, insert_code.
  8. M7 — Secrets node (read_env, secret_get) and the keystore. Land the redactor and handle pattern from day one.
  9. M8 — Inline ghost completions. Optional — gauge after M6.

Each milestone is independently shippable to dev and useful on its own, so we can pause anywhere.

Decisions logged​

  • Conversation persistence — localStorage + summarise (2026-10-04). (Supersedes "none for now", 2026-04-27.) Chats are saved in the renderer's localStorage (frontend/src/js/chat_history.js): up to 30 conversations, capped at ~2 MB total (oldest dropped first), never written to a .weave, the backend or disk files we manage. The chat toolbar has a saved-chat dropdown, New, Summarise and Delete; each conversation remembers its provider/model. This is deliberately per-machine and unsynced — moving it server-side is a later decision (it brings back the redaction question). Summarise / compact: once a chat reaches 20 user+assistant turns or ~40k characters (~10k tokens), a banner suggests continuing from a summary, because every turn resends the whole history (cost grows with length) and models lose track of early details. Summarise asks the selected model for a structured markdown summary (Goal / Key facts & decisions / Code / Open questions) and starts a new chat seeded with it as a system turn (kind: "summary", shown as a collapsible bubble); the original chat is kept. This is the "markdown summary seeds the next session" idea from the earlier decision, done client-side and provider-agnostic. (Anthropic's server-side compaction isn't used: it only applies to Claude and to contexts far larger than a chat panel reaches.)

Open questions​

  • Cost telemetry. Show running token/cost counter in the chat header? Useful for users with their own keys; mandatory for any future hosted plan.
  • Tool approval UX. Auto-approve read_file, prompt for insert_code, never auto-approve anything that touches .env. This is the same model Claude Code uses and it works.
  • Local model story. Ollama is easy via the abstraction; do we ship a "local-only" mode that disables network egress for the paranoid?
  • Chat in the graph editor too? The right panel currently exists only on code_editor.html. Worth mirroring on index.html (graph editor) so users can ask "what does this node do?" while viewing the graph. The same chat_ui.js should drop into both.

Appendix — small unrelated heads-up​

While reading sdk/decorators.py I noticed that register_block explicitly whitelists meta keys when populating NODE_REGISTRY, and the round-2 resizable flag isn't in the whitelist — so it never reaches the frontend via `/api