Tutorial: chat with the assistant, then build RAG in a graph
EdgeWeave has two AI features:
- the Chat assistant in the Code Editor, which knows EdgeWeave's nodes, example workflows and these docs;
- the AI / LLM and AI / RAG nodes, which put language models and retrieval inside your workflows.
Both work with Ollama, which is free, local and needs no key, or with OpenAI, Anthropic Claude, Google Gemini, or any OpenAI-compatible server.
1. Choose a provider
Option A: Ollama (local, no key)
-
Install Ollama and pull a small model:
ollama pull gemma3:4b -
Leave Ollama running. EdgeWeave finds it at
http://localhost:11434. SetOLLAMA_HOSTin.envif yours runs elsewhere.
Short on disk space? Ollama can keep its models on another drive. Set a
user environment variable OLLAMA_MODELS to a folder there, for example
E:\ollama\models, then quit Ollama from the system tray and start it again.
New pulls go to that folder. To keep models you already have, copy
%USERPROFILE%\.ollama\models there first.
Option B: a cloud provider
Add your key to the .env file. In the installed app that's
%APPDATA%\EdgeWeave\.env. The chat panel's hint shows the exact path.
OPENAI_API_KEY=sk-...
# ANTHROPIC_API_KEY=...
# GEMINI_API_KEY=...
EdgeWeave rereads the file before each message, so you don't need to restart. Keys are
never saved in .weave files or exports.
2. Ask the assistant
-
Open Code Editor and pick the Chat tab.
-
Choose a Provider and Model. Providers without a key show as not set up, with a hint saying what to add.
-
Keep 📚 EdgeWeave ticked. The assistant then looks up relevant nodes, demos and docs for each question and lists them under EdgeWeave sources.
-
Ask something concrete, for example:
Which nodes build a choropleth map of US states from a CSV?
Other things to know about the chat panel:
- Stop a long reply at any time.
- Chats are saved on your computer (Saved chats). When a chat gets long, Summarise into new chat carries the gist into a fresh one, which keeps replies cheaper and sharper.
- Paste an error message or a traceback and ask what's wrong.
3. Run the RAG demo
Open ai/rag_demo.weave. It answers a question using a folder of Markdown
notes:
Load Documents → Split Text → Build Vector Index → Retrieve → Format Context ─┐
▲ ▼
question ─────────▶ Prompt Template → LLM Chat
- Click Run. Everything up to Prompt Template works offline: Load Documents should find 5 files, and Retrieve shows the best-matching chunks.
- On LLM Chat, pick your provider (
ollamafor a local model) and optionally a model. Blank means the provider's default. Run again: the answer cites its sources as[1],[2], and so on. - Change the question in the String node at the left and re-run.
Each run of LLM Chat makes a real model call, which may be billed on a cloud provider.
Make retrieval understand meaning
Build Vector Index defaults to tfidf, which matches shared words. For
matching by meaning, pull an embedding model and switch the method to
ollama (or auto, which uses Ollama when the model is there):
ollama pull nomic-embed-text
openai uses OpenAI's embeddings instead. Chat-only models such as gemma3
can't produce embeddings; the node tells you if you pick one.
Use your own documents
Point Load Documents at your own folder and change its pattern, for example
**/*.md, **/*.txt. Then tune Split Text (by paragraphs, Markdown sections,
or fixed-size chunks with overlap) and Retrieve's top_k.
4. Export it
Export to Python works for the whole RAG graph. The script contains the same
functions the nodes run and reads keys from the environment, so set
OPENAI_API_KEY (or use Ollama) before running it.
Next: Write your own node.