What you’ll build
- A chat UI with a conversation list and a message input.
- Streaming responses (tokens appear as the model generates them).
- Persistent history (refresh the page; conversations are still there).
- Tool calling — the bot can search your tasks, create a task, or send an email.
- Provider switching (OpenAI vs. Anthropic) without code changes.
Prerequisites
Step 1 — Project foundation
Use Plan mode for the first prompt. Submit. Review the plan. Approve. vly will prompt forOPENAI_API_KEY. Paste it.
Build takes ~90 seconds.
Step 2 — Test the basic flow
- Open the preview, sign in.
- Click New conversation.
- Type “What’s the capital of France?” and submit.
- Watch tokens stream in. The message should populate live, not wait for the full response.
Step 3 — Add tool calling
Now make the bot take actions. We’ll let it search and create tasks. If you don’t have atasks table from a prior recipe, vly will create one (it’s a small ask).
After build:
- “Find my tasks about the Q2 report” — bot calls
search_tasks, returns results. - “Make a task to follow up next Tuesday” — bot calls
create_task, confirms.
send_email, update_task, etc.) by extending the same pattern.
Step 4 — Provider switching
Let users pick between OpenAI and Anthropic models. vly may need an Anthropic API key — paste it when prompted. After build, switching the dropdown changes which provider runs the next message.Step 5 — Polish
Common refinements:- Markdown rendering
- Copy code blocks
- Regenerate
- Search past conversations
- Token / cost tracking
Going further
Multi-modal (images)
Allow image attachments. Use GPT-4o or Claude vision. Convex file storage for the upload.
Voice in/out
Combine with the voice AI recipe — push-to-talk in, TTS out.
RAG over your docs
Embed your docs with vector search. Add a
search_docs tool that returns relevant chunks.Memory across conversations
Summarize old conversations and prepend the summary as context to new ones.
Common pitfalls
Tokens not streaming smoothly
Tokens not streaming smoothly
The action is buffering. Make sure the streaming loop writes to the database on every chunk (not at the end). Convex’s reactive sync handles the client side.
Tool calls timeout
Tool calls timeout
Convex actions have a 10-min timeout. Long tool chains shouldn’t hit this, but if they do, split into multiple actions and use scheduled functions.
Model returns invalid JSON in tool calls
Model returns invalid JSON in tool calls
OpenAI’s structured output mode (
response_format: { type: "json_object" }) helps. For tool calling, the SDK handles parsing — just catch errors and re-prompt the model.Costs ballooning
Costs ballooning
Default to
gpt-4o-mini for most conversations. Cap conversation history sent to the model. Cache repeated tool results.Related
OpenAI integration
The full OpenAI integration reference.
Anthropic integration
Same surface, Claude models.
Voice AI app
The voice variant.
