Prerequisites
Setup
1
Get your OpenAI API key
- Sign in at platform.openai.com.
- Click API keys → Create new secret key.
- Name it
vly-yourapp(or similar). - Copy the key (starts with
sk-...). You won’t see it again.
2
Connect OpenAI in vly
In a prompt:vly opens the integration setup dialog. Paste your key and confirm. The key is stored encrypted as
OPENAI_API_KEY.3
Use it in a feature
Configuration
string
required
Your OpenAI API key. Server-only. Starts with
sk-....string
Optional. Your OpenAI organization ID, if you want to scope usage to a specific org.
string
default:"gpt-4o-mini"
Default model for generic AI features. Override per-call. Common values:
gpt-4o, gpt-4o-mini, gpt-4-turbo.Code patterns
Chat completion
convex/ai.ts
Streaming chat
For chatbot UIs, stream tokens to the client as they arrive:convex/ai.ts
useQuery(api.chats.get, { id }) and sees tokens stream in real time.
Embeddings (for vector search)
convex/ai.ts
Image generation (DALL-E)
convex/ai.ts
Audio (Whisper STT)
convex/ai.ts
Tool / function calling
For agents that can call your own functions:Cost management
OpenAI bills per token. Control costs by:Use the smallest model that works
gpt-4o-mini is 15× cheaper than gpt-4o and good enough for most summarization, extraction, and simple chat tasks.Cap tokens
Set
max_tokens on every call. Prevents runaway responses and runaway bills.Cache responses
For deterministic prompts (summarize the same text), cache the result in a Convex table keyed by content hash.
Set a usage limit in OpenAI
OpenAI dashboard → Settings → Limits. Set a hard monthly cap; calls fail when exceeded (better than a surprise bill).
Choosing a model
Comparison with other providers
For chatbots, Anthropic’s Claude is often a better choice for nuanced conversations. For everything else, OpenAI is the default.
Troubleshooting
'You exceeded your current quota'
'You exceeded your current quota'
Your OpenAI account has no credit, or you hit your usage limit. Add credit at platform.openai.com/billing.
'The model gpt-4o does not exist or you do not have access'
'The model gpt-4o does not exist or you do not have access'
Some models require tier-1+ access. Check Tier requirements. Most accounts auto-graduate after $5 of usage.
Streaming response 'cuts off'
Streaming response 'cuts off'
max_tokens is set too low, or the model hit its context window. Increase max_tokens, or for very long inputs, use a model with a larger context window.Slow response time
Slow response time
Some prompts are slow regardless of model. Add a loading state in the UI; for chatbots, switch to streaming so users see partial output.
Related
AI chatbot recipe
Complete walkthrough — streaming, history, tool calls.
Image generator recipe
User-facing image generation with DALL-E.
Vector search
Semantic search with OpenAI embeddings + Convex.
