STT (Whisper) + LLM + TTS (ElevenLabs). Push-to-talk in, AI voice out. ~60 minutes.
A voice interface: user speaks; Whisper transcribes; an LLM responds; ElevenLabs speaks the answer.Time: ~60 min. Difficulty: Advanced. Integrations:OpenAI (Whisper, GPT-4o), ElevenLabs.
A voice AI assistant. Push-to-talk button: hold to record, release to send.Whisper transcribes the audio. The transcription goes to GPT-4o; the responsestreams back. ElevenLabs converts the response to speech and plays it.Conversation history preserved per user. Show transcripts of both user and AI.