Add a record button that captures audio, sends it to a transcription model through your server, and drops the text into your input.
Add a mic button so users can speak instead of typing
How it works
- Add a record button: Vovy asks Claude Code to add a mic button that asks for microphone permission and records audio in the browser.
- Send audio to your server: Claude Code uploads the recording to a server function, keeping your API key off the browser.
- Transcribe it: The server function calls OpenAI's transcription endpoint with the audio file and returns plain text, usually within a couple of seconds.
- Handle the rough edges: Claude Code adds a max recording length, a clear error if permission is denied, and a note that Safari needs HTTPS for the mic.
- Test and price it: Vovy records a test phrase, shows the transcript, and puts the cost per minute of audio on a card.
What you provide
- An OpenAI API key
- A backend function
- Your app's input screen
What you get
- A working mic button
- Server-side transcription
- A cost per minute estimate
FAQ
Why not use the browser's built-in speech recognition?
It is free but support and accuracy vary by browser. A transcription API is consistent everywhere.
What does transcription cost?
Priced per minute of audio, typically well under a cent per short voice message.
Does it work in other languages?
Yes. Transcription models handle dozens of languages and usually detect the language on their own.
Related tasks
All tasks