Switch your AI calls to streaming so replies appear token by token, cutting perceived wait time from seconds to under one.
My chatbot makes people wait 10 seconds. Make the answer stream in word by word
How it works
- Measure the current wait: Vovy sends a test message and times it: seconds to the first word versus seconds to the full reply. Streaming fixes the first number.
- Turn on streaming server-side: Vovy asks Claude Code to set stream to true on your API call and pass chunks through as server-sent events, a way to push text as it arrives.
- Render chunks in the UI: Claude Code updates your chat component to append text as each chunk lands, with a stop button and a typing indicator.
- Check your host supports it: Vovy checks your function timeout and that no layer buffers the response, a common reason streaming works locally but not on Vercel.
- Show the difference: A card compares time to first word before and after, with a short screen recording of the new feel.
What you provide
- Your app's chat or AI feature
- Access to your code and host
What you get
- Streaming AI replies
- A stop button and typing state
- Before and after timing
FAQ
Does streaming cost more?
No. You pay for the same tokens either way. It only changes how fast the user sees them.
Why does it stream locally but not in production?
Something between your server and the browser is buffering, or the function times out. Vovy checks both.
Can I stream JSON or structured output?
You can, but partial JSON is not valid until it finishes. Stream text for chat, and wait for the whole result for structured data.
Related tasks
All tasks