Voice Input in the AI Chat

What is this for?

Voice input lets you dictate messages to the AI Assistant instead of typing — handy when your hands are busy, the text is long, or you simply want to phrase your request in natural spoken language. The system automatically recognizes your speech, converts it to text, and places it in the input field. The message is not sent automatically — you review the text and click "Send" yourself.

What It Looks Like in the Interface

In the message input panel (below the chat history), on the left side of the text field there is a microphone button (🎤 icon).

  • Gray microphone — recording is not active. Tap it to start.
  • Red microphone (Stop square) with pulsing animation — voice is being recorded. Tap it again to stop early.
  • A "Voice input" tooltip may appear next to the button when you hover over it.

Step-by-Step: How to Use Voice Input

  1. Open the chat — click the green hexagonal button in the bottom-right corner.
  2. Tap the microphone on the left side of the input field.
  3. Allow microphone access if your browser asks (this only happens the first time you use it).
  4. Speak clearly — the system is listening. You don't need to hold the button.
  5. Recording stops automatically after 1.5 seconds of silence (when you stop talking).
  6. Text appears in the input field — read it through.
  7. Fix any mistakes manually (with your keyboard) if needed.
  8. Tap the "Send" button (paper airplane) or press Enter to send your message to the AI.

Important: The message is not sent automatically after transcription. This is intentional — you stay in control of what goes to the AI.

Things Worth Knowing

  • Auto-stop: The system uses voice activity detection (VAD). If you go silent for 1.5 seconds, the recording stops on its own. No need to press Stop if you've finished your sentence.
  • Manual stop: If you want to stop recording earlier (for example, you changed your mind or are interrupting yourself) — tap the red microphone (Stop) again.
  • Maximum duration: There's a hard safety limit — about 60 seconds. If you speak without pauses longer than that, the recording stops automatically.
  • Language detection: The system automatically detects the language of your speech (Ukrainian, English, Spanish, and others). It works best if you speak one language throughout the entire recording.
  • Sound quality: For best results, speak close to the microphone in a quiet environment. Background noise, music, or other voices may reduce accuracy.
  • Browsers: Works in Chrome, Edge, Firefox, and Safari (modern versions). Requires HTTPS (or localhost).

Possible Errors and How to Fix Them

Situation What You See What to Do
Browser blocked microphone access The browser popup "Allow microphone access?" didn't appear, or you clicked "Block" Click the lock / microphone icon in the browser's address bar → allow microphone for this site → reload the page and try again.
"Could not start recording" (error) A red toast notification saying "Could not start recording" Check that your microphone is connected and not being used by another app (Zoom, Meet, etc.). Reload the page.
"No audio was recorded" A notification saying "No audio was recorded" You tapped Stop before saying anything. Tap the microphone again and speak.
"Could not recognize audio" A notification saying "Could not recognize audio" Try again: speak louder, more clearly, and closer to the microphone. Check your internet connection (recognition happens on the server).
Text was recognized with errors The text in the input field contains mistakes Fix them manually before sending. The system learns over time but isn't perfect.

Tips for Better Recognition

  • Speak in full sentences rather than single words — context helps the model.
  • Pause between thoughts — 1.5 seconds of silence works like a period.
  • Avoid background noise (air conditioning, music, colleagues chatting).
  • Don't shout or whisper — normal conversational volume works best.
  • If you switch languages mid-recording (for example, from Ukrainian to English), it's better to stop the recording, send the current message, and start a new recording for the other language.

Benefits of Voice Input in the AI Assistant

  • Speed: Dictating is 3–4 times faster than typing on a keyboard.
  • Natural phrasing: You say "Create a registration form for a webinar with fields for full name, work email, company name, required experience level dropdown: junior, middle, senior, team lead, and an optional dietary restrictions field", instead of figuring out how to shorten it.
  • Multitasking: You can speak while looking at the form editor, scrolling through the field list, and so on.
  • Accessibility: Helps users with limited hand mobility or vision.

Limitations

  • No auto-send — this is a feature, not a bug. You control the moment of sending.
  • Single-speaker only — if multiple people talk at the same time, quality drops.
  • Requires internet — recognition happens on the server (Whisper API). Does not work offline.
  • No file uploads — you cannot upload an audio file for transcription, only live microphone input.

Example Scenario

You: tap the microphone → "Create a registration form for a developer conference with fields: full name, work email, company name, experience level dropdown: junior, middle, senior, team lead, and an optional dietary restrictions field."

System: after 1.5 seconds of silence the recording stops, text appears in the input field

You: read through, fix "detary" to "dietary", tap "Send"

AI: "Done! The form has been created. Review the draft in the editor and click 'Apply'."


Remember: Voice input is a tool for creating your message text, not for controlling the interface. Commands like "close the chat" or "minimize the window" won't work by voice — use the buttons in the interface.