Voice Input
Voice Input skills allow AI Agents and AI Flows to accept spoken input and convert it into text for further processing. This enables hands-free interactions and voice-driven automation.
Available Skills
Section titled “Available Skills”Google Speech-to-Text (V2)
Section titled “Google Speech-to-Text (V2)”Convert spoken audio into text using Google Cloud Speech-to-Text (V2).
The skill supports accurate speech recognition across multiple languages and can be integrated into AI Agents and workflows to process voice input in real time or from audio recordings.
Common use cases
- Voice-enabled AI assistants
- Transcribe meetings and interviews
- Process voice commands
- Convert audio recordings into text
- Build voice-driven automation workflows
Typical Workflows
Section titled “Typical Workflows”Voice Input skills can be combined with other built-in skills to automate end-to-end voice processing.
Examples include:
- Transcribe a meeting and email the summary to participants.
- Convert customer voice messages into support tickets.
- Process voice commands and execute infrastructure tasks.
- Transcribe audio files and store them in a knowledge base for semantic search.
- Generate AI summaries from recorded conversations.
