From the governance test checklist gap analysis (governance#41). No RFC exists for voice/speech output.
What's implemented (no spec)
- Speak-on-space: the engine detects word boundaries in the output buffer and speaks the completed word via the platform TTS
- Speak control node: a control-mode node that speaks the current buffer
- Per-platform TTS: GTK uses speech-dispatcher (via rust-tts-wrapper), Windows uses SAPI, Android uses Android TTS, Apple uses AVSpeechSynthesizer
- Speech volume/rate controls in Settings on some platforms
What needs specifying
- When does speech fire? (word boundary in the engine? configurable?)
- Can speech be interrupted? (new word supersedes the old? explicit stop?)
- Voice selection (system default? user-configurable per-platform?)
- Rate/volume/pitch controls (engine-managed or platform-managed?)
- What happens in direct mode? (speak the target's text or Dasher's buffer?)
- Integration with communication boards / phrase storage (if/when that exists)
Why this matters
Core AAC: users who can't see the screen (or can't look at it while typing) rely on auditory feedback. This is the voice of the product.
From the governance test checklist gap analysis (governance#41). No RFC exists for voice/speech output.
What's implemented (no spec)
What needs specifying
Why this matters
Core AAC: users who can't see the screen (or can't look at it while typing) rely on auditory feedback. This is the voice of the product.