Skip to content

RFC needed: Voice/speech output (speak-on-space, speak control nodes, per-platform TTS) #42

Description

@willwade

From the governance test checklist gap analysis (governance#41). No RFC exists for voice/speech output.

What's implemented (no spec)

  • Speak-on-space: the engine detects word boundaries in the output buffer and speaks the completed word via the platform TTS
  • Speak control node: a control-mode node that speaks the current buffer
  • Per-platform TTS: GTK uses speech-dispatcher (via rust-tts-wrapper), Windows uses SAPI, Android uses Android TTS, Apple uses AVSpeechSynthesizer
  • Speech volume/rate controls in Settings on some platforms

What needs specifying

  1. When does speech fire? (word boundary in the engine? configurable?)
  2. Can speech be interrupted? (new word supersedes the old? explicit stop?)
  3. Voice selection (system default? user-configurable per-platform?)
  4. Rate/volume/pitch controls (engine-managed or platform-managed?)
  5. What happens in direct mode? (speak the target's text or Dasher's buffer?)
  6. Integration with communication boards / phrase storage (if/when that exists)

Why this matters

Core AAC: users who can't see the screen (or can't look at it while typing) rely on auditory feedback. This is the voice of the product.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions