Important
This is 1 of 2 repos — you need both halves. This repo is the backend add-on (the voice "brain"). It needs the custom Voice PE firmware to connect to it — the stock Home Assistant voice pipeline won't talk to this add-on. You must set up both:
- 🧠 Backend add-on (this repo) — runs inside Home Assistant
- 🔌 Device firmware → xandervanerven/home-assistant-voice-pe (flashed onto the Voice PE)
📖 New here? The full INSTALL guide walks through both halves, step by step.
A Home Assistant add-on that turns a Voice PE
device into a low-latency voice assistant built on OpenAI's Realtime API
(gpt-realtime-2). The device streams microphone audio to this add-on over a
plain WebSocket; the add-on runs the Realtime speech-to-speech session and
controls Home Assistant through the official
Home Assistant MCP Server
integration. STT, TTS and the LLM all run in the Realtime session — there is no
Home Assistant voice_assistant pipeline on the audio path.
Fork of fjfricke/ha-openai-realtime, retargeted at
gpt-realtime-2, the official HA MCP Server, optional web search, and the Voice PE thin-client firmware (a separate repo — see below).
openai_realtime_voice_agent/— the Home Assistant add-on (Python / Pipecat). This is the only thing you install.DOCS.md— full setup: OpenAI key, the Home Assistant MCP connection, recommended settings, web search, all options.CHANGELOG.md— what changed per version.
The device firmware lives in its own repository —
xandervanerven/home-assistant-voice-pe
(a custom va_client ESPHome component, specific to the Voice PE hardware).
- In Home Assistant, open Settings → Add-ons → Add-on store → ⋮ → Repositories
and add
https://github.com/xandervanerven/ha-openai-realtime. - Install OpenAI Realtime 2 Voice Agent. It ships with no prebuilt
image:, so Home Assistant builds it locally on first install (a few minutes on a Pi). - Configure the add-on and flash the companion firmware — see
openai_realtime_voice_agent/DOCS.md.
(An optional GitHub Actions workflow can publish container images to ghcr.io; it isn't needed for a normal local-build install.)
Voice PE (ESP32-S3) ──WS, 16 kHz PCM up──▶ this add-on ──▶ OpenAI Realtime API
va_client firmware ◀──── 24 kHz PCM down── (Pipecat) (gpt-realtime-2)
│ tools
▼
Home Assistant MCP Server
The device does wake-word detection and XMOS audio cleanup locally and is a thin client. Interrupt a reply with the "stop" word or the center button.
- No voice timers or alarms yet — every other Assist action (lights, switches, scenes, climate) and online questions work.
- A brief reconnect about once an hour (OpenAI's 60-minute session cap; the add-on refreshes proactively during a quiet moment, so it rarely interrupts).
- Rarely, the assistant may stop itself on a word in its own reply that sounds like "stop" — just ask again.
- Forked from fjfricke/ha-openai-realtime.
- Built on Pipecat.
- Firmware thin-client design based on maxmaxme/home-assistant-voice-pe (a fork of esphome/home-assistant-voice-pe, Nabu Casa / ESPHome).
- Inspiration from marcinnowak79/home-assistant-voice-pe (gemini-live-proxy).
MIT — see LICENSE.