Cerebras Systems and Hugging Face launched an open-source, full speech-to-speech AI pipeline on July 1, 2026. The system integrates Nvidia’s Parakeet for speech recognition and Google DeepMind’s Gemma 4 31B for reasoning. Alibaba’s Qwen3-TTS provides the voice output.
The pipeline utilizes Cerebras wafer-scale chips to minimize conversational response delays. Google’s Gemma 4 31B model runs approximately 35 times faster on this hardware compared to typical GPU endpoints. This performance increase aims to eliminate multi-second stalls in AI interactions.
Over 9,000 Reachy Mini robots currently use the pipeline in production environments. This deployment confirms the system's viability for real-time conversational AI applications.