Synthesize natural spoken speech from text using high-quality system voices with speed and pitch controls.