Voice speed is part of each voice’s sound: every line is voiced at a specific pace, so the time to decide it is before you generate. The speed panel in the Voices cast bar sets the whole cast at once: one pace for everyone, or the Varied preset, which spreads speeds evenly across the cast so each voice reads a little differently (new projects start on the varied spread, 1.25× to 2×). The samples in the panel are free: play any pace and hear exactly how fast it really is before anything is billed.
Because the spoken pace is baked into the audio, changing it after voices exist re-voices them, and the panel shows what that would cost before you apply anything. Speeds go up to 3×, and there are no per-voice walls: where a voice's native range ends, the studio time-stretches the audio: same pitch, faster speech, the way reader apps play text at 3× and still sound natural. The panel names the voices that use the stretch, and that part of the pace is free to change anytime.
The pace axis in the panel is shaded by what listeners can consciously follow. Two independent lines of research agree on where that sits: comprehension of normal speech holds up to roughly 275 words per minute (about 1.8× of conversational speed), and a 2022 study of 231 students found comprehension essentially unchanged from 1× to 2×, dropping only at 2.5×. Past about 2.5× the fall is steep.
Your program is an easier case than those studies, which is why the shading advises instead of blocking. They measured novel, continuous sentences; you loop short, familiar lines with gaps between them, and inserted gaps alone were measured to cut errors at high speed dramatically. A subliminal also does not depend on you consciously following the words.
The per-voice Speed-up is the manual version of that same pitch-preserving stretch: it plays a finished recording faster at mix time, free to change as often as you like, even after synthesis.