L1.12
Latency and Streaming Experience
Cards in this group · 5
- L1.12.1Generation latency usually sits in the seconds people can feel waiting but will not yet abandon
- L1.12.2Streaming shortens time-to-first-token, not total duration; it changes perception, not fact
- L1.12.3Token-by-token invites users to start judging before the result is complete, so decisions rest on a half-product
- L1.12.4Mid-stream self-correction lets users see intermediate content that will later be overturned
- L1.12.5Long background jobs should not stream; they need progress and permission to leave