L1.12

Latency and Streaming Experience

Cards in this group · 5

  1. L1.12.1Generation latency usually sits in the seconds people can feel waiting but will not yet abandon
  2. L1.12.2Streaming shortens time-to-first-token, not total duration; it changes perception, not fact
  3. L1.12.3Token-by-token invites users to start judging before the result is complete, so decisions rest on a half-product
  4. L1.12.4Mid-stream self-correction lets users see intermediate content that will later be overturned
  5. L1.12.5Long background jobs should not stream; they need progress and permission to leave