Observability and Tail Latency Across Services · seed 1 · A4, ink-friendly. The answer key prints on its own page for grown-ups.

Finding the slow service in a trace

Computing · Computer Systems · ages 23-24
Name ______________________   Date ____________
  1. What does a distributed trace follow?

    • Every request blended together
    • One request across service boundaries
    • Only the slowest service
  2. A trace of one request shows three spans, and one span holds most of the time. What does that tell you?

    • That service holds the latency
    • The request never reached that service
    • The trace mixed up several requests
  3. The mean latency always reveals a slow tail.

    Circle one:   True   False

  4. A request fans out to three services and waits for every reply. The replies take 20, 30, and 400 ms. About how slow is the whole request?

    • About 400 ms
    • About 30 ms
    • About 20 ms
  5. The mean looks fine but the ninety-ninth percentile is bad. What is happening?

    • Every request is slightly slow
    • All the measurements are broken
    • A few requests suffer while most are fast
  6. A slow request's trace shows one very long span among short ones. What do you do?

    • Blame the average of all the spans
    • Blame the service holding the long span
    • Blame the shortest span
  7. A fan-out request is slow, but every service is usually fast. Its trace shows all spans short except one long span. What happened?

    • That service was slow this time, and the request inherited the wait
    • The trace added extra time itself
    • Fast services always cause slow requests
  8. A teammate says the average proves no request is slow. What is wrong with that?

    • Averages are always computed wrong
    • Traces never show slow requests
    • The average can look fine while the tail suffers
LightMySky · lightmysky.comW1-mt_NT5tmumdfA-s1

Answer key

For grown-ups. Fold this page away before handing over the rest.

Finding the slow service in a trace W1-mt_NT5tmumdfA-s1

  1. One request across service boundaries · A trace sticks to one request as it crosses each process boundary.
  2. That service holds the latency · The longest span is where the request spent its time.
  3. False · Fast requests drown out the slow ones, so the mean hides the tail.
  4. About 400 ms · Waiting on all means waiting on the slowest.
  5. A few requests suffer while most are fast · A bad tail with a fine mean means a suffering minority.
  6. Blame the service holding the long span · Line the spans up and blame the longest one.
  7. That service was slow this time, and the request inherited the wait · Fan-out inherits the slowest wait, even from a usually fast service.
  8. The average can look fine while the tail suffers · The average cannot clear the tail; only a percentile can.
Worksheet · LightMySky