Differential Expression from Counts: Dispersion, Fit and Shrinkage · seed 1 · A4, ink-friendly. The answer key prints on its own page for grown-ups.

Reading change out of RNA count tables

Science · Genetics & Evolution · ages 22-24
Name ______________________   Date ____________
  1. What does size factor normalisation correct?

    • The GC content of the reference genome
    • Differences in sequencing depth between samples
    • Batch effects from different lab coats
  2. Why is a normal model wrong for read counts?

    • It allows impossible negative values and misses the mean variance link
    • It always reports exactly zero for every gene
    • It cannot handle numbers larger than one hundred
  3. A tenfold raw ratio built from thirty reads versus three is strong evidence of change.

    Circle one:   True   False

  4. What is dispersion in this analysis?

    • The per gene extra variance beyond Poisson sampling, estimated across similar genes
    • The number of lanes used on the sequencer
    • The length of the longest transcript in the sample
  5. One condition switches on a few extremely abundant transcripts. What happens to standard scaling?

    • It works better because totals grow
    • It breaks, since a few transcripts distort the shares of the rest
    • It converts counts into exact measurements
  6. After shrinkage, a low count gene moves from eightfold to near zero. How do you read it?

    • As a confirmed eightfold discovery
    • As proof the pipeline deleted real biology
    • As no trustworthy evidence of change from weak data
  7. Diagnostic plots show count distributions warped after scaling. What do you do?

    • Publish quickly before anyone notices
    • Double all counts and rerun without scaling
    • Suspect composition effects and distrust comparisons until the failure is fixed
  8. A colleague reports only raw fold changes, topped by a 3-versus-0 gene. What is the flaw?

    • Raw ratios ignore depth, variance, and uncertainty that shrinkage would expose
    • Raw ratios are always smaller than shrunken ones
    • Negative binomial models forbid ranking genes
LightMySky · lightmysky.comW1-mt_nEV9547kSN-s1

Answer key

For grown-ups. Fold this page away before handing over the rest.

Reading change out of RNA count tables W1-mt_nEV9547kSN-s1

  1. Differences in sequencing depth between samples · Deeper samples show higher counts for every gene, so you must scale first.
  2. It allows impossible negative values and misses the mean variance link · Counts are discrete and noisy, so they need a model built for count data.
  3. False · Tiny counts carry huge uncertainty, so the raw ratio is noise dressed as discovery.
  4. The per gene extra variance beyond Poisson sampling, estimated across similar genes · You borrow strength across genes of similar strength to estimate it.
  5. It breaks, since a few transcripts distort the shares of the rest · Size factors assume most genes stay constant, which massive shifts violate.
  6. As no trustworthy evidence of change from weak data · Shrinkage pulls unstable estimates toward zero, and weak genes move the most.
  7. Suspect composition effects and distrust comparisons until the failure is fixed · Distribution plots catch scaling failure before it corrupts your results.
  8. Raw ratios ignore depth, variance, and uncertainty that shrinkage would expose · The published number should state conservatively what the data actually support.
Worksheet · LightMySky