Models of Sequence Evolution and Maximum Likelihood Trees
Estimating a tree by likelihood needs an explicit model of how sites change: a rate matrix, an account of rate variation across sites, and a way to compare candidate models. This stop connects the model chosen to the tree that comes out.
What a learner can do afterwards
- Compare Jukes-Cantor, HKY and the general time-reversible model by the parameters each adds
- Explain what a gamma rate parameter corrects for and what goes wrong when it is left out
- Interpret the likelihood of a tree given an alignment and a stated model
1 · Read
Estimating a tree by likelihood needs an explicit model of how sites change. The simplest model treats all changes as equally likely with equal base frequencies. HKY adds realistic unequal base composition plus a bias toward transitions. The general time reversible model goes furthest, allowing each pair of nucleotides its own rate alongside unequal base frequencies. Each step adds parameters that correct hidden changes, where letters flip back and forth and distant sequences match by chance.
Suppose two fast evolving lineages pile up many overlapping changes. A model without rate variation averages one speed across all sites and underestimates those long branches. A gamma rate parameter fixes this by letting sites evolve at different speeds, which you already met as position-specific thinking in profile models. Leave gamma out and the method can group the two fast lineages together incorrectly, fooled by their coincidental matches.
The likelihood of a tree given an alignment and a model is the probability of the observed alignment under that tree and model, computed by summing over every possible ancestral state at the internal nodes. Maximum likelihood then asks which candidate tree makes the alignment most probable. Because every topology score runs through the model, changing the model changes branch lengths and can crown a different winner.
To defend a published tree, justify its substitution model and show the result holds under richer ones. Read trees from shared nodes, never from the order of names at the tips, and remember rooting choices tell their own stories. A poor model can favour the wrong topology with high apparent confidence, so checking richer models is part of honest inference.
Models price hidden changes, gamma spreads speeds across sites, and likelihood crowns the tree that best explains the alignment.
2 · Watch
Take it off screen
Where it sits
8 questions wait behind this lesson, each with its answer explained. Every answer feeds the sky: stars light as they are learned, and dim when it is time to come back.