1 of 45

CLADISTICS

Phylogenetic systematics

ODWS Paul Billiet 2011

2 of 45

The basic assumption

  • All life on Earth shares a common origin
  • Therefore, two different organisms will share a common ancestor

ODWS Paul Billiet 2011

3 of 45

Distant cousins

  • Merlin is clearly a cat and I am a human
  • We share a common ancestry that can be seen in our anatomy

ODWS Paul Billiet 2011

4 of 45

Vertebrates

  • Both Merlin and I have, a skull followed by a vertebral column, paired sense organs, a tail that continues on beyond the anus
  • All vertebrates have these, they must have a shared ancestor

Silky shark

Carcharhinus falciformis

ODWS Paul Billiet 2011

5 of 45

Tetrapods

  • Merlin and I both have jaws with teeth and two pairs of limbs
  • We share these features with a more select group of vertebrates called tetrapods

Common frog

Rana temporaria

ODWS Paul Billiet 2011

6 of 45

Amniotes

  • When we were embryos both Merlin and I were protected by membranes
  • One is called the amnion that is a feature of many terrestrial vertebrate animals

ODWS Paul Billiet 2011

7 of 45

Mammals

  • Both Merlin and I have: �hair, �we are endothermic, �we have jaws that connect to the skull in a particular way, �we suckled milk when were young, �we have a diaphragm between our thorax and our abdomen
  • We are mammals

ODWS Paul Billiet 2011

8 of 45

Eutherians

  • Merlin and I spent the early parts of our life in a womb supported by a placenta
  • We are eutherian mammals

ODWS Paul Billiet 2011

9 of 45

Merlin’s relatedness to me

ODWS Paul Billiet 2011

10 of 45

What we know and what we don’t know

  • We know that Merlin and I shared a common ancestor
  • We do not know:�when�where
  • We have some ideas on what it might have looked like
  • We do not know how we came to be the way we are

ODWS Paul Billiet 2011

11 of 45

Adding in another cousin

  • Soup is another cat-like animal
  • Soup shares more features with Merlin than I do

ODWS Paul Billiet 2011

12 of 45

An extended family: �Merlin, Soup and I

ODWS Paul Billiet 2011

13 of 45

An alternative view

  • There is more than one way we three could be related

ODWS Paul Billiet 2011

14 of 45

Cladograms and clades

  • These diagrams are called cladograms
  • Comes from the Greek word meaning a branch
  • Each branch point or node represents a common ancestor
  • The branches above a node represent a clade
  • All the organisms in a clade share a number of features

ODWS Paul Billiet 2011

15 of 45

Common sense v Science

  • These cladograms suggest that there may be different ways of obtaining the same result
  • Common sense would suggest that the first cladogram is the correct approach
  • Common sense is not objective
  • Common sense is not scientific

ODWS Paul Billiet 2011

16 of 45

Cladistics

  • Cladograms belong to a method of taxonomy called cladistics �(aka phylogenetic systematics)
  • Cladistics has become an accepted way of classifying organisms
  • It permits hypothesis of relatedness to be tested
  • It uses the the principle of Occum’s razor to decide which is the most plausible hypothesis

ODWS Paul Billiet 2011

17 of 45

Occum’s razor

  • Occum’s razor states that if there are two or more conflicting hypotheses to explain a phenomenon the simplest is chosen as the working hypothesis
  • This is called The Principle of Parsimony
  • This does not mean that it is the right hypothesis
  • It still needs to be tested
  • All hypotheses are provisional

ODWS Paul Billiet 2011

18 of 45

The most parsimonious route

  • The cladogram on the left implies that cat-like features evolved only once in the clade containing Soup and Merlin
  • The one on the right implies that they evolved twice independently
  • So it seems from first analysis that the first cladogram is the one to retain…
  • … for the moment

ODWS Paul Billiet 2011

19 of 45

An alternative hypothesis

  • Evolution is not just about gaining new characters it is also involves losing characters
  • Suppose that the ancestors of humans and cats were all cat-like…
  • …and these characters were lost just once during the evolution towards me as shown on the right
  • This hypothesis is just as parsimonious as the first

ODWS Paul Billiet 2011

20 of 45

How do we resolve the problem?

  • The two hypotheses can be tested using a fourth organism
  • This organism has to be clearly unrelated to the rest of the group
  • e.g. An animal that is not a eutherian mammal
  • This is called an outgroup and the test is called an outgroup comparison
  • Enter Albert…

ODWS Paul Billiet 2011

21 of 45

Albert is not a eutherian mammal

ODWS Paul Billiet 2011

22 of 45

Two cladograms are possible

  • The cladogram on the left requires cat-like features to have evolved just once on the branch to Merlin and Soup

ODWS Paul Billiet 2011

23 of 45

Two cladograms are possible

  • The one on the right requires either:�that cat-like features evolved twice independently to Merlin and Soup
  • Or:�Cat-like features evolved once in the common ancestor of Merlin, Soup and myself …
  • … AND was then lost in the evolution of myself

ODWS Paul Billiet 2011

24 of 45

Applying Occum’s razor

  • Hence the cladogram on the left offers the simplest (most parsimonious) route

ODWS Paul Billiet 2011

25 of 45

The power of cladistics

  • Cladistics tests all possible hypotheses objectively
  • It can lead to some surprising conclusions

ODWS Paul Billiet 2011

26 of 45

Cladogram of birds and dinosaurs

Node

ODWS Paul Billiet 2011

27 of 45

What is Phylogenetic tree�

  • A phylogenetic tree is a branching diagram or tree showing the evolutionary relationships among various biological species or other entities that are believed to have a common ancestor
  • The branches in a phylogenetic tree represent the lineages of organisms or sequences being studied, and the points where branches split represent common ancestors.
  • The length of the branches typically represents the amount of evolutionary change that has occurred since the split from the common ancestor, although this can vary depending on the type of data being used to construct the tree.
  • Phylogenetic trees are constructed based on various types of data, including morphological traits, genetic sequences (such as DNA or protein sequences), and sometimes even behavioral characteristics.
  • These trees help scientists understand the evolutionary relationships between different species and can provide insights into their shared ancestry, divergence times, and patterns of evolution. Phylogenetic trees are widely used in fields such as evolutionary biology, systematics, ecology, and molecular biology.

28 of 45

29 of 45

  • Rooted trees are trees that have a specified root node, which represents the common ancestor of all the organisms in the tree.

  • Types of Phylogenetic Tree

There are several different types of phylogenetic trees. They can be classified as:

On the basis of the presence or absence of a common root

Unrooted trees do not have a specified root node and show only the branching pattern of the evolutionary relationships among taxa or OTUs, without any information about their common ancestor

30 of 45

On the basis of topology

  • Cladogram is a type of phylogenetic tree that displays only the branching pattern of evolutionary relationships among organisms. Cladograms are unscaled, which means that the branch lengths do not reflect the amount of evolutionary divergence between taxa or operational taxonomic units (OTUs).

Phylogram is a type of phylogenetic tree that represents the evolutionary relationships among organisms by showing both the branching pattern and the amount of evolutionary divergence. Phylograms are scaled, which means that the branch lengths are proportional to the amount of evolutionary divergence.

31 of 45

Phylogenetic Tree Construction Steps

  • �Phylogenetic tree construction is a complex process that involves several steps:
  • 1. Selection of molecular marker
  • 2. Multiple sequence alignment
  • 3. Selection of a model of evolution
  • 4. Construction of the phylogenetic tree
  • 5. Assessment of the reliability of the tree����

32 of 45

1. Selection of molecular marker

  • The first step in constructing a phylogenetic tree is to choose the appropriate molecular marker. 
  • The choice of molecular marker depends on the characteristics of the sequences and the purpose of the study. Either nucleotide or protein sequence data can be used. 
  • For closely related organisms, nucleotide sequences are preferable, while for more divergent groups, slowly evolving nucleotide sequences or protein sequences may be used. 
  • Protein sequences are preferred over nucleotide sequences in many cases because they are more conserved and allow for more sensitive alignment due to having more characters. 
  • Although protein sequences offer several benefits for phylogenetic analysis, DNA sequences can also provide valuable information in certain instances, especially when dealing with closely related sequences.

33 of 45

2. Multiple sequence alignment

  • After the selection of molecular markers, the next step is to align the sequences from different species. 
  • This is the most important step because the accuracy of the resulting phylogenetic tree depends on the quality of the alignment.
  • Alignment programs such as T-Coffee can be used. 
  • Gblocks is one of the automatic programs that can help improve alignment by eliminating poorly aligned positions and divergent regions.

34 of 45

3. Selection of a model of evolution

  • The third step of phylogenetic tree construction is the selection of an appropriate evolutionary model. 
  • Evolutionary (or substitution) models are statistical models that describe the substitution and divergence of sequences over time. 
  • There are several substitution models available for both nucleotide and amino acids.
  • Two commonly used substitution models for nucleotides are the Jukes-Cantor (JC) model and Kimura’s two-parameter model.
  • There are also many amino acid substitution models. The most commonly used ones are the Dayhoff model (PAM) and the Jones-Taylor-Thornton (JTT) model. 

35 of 45

4. Construction of the phylogenetic tree

  • The next step is the construction of the phylogenetic tree. 
  • The two main methods for constructing phylogenetic trees are distance-based and character-based methods. 
  • Distance-based methods rely on computing the amount of dissimilarity between sequences, while character-based methods use molecular sequences from individual taxa to trace the character states of the common ancestor. 

36 of 45

5. Assessment of the reliability of the tree

  • The final step involves assessing the reliability of the phylogenetic tree. This can be done by a statistical method called bootstrapping which is used to assess the reliability of a phylogenetic tree’s topology.
  • It involves repeatedly resampling the initial sequence data to generate multiple subsets of derived sequences, referred to as bootstrap samples.
  • These samples are then used to construct a new phylogenetic tree using the same method as the original tree.
  • Interior branches that are accurately predicted by the new tree are assigned a value of 1. This process is repeated numerous times, and the percentage of times each interior branch receives a value of 1 is calculated as the bootstrap value or confidence value. 
  • A bootstrap value of 95 or more is generally considered to indicate an accurate topology, and these values are expressed as percentages on the branches of the phylogenetic tree.
  • Besides bootstrapping, other resampling strategies like Jackknifing and Bayesian Simulation can also be used.

37 of 45

38 of 45

39 of 45

Multiple sequence alignment

40 of 45

Why we do multiple alignments?

  • Multiple nucleotide or amino sequence alignment techniques are usually performed to fit one of the following scopes :
  • In order to characterize protein families, identify shared regions of homology in a multiple sequence alignment; (this happens generally when a sequence search revealed homologies to several sequences)
  • Determination of the consensus sequence of several aligned sequences.

41 of 45

Why we do multiple alignments?

  • Help prediction of the secondary and tertiary structures of new sequences;
  • Preliminary step in molecular evolution analysis using Phylogenetic methods for constructing phylogenetic trees.

42 of 45

An example of Multiple Alignment

VTISCTGSSSNIGAG-NHVKWYQQLPG

VTISCTGTSSNIGS--ITVNWYQQLPG

LRLSCSSSGFIFSS--YAMYWVRQAPG

LSLTCTVSGTSFDD--YYSTWVRQPPG

PEVTCVVVDVSHEDPQVKFNWYVDG--

ATLVCLISDFYPGA--VTVAWKADS--

AALGCLVKDYFPEP--VTVSWNSG---

VSLTCLVKGFYPSD--IAVEWWSNG--

43 of 45

Multiple Alignment Method

  • The most practical and widely used method in multiple sequence alignment is the hierarchical extensions of pairwise alignment methods.
  • The principal is that multiple alignments is achieved by successive application of pairwise methods.

44 of 45

Multiple Alignment Method

  • The steps are summarized as follows:
  • Compare all sequences pairwise.
  • Perform cluster analysis on the pairwise data to generate a hierarchy for alignment. This may be in the form of a binary tree or a simple ordering
  • Build the multiple alignment by first aligning the most similar pair of sequences, then the next most similar pair and so on. Once an alignment of two sequences has been made, then this is fixed. Thus for a set of sequences A, B, C, D having aligned A with C and B with D the alignment of A, B, C, D is obtained by comparing the alignments of A and C with that of B and D using averaged scores at each aligned position.

45 of 45

Steps in Multiple Alignment