# John Harshman: The Phylogeny of Crocodiles

**URL:** <https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057>\
**Category:** Office Hours\
**Tags:** Science\
**Created:** [July 30, 2019, 3:20pm UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057 "2019-07-30T15:20:51Z")\
**Posts on this page:** 12\
**Page:** 3

<div class="post-metadata">

**Author:** ![swamidass](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/swamidass/32/3_2.png) [@swamidass](https://discourse.peacefulscience.org/u/swamidass)\
**Post date:** [August 9, 2019, 3:39am UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/42 "2019-08-09T03:39:59Z")

</div>

> [@John\_Harshman](#):
>
> > Can measure and report the number of homoplastic mutations?
> 
> I could report the consistency index; that’s easy enough. Why?  
> Permuted: CI .5337; regular: CI .9531.

Please explain CI and how it relates to the number of homoplastic mutations?

---

<div class="post-metadata">

**Author:** ![John\_Harshman](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/john_harshman/32/1121_2.png) [@John\_Harshman](https://discourse.peacefulscience.org/u/John_Harshman)\
**Post date:** [August 9, 2019, 3:50am UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/43 "2019-08-09T03:50:55Z")

</div>

> [@swamidass](#):
>
> Please explain CI

CI is a property of data mapped onto a tree. It’s the ratio of the minimum possible number of transformations for each character on any tree to the number of transformations as mapped onto the current tree. The minimum possible number of transformations is the number of observed states minus one. Most of the additional transformations above the minimum would be considered homoplasy. CI is from 0 to 1; higher CI means less homoplasy. A CI of .95 means very little implied homoplasy; .53, quite a lot.

---

<div class="post-metadata">

**Author:** ![swamidass](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/swamidass/32/3_2.png) [@swamidass](https://discourse.peacefulscience.org/u/swamidass)\
**Post date:** [August 9, 2019, 4:00am UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/44 "2019-08-09T04:00:49Z")

</div>

Can you please work that out with a small example?

---

<div class="post-metadata">

**Author:** ![John\_Harshman](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/john_harshman/32/1121_2.png) [@John\_Harshman](https://discourse.peacefulscience.org/u/John_Harshman)\
**Post date:** [August 9, 2019, 1:18pm UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/45 "2019-08-09T13:18:40Z")

</div>

> [@swamidass](#):
>
> small example?

Sure. Suppose you have a site at which there are two states, say T and C, with some species having T and others C. Further suppose that you know the root state to be C. (That isn’t necessary and doesn’t affect the calculations, but it makes things easier to explain.) The minimum number of changes to explain the distribution is one, from C to T. Now suppose you have a tree that best fits all the data, and on this tree the species with T do not form a single group. We may have to suppose that the change from C to T happened twice independently in different parts of the tree, or we may have to suppose that C changed to T and, at some point, back to C. In either case, there are two changes necessary to explain the distribution of C and T over that tree. The consistency index for that site would be 1/2.

---

<div class="post-metadata">

**Author:** ![Rumraket](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/rumraket/32/9328_2.png) [@Rumraket](https://discourse.peacefulscience.org/u/Rumraket)\
**Post date:** [August 9, 2019, 1:54pm UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/46 "2019-08-09T13:54:41Z")

</div>

That was excellent. What program does one use to obtain a CI for a data set?

---

<div class="post-metadata">

**Author:** ![John\_Harshman](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/john_harshman/32/1121_2.png) [@John\_Harshman](https://discourse.peacefulscience.org/u/John_Harshman)\
**Post date:** [August 9, 2019, 2:18pm UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/47 "2019-08-09T14:18:53Z")

</div>

> [@Rumraket](#):
>
> What program does one use to obtain a CI for a data set?

I used PAUP. Note that the CI isn’t for the data set; it’s for the data set combined with some particular tree.

---

<div class="post-metadata">

**Author:** ![Rumraket](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/rumraket/32/9328_2.png) [@Rumraket](https://discourse.peacefulscience.org/u/Rumraket)\
**Post date:** [August 9, 2019, 2:42pm UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/48 "2019-08-09T14:42:41Z")

</div>

> [@John\_Harshman](#):
>
> Note that the CI isn’t for the data set; it’s for the data set combined with some particular tree.

Thanks, I understand.

---

<div class="post-metadata">

**Author:** ![davecarlson](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/davecarlson/32/2139_2.png) [@davecarlson](https://discourse.peacefulscience.org/u/davecarlson)\
**Post date:** [August 9, 2019, 4:44pm UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/49 "2019-08-09T16:44:50Z")

</div>

> [@davecarlson](#):
>
> I think PAML’s evolver package can do this. When I get a little time, I’ll try working on this.

Okay, here is a follow up.

I used the [T-REX webserver](http://www.trex.uqam.ca/) to generate an arbitrary (but resolved) tree with 10 taxa.  
I then used [PAML’s](http://abacus.gene.ucl.ac.uk/software/paml.html) evolver package to simulate 10,000 bp of sequence along each branch of the tree. Then, I used [IQ-tree](http://www.iqtree.org/) to perform a Maximum Likelihood phylogeny inference with 100 bootstrap replicates.

Here is the tree (with [midpoint rooting](https://www.mun.ca/biology/scarr/Panda_midpoint_rooting.html) for easy visualization):

 ![nonRand](https://us1.discourse-cdn.com/flex016/uploads/peacefulscience/original/2X/b/b713c5b6a2b92337706f6a5eec3df72ae056759a.png)

As expected, the tree is completely resolved with 100% bootstrap scores.

Next, I repeated the same procedure but used a completely unresolved tree (i.e., a polytomy) to simulate the DNA sequences. Here is the tree resulting from that data set:

 ![Rand](https://us1.discourse-cdn.com/flex016/uploads/peacefulscience/original/2X/1/158f541389ddd6ccf043d972936911dc53d9fb79.png)

This time the internal branches are all very short, and the low bootstrap scores suggest that all the relationships are very uncertain. Again, as expected.

Edit: fixed some wording

---

<div class="post-metadata">

**Author:** ![Rumraket](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/rumraket/32/9328_2.png) [@Rumraket](https://discourse.peacefulscience.org/u/Rumraket)\
**Post date:** [August 9, 2019, 4:51pm UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/50 "2019-08-09T16:51:42Z")

</div>

What would be the CI values you get from each of those two trees?

---

<div class="post-metadata">

**Author:** ![davecarlson](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/davecarlson/32/2139_2.png) [@davecarlson](https://discourse.peacefulscience.org/u/davecarlson)\
**Post date:** [August 9, 2019, 5:00pm UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/51 "2019-08-09T17:00:12Z")

</div>

I don’t currently have PAUP installed and don’t know which other programs estimate CI, but if I get a chance, I’ll install PAUP and try to check.

---

<div class="post-metadata">

**Author:** ![Mercer](https://avatars.discourse-cdn.com/v4/letter/m/e274bd/32.png) [@Mercer](https://discourse.peacefulscience.org/u/Mercer)\
**Post date:** [August 9, 2019, 7:37pm UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/52 "2019-08-09T19:37:27Z")

</div>

> [@John\_Harshman](#):
>
> Sure. Suppose you have a site at which there are two states, say T and C, with some species having T and others C. Further suppose that you know the root state to be C. (That isn’t necessary and doesn’t affect the calculations, but it makes things easier to explain.) The minimum number of changes to explain the distribution is one, from C to T. Now suppose you have a tree that best fits all the data, and on this tree the species with T do not form a single group. We may have to suppose that the change from C to T happened twice independently in different parts of the tree, or we may have to suppose that C changed to T and, at some point, back to C. In either case, there are two changes necessary to explain the distribution of C and T over that tree.

For laypeople, it might be clearer to simply point out that we would mis-score 2 mutations (a mutation and a reversion to the original base or amino-acid residue) as 0 mutations if we have not sampled a species with the mutation. This is noise, but it is washed out by the enormous amount of data we can collect.

---

<div class="post-metadata">

**Author:** ![Joe\_Felsenstein](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.peacefulscience.org/joe_felsenstein/32/8176_2.png) [@Joe\_Felsenstein](https://discourse.peacefulscience.org/u/Joe_Felsenstein)\
**Post date:** [August 10, 2019, 1:46am UTC](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057/53 "2019-08-10T01:46:18Z")

</div>

Glad to hear it seemed cool. It is fairly old right now (the latest release about a decade old) but I am working on a new version with Java interfaces for the programs. I do like to think that it has the best documentation in the (phylogeny) industry.

[Previous page](https://discourse.peacefulscience.org/t/john-harshman-the-phylogeny-of-crocodiles/7057.md?page=2)
