Response to Will Duffy | Phylogenetic systematics

Erica (@GutsickGibbon) is still continuing with her live-stream series teaching Will Duffy about evolution. I missed a couple of streams since the last time I posted a response to Duffy, but I did watch the most recent one.

During the segment 05:42 - 36:40, Will Duffy presents his rebuttals to Erika regarding her presentation of the previous stream, which specifically concerned Tetrapod evolution. However, Duffy’s his main issue seemed to be cladistics, specifically phylogenetic systematics. There are a few unrelated claims that Duffy makes about cosmology at the beginning, but this post will focus on that topic.

Duffy says he reached out to Rebekah Davis and Dr. Cornelius Hunter to notify him about some papers on the subject. I have seen what happens when Rebekah and Hunter discuss a science paper together, so this did not inspire confidence in me. Sure enough, the arguments he makes are very flawed. In particular, the mistakes he makes in his final argument, during the segment 34:50 - 36:02, prompted me to write a comment under the livestream to explain how Duffy gets (practically) everything wrong about the figure he shows. But here I will go over into more details about this argument, and others.

1. Cladistics vs. Phenetics

Concerning segment 19:10 - 30:56

First, some of the good stuff. This was the very first time I have seen a creationist who did a descend summary on the difference between cladistics and phenetics. Will Duffy did some work to understand what is meant by phrases like “birds are dinosaurs” and “whales and snakes are tetrapods” since he was confused about, considering that the common traits used to define the category (e.g. four limbs of tetrapods) do not apply to some of its members (e.g. whales and snakes). He came across the destinction between phenetics and cladistics and the whole debate between these during the 1960s to the 1980s.

Duffy gets the basics right. Phenetics is based on numerical taxonomy, basically the goal is to comprehensively analyze all traits of a given set of organisms and determining their overall similarities shared between them. The animals are classified according to these similarities, with less inclusive taxa defined by greater degrees of similarities, while more inclusive taxa are defined with fewer shared traits. Cladistics goes one step further. They use statistical analyses from phylogenetic systematics to determine which traits are likely homologous (shared via common descent) as opposed to homoplasies (convergent evolution). The homologous traits, specifically synapomorphies (shared derived traits), are used to determine the monophyletic taxa, i.e. clades.

That is good, but Will gets some of the details a bit off. He shows Linnaean taxonomy on his slide as being part of phenetics, but it is not. Linnaean taxonomy was - more or less - based on arbitrary judgements (no numeric or statistic analysis), which at the time was the only practical option. Sokal and Sneath devised phenetics, based on numerical taxonomy, in order to make the discipline objective.

Duffy also gets one thing very wrong. Duffy claims that snakes are not tetrapods according to phenetics. However, despite not having four limbs… the overall morphology of snakes is still more similar to that of other tetrapods than any non-tetrapod taxon. More specifically, snakes are overall most similar to reptiles. So phenetics would still classify these as such as well.

2. Why Cladistics wins

The pheneticists explicit goal was to devise a system of classification that provided the most information content (or naturalness) regarding the diversity of life. Basically, how can we construct the simplest taxonomy can efficiently summarize, and be predictive, of all the diverse forms, with minimal redundancy or loss of data. Pheneticists maintained that this was the primary goal of taxonomy, and that phylogeny had no place in it. However, in a series of papers published from 1977 to 1983 (e.g. this one), James S. Farris used information theory to show that cladistics maintains more information in taxonomy than phenetics. This is mainly due to the fact that phenetics, by averaging out characters as overall similarities, necessarily destroys data, while cladistics maximizes information. For detailed explanations, see here (skip to ‘INFORMATIVENESS - PREFERENCE OF CLADISTICS OVER PHENETICS’).

In other words, cladistics is a better method of taxonomy than phenetics, even according to the criteria espoused by the pheneticists. The fact that cladistics maximizes information in taxonomic classification is strong evidence for shared ancestry.

What this also means is that the very issue Will brought up - regarding situations where cladistics seperates seemingly similar organisms (e.g. falcons and hawks) and unites seemingly dissimilar organisms (e.g. hippos and whales) - is only made WORSE under phenetics. Ironically, Duffy’s argument here is an argument FOR cladistics. Since phenetics destroys information by averaging out morphological data, it more often leads to situations where organisms are separated into two taxa, such that the actual character distributions strongly overlap between the two taxa; i.e. phenetics is more likely to seperate similar organisms into unnatural groups. Expressing this point in cladistic terminology, phenetics has a higher tendency to erroneously construct paraphyletic (or polyphyletic) taxa.

To me, the best and most intuitive illustration of this point is provided by this video below. It is also one of my personal favorite YT videos. It shows how classifying birds - not only as dinosaurs - but also as reptiles, minimizes the the aforementioned contradictions that phenetics tends to produce.

3. Circular definitions?

Concerning segment 14:05 to 19:10

Will Duffy here gets stuck on how to him the definitions of cladisitcs are circular. One example he gives in paritcular is the first sentence on a wikipedia page.

Duffy is right that this definition is circular - as in - it does does not informatively define what terapods and the group tetrapoda are. However, Duffy is also not fair to the wikipedia article, since the very first subsection (Definitions) discusses in detail precisely how Tetrapoda is defined as a group, and it also includes some discussion about competing definitions, which are not circular.

Having said that, I took the initiative to update the introductory page by giving a brief and informative definition for Tetrapoda.

But why do we include limbless things within a group that is characterized by limbs? Doesn’t that seem contradictory? Well, the alternative to this would be coming up with a taxon that is described as the “lack of limbs” to include snakes (which are reptiles), but that would also end up erroneously lumping other limbless and elongated forms, such as caecilian (which are amphibians), and perhaps all things that are worm-shaped like annelids (which are not even vertebrates). What you end up with is a group of - “WORMS” - all of which may have silhouettes that look similar superficially, but are still very different from each other in many more details.

Fun trivia, before its usage was restricted to invertebrates, the term ‘worm’ (or ‘wyrm’ in old-English) was actually used in this manner, even commonly used to refer to large serpents, including those of mythology like the world serpent, Jörmungandr.

As explained in the previous section, even phenetics does not do this. Since it classify things by their overal (average) similarity, such that snakes are still tetrapods even under that system. However, cladistics fares better, which was also covered in the former section.

4. Much ado about statistics

Concerning segment 30:56 - 36:02

First, I have to make a general point about statistics. Suppose someone named Mike has a hypothesis about the relationship between diet and height. He claims that people who consume 1200 calories per day grow to be 1.5 meters, and those 2400 calories per day grow to be 1.7 meters. He looks at 2000 people, 1000 consuming 1200 calories and 1000 consuming 2400 calories. In an ideal scenario, where calories is the ONLY factor determining height, this is what he would expect to see.

I made these figures in R with the hist() function.

Within each cohort, the heights are perfectly uniform, no exceptions. However, the real world of course does not work this way. Diet is not the only variable that influences height. There are confounding variables involved. Let’s assume for simplicity that all confounding variables sum up to be a random factor, causing the heights to have a normal distribution with a standard deviation of 0.1 in both groups. This means that the majority (~68%) of all measured heights will tend to fall within the ± 0.1 range away from the predicted values: 1.4 - 1.6 meters for 1200 calories and 1.6 - 1.8 meters for 2400 calories. As a result, the histograms will look like this.

I generated the data in R with the rnorm() function.

Now, things seem a lot messier. The heights measured in both groups overlap one another between 1.40 and 1.85 meters. Specifically, ~15% of those who consume 1200 calories per day have heights that are closer to 1.7 meters than 1.5 meters, and ~2% exceed the average 1.7 meters of the group that consume 2400 calories group. The same is true vice versa: ~15% of the 2400 calories group have heights closer to 1.5 meters than 1.7 meters, and ~2% are shorter than the average 1.5 meters of the 1200 calories group.

But Mike is not deterred. He is confident that the two groups support his hypothesis. He demonstrate his confidence with a statistical test. In this case, a two-sample t.test is used. We can do this easily in R with the t.test() function. The p-value it gives is < 2.2e-16, which is actually the lowest p-value that R can show. The real value is closer to < 2.2e-307 That is the probability that the observed differences in the height distributions can be explained by random chance. In other words, we can be confident with a certainty of OVER 99.99999… [followed by another 300 nines] percent that the mean difference between the groups are not random. R also provides a 95% confidence interval for the mean difference between the two groups, which is 0.193 - 0.210, which includes the predicted mean difference of 0.2.

Now… what if someone called Darren comes along and argues that the conclusion made by Mike is completely bogus. He points out that ~15% of the data is incongruent with Mike’s hypothesis, and that the t.test is just a sneaky trick to sweep all the incongruences under the rug. You might think that this is a silly argument, especially considering the statistical probabilities shown previously, and you would be correct. However, this is exactly what creationists are doing when they point to incongruences in phylogenetic systematics… which is deeply entrenched in statistics. This is especially the case when they point to examples where results differ between morphological- and genetic-based data, or when they point to incongruences within morphological data, such as homoplasies; i.e. traits shared by convergent evolution. Will Duffy also does this throughout the last segment of his slide show from 30:56 - 36:02.

The fact is that phylogenetic systematics - like almost all other statistical disciplines - will be subject to confounding variables, such that incongruences are actually expected to occur. The point remains that - despite the confounding variables - the morphological and genetic data is statistically overwhelmingly in favor of common ancestry. In the following section #5, I will cover a very good example that illustrates this fact.

5. Completely misunderstanding a graph

Skipping to the segment 34:50 - 36:02 when Duffy shows the image below from this 1991 paper.

Better screenshot of the graph here:

The figure shows how the number of taxa included in a dataset affects the resulting consistency index. Exactly what this means will be made clear in the following paragraphs.

First some (relatively) minor corrections. Will says (emphasis mine):

Here is a chart from that paper and I’m going to do my best to explain it, and I hope I don’t get this wrong, but I don’t think I will. A consistency index of 1.0, which is that red line at the top, means essentially that the morphological data and the genetic data match. That’s kind of like a perfect match.

That’s not the case. The paper only deals with morphological (character) data. What the consistency index notates is how good a given dataset (all the characters states of a given number of taxa) is explained by the corresponding phylogenetic tree. In an ideal scenario, the number of evolutionary steps (character changes) needed according to the phylogenetic tree should be the minimal theoretical number required to explain the data. However, data can be confounded by homoplasies (shared traits due to convergent evolution). For example, if the theoretical minimum number of evolutionary steps required is 100, but the actual number is 200 (an extra 100 changes due to homoplasies) the consistency index is 100/200=0.5.

Will further states (emphasis mine):

Now these two lines that you see going down, those are two random data sets that they used for purposes of looking at… what is predicted is the red line… and then what’s just purely random data.

Another relatively minor correction. Those dotted lines, that Will points at with a red arrow, these represent the 95% confidence interval, based on 30 randomized datasets. Basically, if you used a randomized morphological dataset, the CI value produced from it will fall between the two dotted lines with a probability of 95%. Lets call this the “zone of chance”.

Here is where Will Duffy goes completely off the rails.

If the pattern of morphological traits is constrained by phylogeny, then the results of real datasets should fall OUTSIDE and ABOVE the zone of chance. Those open triangles, open dots, and solid dots represent the results of 75 real datasets. The citations of these are provided in the caption of the figure. As you can see for yourself (except for 3 points) all of them are outside and above the zone of chance. The probability that one data point falls outside and above of the zone of chance due to randomness is 2.5%. The probability that ≥72 out of 75 datasets do this at random is less than… 1 in 10^109… which is beyond astronomical… literally… you are more likely to succeed at picking the exact same atom twice at random from the observable universe, with a probability of 1 in 10^80. And bear in mind that this is just asking whether the datasets fall outside and above the zone of chance… we are not even asking how far away from the zone of chance they are, which will lower these odds even further. This means that those datasets significantly support their corresponding phylogenetic trees. The figure shows statistically significant evidence FOR common ancestry.

Duffy gets the figure completely backward, claiming it contradicts common ancestry (emphasis mine):

You’ll notice… the number of taxa that they look… that they looked at, which goes here along the x-axis. The more taxa that they looked at and compared, the more it showed closer to random versus the top line of 1.0.

So, only when they used five taxa did they get a result that matched. As soon as they increase that, all the results more closely aligned with the random data set versus the prediction.

Here he makes several erroneous assumptions.

First, he assumes that, according to common ancestry, the datapoints should follow the horizontal red line he added in his slide. He expects that all points should have CI = 1.0 - or at the very least - he expects that the points are positioned visually closer to 1.0 relative to the zone of chance. Why is this assumption wrong? As we have established in former section #4, phylogenetic systematics is a statistical science which always deals with confounding variables. This means we will never expect a 100% perfect fit, since that would entail perfect data unaffected by confounding variables. In this case, the key confounder is homoplasy. There are others confoundes that are encountered in plylogenetic systematics, which I will leave aside. This means we do NOT expect that the data follows the red line (CI = 1.0) that Duffy added to the graph.

We also don’t expect the data to be positioned visually closer to the red line relative to the zone of chance, because that wrongly assumes that the effects of the variables are linear. If you increase the number of taxa, the number of possible homoplasies increases - by a far greater rate - than the number of possible homologies. Why? Let’s think about one character for which there are two character states: 0 and 1. If the character state indicates only homology (zero homoplasy), the character state only changes once from an ancestral (0) to a derived (1) state in the phylogenetic tree. This would remain true no matter how many taxa you include in the phylogeny. However, with regard to a character that is confounded by homoplasy, there are many possible ways in a phylogenetic tree where the character state can shift from 0 to 1 (or go back from 1 to 0) two times independently. And it is possible for this to happen more than twice. Thus, with increasingly larger phylogenies (correspoding to more taxa), the probability that a given character is confounded by homoplasies increases. Ergo, the consistency index - predictably - tends to go down when you include more taxa in your datasets. As you may have noticed, the zone of chance is also not linear, going down fast when more taxa are included, precisely for the same reason.

To illustrate this further, let’s bring back the example I gave regarding Mike’s hypothesis on the relationship between diet and height in former section #4. Here below, I plot the datasets of both groups into the same histograms.

  • The left histogram shows the effect of diet under an ideal, unrealistic scenario. Diet is the only factor, and there are NO confounding variables. This is analogous to the red line (CI = 1.0) that Duffy added to the graph above.
  • The middle graph represents a realistic scenario, when we introduce a random confounding variable, which is analogous to the 75 points based on real datasets in graph above.
  • The right graph is what happens when we remove the dietary effect, such that the expected average is 1.6 meters, leaving the random effect as the only variable. This is analogous to the zone of chance in the graph above.

Hold up… wait a minute!? The realistic data (middle plot) looks far more like the randomized data (right plot) than the data that is predicted from diet as the only factor (left plot). But this does not tell you anything statistically relevant. Thinking otherwise is the very mistake Will Duffy makes here.

Another error that Duffy makes with this figure is when he highlights the only point that has a CI equal to one, saying (emphasis mine):

Only when they used five taxa did they get a result that matched. As soon as they increase that, all the results more closely aligned with the random data set versus the prediction.

He says this as if to claim that only datasets with a low number of taxa are consistent with common ancestry, but the graph shows the very opposite. Datasets with low taxa numbers are at risk of falling within the zone of chance. See those three dots inside the dotted lines, which are all the way on the left-hand side of the graph. Datasets with higher taxa numbers consistently fall outside the zone of chance. Thus, datasets that include more taxa are more reliable and more supportive of common ancestry. NOT less.

VERY IMPORTANT CAVEAT: I am not blaming Will Duffy personally for getting this figure so wrong. I can see how a laymen can look at that graph and come to such a conclusion that seems intuitive to them. Statistics is not intuitive. However, I am not sure if it was Dr. Cornelius Hunter or Rebekah who notified Will about this paper and the graph. If it was Dr. Hunter… then… yeah… I would certainly blame him. I definitely expect him to know better.

6. Reading what the paper actually says

Concerning segment 30:56 - 34:50

Before Duffy shows the graph discussed in section #5, he goes through several papers (provided to him by Rebekah and Dr. Hunter) in rapid succession. I am going to wrap this up as quickly as possible, since this post is already very long. So I won’t cover every paper he mentioned.

Here is the first paper that Duffy shows, along with a quote next to it. The full paragraph reads as follows (highlighting sentences that are omitted in Duffy’s slide):

Obtaining additional DNA specimens of the two New Zealand species of Thaumledone, T. marshalli and T. zeiss will greatly aid in our understanding of the evolutionary history of the genus. Morphometric analysis fails to separate the two New Zealand species from one another and from T. peninsulae and it is unlikely that such data would prove phylogenetically useful, even if they provided separation among species, since there was no congruence between morphological and molecular matrices for the Southern Ocean species. Codeable morphological characters are scarce and many (e.g. the size of the salivary glands) seem to reflect environmental influences (e.g. habitat depth) rather than evolutionary history. The addition of molecular sequences for these taxa is therefore essential to determine the origins and the genetic divergence within the genus. Bearing in mind the population level differences seen over relatively small distances in this study, it is also likely that future trawling efforts in the Southern Ocean, particularly on the slope waters of sub-Antarctic islands, will discover further species of Thaumeledone.

To add more clarity, the paper concerns five species of the genus *Thaumledone - *a genus of small, deep-sea benthic octopuses which mostly live in the cold, deep waters of the Southern Hemisphere. The five species are:

  • Two New Zealand species, for which they had only morphological data: T. marshalli and T. zeiss
  • Three Southern Ocean species, for which they had both morphologic and genetic data: T. peninsulae, T. gunteri, and T. rotunda

As Duffy’s quote shows, there was no congruence between the morphologic and genetic data, which sounds worse than it actually is, because when you are only dealing with 3 taxa, there are only 3 possible rooted trees you can consturct, each of which will be 100% incongruent with the other two (sharing zero nodes). In simpler terms, with 3 taxa and rooted trees, two datasets can either be 100% congruent or 0% congruent. There is no other possibility.

More importantly, the authors noted good reasons to doubt the usefulness of the morphological data that they had available. First and foremost, the morphological was rather limited. They had three data matrices:

  • 16 morphological variables for the three Southern Ocean species (9 individuals total)
  • 11 morphological variables for the three Southern Ocean species (21 individuals total)
  • 11 morphological variables for all five species (29 individuals total).

That is really poor. The modern standard lies between 100 to 500 morphological variables. The authors in particular note (even in the quote that Duffy shows) that the morphological data was not sufficient enough to even delineate the individuals of the New Zealand species and T. peninsulae into separate species. Statistical incongruences between one limited dataset (morhphology) and a more comprehensive data set (genetics) is not surprising.

The second reason the authors doubt the usefulness of the morphological data is that the limited morphological data is also likely confounded. Morphological data suggested that, among the southern ocean species, T. rotunda was the odd-one out, but genetics very convincingly (with Bayesian support values of 1.00, and maximum likelihood bootstrap value of 99.5) shows that T. peninsulae is the odd-one out. The authors noted that T. rotunda is found in much deeper waters with a likely circumpolar distribution, while the other two species are found in the same geographical location at similar depths. Thus, the morphological oddness of T. rotunda is likely due to it being adapted to a very different environment, while the other two species have either evolved convergently, or they have simply conserved the same morphology as their common ancestor. Quoting from the authors directly:

Despite some morphological similarities between T. peninsulae and T. gunteri, the molecular phylogenetics show a close sister taxa relationship between T. gunteri and T. rotunda. This suggests that at least some of the morphological features unique to T. rotunda may have evolved in conjunction with its distribution in deeper waters. For example the reduction in posterior salivary gland size (associated with paralysis of prey) may be due to the greater propensity of small and soft bodied prey items in the deep sea than in shallower depths (Voss 1988), and thus a reduced requirement for toxic agents to subdue these prey items. Reduction of such features in T. rotunda is consistent with the concept that reductions and losses of characters in many deepwater octopods has occurred convergently (Voight 1993).

FIN

4 Likes

Thanks for this. I also contacted Will after the latest stream and I’ve been trying to walk him through the evidence provided by ancestral sequence reconstruction, which is something that initially helped convince me about common descent. I’ll email him about this thread, hopefully he will read through your response, since he’s been very open minded.

2 Likes

I think a mistake Will is doing is to go to Rebekah and Cornelius Hunter and judging them to be honest and qualified to explain these things to him. He doesn’t have the knowledge to be able to make such an evaluation.

In my estimation Will has a small blindspot for apparent sincerity, and he appears to put way too much scientific value on if people can tell personal stories. Or if they can, occasionally, admit to being wrong about some things.

People can of course sincerely believe they are in the right and yet be mistaken, or even delude themselves to various degrees, and heck, they can even put on a facade of sincerity.

In my experience almost everyone will, at one point or another, admit to mistakes and being wrong. They will even some times do so deliberately to generate trust, as long as admitting that mistake does not appear to cast any significant amount of doubt on what they believe overall.

It is one thing to say “I was wrong about this particular piece of evidence”, but another entirely to admit you’re fundamentally wrong about the overall point of contention.

It is a mistake to put so much stock in these token admissions of mistake, that you then infer you can generally trust the person you’re talking to. The sad truth is you can’t make that inference.

Obviously this goes for everyone on both sides of any contended topic including, of course, the creation-evolution debate.

4 Likes

Couple of points about Will Duffy’s presentation.

First, he’s confused about the difference between cladistics and phenetics. The second term is ambiguous in meaning, and he seems to take a restrictive view, tying it closely to the criterion of parsimony. A more common usage these days refers not to a tree-selection criterion but to the practice of matching classification to phylogeny. But even under the restrictive definition, he’s wrong. Phenetics and cladistics are both numerical taxonomy, though only cladistics is phylogenetic, so he got that bit right. “Observable traits” are different from “shared derived traits” only in retrospect, after the phylogenetic analysis is complete. “Quantitative” is not in conflict with “parsimonious”. Nor is Linnean taxonomy inherently phenetic. It’s either entirely intuitive or relies on limited sets of characters viewed by the taxonomist as important. In any case it doesn’t rely on overall similarity. Finally, phenetics is not necessarily attached to morphological characters while cladistics is not necessarily attached to molecular characters. All methods can be applied to any data type. We should also distinguish between tree-building and classification; the latter is subsequent to the former, true both in phenetic and phylogenetic treatments.

Second, he’s very confused about taxon definitions. A tetrapod can indeed be defined as a member of Tetrapoda. That’s not circular; it just passes the need for a definition to the taxon Tetrapoda. And this definition has always been more or less phylogenetic. These days the definition is explicitly phylogenetic and relies on reference to a tree. Snakes and whales have been considered tetrapods since the term was first used. Etymology is not destiny, and he should stop harping on it.

Same with Diapsida. Yes, turtles are diapsids because they are part of that clade, phylogenetically defined. Note that this was first proposed based on morphological characters — see the work of Olivier Rieppel. (My favorite such character is the hooked 5th metatarsal.) Anyway, this turns out not to be a conflict between molecular and morphological data.

That’s as far as I got, but so far it’s not impressive.

3 Likes

Good clarifications.

After explaining that Linnaean taxonomy is not really phenetics…

…I also wanted to comment on how weirdly he dichotomized “observable traits” and “shared derived traits”, but I did not want to get bogged down into explaining all the details and move on to issues I found more pressing.

I do think it is circular… IF the taxon is not defined informatively. If we just define a ‘tetrapod’ as any member of a group called ‘tetrapods’ or ‘tetrapoda’, then we are simply stating that all tetrapods are tetrapods. It is indeed circular.

Of course, it is very easy to find a basic definition that is not circular, and we can also go into the difference between apomorphy-based, node-based, or stem-based definitions in cladistics, which are not circular either.

I haven’t watched the video, and don’t think I will. But from your response, it seems to me Duffy is making a mistake that that many creationists do. They fail to understand that common ancestry is concluded from the fact that the branching tree pattern follows as a consequence from whatever system is used to classify organisms. They get hung up on whether and how it can be determined that shared characteristics, taken on their own, result from common ancestry rather than from “common designer” or some other hypothesis.

It may be true that, if a small number of organisms share a small number of attributes, there may be a reasonable chance that this could result from causes other than shared ancestry. But when the nested hierarchy is repeatedly confirmed as ever greater numbers of attributes of ever greater numbers of species are analyzed, it becomes perverse to deny there could be any explanation other than common ancestry.

3 Likes

Continuing through the video, and few minutes on:

Duffy claims that evolution makes two predictions: “Organisms with similarities in phenotype will have similarities in genotype” and “organisms without similarities in phenotype will not have similarities in genotype”. Unfortunately, these are too poorly stated to have any real meaning. His prime example of the first is chimps and bonobos, which are nearly identical in both phenotype and genotype. His example of the second is a fruit fly and (I think) a jellyfish. But of course they do have similarities in both phenotype and genotype; again, without quantification of the expected degree of similarity the “prediction” is useless.

Then he gets into pairs that he thinks contradict the explanation, the first being hawks and falcons. They have phenotypic similarities but, he supposes, no genetic similarities. But what he really means, apparently, is that they are genetically more similar to some other birds than to each other, because of course they do have considerable genetic similarity. And he doesn’t even really mean that, since his index of genetic similarity is phylogenetic relationships, which are not determined by genetic similarity. Falcons and hawks are not either other’s closest relatives, but they are relatives, and much closer to each other than to any birds outside their common clade Telluraves. And neither are falcons and hawks determined to be phenotypically similar in any rigorous way. They just look similar in certain noticeable features. But of course they’re different in many other features.

The next supposed contradiction of evolutionary expectation is whales vs. hippos, which supposedly have genetic similarities but no phenotypic similarities. And yet they have a great many phenotypic similarities: that’s how we have long known that they’re both mammals, and more recently, based on a number of fossil whales with diagnostic double-pulley astragali, that they’re both artiodactyls. Once again, the claim is vacuous.

If we can draw anything coherent from these examples, it would seem to be a claim that the degree of phenotypic similarity should be closely correlated to the degree of genetic similarity. But there is no reason, given our understanding of evolution, to expect such a thing. The simplest model under which that would be expected would be a combination of absolute molecular and morphological clocks, i.e. a constant rate of evolution in both systems. Needless to say, no such thing exists or would be expected. Rates of molecular evolution change over time, and rates of phenotypic evolution even more so. Duffy raises a very confused strawman.

Next, he makes the claim “genetics is producing very different hierarchies than phenetic classification”. This is confused on many levels. First, he confuses phenetics with classical Linnean classification, which is not the same at all. Very few taxa have been assessed by phenetic methods, because phenetics was a brief fad adopted by few systematists. Second, he confuses phenetics (a set of methods) with morphology (data). Phenetic methods can be used with any data, and morphological data can be analyzed using any method. The most common such method, at least for the past 50 years or so, is maximum parsimony, i.e. what he means by “cladistics”. Now sometimes he’s comparing molecular phylogenies to cladistic analyses of morphological data, and sometimes he’s compariing them to traditional classifications, which were neither cladistic nor phenetic. Much of the difference between modern cladistic and traditional classification is of philosophy. It’s been known at least since the 1860s that birds are nested within reptiles, but while traditional classification allowed the dismemberment of clades, these days we don’t. Changes in classification that reflect changes in underlying ideas of phylogeny are less common than those that reflect this change in what can be called a taxon.

Still, if we can make sense of his poorly stated claim, it’s about the conflict between morphological and molecular phylogenies. And those definitely exist, though they are less prevalent than he supposes.

Some advice: if Duffy wants to know about phylogenetics, he should ask a phylogeneticist rather than a couple of creationists.

But he got a bunch of papers showing differences between some morphological and molecular phylogenies. Two of them, oddly enough, are publications of mine, and he quote-mines me to exaggerate the difference between phylogenies. I suspect the other examples are similar. And I suspect both that the sample is biased and that differences among trees are not as great as implied. Incongruities tend to be between trees that differ only slightly and are much more similar than expected by chance. This is expected if the data are noisy and/or inconclusive, as is the case with quite a few morphological analyses. And of course many traditional classifications aren’t based on any sort of rigorous analysis of data.

1 Like

To be fair, people working on “cladistics” have not done a good job explaining the issues. They tend to conflate classification with reconstruction of phylogenies (it is important to realize that those are two different tasks). There are many definitions of cladistics, most of which declare that it is using monophyletic classification, but then continue by insisting that inference of phylogenies be done by nested synapomorphies, or by parsimony, or by some other means. And every advocate of cladistics is sure that their definition is the one everyone else uses. It’s a complete mess. In addition many of them call use of distance matrix methods “phenetics”, even though that term describes a position on classification, not how phylogenies are to be inferred.

Will Duffy could not have chosen a discussion in which the advocates of different positions do a worse job of explaining them.

I don’t recall any actual use of “Cladistics” by scientists in the past 30 years or so except as the title of the journal of the Hennig Society. There’s “cladistic classification”, though I think that’s largely been superseded by “phylogenetic classification”.

From two creationists. That he judged to both be honest because they’ve admitted to having made mistakes before.

I’ve left comments on mutiple of Rebekah’s videos explaining her many misconceptions about incongruent phylogenetic trees, since. Many of these comments she either ignored entirely, or failed to comprehend.

Given that she’s re-telling the same basic mistakes to Will now (the mistake being that she thinks she can basically dismiss phylogenetics entirely on the basis that different trees will be incongruent to some degree or another), I stopped posting on her vidoes because it seemed a waste of time. I ended up concluding that she’s either dishonest or suffering some severe case of Morton’s Daemon.

1 Like

To elaborate: he cites my phylogenetic analysis of Crocodylia. As stated therein, previous phylogenetic analyses of morphology had disagreed with previous such analyses of molecular data (some of which, incidentally, had used methods that might be judged “phenetic”), and also with my analyses. But the difference involved the placement of only one species. Trees were otherwise identical. Recent analyses of morphology have also come to agree with the molecular tree.

And he cites my phylogenetic analysis of Palaeognathae. There had in fact been very few previous analyses using morphology, and one of them had found the same result I did. Other previous molecular analyses had suffered from problems of rooting, long-branch attraction, and in one case of forcing the analysis to conform to an a priori tree. And again, though the result was surprising, it involved a disagreement regarding only one or two species (depending on how many species you divide ostriches into). And again, some recent analyses of morphology have come to agree with our result.

Meant to include links:

https://www.researchgate.net/publication/10734528_True_and_False_Gharials_A_Nuclear_Gene_Phylogeny_of_Crocodylia

https://www.researchgate.net/publication/23232033_Phylogenomic_evidence_for_multiple_losses_of_flight_in_ratite_birds

Probably useless. Creationists generally don’t want to hear and are at any rate suspicious of anything they get from satanic evilutionists. But the sample is inherently biased. There’s not much reason to publish “we got the same result as everyone else has” or, even before that, to revisit well-supported conclusions. You don’t get much out of showing that birds are monophyletic or showing that whales are mammals. And so the literature is enriched in conflicts and surprises beyond what a random sample of the tree of life would show.

Let me expand on that from personal experience. Consider the two pubs of mine that Will mentoned. The one on crocodylians came from a data set originally intended only to provide an outgroup for a bird study. When we noticed that it resolved an existing controversy in crocodylian systematics we published it separately; otherwise it would at most have been a sentence in another paper, probably just in the methods section. The one on palaeognaths originated from a big project (“Early Bird”) on all birds, which again had a crocodylian outgroup that merited only a sentence or two. It’s a separate paper solely because it found a novel arrangement of taxa suggested only once previously, and I thought this little bit of the tree deserved analysis in great detail and a publication all its own.

2 Likes

There is a wonderful resource called Orthomam which is for mammal genes. It contains all fully-sequenced mammals, and alignments of DNA for each gene from each of those mammals. Currently, 15,868 alignments from 190 species. They also provide protein sequence alignments. And, for each of the 15,868, phylogenies for all species that have that locus.

Using their interface, one can download the trees for different loci of your choice and compare them. Or you can download the alignments and make trees yourself. This should go some distance toward correcting for the overpublication of discrepancies between trees.

Link: https://orthomam.mbb.cnrs.fr/

2 Likes

Unfortunately it’s for protein-coding exons only. It would be nice if they could include at least the introns. Better reflection of phylogeny. [Correction: I wasn’t reading the tree right; Atlantogenata is there just fine. Never mind.]

Well, the problem with including introns is aligning them. Big, big problem. As far as which groups are in “the tree”, you are looking at the overall tree inferred from all exon sequences. The interesting thing one can do with Orthomam is to make trees for individual loci, and for subsets of the 190 species. And seeing whether the amounts of discordance between the trees from different loci are anywhere near as large as claimed by creationists.

1 Like

Not so big as you might think. At least, you can align introns for all birds well enough.

Exactly. I think omitting introns makes it a great tool for lay illustrations.

I have followed her videos for a while, but I did not notice your comments. Or perhaps you are using a different name on YT.

Under which videos did you comment?