Can someone explain to me just what the hell information entropy is a measure of in a way where it’s relevant to, like, anything? Who gives a crap if SIE goes up or down? What does that matter?
Shannon information entropy is a measure of… what exactly? What does it have to do with life? I get the distinct impression someone is trying to claim it’s a measure of fitness, or functional information. But such a claim hasn’t met it’s burden of proof. What am I missing?
Also, doesn’t the 2nd law of thermodynamics say that the total entropy of an isolated system always goes up? And this is the whole “things are running down” sort of thing that originally inspired the 2LOT arguments against evolution? And the “running down” thing is because the total entropy of the universe increases. As things are running down, entropy increases.
But the shannon information entropy of SARS-Cov2 goes DOWN, in contrast to the thermodynamic entropy of an isolated system, which is going up.
So the information entropy, which is a measure of… randomness? - Is going down, which means it’s becoming less random with time? The opposite of information “heat death”?
That’s my impression as well. They keep talking about entropy decreasing which would be backwards from our expectations in a closed thermodynamic system.
I’m also getting an impression that it is measuring sequence variation. In my own analyses of the SARS-CoV-2 data there is a big drop off in sequence variation within the population as Delta and Omicron swept through, especially with Omicron. A pair-wise comparison within any 1 month window will show an increase of variation for a time, and then a quick drop off as a new strain sweeps through.
From my own poor undestanding of thermodynamics, a lack of variation would seem to be analogous to a highly order system with entropy increasing as more sites are mutated.
No. The information entropy is measured for each individual sequence independently.
I got involved in this discussion precisely to differentiate those possibilities (information entropy of a sequence vs information entropy of a population) for myself. I was able to precisely reproduce the information entropy values in the table in the paper just be calculating the entropy the distribution of nucleotides in a given sequence. (It is seriously ~5 lines of R code to reproduce the table and the two figures.)
What does it all “mean”? The SARS-CoV-2 reference genome from 2019 does not have a perfectly uniform distribution of nucleotides; it is AU enriched. That’s why the information entropy of that sequence is slightly less than 2, the theoretical maximum for a fully uniform distribution. Over time, it has tended towards a slightly more AU enriched state, taking the distribution further from perfectly uniform and lowering the information entropy slightly more.
That could be because the mutation probabilities in the human host are a little different than the previous host, so that the steady state of the Markov process is skewed more towards AU in humans. That could be because the codon usage in humans favors AU a little more than the previous host’s codon usage. That could be because the adaptive mutations for evading human immunity happen to involve As and Us. That could be because the poly A tail of the RNA genome is getting longer. Some combination of all of the above is also possible.
Now, we could also measure the information entropy of the population. Then we would expect what you describe: selection would decrease sequence diversity and decrease information entropy, while mutation would increase sequence diversity and increase information entropy of the population. But that’s not what was done in the paper.
Let’s sort this out, because Vopson isn’t even looking at sequence information. He is taking the sequence as a sample of single nucleotides. I’m pretty sure we are thinking the same thing, but the terminology here here is dreadful/
That’s interesting. However, immediately after the emergence of this initial Omicron variants, a quite impressive drop in entropy is observed. Do you how this drop could be explain?
It seems to me that no such viruses with entropy at or near the original value is observed for the most recent sequences. Do you agree?
Gil, don’t let the graph axis fool you - the drop is in the 3rd and 4th decimal place. This is trivial, especially since the SIE is so close to the theoretical maximum. I agree it is interesting, but as far as I can tell it is biologically irrelevant.
Made it more visible here. The range between min and max entropy seems to expand and contract over time, but it’s true that there’s almost always some few sequences near the initial value.
On a hunch this mysterious trend might just be random, I set up a simple simulation on a spreadsheet. I started with 40 “nucleotides” numbered 0,1,2,3, with 10 of each. From that I calculate the SIE in the same way as Vopson, which always starts at the maximum value of 2.0.
Next I randomly select on nucleotide and mutate it randomly to a different value, recalculate SIE. Rinse and repeat.
Plotted here is an SIE random walk sequence for some 400+ sequential mutations. Patterns vary, but this example is fairly common, showing a series of downward trends followed by jumps upwards. Later (maybe tomorrow) I’ll put this up on Google Drive for everyone to try.
I went through several possible explanations a few posts above. And as @Dan_Eastwood pointed out, this is a small change (0.03%) and consistent with a random walk. Although ‘quite impressive’ is in the eye of the beholder, I suppose.
Nope, although I can understand why it is not apparent on my chart. @Rumraket did a nice job improving the visibility (thanks!), and in case that is not clear enough here’s a chart of the weekly max with a horizontal line at the value of the reference sequence from the paper.
The surprising result obtained here has massive implications for future developments in genomic research, evolutionary biology, computing, big data, physics, and cosmology.
Implication for cosmology from SARS-CoV-2 variants??? I’m starting to settle on woo.
Two caveats.
The reference RNA sequence of the SARS-CoV-2, representing a sample of the virus collected early in the pandemic in Wuhan, China in December 2019 has 29 903 nucleotides, so N = 29 903. For this reference sequence, we computed the Shannon information entropy using relation
One, if one extrapolates backwards, do we find the sequence as it existed prior to the species jump? How does the Wuhan alpha case constitute a baseline? Is it arbitrary?
The observed correlation between the information entropy and the time dynamics of the genetic mutations is truly unique, because it reconfirms the second law of infodynamics, but it also points to a possible deterministic approach to genetic mutations, currently believed to be just random events. The existence of an entopic force that governs genetic mutations instead of randomness is very powerful and it could lead to the future development of predictive algorithms for genetic mutations before they occur.
[Mod edit to add blockquote just above - Dan]
Two, it is likely near all the possible genetic mutations have probably already occurred, just not together with other favorable mutations or in somebody who transmitted it. Full infection means entails a viral swarm even with a dominant strain. At least one successful prediction of a specific mutation has been made on the basis of selection, not some second law of, uh, infodynamics.
This is mutation only? A population that had mutation making new alleles that then were either beneficial or deleterious would occasionally show “selective sweeps” that greatly reduce variability in the population. It then might build up again, gradually.
Yes, mutation only. SIE isn’t measuring sequence variability, but my intuition is that selective sweeps should perturb the SIE random walk. Still no useful interpretation here.
Because it suggests that the genetic changes that has accumulated with time in the SARS2 and H1N1 genomes are more a product of thermodynamics than selection.
Maybe he thinks AU bases are invisibly deleterious and so the putative slow creep of AU bases in SARS-Cov2, and in H1N1, he takes to imply they’re suffering “Genetic Entropy”?