There’s a time or two when I nearly kicked the bucket.
In my day the wastebaskets that collected the punched holes from the punch card stations and the paper-tape rolls were called the confetti bins. (At least some of the computer staff called them that. Students would occasionally raid them for use in parades or party pranks.)
For example: there are many information storage devices that depend solely on the positioning of their components (abacuses, cricket scoreboards, fridge magnet letters, etch-a-sketches, xkcd’s rocks in the desert, etc), meaning that the mass of a bead on a wire not only differs depending on its position on the wire, but also (since people use and erase abaci in different ways) depending on who’s looking at it and from which side.
There are ways of retrieving data from ‘erased’ magnetic media. The only way to completely eradicate data from magnetic media is to overwrite it multiple times with random values. Since those random values are themselves in a sense information, there isn’t really any such thing as an ‘erased state’ which doesn’t contain information. The same is effectively true of many other information storage devices where the ‘erased state’ is just one possible arrangement out of many.
The number of bits of information stored on a device can vary depending on external factors such as encoding methods, encryption, compression, language etc. Since it is possible for the exact same arrangement of data to have different meanings and hence different numbers of bits depending on both context and the observer, a data storage device would have to have many different mass values all at the same time.
There are data storage devices such as punch cards and ticker tape that decrease in mass when data is stored on them, and increase in mass when that data is erased (by taping over the holes).
Having said that, there is a genuine issue of the minimum energy required to make a change to stored information. But although that could imply a change in mass, (AIUI) that change would be a loss through heat dissipation whether adding or erasing information. Though I am not a particle physics expert either.
Note that the decrease in SIE observed by the authors for SARS2 was also observed by Sanford and Carter for H1N1. So, as the authors argue, it may very well be a general law that the SIE of genetic systems tend to decrease with time. Given the phenomenon of mutational biases, this isn’t surprising.
OR it could be an expected effect in (for example) viruses that have recently jumped species. I suspect the answer may exist in the literature, somewhere.
So, as the authors argue, it may very well be a general law that the SIE of genetic systems tend to decrease with time.
Recall this is the information in the distribution of single nucleotides, NOT the information in sequences of nucleotides. This is a very different proposition, and IIRC not at all what Sanford and Carter propose.
Where did they measure SIE? The impression I got from Sanford’s work is that he predicted a decrease in fitness to the point of extinction. I don’t remember him using any quantitative measurement of information.
I note that they did nothing of the sort, as they deceptively omitted the vast majority of H1N1 data and grossly misrepresented the H1N1 subtype as a lineage. Once again, IDcreationist misrepresentations are so basic that they can be seen by anyone with a Wikipedia-level understanding.
This misrepresentation is ludicrous because isolates of the H1N1 subtype only necessarily are related in two of the eight genome segments; therefore, they cannot possibly be a strain or a lineage. Perceptive readers can figure out which, but I will highlight the two for those who cannot:
The entire Influenza A virus genome is 13,588 bases long and is contained on eight RNA segments that code for at least 10 but up to 14 proteins, depending on the strain. The relevance or presence of alternate gene products can vary:[29]
Segment 1 encodes RNA polymerase subunit (PB2).
Segment 2 encodes RNA polymerase subunit (PB1) and the PB1-F2 protein, which induces cell death, by using different reading frames from the same RNA segment.
Segment 3 encodes RNA polymerase subunit ¶ and the PA-X protein, which has a role in host transcription shutoff.[30]
Segment 4 encodes HA (hemagglutinin). About 500 molecules of hemagglutinin are needed to make one virion. HA determines the extent and severity of a viral infection in a host organism.
Segment 5 encodes NP, which is a nucleoprotein.
Segment 6 encodes NA (neuraminidase). About 100 molecules of neuraminidase are needed to make one virion.
Segment 7 encodes two matrix proteins (M1 and M2) by using different reading frames from the same RNA segment. About 3,000 matrix protein molecules are needed to make one virion.
Segment 8 encodes two distinct non-structural proteins (NS1 and NEP) by using different reading frames from the same RNA segment.
Not only codon bias, but also the relative percentage changes in the four nucleotides with time. See fig3 of the piece below:
Clearly, the change in nucleotide composition within the H1N1 genome, with the increase frequency of A and U and the corresponding decrease of G and C reflects an overall decrease if SIE with time. And as the authors argue, the genetic changes that has accumulated with time in the H1N1 genome is more a product of thermodynamics than selection.
I can think of plausible reasons, but I don’t know. I don’t think you know either. I was just glancing at a paper about the accumulation of “host specific” mutations, which is just the sort of thing which MIGHT cause the trend after jumping species. Without a better understanding it’s not appropriate to draw conclusions one way or the other.
You seem to be under the impression these are two different things. It’s two ways of saying the same thing.
As far as I am aware (and I would happily be corrected) the relative frequency of single nucleotides (ATCG) is not biologically relevant. so long as the proportions stay in a range of 0.3 to 0.7 there is little loss (<16%) in terms of coding efficiency for long strings of codons. Certainly nothing that can’t be made up with a slightly longer string.
IF mutational bias were driving these proportions to 0 or 1, THEN they should have reached that state (0 or 1) long ago. I suspect there are other mechanisms at work which prevent this from happening, and it’s probably something very basic, if I only knew what to look for.
I was reluctant to agree to this based on just the artificial sample I limited myself to previously. So I took the next step of download all available full genome sequences of SARS-CoV-2 from NCBI. After filtering ones with absent/unclear collection dates and ones with sequence characters other than A, C, G, and T, that yields just shy of 800,000 sequences. Here’s the entropy vs time chart:
The light gray points are the individual sequences. There are so many of them that it’s impossible to assess density that might impact trends. So the darker grey circles are averages over individual Pango lineages, scaled to the number of sequences. The red line is a trend line if we require a linear model; the blue line is a generalized additive model fit. That line makes it more clear that while there is an overall downward trend, the initial Omicron variants (early 2022) represented an increase of entropy. So it’s not clear to me that decreases are inevitable. Further, as you can see from the light dots, at any given time there are always viruses with entropy at or near the original value. So if there is any sort of detrimental impact of decreasing entropy in the distribution of nucleotides, there are options with higher entropy available to selection.
OK, you are right on that point; it’s not the same thing.
BUT the OP paper is still measuring the distribution of nucleotides, not codons so the part you are right about just isn’t relevant.
NOW take a look at sentences 2 and 3 from the paper you just cited …
Different species have consistent and characteristic codon biases. Codon bias varies not only with species, family or group within kingdom, but also between the genes within an organism. Codon usage bias has evolved through mutation, natural selection, and genetic drift in various organisms.
This is just the sort of thing I thought we might find if we looked.
@Giltil I urge you to take a hard look at exactly what it is you are right about. Based on the article you cited, some mutations will change the third codon to a more stable configuration, decreasing SIE without changing the amino acid, and therefore not changing the genetic information encoded.
@AndyWalsh Extrapolating the linear trend backwards (Yes I know that is a bad idea) then a little more than 100 ago 2 bits encoded more than 2 bits of information. This isn’t possible. We also should not observe nearly 2 bits of information if this is a long term trend. Therefore, this can’t be a long term trend.
Yes, I’d agree that since the values are this close to the analytical maximum, this can’t be a long-term trend. Of course, to different people that might imply widely different conclusions, especially in the case of this virus.
I was poking around the NCBI virus database a little to see what options there might be for looking at the trends in other viruses. As you might imagine, SARS-CoV-2 is vastly overrepresented, accounting for ~60% of all the virus sequences available. And if we want to look for longer-term patterns, we’re getting into eras of different sequencing technology and then into sequencing of preserved samples. At some point, a more careful analysis is required even for casual exploration. Not sure if I’m up for that.