Uncommon Descent Serving The Intelligent Design Community

Human and chimp DNA: They really are about 98% similar

Categories
Intelligent Design
Share
Facebook
Twitter/X
LinkedIn
Flipboard
Print
Email

A few days ago, scientist and young-earth creationist Dr. Jay Wile wrote a post on his Proslogion blog, in which he reported that Dr. Jeff Tomkins had abandoned his claim that human and chimpanzee DNA are only about 70% similar, in favor of a revised figure of 88%. But even that figure is too low, according to the man who spotted the original flaw in Dr. Tomkins’s work.

Dr. Wile reports:

More than two years ago, Dr. Jeffrey P. Tomkins, a former director of the Clemson University Genomics Institute, performed a detailed, chromosome-by-chromosome comparison of human and chimpanzee DNA using a widely-recognized computer program known as BLAST. His analysis indicated that, on average, human and chimpanzee DNA are only about 70% similar. This is far, far, below the 95-99% numbers that are commonly cited by evolutionists, so once I read the study, I wrote a summary of it. Well, Dr. Tomkins has done a new study, and it invalidates the one he did two years ago.

The new study was done because last year, a computer programmer of financial trading algorithms (Glenn Williamson) discovered a bug in the BLAST algorithm that Tomkins used. This bug caused the program to ignore certain matches that should have been identified, which led to an artificially low similarity between the two genomes.

Here is what Glenn Williamson has to say about himself:

Yeah – 36 year old, stay-at-home father of four – including triplets, ha! 🙂

I don’t have any formal qualifications in genetics, or anything biological for that matter. I have a bachelors degree in computing science (i.e. programming) from the University of Technology in Sydney. Started my career as a programmer, but transitioned into derivatives trading, which is a lot more fun…

And for what it’s worth, I believe that my paper is more of a computing science paper than a genetics paper. It’s more my area of expertise than Jeff Tomkins’ area.

Glenn Williamson’s detailed takedown of Dr. Tomkins’s 70% similarity figure can be accessed here. Dr. Tomkins claims he submitted his paper to the creationist publication, Answers Research Journal, but it was never published. Here’s an excerpt from the paper (emphasis mine – VJT):

In this paper I carefully reproduce a subset of Dr Tomkins’ results, and show clearly and unambiguously that Dr Tomkins has fallen victim to a serious bug in the software used to obtain his results. It is this bug that causes Dr Tomkins to report the erroneous figure of 70% similarity. After correcting for both the effects of this bug and some non-trivial errors in Dr Tomkins’ methodology, I report an overall similarity of 96.90% with a standard error of ±0.21%. This figure includes indels, and the result is largely in line with the secular scientific consensus.

What happened next? Dr. Wile takes up the story:

As a result, Dr. Tomkins redid his study, using the one version of BLAST that did not contain the bug. His results are shown above… The overall similarity between the human and chimpanzee genomes was 88%.

In an update at the top of his post, Dr. Wile now admits to having cold feet, even about the revised 88% figure:

Based on comments below by Glenn (who is mentioned in the article) and Aceofspades25, there are questions regarding the analysis used in Dr. Tomkins’s study, upon which this article is based. Until Dr. Tomkins addresses these questions, it is best to be skeptical of his 88% similarity figure.

So what was wrong with Dr. Tomkins’s new study? I’ll let Glenn Williamson explain (emphasis mine – VJT):

October 16, 2015 4:22 pm

As I’ve said many times, if there is a single base pair indel in the middle of a 300bp sequence, Tomkins will say this is a 50% match.

Tomkins is most certainly aware of this, yet he chose to publish it. I think that says pretty much everything.

Another commenter named Aceofspades25 has this to add (emphases mine – VJT):

October 16, 2015 2:13 pm

The other obvious thing that Thompkins hasn’t dealt with in his BLASTN analysis, I talk about here.

There are a few cases where no match will be found because this entire sequence appears de-novo in Chimpanzees as the result of a single mutation (e.g. a novel transposable element – see here) or because humans have had a large deletion which other primates don’t. Deletions like this also likely occurred in a single mutation – see here

Thompkins (sic) would count both of these as being a 0% match (or 600 effective mutations if the sequences he was searching for were 300bp each). In reality, these probably represent just 2 mutations.

I’ll let Glenn Williamson have the final word (emphases mine – VJT):

Thanks Ace, for letting me know about this post. I reiterate here a few things about my (unpublished!) paper, and about Tomkins’ new paper.

The first thing is that he uses the “ungapped” parameter in his BLAST comparisons. As I’ve written in a few other places now, using this parameter, and calculating results in the way that Tomkins does is entirely disingenuous. If you are comparing two 300bp sequences, and one of those sequences has a single indel smack bang in the middle, Tomkins counts this as the sequences being only 50% identical.

I’ve told him at least twice that he cannot use ungapped and then calculate the result in this way. He can do one of two things:

1. Use ungapped, which ignores indels and therefore he can only report the substitution rate. (emphasis mine – VJT)If he did this, he would get a result of around 98.8%.

2. Allow gaps, and – this is what he fails to mention in his paper – get a result of around 96.9%. And this is using a very conservative method of calculation as well, since it counts a 50bp indel as having the same weight as 50 individual mutations. If you counted a 50bp indel as a single event (which it probably was), then the overall result would be pushed up towards 98%, which is the figure usually thrown around anyway.

In a comment dated 14 August 2015 (at 03:44) on an article titled, Chimp and Human DNA vs “Sophisticated Nonsense”! on a blog called Marmotism, Glenn Williamson adds:

I’ve actually written a paper on Tomkins’ 70% result, and have _ATTEMPTED_ to get it published in Answers Research Journal. Obviously they are not having a bar of it – Tomkins is the sole peer-reviewer, and he is currently refusing to provide any critique of my work – he has been silent for 8 months, while the ball is in his court ..

See the paper here:

https://www.dropbox.com/sh/dm2lgg0l93sjayv/AAATnWSJdER53EYEYZvcgiwma?dl=0

It’s the two PDFs ..

Dr. Tomkins’s latest article in Answers Research Journal (October 7, 2015) acknowledges Williamson’s work in a single sentence:

As of 2013, the issue of overall genome similarity between chimpanzee and humans seemed to be about 70% based on five different reports, three of which were based on actual data analyses. However, in 2014 , a computer programmer of financial trading algorithms discovered an apparent bug in the BLASTN algorithm and notified this author of the situation (Glenn Williamson, Tibra Capital, personal communication).

So, a submitted article only counts as a “personal communication”? Perhaps Dr. Tomkins needs to be a little more up-front about giving credit where credit is due, and acknowledging his mistakes. At any rate, the ball is definitely in his court, and his latest 88% similarity figure warrants skepticism. I have to say that Dr. Tomkins’s methodology sounds rather suspicious to me.

What do readers think?

Comments
Not looking to join, I was just curious what you'd say. I know my point didn't really add anything to the topic at hand as you guys are comparing different species. But anyways gametes always have half the genome because they are produced via meiosis if that was what you weren't sure about.Alicia Cartelli
November 6, 2015
November
11
Nov
6
06
2015
01:15 PM
1
01
15
PM
PDT
4. Finally, some taxonomically restricted cases of varying c-values (like onions) may be legitimate cases of junk through runaway transposon duplication. Perhaps these clades are simply more prone to transposition. However, I don’t think these can be used to then argue that most genomes in most organisms are junk. Casey Luskin has been trying out this line of argument also, hoping that genome size variation can be brushed off as a few exceptional groups with bloated genomes. The problem is that it relies on wishful thinking rather than facts. Genome size variation is ubiquitous, and the size variations are over several orders of magnitude. There are plants and animals with small genomes and little that looks like junk (e.g. fruit flies and Arabidopsis), but they are exceptions. Unnecessarily huge genomes, with pretty huge variations in genome size, are normal and very common. Download T. Ryan Gregory's genome size databases and plot the variation. I have. It's everywhere -- yet virtually all of these animals and plants have roughly the same number of genes, ~15000-30000.NickMatzke_UD
November 6, 2015
November
11
Nov
6
06
2015
01:00 PM
1
01
00
PM
PDT
@Alicia Cartelli, Welcome to our friendly debate : ) Larger cells could either increase transcription and copy number. But having increased copy number allows a faster response to changes. I use the term "copy number" extremely loosely, as not all copies need be the same sequence, per the principles of redundancy I outlined above. Granted this is speculative. But I think we have enough design-compatible reasons for varying C-values that we can't say junk DNA is the only possible explanation for varying C-values. As you know, sperm and oocytes have the same amount of genetic information because cells within the same organism have the same haploid genome size. Perhaps due to a constraint of cell differentiation itself?--I'm not up to speed on those details.JoeCoder
November 6, 2015
November
11
Nov
6
06
2015
12:41 PM
12
12
41
PM
PDT
I reject common descent. I also question the ages of life on earth, due to carbon-14 and soft tissue being found throughout the fossil record. Genetic entropy adds another point to this argument. But the ages of the rocks around the fossils argues against it. So I am unable to argue for a definitive position one way or the other, which is why I've avoided that topic so far. I think the earth itself and universe are billions of years old and do not question the conventional dates there. But saying that most DNA must be functionally neutral because of evolution, requires genomes to be created by unguided evolution. Behe for example, who accepts common descent, still thinks an intelligence is necessary to inject larges amounts of information. As I said above, I think junk DNA is a good explanation for taxonomically restricted cases of large C-value differences. Some onion genomes probably are mostly junk DNA. But that does not mean that most DNA in most organisms is junk. I also fully agree that larger genomes are more likely to accumulate excess DNA over time, due to weaker selection and because a deletion is more likely to be deleterious than an insertion. My point is that there are equally good explanations for C-value differences under design. Therefore C-value variation cannot be used to argue for junk DNA. One thing to keep in mind is that much functional redundancy is not due to duplications. From a paper in yeast:
I used functional genomics data from the yeast Saccharomyces cerevisiae to test the hypotheses related to the following: if gene duplications are mostly responsible for robustness, then a correlation is expected between the similarity of two duplicated genes and the effect of mutations in one of these genes. My results demonstrate that interactions among unrelated genes are the major cause of robustness against mutations.
This is good design, because having backup systems that operate in very different ways makes them less prone to the factors that made primary systems fail. So because of their uniqueness, these redundant systems should also be counted among the functional nucleotides that selection must maintain. They can't be easily replaced simply through more duplications of existing systems.JoeCoder
November 6, 2015
November
11
Nov
6
06
2015
12:14 PM
12
12
14
PM
PDT
"There seems to be a good correlation between cell size and genome size. The idea being that larger cells require more RNA’s." That doesn't make any sense, if more RNAs are needed, cells would increase transcription, not genome size. (Usually, I guess polytene chromosomes in Drosophila would be an exception) The smallest and largest human cells are the sperm and oocyte respectively, and both have the same amount of genetic information (half that of a typical cell.)Alicia Cartelli
November 6, 2015
November
11
Nov
6
06
2015
11:59 AM
11
11
59
AM
PDT
JoeCoder: That depends on the assumption that evolution is what created our genomes, which is the very thing I am contesting here. You are? Do you accept common descent? JoeCoder: More cell types would generally require more information. So onions have more cell types than humans. JoeCoder: There seems to be a good correlation between cell size and genome size. The idea being that larger cells require more RNA’s. Mutations that might have been deleterious are no longer so. Redundancy, such as polyploidism, results in huge flexibility in response to mutation. JoeCoder: The argument could be rephrased as “Most of the genome bust be functionally neutral because evolution could create no better”. Fast reproducing organisms, such as bacteria, have more compact genomes. Slow reproducing organisms, such as mammals, tend to accumulate genomic material, which can also have an evolutionary advantage. Retroviruses may be an important component of this process. JoeCoder: This paper in the Journal of Creation estimates between 1200 and 1.5 million years to extinction based on the parameters used. It doesn’t look like something easy to pin down. From a thousand to a million years. Quite the range. In any case, the average lifespan of a species is a few million years, but it's not as if they don't leave descendants — they often do.Zachriel
November 6, 2015
November
11
Nov
6
06
2015
11:46 AM
11
11
46
AM
PDT
Zachriel: Pheasant and Mattick.. are including spacers, as well as other functions which are not sequence specific or otherwise relaxed in terms of sequence. Correct. But this is still strong evidence against the claim that only 1 to 2% is nucleotide-specific functional. The nucleic acids that make up RNA products connect to each other in very specific ways, which force RNA molecules to twist and loop into a variety of complicated 3D structures. John Mattick was quoted in a press release from one of the papers I cited above:
We believe that RNA structures probably operate in a similar way to proteins, which are composed of structural domains that assemble together to give the protein a function.
Even if only one fourth of the 85% transcribed RNA nucleotides are specific, that's still a number 10 to 20 times higher than the 1 to 2% we were talking about above, and twice as much as the number the Mendel team is using. Zachriel: the C-value paradox, a.k.a. onion text. I've read much about what Dr. Moran, Graur, and others have written about C-values and I don't find it convincing. I think there are good reasons for varying C-values that are still entirely compatible with design. 1. On a very large scale, genome size is correlated with number of cell types. See figure 2B from this paper. More cell types would generally require more information. Although note that they only plot protostomia and deuterostomia animals as two points and we can't see differences between individual clades. 2. Take a look at this composite image I put together from this angiosperm paper and also a paper from T. Ryan Gregory in 2001. There seems to be a good correlation between cell size and genome size. The idea being that larger cells require more RNA's. One textbook makes an analogy:
The situation is like that of a car factory aiming for a steady output of cars: engines, wheels and doors must be made at the same rate; if overall output is to be increased the number of each must be increased by the same proportion. Moreover, if each robot, machine tool, and operative is already working at maximal rates, one can increase output only by increasing the number of assembly lines
Although I'm sure it's more complex than that, since I don't think the number of protein coding genes increases. I have also seen papers reporting correlations between c-value and longevity, but I am hesitant to cite them because I have not yet read them. It makes sense that more DNA would bring more redundant systems against failure, leading to longevity. 3. Genome size differences may represent tradeoffs between different forms of data storage. One ENCODE researcher commented on reddit:
Organism introduce genetic variation in different ways. For instance, in Drosophila, a 100 kilobase gene (DSCAM) encode thousands of different proteins through a complex alternate splicing mechanism. One could envision copying each of these transcripts - without the alternate splicing - into the genome thus increasing the size of the genome by 10 million bases, or roughly 10%, but not changing the complexity at all.
This is similar to how our own compression algorithms operate--frequently used sequences are stored only once and re-referenced, as opposed to the same information being stored multiple times, at perhaps higher fidelity. A png image is not 90% junk just because it's 10 times larger than an equivalent lossy jpeg. 4. Finally, some taxonomically restricted cases of varying c-values (like onions) may be legitimate cases of junk through runaway transposon duplication. Perhaps these clades are simply more prone to transposition. However, I don't think these can be used to then argue that most genomes in most organisms are junk. Our own genomes are very close to the same size as all other mammals, so I don't think we or they are subject to excess transposition. Zachriel: [a highly neutral genome is] also supported by many experiments showing consistency in rates of neutral evolution in many different taxa and many different parts of the genome That depends on the assumption that evolution is what created our genomes, which is the very thing I am contesting here. The argument could be rephrased as "Most of the genome bust be functionally neutral because evolution could create no better". Zachriel: Sanford explicitly claims Mendel’s Accountant shows that the history of biological descent cannot be measured in millions of years. Perhaps Sanford is overstating his case and not taking redundancy into account? This paper in the Journal of Creation estimates between 1200 and 1.5 million years to extinction based on the parameters used. It doesn't look like something easy to pin down.JoeCoder
November 6, 2015
November
11
Nov
6
06
2015
10:20 AM
10
10
20
AM
PDT
JoeCoder: Pheasant and Mattick Looking at Pheasant and Mattick, they are including spacers, as well as other functions which are not sequence specific or otherwise relaxed in terms of sequence. The general claim that most of the genome is non-functional in terms of sequence is supported by the C-value paradox, a.k.a. onion text. It's also supported by many experiments showing consistency in rates of neutral evolution in many different taxa and many different parts of the genome. Of course, much of the genome may have function which is relaxed in terms of sequence. JoeCoder: Given the levels of redundancy we see in genomes, and that the worst mutations are still selected against, I don’t think thousands of generations is too many. Unless you are able to quantify it? Sanford explicitly claims Mendel's Accountant shows that the history of biological descent cannot be measured in millions of years.Zachriel
November 6, 2015
November
11
Nov
6
06
2015
07:16 AM
7
07
16
AM
PDT
Larry Moran's view requires that mutations falling within 98% of the genome be functionally neutral. I don't see how that's possible given recent research. A few examples: 1. Pheasant and Mattick, 2007 estimated that "the functional portion of the genome may exceed 20%" based on conserved sequences and those under purifying selection. 2. ENCODE 2012 estimated "that at a minimum 20% (17% from protein binding and 2.9% protein coding gene exons) of the genome participates in... specific functions, with the likely figure significantly higher" 3. ENCODE 2014 estimated that "12-15%" of DNA is evolutionarily conserved. 4. Martin Smith et al, 2013 estimated that at least 13.6 - 30% of DNA is conserved in its RNA structure. This means that animals with different DNA sequences still produce RNA molecules that have the same shape. They report of their conserved RNA's that "88% of which fall outside any known sequence-constrained element, suggesting that a large proportion of the mammalian genome is functional." 5. Additionally, we know at least 85.2% of the genome is transcribed, and John Mattick says there are usually functional consequences in disrupting those transcripts:
where tested, these noncoding RNAs usually show evidence of biological function in different developmental and disease contexts, with, by our estimate, hundreds of validated cases already published and many more en route, which is a big enough subset to draw broader conclusions about the likely functionality of the rest.
Granted conserved sequences rely on the assumption of common descent, but that and protein binding alone would make the deleterious rate at least 20 mutations per generation. And all of those are lower-bound measurements of the number of non-netural nucleotides. I don't see how it's reasonable to believe only 1 to 2% of nucleotides are non-neutral? Zachriel: humans have been around for thousands of generations and are still happily making babies by the millions. Supposing you agreed genetic entropy were true, how many generations do you think would be too many before humans go extinct? Given the levels of redundancy we see in genomes, and that the worst mutations are still selected against, I don't think thousands of generations is too many. Unless you are able to quantify it?JoeCoder
November 5, 2015
November
11
Nov
5
05
2015
12:08 PM
12
12
08
PM
PDT
JoeCoder: In using a deleterious rate of 10, the Mendel authors are already assuming about 90% of mutations are neutral. You cited Larry Moran. He suggests "there are fewer than 2 detrimental mutations per generation and this is an acceptable genetic load." http://sandwalk.blogspot.com/2014/04/a-creationist-tries-to-understand.html JoeCoder: I don’t think there are any reasons left for me to believe that Mendel is not an accurate-enough simulation of genetic load in humans. Sure there are. We have strong evidence that humans have been around for thousands of generations and are still happily making babies by the millions.Zachriel
November 5, 2015
November
11
Nov
5
05
2015
08:07 AM
8
08
07
AM
PDT
Zachriel: it’s presented as an accurate model of biological evolution. That’s where it fails. In using a deleterious rate of 10, the Mendel authors are already assuming about 90% of mutations are neutral. I don't think there are any reasons left for me to believe that Mendel is not an accurate-enough simulation of genetic load in humans. Do you have any other ideas?JoeCoder
November 5, 2015
November
11
Nov
5
05
2015
07:55 AM
7
07
55
AM
PDT
JoeCoder: Granted 20,000 still isn’t a very large population, but that graph makes it look like there won’t be much benefit to going larger and computational resources make it difficult. The effective population size of humans is only about 100,000 due to a bottleneck in the recent past. JoeCoder: Do you likewise disagree with Larry Moran’s view that anything more than a small number of deleterious mutations will lead to extinction? Sure. However, most mutations are neutral. JoeCoder: Mendel’s Accountant is a simulation and does not model reality perfectly. That's fine. All models are wrong, but some are useful. If it were presented as a general model, then it might have some value, but it's presented as an accurate model of biological evolution. That's where it fails.Zachriel
November 5, 2015
November
11
Nov
5
05
2015
07:47 AM
7
07
47
AM
PDT
As you know, John Sanford is a young earth creationist and uses genetic entropy as an argument against long ages. However, I'm not sure genetic entropy would prevent humans from existing for thousands of generations. Genomes are very redundant with one system kicking in to compensate for when another is knocked out. The ENCODE team noted the difficulty this caused in testing for functions via knockout:
Loss-of-function tests can also be buffered by functional redundancy, such that double or triple disruptions are required for a phenotypic consequence. Consistent with redundant, contextual, or subtle functions, the deletion of large and highly conserved genomic segments sometimes has no discernible organismal phenotype, and seemingly debilitating mutations in genes thought to be indispensible have been found in the human population
Zachriel: Fitness decline can certainly occur, but typically with much smaller population sizes than claimed by Sanford. Take a look at Figure 9 from the 2011 Can Purifying Natural Selection Preserve Biological Information paper. The Mendel authors write:
Figure 9 shows the effect of population size on percent retention after 10,000 generations. Within this limited amount of time, there was only a trivial advantage in having population sizes greater than 5,000. With a population size of 5,000, the rate of mutation accumulation was 89.38%. Doubling the population size to 10,000 resulted in 89.05% accumulation, and doubling the population size again to 20,000 resulted in no further improvement (89.05% accumulation). It is clear that the advantage of larger population size beyond 1000 is only realized in deep time, which seems to imply the need for some type of very long-term selection equilibrium, which may be conceptually problematic.
Granted 20,000 still isn't a very large population, but that graph makes it look like there won't be much benefit to going larger and computational resources make it difficult. We can try simulating larger populations if you think it would help. Zachriel: There’s clearly something wrong with his model, actually many things. Do you likewise disagree with Larry Moran's view that anything more than a small number of deleterious mutations will lead to extinction? You're departing from the majority of biologists knowledgeable in population genetics who argue against ID. That's OK because I sometimes also take minority positions. But I do want to point that out. Mendel's Accountant is a simulation and does not model reality perfectly. But if you want to win me to your position then I will need a reason why the results are not reliable enough.JoeCoder
November 5, 2015
November
11
Nov
5
05
2015
07:34 AM
7
07
34
AM
PDT
JoeCoder: The inevitability of fitness decline also seems consistent with the claims of Susumu Ohno, Larry Moran, and others per comment #191 Fitness decline can certainly occur, but typically with much smaller population sizes than claimed by Sanford. As already pointed out, humans have been around for thousands of generations, and are happily fecund. There's clearly something wrong with his model, actually many things.Zachriel
November 5, 2015
November
11
Nov
5
05
2015
06:48 AM
6
06
48
AM
PDT
Zachriel: Fitness in Mendel’s Accountant doesn’t necessarily translate into the selection coefficient. A selection coefficient of 0.1 means the organism will, on average, leave 10% more fertile offspring. (The use of sign may vary with context.) Right, but I don't think this will have any significant effect on the outcome when simulating the consequences of deleterious load. Mendel also has a "Fitness-dependent fecundity decline" option in the advanced selection parameters, but I have not studied it. Zachriel: http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2607448/ Ah, I recognize that paper from having it saved in my notes with a "need to read" tag. All the more reason I should do so. Thank you nonetheless. I will have to think more about the model for accidental death. I think you may be right that there is not a need for a "dumb luck" coefficient since fitness is equivalent to number of offspring. Still, per comment #214 I don't see a reason to doubt the results of Mendel's Accountant. The inevitability of fitness decline also seems consistent with the claims of Susumu Ohno, Larry Moran, and others per comment #191. That's one of the main reasons Larry Moran argues that so much DNA is junk--so that the deleterious mutation rate can be decreased far enough to avoid the problem.JoeCoder
November 4, 2015
November
11
Nov
4
04
2015
04:25 PM
4
04
25
PM
PDT
JoeCoder: I used probability selection and seeded a population of 2000 with 5 beneficial mutations each having a selective advantage of 0.1. Fitness in Mendel's Accountant doesn't necessarily translate into the selection coefficient. A selection coefficient of 0.1 means the organism will, on average, leave 10% more fertile offspring. (The use of sign may vary with context.) JoeCoder: Also, do you have a link so I can read how the 2s(Ne/N) formula is derived? Here's a review. If you want more details, you can follow the citations to Kimura's original diffusion model. http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2607448/ JoeCoder: I’m surprised it doesn’t have a “dumb luck” coefficient in there somewhere, that would vary according to how efficient selection is in various populations. Take the simple model of accidental death. Because death is random, as long as there are a substantial number of survivors, the survivors will have variations similar in proportion to the larger population. The survivors will still be subject to selection, and those survivors with a selective advantage will still tend to leave more offspring.Zachriel
November 4, 2015
November
11
Nov
4
04
2015
04:07 PM
4
04
07
PM
PDT
Zachriel, it looks like we have three options: 1. The formula you provided in comment #209 is wrong. 2. I somehow made a mistake in the parameters when running Mendel's Accountant in comment #136. I don't think I did. 3. There never was any bug in Mendel's Accountant to begin with, even back in 2009 when you first looked at it. Thoughts?JoeCoder
November 4, 2015
November
11
Nov
4
04
2015
04:01 PM
4
04
01
PM
PDT
Zachriel, https://uncommondescent.com/genetics/paul-giem-on-overlapping-genetic-codes/ Comments 48, 56, 58, 60 Does that refresh your memory?Mung
November 4, 2015
November
11
Nov
4
04
2015
03:54 PM
3
03
54
PM
PDT
Mung: Well, to be honest I was sort of hoping we could avoid all that. Sorry. Don't want to be a bother. We have looked at the code briefly in the discussion with JoeCoder. We have not studied the code in detail in several years. As there is a new version, some of the code is undoubtedly changed. However, the code concerning probability selection and division by random(0,1) is still there.Zachriel
November 4, 2015
November
11
Nov
4
04
2015
03:52 PM
3
03
52
PM
PDT
Mung:
Zachriel, are you intentionally ignoring my question @196? Did you or did you not tell me that the software had not changed? And how did you know that if you haven’t looked at it since 2009?
Zachriel:
Please provide a reference.
Well, to be honest I was sort of hoping we could avoid all that. But I do enjoy a challenge. May take some digging. We can start here though. Zachriel December 1, 2014 at 10:03 am:
Mendel’s Accountant has a demonstrable flaw which virtually eliminates the known empirical effects of selection.
Zachriel now:
We haven’t looked at the code in several years (2009).
Now even if you last looked at the code at the end of 2009, that's still five years.Mung
November 4, 2015
November
11
Nov
4
04
2015
03:36 PM
3
03
36
PM
PDT
Zachriel: If a variant has a selective advantage of 0.1, then it will fix 20% of the time. In my comment #136 above, I used probability selection and seeded a population of 2000 with 5 beneficial mutations each having a selective advantage of 0.1. The mutation rate was 0. 2 out of 5 of those beneficials fixed--40%. Doesn't that make even probability selection in Mendel's accountant twice as generous as your formula? Also, do you have a link so I can read how the 2s(Ne/N) formula is derived? I'm surprised it doesn't have a "dumb luck" coefficient in there somewhere, that would vary according to how efficient selection is in various populations. Maybe that somehow cancels out but I don't see how?JoeCoder
November 4, 2015
November
11
Nov
4
04
2015
03:14 PM
3
03
14
PM
PDT
JoeCoder: What is the correct rate? In a large population, the chance of fixation for a new beneficial mutation is about 2s.* If a variant has a selective advantage of 0.1, then it will fix 20% of the time. Neutral or near neutral variants will fix at a rate equal to their proportion in the population. If there is a new mutant in a population of a thousand, then the chance of fixation is one in a thousand. With probability selection, the result tends to resemble the latter rather than the former. Sure, luck is a factor, but in the long run, selection does play out. * 2s(Ne/N), where s is the selection coefficient, N is the population size, and Ne is the effective population size. Effective population varies, but is important in organisms with harem breeding strategies, or organisms that have experienced population bottlenecks.Zachriel
November 4, 2015
November
11
Nov
4
04
2015
02:30 PM
2
02
30
PM
PDT
That figure doesn’t include all reproduction, which includes miscarriages, stillborns, and those who die before their own chance to reproduce.
Right, I'm only counting those who survive and reproduce because that's how you calculate population growth. Assume that each woman has on average 6 offpsring (or 12 if you count failed pregnancies) and either 2.0 or 2.23 survive to reproduce, based on the strength of selection.
one wouldn’t expect 1% of the population to produce a hundred times the offspring by accident.
It makes them 100 times more likely to survive to the point of producing offspring at all. Not produce 100 times more offspring. So it's only used for ranking.
probability selection, small but significant selective advantages are lost at a rate higher than expected based on population genetics
What is the correct rate? I've been asking you for it through the last several comments now. I don't know it either but you are the one claiming Mendel is using the wrong rate : )JoeCoder
November 4, 2015
November
11
Nov
4
04
2015
02:04 PM
2
02
04
PM
PDT
Mung: Zachriel, are you intentionally ignoring my question @196? Please provide a reference. JoeCoder: With a parameter of 0, the distribution is such that 50% are less than 2, 90% of them are less than 10, and 99% of them are less than 100. If working fitness includes reproductive potential, then there's the obvious problem that one wouldn't expect 1% of the population to produce a hundred times the offspring by accident. Sure it could happen — all the superior stags got wiped out by a lava flow —, but it's very unlikely. If working fitness doesn't include reproductive potential, then the working fitness is only used for ranking. Even then, with probability selection, small but significant selective advantages are lost at a rate higher than expected based on population genetics. JoeCoder: Maybe he woos women with good poetry? That would be a superior phenotype then. JoeCoder: I don’t thinks 2.23 versus 2.0 reproducing offspring is a very big difference in terms of selection. That figure doesn't include all reproduction, which includes miscarriages, stillborns, and those who die before their own chance to reproduce. (About half of human conceptions end in miscarriage, often due to defects, the first step in natural selection, so the reproductive rate would be at least twice the population increase.) In any case, a low reproductive rate allows deleterious mutations to accumulate faster. Humans seem to be more than healthy enough to continue their existence into the foreseeable future. This is contrary to the claim made by Sanford. Truncated selection seems to work reasonably well, but because the method of randomization and the calculation of working fitness is non-biological, it's hard to make sense of it. You can't break it down to compare it to real examples. For instance, fitness doesn't seem equivalent to selection coefficient.Zachriel
November 4, 2015
November
11
Nov
4
04
2015
01:40 PM
1
01
40
PM
PDT
Zachriel, are you intentionally ignoring my question @196? Did you or did you not tell me that the software had not changed? And how did you know that if you haven't looked at it since 2009?Mung
November 4, 2015
November
11
Nov
4
04
2015
10:44 AM
10
10
44
AM
PDT
It also creates a strange skew of working fitness from small to infinity
With a parameter of 0, the distribution is such that 50% are less than 2, 90% of them are less than 10, and 99% of them are less than 100. randonmum() has to return a value less than 0.01 to go above 100. This seems realistic because every now and then super-unfit guy still has lots of kids. Maybe he woos women with good poetry? Mendel realistically makes the rareness of this event proportional to his unfitness. So it's a realistic distribution. I don't know if it's an absolutely correct distribution or how to determine that. I also don't know how to determine whether a parameter of 0 or 0.5 is more biologically realistic.
Yet humans survived a bottleneck with very low population, and then survived rapid expansion.
Under the genetic entropy model being proposed, by extrapolation humans would've had far fewer deleterious alleles back at that bottleneck. Thus the negative effect of such a bottleneck would be much less. Also how rapid of an expansion are you suggesting? Even the YECs who want to fit everything within only 4350 yeas need need 2.23 reproducing offspring per female to go from 6 people to a world population of 1.7 billion in the year 1900AD: N1 = N0 * e ^ (r*t) 1.7 billion people = 6 people * e^(r * 170 generations) r = 0.115 2r = 0.230, since you need two people to reproduce While 2 reproducing offspring per female is required to maintain constant population size. I don't thinks 2.23 versus 2.0 reproducing offspring is a very big difference in terms of selection.JoeCoder
November 4, 2015
November
11
Nov
4
04
2015
10:19 AM
10
10
19
AM
PDT
JoeCoder: Small variations in selection coefficients SHOULD be the most diluted (per Kimura), so I see no problem there. Sure, but even strong selection is strongly diluted when working fitness is derived by division with a random number between zero and one. It also creates a strange skew of working fitness, from small to infinity. JoeCoder: So I think grown is expected any time selective pressures decrease faster than the deleterious load increases.
J Sandord (2008): "When biologically realistic parameters are selected, Mendel shows consistently that genetic deterioration is an inevitable outcome of the processes of mutation and natural selection. The primary reason is that most deleterious mutations are too subtle to be detected and eliminated by natural selection and therefore accumulate steadily generation after generation and inexorably degrade fitness."
Yet humans survived a bottleneck with very low population, and then survived rapid expansion.Zachriel
November 4, 2015
November
11
Nov
4
04
2015
09:53 AM
9
09
53
AM
PDT
Division by random(1) is ill-considered, and doesn’t properly represent randomness in nature. As important, it dilutes most of the signal from selection, especially with small variations in selection coefficients.
I think we need a way to quantify what amount of randomness is biologically realistic. Without that I don't see us advancing past this point. I see selection as pretty random to begin with, but I am open to having my view amended in light of data. Small variations in selection coefficients SHOULD be the most diluted (per Kimura), so I see no problem there. I agree selection can keep bacterial genomes more efficient. I'm just sharing the data I've come across.
That paper led to claims that the human genome was melting down, which is odd considering their rapid expansion in population over the last few centuries.
Populations expand when the strength of selection is decreased. In the last few thousand years technological progress (beginning with farming) has allowed us to decrease our own selective pressures. So I think grown is expected any time selective pressures decrease faster than the deleterious load increases. Although decreased selective pressures also lead to faster deleterious accumulation. Out of curiosity, who is "we"? I don't care about names, I value my privacy, and I have no desire to dox you. An answer like "me and my flatmate" would be fine.JoeCoder
November 4, 2015
November
11
Nov
4
04
2015
08:33 AM
8
08
33
AM
PDT
JoeCoder: Mice and rabbits have less time and fewer cell divisions between generation, which likely gives them a lower mutation rate. Yes. Mammals have about the same mutation rate per year, not per generation. See Kumar & Subramanian, Mutation rates in mammalian genomes, PNAS 2002. JoeCoder: Generally it seems like the larger animals are the most prone to extinction. Sure, but that's usually due to other causes, at least until they reach too low a population to be sustainable. JoeCoder: "We used a bacterial system in which the fitness effects of a large number of defined single mutations in two ribosomal proteins were measured with high sensitivity… most mutations (120 out of 126) are weakly deleterious and the remaining ones are potentially neutral." Bacteria tend to have highly optimized genomes. The selection coefficient of these mutations are on the order of -0.005. However, consider that once these mutations have occurred, assuming they aren't purged by selection, they likely will undone by a future mutation, in which case they would be beneficial. In other words, a natural population will exhibit significant variation, with traits drifting about the maximum for each trait. JoeCoder: The 2007 paper from the Mendel team I linked above used a beneficial rate of 1% which I think is unrealistically high. But that paper used probability selection (the mostly random one). Our own efforts were in response to the 2007 paper. Division by random(1) is ill-considered, and doesn't properly represent randomness in nature. As important, it dilutes most of the signal from selection, especially with small variations in selection coefficients. That paper led to claims that the human genome was melting down, which is odd considering their rapid expansion in population over the last few centuries. Humans seem rather fecund, and hear-tell, they have resorted to artificial means to limit their reproduction. We're running some experiments with 0.5 partial truncation...Zachriel
November 4, 2015
November
11
Nov
4
04
2015
08:04 AM
8
08
04
AM
PDT
thank youMung
November 3, 2015
November
11
Nov
3
03
2015
07:47 PM
7
07
47
PM
PDT
1 2 3 4 9

Leave a Reply