Uncommon Descent Serving The Intelligent Design Community

Functional information defined

Categories
Intelligent Design
Share
Facebook
Twitter/X
LinkedIn
Flipboard
Print
Email

What is function? What is functional information? Can it be measured?

Let’s try to clarify those points a little.

Function is often a controversial concept. It is one of those things that everybody apparently understands, but nobody dares to define. So it happens that, as soon as you try to use that concept in some reasoning, your kind interlocutor immediately stops you at the beginning, with the following smart request: “Yes, but what is function? How can you define it?

So, I will try to define it.

A premise. As we are not debating philosophy, but empirical science, we need to remain adherent to what can be observed. So, in defining function, we must stick to what can be observed: objects and events, in a word facts.

That’s what I will do.

But as usual I will include, in my list of observables, conscious beings, and in particular humans. And all the observable processes which take place in their consciousness, including the subjective experiences of understanding and purpose. Those things cannot be defined other than as specific experiences which happen in a conscious being, and which we all understand because we observe them in ourselves.

That said, I will try to begin introducing two slightly different, but connected, concepts:

a) A function (for an object)

b) A functionality (in a material object)

I define a function for an object as follows:

a) If a conscious observer connects some observed object to some possible desired result which can be obtained using the object in a context, then we say that the conscious observer conceives of a function for that object.

b) If an object can objectively be used by a conscious observer to obtain some specific desired result in a certain context, according to the conceived function, then we say that the object has objective functionality, referred to the specific conceived function.

The purpose of this distinction should be clear, but I will state it explicitly just the same: a function is a conception of a conscious being, it does not exist  in the material world outside of us, but it does exist in our subjective experience. Objective functionalities, instead, are properties of material objects. But we need a conscious observer to connect an objective functionality to a consciously defined function.

Let’s make an example.

I am a conscious observer. At the beach, I see various stones. In my consciousness, I represent the desire to use a stone as a chopping tool to obtain a specific result (to chop some kind of food). And I choose one particular stone which seems to be good for that.

So we have:

a) The function: chopping food as desired. This is a conscious representation in the observer, connecting a specific stone to the desired result. The function is not in the stone, but in the observer’s consciousness.

b) The functionality in the chosen stone: that stone can be used to obtain the desired result.

So, what makes that stone “good” to obtain the result? Its properties.

First of all, being a stone. Then, being in some range of dimensions and form and hardness. Not every stone will do. If it is too big, or too small, or with the wrong form, etc., it cannot be used for my purpose.

But many of them will be good.

So, let’s imagine that we have 10^6 stones on that beach, and that we try to use each of them to chop some definite food, and we classify each stone for a binary result: good – not good, defining objectively how much and how well the food must be chopped to give a “good” result. And we count the good stones.

I call the total number of stones: the Search space.

I call the total number of good stones: the Target space

I call –log2 of the ratio Target space/Search space:  Functionally Specified Information (FSI) for that function in the system of all the stones I can find in that beach. It is expressed in bits, because we take -log2 of the number.

So, for example, if 10^4 stones on the beach are good, the FSI for that function in that system is –log2 of 10^-2, that is  6,64386 bits.

What does that mean? It means that one stone out of 100 is good, in the sense we have defined, and if we choose randomly one stone in that beach we have a probability to find a good stone of 0.01 (2^-6,64386).

I hope that is clear.

So, the general definitions:

c) Specification. Given a well defined set of objects (the search space), we call “specification”, in relation to that set, any explicit objective rule that can divide the set in two non overlapping subsets:  the “specified” subset (target space) and the “non specified” subset.  IOWs, a specification is any well defined rule which generates a binary partition in a well defined set of objects.

d) Functional Specification. It is a special form of specification (in the sense defined above), where the rule that specifies is of the following type:  “The specified subset in this well defined set of objects includes all the objects in the set which can implement the following, well defined function…” .  IOWs, a functional specification is any well defined rule which generates a binary partition in a well defined set of objects using a function defined as in a) and verifying if the functionality, defined as in b), is present in each object of the set.

It should be clear that functional specification is a definite subset of specification. Other properties, different from function, can in principle be used  to specify. But for our purposes we will stick to functional specification, as defined here.

e) The ratio Target space/Search space  expresses the probability of getting an object from the search space by one random search attempt, in a system where each object has the same probability of being found by a random search (that is, a system with an uniform probability of finding those objects).

f) The Functionally Specified  Information  (FSI)  in bits is simply –log2 of that number. Please, note that I  imply  no specific  meaning of the word “information” here. We could call it any other way. What I mean is exactly what I have defined, and nothing more.

One last step. FSI is a continuous numerical value, different for each function and system.  But it is possible to categorize  the concept in order to have a binary variable (yes/no) for each function in a system.

So, we define a threshold (for some specific  system of objects). Let’s say 30 bits.  We compute different values of FSI for many different functions which can be conceived for the objects in that system. We say that those functions which have a value of FSI above the threshold we have chosen (for example, more than 30 bits) are complex. I will not discuss here how the threshold is chosen, because that is part of the application of these concepts to the design inference, which will be the object of another post.

g) Functionally Specified Complex Information is therefore a binary property defined for a function in a system by a threshold. A function, in a specific system, can be “complex” (having  FSI above the threshold). In that case, we say that the function implicates FSCI in that system, and if an object observed in that system implements that function we say that the object exhibits FSCI.

h) Finally, if the function for which we use our objects is linked to a digital sequence which can be read in the object, we simply speak of digital FSCI: dFSCI.

So, FSI is a subset of SI, and dFSI is a subset of FSI. Each of these can be expressed in categorical form (complex/non complex).

Some final notes:

1) In this post, I have said nothing about design. I will discuss in a future post how these concepts can be used for a design inference, and why dFSCI is the most useful concept to infer design for biological information.

2) As you can see, I have strictly avoided to discuss what information is or is not. I have used the word for a specific definition, with no general implications at all.

3) Different functionalities for different functions can be defined for the same object or set of objects. Each function will have different values of FSI. For example, a tablet computer can certainly be used as a paperweight. It can also be used to make complex computations. So, the same object has different functionalities. Obviously, the FSI will be very different for the two functions: very low for the paperweight function (any object in that range of dimensions and weight will do), and very high for the computational function (it’s not so easy to find a material object that can work as a computer).

4) Although I have used a conscious observer to define function, there is no subjectivity in the procedures. The conscious observer can define any possible function he likes. He is absolutely free. But he has to define objectively the function, and how to measure the functionality, so that everyone can objectively verify the measurement. So, there is no subjectivity in the measurements, but each measurement is referred to a specific function, objectively defined by a subject.

Comments
Piotr: You quote the following paper: "Protein superfamily evolution and the last universal common ancestor (LUCA)." You are very kind. Here is the abstract:
By exploiting three-dimensional structure comparison, which is more sensitive than conventional sequence-based methods for detecting remote homology, we have identified a set of 140 ancestral protein domains using very restrictive criteria to minimize the potential error introduced by horizontal gene transfer. These domains are highly likely to have been present in the Last Universal Common Ancestor (LUCA) based on their universality in almost all of 114 completed prokaryotic (Bacteria and Archaea) and eukaryotic genomes. Functional analysis of these ancestral domains reveals a genetically complex LUCA with practically all the essential functional systems present in extant organisms, supporting the theory that life achieved its modern cellular status much before the main kingdom separation (Doolittle 2000). In addition, we have calculated different estimations of the genetic and functional versatility of all the superfamilies and functional groups in the prokaryote subsample. These estimations reveal that some ancestral superfamilies have been more versatile than others during evolution allowing more genetic and functional variation. Furthermore, the differences in genetic versatility between protein families are more attributable to their functional nature rather than the time that they have been evolving. These differences in tolerance to mutation suggest that some protein families have eroded their phylogenetic signal faster than others, hiding in many cases, their ancestral origin and suggesting that the calculation of 140 ancestral domains is probably an underestimate.
Emphasis mine. No comment. Please, quote that kind of paper more often. :)gpuccio
May 9, 2014
May
05
May
9
09
2014
02:15 PM
2
02
15
PM
PDT
Piotr: OK, you are a fighter. I like that. And your arguments are better, again. More technical, more interesting. And they can be discussed. So, let's fight! One point at a time. Let's start with this bizarre statement (at #184). Ti my question (what am I missing?) you answer:
The simple fact that the homologies are between members of the same superfamily, not between bacterial and archeal superfamilies themselves. Its the common ancestor of the whole superfamily that can be traced back to Luca, not the individal proteins.
But that is simply not true. I will give you facts. Let's take again our old friend, the beta subunit of ATP synthase. I have blasted the E. coli protein against the form in Methanosarcina barkeri, an archea that produces methane. This is the result: E. coli: length 460 AAs. Methanosarcina barkeri: Identities: 225/455(49%) Positives: 313/455(68%) Expect: 1e-153 Now, these are not abstractions. They are not myths. They are real proteins, indeed protein subunits of one of the most important and complex molecular machine we know of. Most of the moleculte (approximately AAs 74 - 345) represents a single domain which is part of the RecA-like NTPases. Now, as you can see: a) these two real proteins share a very high homology osf sequence, structure and function between archaea and bacteria. The sequence is mainly the same, as is the structure and function. Half of the sequence is identical. How do you explain that? It is simpple: this molecule was already there in LUCA, and it had similar sequence, and the same structure and function, in LUCA. So, why do you say: "the homologies are between members of the same superfamily, not between bacterial and archeal superfamilies themselves"? It is simply not true. The 900 - 1000 superfamilies which were already in LUCA were already in LUCA. That nmeans that the proteins which are part of those superfamilies share high homology, and similar structure and function, not only among themselves in in species, but also between archea and bacteria. That's how we know that they were already present in LUCA. Please, look at this very good paper: "The Evolutionary History of Protein Domains Viewed by Species Phylogeny" http://www.plosone.org/article/info%3Adoi%2F10.1371%2Fjournal.pone.0008378 A quote: "Table 1 lists the predicted number of domains and domain combinations originated in the major lineages of the tree of life. 1984 domains (at the family level) are predicted to be in the root of the tree (with the ratio Rhgt = 12), accounting for more than half of the total domains (3464 families in SCOP 1.73)."gpuccio
May 9, 2014
May
05
May
9
09
2014
02:09 PM
2
02
09
PM
PDT
Gpuccio, @164:
So, let’s talk about duplicated genes. What do you mean here? A gene duplicated and inactivated? Or a gene duplicated which goes on with its function?
For simplicity, let's restrict our discussion to a gene whose paralogue is initially identical to the original and remains "functional" (that is, gets transcribed and translated into a functional product). Of course in organisms with vast effective population (especially prokaryotes) the cost of maintaining a redundant copy may lead to counterselection. If, however, the effective population is relatively small, redundancy may be (nearly) neutral and duplicated genes may be retained in the genetic pool for a long time. Another possibility is that duplication itself is beneficial (like that multiplied amylase gene in domestic dogs, see Rhampton7 @175), in which case it will be maintained by positive selection. I'd like to concentrate on those cases when duplication is practically neutral.
If you mean a gene which is not functional, the situation is even worse. Either the second gene is really useful, and then it is subject to negative selection: it will remain in iots “hole”.
??? -- I do not understand this paragraph as it stands. Did you mean "a gene which is functional"? And what exactly is the second sentence trying to convey? Why "negative selection"? Under the standard meaning of the term, it isn't something that affects "really useful" genes. I can only guess that you are indirectly (and confusingly) referring to "Ohno's dilemma": if the retention of both copies is beneficial, it restricts their ability to diverge. If that's what you mean, I agree.
Or the second gene is not useful, and it can change its sequence. Then, it will lose its original functionalities very early (a single stop codon will be enough), and become like the inactivated gene.
Nope. You have smuggled in a false tacit assumption (after Behe?): gene = protein = function. It's actually gene -> protein(s) and protein -> {function_1, function_2, ... function_n}. The consequences of a mutation may damage one of the functions while leaving others unaffected. This opens many interesting possibilities.
Proteins are isolated functional islands. You cannot go from one island to the other. You will be immediately lost in the ocean of non functionality.
This is a myth.
It’s not a case that we have no single example of the transition from one protein superfamily to another one, from one structure and function to another completely different structure and function.
When interpreted literally, it seems to mean that have examples of such transitions. Which is of course just fine as far as I'm concerned, but I suspect that the sentence has mutated via double negation into the opposite of what youy had intended to say.Piotr
May 9, 2014
May
05
May
9
09
2014
10:42 AM
10
10
42
AM
PDT
Gpuccio @181
The simple point is that the facts we observe suggest that LUCA is also FUCA. Ideology, and only ideology, suggests otherwise.
LUCA is nothing of the kind. It's just the point in the past where all the genealogies of modern gene families finally coalesce. It was no more the first organism than "mtEve" was the first female. It wasn't the first prokaryotic cellular organism either, or the only form of life on young Earth. If we ignore horizontal transfer, LUCA is technically the most recent common ancestor of Bacteria and Archaea, but since HT was surely widespread among early prokaryotes, the actual point of coalescence must be a bit deeper.
From Wikipedia: “The LUA is estimated to have lived some 3.5 to 3.8 billion years ago (sometime in the Paleoarchean era).[3][4] The earliest evidences for life on Earth are graphite found to be biogenic in 3.7 billion-year-old metasedimentary rocks discovered in Western Greenland[5] and microbial mat fossils found in 3.48 billion-year-old sandstone discovered in Western Australia.[6][7]”
Some of those fossil traces of life (if their interpretation is correct may well be older than LUCA.
So, if those estimates are credible, the window ofr “OOL before LUCA”, if it ever existed, is really small. In that window (let’s say 200 – 300 Ma at best) not only OOL must have taken place, by some magical mechanism out of thin air, but also about half pf protein superfamilies must have appeared, with their sequence, structure and function (see previous post).
I wouldn't say that the 3.5 Ga estimate of the age of LUCA is impossible (though I've seen more modest estimates, of the order of 2.9 Ga, in recent literature). It still gives us a few hundred million years for the first chemical replicators to evolve into prokaryotic life more or less as we know it. It's an interval roughly equal to that between the end of the Carboniferous and now. You seem to think that for a few billion years the "designer" has not really been creating anything out of thin air. He's only been using small-scale magic to add some intelligent organisation to elements already present, like when he turns small ORFs into genes in fruit flies. He must have a soft spot for fruit flies; he's still churning out orphan genes by the hundred for the 1500 Drosphila species, at such a rate that many of them are still segregating in D. melanogaster, for example. Perhaps the designer is the Lord of the Flies? But when we get to the root of life, you have to assume the magical creation of a complete LUCA out of thin air. The designer even had to create "superfamilies of genes" looking as if they consisted of homologues (but they can't be true homologues if LUCA was Generation Zero; the designer simply made them look related).Piotr
May 9, 2014
May
05
May
9
09
2014
09:55 AM
9
09
55
AM
PDT
Gpuccio:
Wait a moment. If disfferent proteins in different species are part of the same superfamily, that means that they share sequence homologies, similar structure and similar function.
Ranea et al. 2006:
"However, even though there may be conservation of very general functional or molecular mechanisms (see Todd et al. 2001), some ancestral superfamily domains show such high functional diversification that it is difficult to define a concrete function for them. For example, the ATP-loop superfamily has representatives in 230 different orthologue clusters divided into 19 different COG functional subcategories. Although the ATP-loop domain is mainly represented in metabolic pathways, this domain is also involved in disparate functional roles."
Gpuccio:
If the superfamily is among those which were in LUCA, it means that those homologies, structure and function are already observable both in acrhea and in bacteria (and therefore precede the archea bacteria divergence). That means that the sequence, the structure and the function (therefore, the superfamily) were already in LUCA. What am I missing?
The simple fact that the homologies are between members of the same superfamily, not between bacterial and archeal superfamilies themselves. Its the common ancestor of the whole superfamily that can be traced back to Luca, not the individal proteins. More later.Piotr
May 9, 2014
May
05
May
9
09
2014
06:45 AM
6
06
45
AM
PDT
rhampton7: Thank you for the contributions and the links. I will look at them, and cone back to you.gpuccio
May 9, 2014
May
05
May
9
09
2014
04:31 AM
4
04
31
AM
PDT
Piotr:
The controversy is not about whether the genes encoding for the L and M opsins have a common origin but whether trichromacy arose once or independently twice in catarrhines and platyrrhines.
Let's admit that the genes have a common origin (I should check the homologies, but I have not the time now). And so? That would simply be an argument about common descent, which I accept and defend. It tells us nothing about the mechanism, unless we analyze the complexity of the functional transition. Which I have not the time to do now, but can certainly be done.gpuccio
May 9, 2014
May
05
May
9
09
2014
04:22 AM
4
04
22
AM
PDT
Piotr:
LUCA was not the first living thing on Earth but the last universal common ancestor, itself the product of a long prehistory — presumably hundreds of millions of years — which can’t be reconstructed by comparing the genomes of its descendants. It was already pretty “modern” in terms of its genomic status. I have no idea how its genes originated.
I agree with you. You have no idea. The simple point is that the facts we observe suggest that LUCA is also FUCA. Ideology, and only ideology, suggests otherwise. From Wikipedia: "The LUA is estimated to have lived some 3.5 to 3.8 billion years ago (sometime in the Paleoarchean era).[3][4] The earliest evidences for life on Earth are graphite found to be biogenic in 3.7 billion-year-old metasedimentary rocks discovered in Western Greenland[5] and microbial mat fossils found in 3.48 billion-year-old sandstone discovered in Western Australia.[6][7]" From Wikipedia: "The age of the Earth is 4.54 ± 0.05 billion year" "4100–3800 Ma Late Heavy Bombardment: extended barrage of impact events upon the inner planets by meteoroids. Thermal flux from widespread hydrothermal activity during the LHB may have been conducive to life's emergence and early diversification.[5] 3900–2500 Ma Cells resembling prokaryotes appear.[6] These first organisms are chemoautotrophs: they use carbon dioxide as a carbon source and oxidize inorganic materials to extract energy. Later, prokaryotes evolve glycolysis, a set of chemical reactions that free the energy of organic molecules such as glucose and store it in the chemical bonds of ATP. Glycolysis (and ATP) continue to be used in almost all organisms, unchanged, to this day.[7][8]" So, if those estimates are credible, the window ofr "OOL before LUCA", if it ever existed, is really small. In that window (let's say 200 - 300 Ma at best) not only OOL must have taken place, by some magical mechanism out of thin air, but also about half pf protein superfamilies must have appeared, with their sequence, structure and function (see previous post). For OOL we have no non design theory which even starts to be credible. For the origin of superfamiles, both at OOL and after it, we have no credible non design theory. Ah, but I forgot. These are God of the gaps arguments, not scientific reasoning.gpuccio
May 9, 2014
May
05
May
9
09
2014
04:18 AM
4
04
18
AM
PDT
Piotr:
Now this is getting really bizarre. The ancestry of a few hundred superfamilies can be traced back to LUCA, which doesn’t mean that they were present in LUCA as superfamilies.
Wait a moment. If disfferent proteins in different species are part of the same superfamily, that means that they share sequence homologies, similar structure and similar function. If the superfamily is among those which were in LUCA, it means that those homologies, structure and function are already observable both in acrhea and in bacteria (and therefore precede the archea bacteria divergence). That means that the sequence, the structure and the function (therefore, the superfamily) were already in LUCA. What am I missing?gpuccio
May 9, 2014
May
05
May
9
09
2014
04:04 AM
4
04
04
AM
PDT
Piotr:
ORFs can originate in various ways. I never said they had to have homologues in more distant relatives (such homologues, even if they existed, would be hard to recognise anyway, given the rather small size of those “proto-genes”).
If they were old genes inactivated and only slightly changed, like in your "extreme" example, they should have those homologues.gpuccio
May 9, 2014
May
05
May
9
09
2014
04:01 AM
4
04
01
AM
PDT
Piotr:
It’s enough if de novo genes have clear non-coding orthologues (primarily, unexpressed small ORFs) in closely related species. It shows that they were not magically created by the designer out of thin air.
As you should have understood (and have understood), I don't think that the designer creates things magically out of thin air. I have clarified many times, even to you, that IMO the designer guides the evolution of sequences adding functional information, probably mainly through guided mutations and guided transposon activity. What has it to do with "creating things magically out of thin air"? And, as I have explained, that view is not only perfectly compatible with de novo genes having homologies with previous non coding regions, but indeed supported by it.gpuccio
May 9, 2014
May
05
May
9
09
2014
03:55 AM
3
03
55
AM
PDT
Piotr:
No, I gave an extreme example to show that “junk” can be recycled. I was only talking about bias in favour of potentially functional products, not about a universal mechanism generating all de novo genes.
And:
No, I don’t buy the idea of “the language gene”; I’m not that naive. You tend to read too much into what I say.
Well, I tend to read, in what you say, arguments against my arguments. I am happy to know you were only digressing. :)gpuccio
May 9, 2014
May
05
May
9
09
2014
03:31 AM
3
03
31
AM
PDT
Piotr: First of all, I want to say that I appreciate very much your contributions here. As I have already said, you are a very good adversary, and you have given me many chances to detail important aspects of what I think. I thank you for that. So, please don't consider the vehemence of intellectual confrontation (which I really like) with any lack of respect for you and your ideas. That said, I must also say, unfortunately, that it is true, for what it's worth, that I have noticed some worsening of your arguments in your last posts. If you allow me the joke, it seems that the quality of your arguments is "decaying beyond recognition", like some unlucky pseudogene. :) For example, frankly I would not have expected, from you, the backpedaling to the God of the gaps "argument". I will not comment on that. I suppose that's what happens when intelligent persons are for some reason committed to defend not so intelligent ideas. However, this thread, and the good discussion we have had here, should be answer enough. So, let's go back to vehement (and interesting, I hope) intellectual confrontation, and let's discuss what can be discussed. There is not much of that, but it's better than nothing. In next thread. (By the way, I am not intentionally trying to increase the number of comments in my thread by fragmenting my answers. Not too much, at least. After all, some of my answers here are still very long. Let's say that the shorter ones, like this one, can be considered "de novo posts". :) )gpuccio
May 9, 2014
May
05
May
9
09
2014
03:26 AM
3
03
26
AM
PDT
gpuccio, As a follow-up to my last post. I came across the Comparative Genomics Lab (Tomas Marques-Bonet), and he is investigating Canid evolution among other lines of research. He co-authored the paper, Genome Sequencing Highlights Genes Under Selection and the Dynamic Early History of Dogs, in which "high-quality genome sequences from three gray wolves, one from each of the three putative centers of dog domestication, two basal dog lineages (Basenji and Dingo) and a golden jackal" were compared. The results were used to reconstruct their evolutionary history. A common criticism of ID is that such an endeavor starts with the unchallenged assumption that evolution (RV + NS) must have been responsible for these changes. It's precisely this assumption that can be tested with the kind of experiment I have suggested.rhampton7
May 8, 2014
May
05
May
8
08
2014
06:16 PM
6
06
16
PM
PDT
So to expand further on my comment @173. The universe of possibilities is incredibly immense. The ability to sample that universe of possibilities is absolutely miniscule in comparison. In spite of these facts we have functional proteins. So the sampling process must be miraculous, finding function in a sea of non-function. OR Function must be ubiquitous, yet another miracle. Pick your poison. Neither can be explained by materialist theories.Mung
May 8, 2014
May
05
May
8
08
2014
05:56 PM
5
05
56
PM
PDT
Piotr:
The universe of aperiodic polymers (and their potential functions) is vast.
And the ability to sample that vast universe so small.
It isn’t the way intelligent designers of the only kind known to me (humans) do their job.
Precisely the point!Mung
May 8, 2014
May
05
May
8
08
2014
05:39 PM
5
05
39
PM
PDT
"Design" doesn't explain anything but "faulty design" sure explains alot!Mung
May 8, 2014
May
05
May
8
08
2014
05:34 PM
5
05
34
PM
PDT
#164: There's so much wrong there that I'll have to leave it till some time tomorrow. #165: The controversy is not about whether the genes encoding for the L and M opsins have a common origin but whether trichromacy arose once or independently twice in catarrhines and platyrrhines.Piotr
May 8, 2014
May
05
May
8
08
2014
05:10 PM
5
05
10
PM
PDT
A complex process must be regulated by complex procedures. You seem to believe that two aminoacids can evolve language and abstract thought.
No, I don't buy the idea of "the language gene"; I'm not that naive. You tend to read too much into what I say. "Importantly contributed to X" is a totally different thing from "was solely responsible for X". I said nothing about abstract thought or language processing. The linguistically relevant effect of the "human variant" of FOXP2 more likely consisted in improving the neuromuscular control of the organs of speech in archaic humans. Still, pretty dramatic if you compare the degree of control humans have over their active articulators with that in chimps.Piotr
May 8, 2014
May
05
May
8
08
2014
04:50 PM
4
04
50
PM
PDT
gpuccio, It seems to me there ought to be a way to scan through published genomes and identify regions that are 1) unique to specific genera and/or species and that 2) are at least X bits (presumably 500) in length. Granted, that wouldn't tell you anything about functionality, but it would substantially narrow the information to be examined. I suggest starting with Canidae since the genomes of many members have already been assembled. The results could be used to create an ID-based cladogram demonstrating where micro-evolution could have been responsible for changes within Canidae and where design was required. While being accessible to the laymen, this experiment would provide ID a means to test its predictions about random variation (and possible natural selection). For example, if such an examination found a complex, functional region of the genome in bloundhounds, but not in gray wolves, then it would refute a premise about the abilities of RV. On the other hand, if complex, functional regions were found in only some of the genera currently defined by science, then it would revolutionize the methods used and the linkages assigned within biological classification. Quite a coup for ID. Best of all, this task should be well within the capabilities of the ID community to quickly accomplish at minimal expense.rhampton7
May 8, 2014
May
05
May
8
08
2014
04:47 PM
4
04
47
PM
PDT
Gpuccio:
Or are you suggesting that all the functional genes which emerged after OOL were already present at OOL, but became inactive pseudogenes in all species, only to be reactivated occasionally in a new species?
No, I gave an extreme example to show that "junk" can be recycled. I was only talking about bias in favour of potentially functional products, not about a universal mechanism generating all de novo genes. Sooner or later pseudogenes and other kinds of ex-functional DNA (e.g. retroviral sequences) decay beyond recognition, but at least those pseudogenised recently can be recruited for some tasks thanks to their residual, potentially functional features.
Unfortunately, such a bizarre and ad hoc theory cannot be true, because the non coding regions which generate new protein coding genes have no homologue not only in coding DNS, but also in non coding DNA of distant species. They appear in direct ancestors (for example, in primates for new human genes) and become ORFs in the final species.
It's enough if de novo genes have clear non-coding orthologues (primarily, unexpressed small ORFs) in closely related species. It shows that they were not magically created by the designer out of thin air. http://www.sciencemag.org/content/343/6172/769 ORFs can originate in various ways. I never said they had to have homologues in more distant relatives (such homologues, even if they existed, would be hard to recognise anyway, given the rather small size of those "proto-genes").
Moreover, you should then explain how at OOL not only the 800 – 900 superfamilies which were already present in LUCA and have persisted up to now were generated, but also how the other 800 – 900 superfamilies which appear in the rest of natural history were generated at OOL, then disappeared, then occasionally reappeared by a single aminoacid mutation (form what?)…
Now this is getting really bizarre. The ancestry of a few hundred superfamilies can be traced back to LUCA, which doesn't mean that they were present in LUCA as superfamilies. LUCA was not the first living thing on Earth but the last universal common ancestor, itself the product of a long prehistory -- presumably hundreds of millions of years -- which can't be reconstructed by comparing the genomes of its descendants. It was already pretty "modern" in terms of its genomic status. I have no idea how its genes originated. Saying, without a shred of positive evidence, that "the designer did it" solves no problems. It's God of the Gaps, as usual.Piotr
May 8, 2014
May
05
May
8
08
2014
04:27 PM
4
04
27
PM
PDT
rhampton7: "Do you know if there is any genetic evidence to support your concept of macroevolution to the application of Biological classification?" My approach is entirely molecular. I am not an expert of biological classifications. I don't believe that the simple observation of the phenotype can help us in understanding the mechanisms of generation of functional information. Information is in the molecules, not in the phenotype. I don't think that my use of "macroevolution" is so different form the usual. "Microevolution" is the term for those few cases where a small molecular difference (1-2 AAs) explains a reproductive advantage under very selective pressure (see again, for example, antibiotic resitance in its simple forms, or the emergence of nylonase). "Macroevolution, at the molecular level, is therefore the emergence of a complex function by a complex sequence variation. I believe that probably the emergence of new species, of new body plans, of different organs or systems, almost always need the emergence of many coordinated new proteins, and of a lot of new regulatory procedures. We have good examples of that. For example, the emergence of the adaptive immune system in jawed vertebrates requires a lot of new molecular tools, first of all the very complex RAG1 and RAG2 proteins. So, there is no true difference between the two approaches, but true reasonings about the causal mechanism of the transitions are possible only at the molecular level.gpuccio
May 8, 2014
May
05
May
8
08
2014
03:29 PM
3
03
29
PM
PDT
Piotr: You have not understood my objection to Szostac's paper. Here is the abstract
Functional primordial proteins presumably originated from random sequences, but it is not known how frequently functional, or even folded, proteins occur in collections of random sequences. Here we have used in vitro selection of messenger RNA displayed proteins, in which each protein is covalently linked through its carboxy terminus to the 3' end of its encoding mRNA1, to sample a large number of distinct random sequences. Starting from a library of 6 times 1012 proteins each containing 80 contiguous random amino acids, we selected functional proteins by enriching for those that bind to ATP. This selection yielded four new ATP-binding proteins that appear to be unrelated to each other or to anything found in the current databases of biological proteins. The frequency of occurrence of functional proteins in random-sequence libraries appears to be similar to that observed for equivalent RNA libraries2, 3.
The fact that his final protein was not really biologically functional (least of all naturally selectable) is just an aside. The real problem is that the final protein, the so called "functional protein", was not in the original random library. It was evolved by intelligent selection, by a design process. From the paper:
Because protein sequences with speci®c functions are expected to be quite rare in protein sequence space, we prepared a DNA library of 4 * 10^14 independently generated random sequences. This DNA library was specifically constructed to avoid stop codons and frameshift mutations4, and was designed for use in mRNA display selections. This DNA library was then used to generate 6 * 10^12 puri®ed non-redundant random proteins that were used as the input into the first selection step. ... Successive rounds of in vitro selection and ampli®cation were performed starting with this random-sequence library. In each round the mRNA-displayed proteins were incubated with immobilized ATP, washed and eluted with free ATP. The eluted fractions were collected and ampli®ed by polymerase chain reaction (PCR); this DNA was then used to generate a new library of mRNAdisplayed proteins, enriched in sequences that bind ATP, for input into the next round of selection. ... We cloned and sequenced 24 individual library members, which showed that the population was now dominated by 4 families of ATP-binding proteins (Fig. 3a). These families show no sequence relationship to each other or to any known biological protein. The members of each family are closely related, indicating that each family is descended from a single ancestral molecule, which was one of the original random sequences. ... One possible explanation for this low level of ATP-binding is conformational heterogeneity, possibly reflecting inefficient folding of these primordial protein sequences. In an effort to increase the proportion of these proteins that fold into an ATP-binding conformation, we mutagenized the library and carried out further rounds of in vitro selection and amplification.
Emphasis mine. IOWs, they organized a process of artificial, intelligent evolution and selection to transform original sequences with low level ATP-binding into a final protein with strong ATP binding. So, the final "functional" protein /which was after all not really functional) was not in the original random library. It was engineered, exploiting some low ATP binding capacity (which can be considered more a biochemical property which is not surprising in a few sequences of a vast random library) that a function, however useless. So, the only legitimate conclusion we can get from that paper is that, if "Functional primordial proteins presumably originated from random sequences", as the abstract states, they did it through a process of intelligent design, including intentional mutation and selection.gpuccio
May 8, 2014
May
05
May
8
08
2014
03:19 PM
3
03
19
PM
PDT
Piotr: From Wikipedia:
Evolution of color vision in primates ... Hypotheses Some evolutionary biologists believe that the L and M photopigments of New World and Old World primates had a common evolutionary origin; molecular studies demonstrate that the spectral tuning (response of a photopigment to a specific wavelength of light) of the three pigments in both sub-orders is the same.[7] There are two popular hypotheses that explain the evolution of the primate vision differences from this common origin. Polymorphism The first hypothesis is that the two-gene (M and L) system of the catarrhine primates evolved from a crossing-over mechanism. Unequal crossing over between the chromosomes carrying alleles for L and M variants could have resulted in a separate L and M gene located on a single X chromosome.[5] This hypothesis requires that the evolution of the polymorphic system of the platyrrhine pre-dates the separation of the Old World and New World monkeys.[8] This hypothesis proposes that this crossing-over event occurred in a heterozygous catarrhine female sometime after the platyrrhine/catarrhine divergence.[4] Following the crossing-over, any male and female progeny receiving at least one X chromosome with both M and L genes would be trichromats. Single M or L gene X chromosomes would subsequently be lost from the catarrhine gene pool, assuring routine trichromacy. Gene duplication The alternate hypothesis is that opsin polymorphism arose in platyrrhines after they diverged from catarrhines. By this hypothesis, a single X-opsin allele was duplicated in catarrhines and catarrhine M and L opsins diverged later by mutations affecting one gene duplicate but not the other. Platyrrhine M and L opsins would have evolved by a parallel process, acting on the single opsin gene present to create multiple alleles. Geneticists use the "molecular clocks" technique to determine an evolutionary sequence of events. It deduces elapsed time from a number of minor differences in DNA sequences.[9][10] Nucleotide sequencing of opsin genes suggests that the genetic divergence between New World primate opsin alleles (2.6%) is considerably smaller than the divergence between Old World primate genes (6.1%).[8] Hence, the New World primate color vision alleles are likely to have arisen after Old World gene duplication.[4] It is also proposed that the polymorphism in the opsin gene might have arisen independently through point mutation on one or more occasions,[4] and that the spectral tuning similarities are due to convergent evolution.Despite the homogenization of genes in the New World monkeys, there has been a preservation of trichromacy in the heterozygous females suggesting that the critical amino acid that define these alleles have been maintained.[11]
Your statement: "We and our primate relatives owe our trichromatic vision to a rather trivial duplication event followed by divergence between the original gene and its copy. I hope you don’t deny the possibility of such a process. " Is it really necessary to "deny" anything? Especially a "possibility"? This is wishful thinking. Even is parts of these hypotheses were true, there is no rigorous eveluation of the mechanisms, of the complexity of the supposed transitions, of the paths, and so on. It's not my habit to deny or not deny the "possibility" of vague explanations.gpuccio
May 8, 2014
May
05
May
8
08
2014
02:52 PM
2
02
52
PM
PDT
Piotr:
If a gene gets duplicated, and one copy is free to drift away from its original role, the effect can be quite dramatic.
So, let's talk about duplicated genes. What do you mean here? A gene duplicated and inactivated? Or a gene duplicated which goes on with its function? If you mean the inactivated gene, it is not different from any other non coding region. Any mutation is neutral by definition. It is probably no more translated. It can "go wherever it likes". That is, nowhere complex and useful. The only effect is neutral variation in a non coding region. IOWs, random sequences, which remain non coding and useless. I can see no drama here. If you mean a gene which is not functional, the situation is even worse. Either the second gene is really useful, and then it is subject to negative selection: it will remain in iots "hole". Or the second gene is not useful, and it can change its sequence. Then, it will lose its original functionalities very early (a single stop codon will be enough), and become like the inactivated gene. End of the story. No drama. Proteins are isolated functional islands. You cannot go from one island to the other. You will be immediately lost in the ocean of non functionality. It's not a case that we have no single example of the transition from one protein superfamily to another one, from one structure and function to another completely different structure and function.gpuccio
May 8, 2014
May
05
May
8
08
2014
02:40 PM
2
02
40
PM
PDT
Piotr:
For example, small changes in the phylogenetically old and highly conserved FOXP2 gene are believed to have contributed importantly to the development of speech in humans. The FOXP2 protein regulates the expression of numerous other important genes (which is precisely the reason why it’s so conserved). We differ from chimps by two amino acids in the protein, due to two non-synonymous substitutions in our lineage. We share one of those point mutations with Carnivora; only the other is really unique to humans (including neanderthals and denisovans), but I wouldn’t describe the effect as minor.
You are making a huge error here. Yoiu start saying: "For example, small changes in the phylogenetically old and highly conserved FOXP2 gene are believed to have contributed importantly to the development of speech in humans." In what sense they "are believed"? From Wikipedia: "Some researchers have speculated that the two amino acid differences between chimps and humans led to the evolution of language in humans.[13] Others, however, have been unable to find a clear association between species with learned vocalizations and similar mutations in FOXP2.[27][28] Insertion of both human mutations into mice, whose version of FOXP2 otherwise differs from the human and chimpanzee versions in only one additional base pair, causes changes in vocalizations as well as other behavioral changes, such as a reduction in exploratory tendencies; a reduction in dopamine levels and changes in the morphology of certain nerve cells are also observed.[15] It may also be, based on general observations of development and songbird results, that any difference between humans and non-humans would be due to regulatory sequence divergence (affecting where and when FOXP2 is expressed) rather than the two amino acid differences mentioned above." Now, don't say that I am quote mining because I don't quote the whole article. My point is simply that what is believed on this point is very controversial. Why?. Because the whole methodology is wrong. FOXP2 is a transcription factor. It has, certainly, important regulatory functions, and it is implied in speech regulation. That's OK. In different species, those regulatory functions will vary, and that can easily explain the differences in humans, and the fact that they are conserved and functional. But how can you leap from that "to amino acid differences between chimps and humans led to the evolution of language in humans"? This is completely unwarranted. Humans have a different brain, and language is primarily an abstract function of the brain. Huge differences between the human brain and the chimp brain can explain the simple fact that humans develop an abstract language, and chimps don't. An error that is made often is that, if one can prove that a final effector is implied in a complex process, for example by showing that a knockout of that effector compromises the process, that one feel authorized that the final effector is the cause of the whole process. That is obviously not true. Transcription factors are key final effectors of the transcription regulation. But the true question is: what regulates the regulators? A complex process must be regulated by complex procedures. You seem to believe that two aminoacids can evolve language and abstract thought. The effect is not minor. But you are wrong about the cause.gpuccio
May 8, 2014
May
05
May
8
08
2014
02:30 PM
2
02
30
PM
PDT
Piotr:
Drift can affect the emergence of functions by allowing populations (especially small ones) to “escape” from local peaks of fitness.
That is irrelevant. Inactivated pseudogenes have no fitness value at all. They can go wherever they like. Non coding DNA, if at least in part it is non functional, has no fitness value at all. It can go wherever it likes. But the probabilistic barriers exclude that a sequence which can "go wherever it likes" will ever reach some place useful (in a complex way). Only functional genes are destined to remain in "peaks of fitness" (which should be more correctly be called "holes of fitness"), as the rugged landscape paper tells us. And the same paper tells us that no drift can help them to get out of the hole, once they have fallen there.gpuccio
May 8, 2014
May
05
May
8
08
2014
02:11 PM
2
02
11
PM
PDT
Piotr:
An algorithm is a step-by-step computation procedure. Neither drift nor selection are “algorithms” in the ordinary sense of the word. They are aspects of the evolution of populations, which is not an algorithm either. It can be modelled algorithmically (to an approximation), which justifies the metaphor but doesn’t make the evolutionary process a sequence of calculations. The orbital motions of planets can also be so modelled, but nobody calls them algorithms.
I don't understand. If you can compute something, you can do it by a computer, automatically. And you do that by algorithms. If you can model drift, NS, or the orbit of planets by a computer program, you can do that by algorithms. What do you mean here? Algorithms can include random processes, and model them by the laws of probability. But they use laws of necessity to compute the probabilities. Drift is an algorithm because it explains (and computes) how a population will evolve in some circumstances, and making some assumptions. What a model of drift cannot tell you is which gene exactly will be fixed, and which will not. That's because the particular gene which is fixed by drift is "selected" randomly. But the process of fixation is mostly algorithmic, and follows the laws of necessity, mathematical laws. The same is true, even more, for NS. NS can be modeled, and is modeled, by certain assumptions. The difference with drift is that which gene will be fixed, in a NS model, is not random, but depends critically on the ability of the gene to contribute to reproduction (or to be a hindrance to it). Therefore, the relationship between the process of fixation and the selection of the gene to which the process applies is not random, in NS, but follows mathematical rules. Even if probabilistic variables can add to the final result, the cause effect relationship between the gene which is fixed and it reproductive advantage is well defined and measured in the model. It is a strong necessity relationship, modified by other variables. In drift, on the other hand, we have no idea of which particular gene will be fixed. That is the difference.gpuccio
May 8, 2014
May
05
May
8
08
2014
02:05 PM
2
02
05
PM
PDT
Piotr:
Let’s imagine that a point mutation disables a functional gene and turns it into a pseudogene — one type of “junk DNA”. A reverse point mutation (which is not astronomically unlikely) may then restore its functionality, “creating” a complete gene out of junk (but not random junk). Such latent functionality is invisible to natural selection, but is important if you want to calculate the likelihood of the emergence of function. It may increase the odds of getting something functional by God knows how many orders of magnitude.
This is really funny. It is simply important to calculate the likelihood of the reactivation of an existing function, not certainly the emergence of a new function. You see, "such latent functionality" would be "invisible to natural selection", but not to science. We would easily find an extremely high homology between the pseudogene and the existing functional gene in the general proteome. That's exactly how we say that a pseudogene is a pseudogene, and not generically another kind of non coding DNA. Or are you suggesting that all the functional genes which emerged after OOL were already present at OOL, but became inactive pseudogenes in all species, only to be reactivated occasionally in a new species? Unfortunately, such a bizarre and ad hoc theory cannot be true, because the non coding regions which generate new protein coding genes have no homologue not only in coding DNS, but also in non coding DNA of distant species. They appear in direct ancestors (for example, in primates for new human genes) and become ORFs in the final species. Moreover, you should then explain how at OOL not only the 800 - 900 superfamilies which were already present in LUCA and have persisted up to now were generated, but also how the other 800 - 900 superfamilies which appear in the rest of natural history were generated at OOL, then disappeared, then occasionally reappeared by a single aminoacid mutation (form what?)... If that is the meaning of: "Besides, the structure of the genome reflects its history in ways that may lead to a pro-functionality bias.", then now I understand why I did not understand. More in next post.gpuccio
May 8, 2014
May
05
May
8
08
2014
01:49 PM
1
01
49
PM
PDT
Piotr:
There’s randomness and randomness. Mutations, recombinations, frameshifts, deletions, inversions, etc. are “random” with regard to the adaptive value of their consequences. But they are not equally probable. Besides, the structure of the genome reflects its history in ways that may lead to a pro-functionality bias.
There's discourse and discourse. There are discourse which are clear and try to express truth, and there are discourses which are vague, confusing and obfuscating. I don't know why, but the tome of your dioscourses seems definitely worse in your last posts. "There’s randomness and randomness."??? What does that mean? I offer you a simple definition of a random system: a system is random if we cannot describe its evolution by necessity laws, but still we can describe it to a certain point by probability distributions. That is simple and clear. "Mutations, recombinations, frameshifts, deletions, inversions, etc. are “random” with regard to the adaptive value of their consequences."??? What does that mean? Mutations are random because we have no way to describe deterministically how and when they happen. But we can describe the general trend of their occurrence probabilistically. That's all. The "adaptive value of their consequences" has nothing to do with that. The same is true for "recombinations, frameshifts, deletions, inversions, etc.". "But they are not equally probable." And so? Unequal probabilities do not make a random system less random. This is an error made by many who do not understand the basics of statistics. Only in systems with an uniform probability distribution the events have the same probability. If the system is described by other types of probability distributions, the probability of the events will be very different (see, for instance, the many natural systems well described by the normal distribution). But random systems they are, just the same. "Besides, the structure of the genome reflects its history in ways that may lead to a pro-functionality bias." I could excuse this statement if it were in Polish, which I do not understand. But in english? What does it mean? More in next post.gpuccio
May 8, 2014
May
05
May
8
08
2014
01:37 PM
1
01
37
PM
PDT
1 2 3 4 5 6 … 10