This blog has been created partly as a companion to Chemistry for the Biosciences, the textbook that I co-author with Tony Bradshaw, and to act as an archive of posts I write for other sites (particularly the OUPblog). Like the book itself, it explores how life on the scale of atoms and molecules has an impact on biology - at the scale of cells, tissues, and organisms - and seeks to demystify a range of biological and chemical concepts.

The blog's name takes as its inspiration the cover of the first edition of Chemistry for the Biosciences, which depicts a gecko seemingly clinging to its surface. To find out what links geckos to chemistry, read this.



Wednesday, 8 June 2011

What is a gene mutation?


In my last three posts I’ve introduced you to the world of biological information, taking you from the storage of biological information in libraries called genomes, which house information in individual books called chromosomes (themselves divided into chapters called genes), to the way the cell makes use of that stored information to manufacture the molecular machines called proteins.

But what happens when the storage of information goes wrong? If we’re reading a recipe and that recipe contains a mistake, chances are that the end-result of our culinary endeavour won’t end up as it should. And so it is at the level of cells. If the information the cell is using is somehow wrong, the end result will also be wrong – sometimes with catastrophic results.

I’ve mentioned in previous posts how biological information is captured by the sequence of the building block ‘letters’ from which DNA is constructed. The sequence of letters is ultimately deciphered by a molecular machine called the ribosome, which reads the sequence of letters in sets of three, and uses each trio to determine which amino acid – the building block of proteins – should be used next in its mission to construct a particular protein. It should come as no surprise that, if the recipe for the protein is changed – if the sequence of DNA ‘letters’ is altered – the protein that is manufactured will probably contain errors as a result. And if a protein contains errors, it won’t be able to function correctly, just as flat-packed furniture will end up being decidedly wobbly if you construct it from the wrong parts.

Imagine a snippet of DNA has the sequence GGTGCTAAG. The ribosome would ‘read’ this sequence, and would use it as the recipe for building a chain of three amino acids: Glycine-Alanine-Lysine. Now imagine that we alter just one letter in our original sequence so that it becomes GGTCCTAAG. All we’ve done is swap a G for a C at the fourth position in the DNA sequence. However, this change is sufficient to affect the composition of the protein that is produced when the sequence is deciphered: the ribosome will now build a chain with the composition Glycine-Proline-Lysine. 

Surely such a small change won’t actually cause significant problems in a cell, though. Right? Wrong. Amazingly (and perhaps unnervingly) the tiniest error can have really quite significant consequences.

Let’s take just one example. Sickle cell anaemia is a condition that affects the red blood cells of humans.  Red blood cells fulfil the essential role of transporting oxygen from our lungs to all the living cells of our body: they continually circulate through our arteries and veins, shuttling oxygen from one place to another. A healthy red blood cell looks a bit like a ring doughnut (though it doesn’t actually have a hole right through the middle); by contrast, the red blood cells of individuals with sickle cell anaemia become warped into crescent-like shapes (like a sickle, the grass-cutting tool, after which the disease is named). These sickle cells no longer pass freely through our arteries and veins. Instead, they tend to get entangled with each other. As a result, the flow of oxygen round the body is impeded, and the individual afflicted with the disease can suffer breathlessness, dizziness, and sudden pain throughout the body as a result.

So what has this to do with changes in the sequence of a gene? Almost unbelievably, this debilitating disease is caused by just a single error in the sequence of a particular gene. The gene in question (one particular ‘chapter’ in one of our chromosomes, the genetic ‘books’ that make up our genome ‘library’) is constructed from a total of 625 building block letters. Yet, a change in just one of those letters – from the letter A to T – results in the amino acid valine being added in place of the amino acid glutamic acid at a particular point in the protein haemoglobin - the part of the red blood cell to which oxygen actually attaches - as it is constructed from the recipe that the gene spells out. This change is all that’s needed to affect the structure of the haemoglobin, which, in turn, affects the shape of the red blood cell in which it is found, causing it to adopt the distorted sickle shape.

Why do such errors happen in the first place? In short, because cells have to make copies of their genomes – and no process of copying is completely error-free. Almost every cell in our body must possess a full library of biological information – a complete copy of our genome. So, every time a cell divides to produce two daughter cells (when our skin cells divide to repair a cut or graze, or the cells of our stomach lining are replenished, for example) it must first make a copy of its genome so that each daughter cell ends up with a full genome. But as the cell copies its genome – literally letter by letter – mistakes creep in. (If you were to re-type this post without using the delete key as you typed, how many errors do you think there would be in the end result? I’m writing this on a laptop that’s just a few months old, and I can already see that the delete key is the most worn key on the keyboard!) These mistakes are what we call mutations.

Fortunately for us, however, the introduction of errors during the genome-copying process is a rare event: the cells of our bodies do a truly remarkable job of keeping things error-free. Indeed, they copy DNA with an accuracy of 99.999% - this means that only one error is introduced for every 100,000 letters that are copied. That’s no mean feat - and is thanks to a particular molecular ‘editor’ that proof-reads the DNA as it is being copied, spotting errors and correcting them as it reads.

Hiccups in the copying of DNA by our cells are not the only cause of gene mutations: various environmental factors can cause them too. The most prevalent of these is ultra-violet light, a natural component of sunlight – but a component that can cause unwanted changes in the chemical structure of DNA when it comes into contact with our cells. The effect of ultra-violet light on our DNA is the reason why we’re encouraged to wear sunscreens when we’re out basking in the sun: you might think a tan will make you look healthy, but the cost might be the creation of an error in the DNA of some of your skin cells that triggers the formation of a skin tumour. Is that a price worth paying for a few weeks of looking bronzed?

Despite this, gene mutation itself isn’t necessarily a bad thing. Indeed, if mutations weren’t to occur, we’d not see around us the remarkable diversity of life that we do. For gene mutation is the molecular ‘fuel’ for the process of evolution, as I’ll explore in a future post.

Sunday, 1 May 2011

How is the information in a gene used by a cell?

In my last two posts I’ve introduced the notion that DNA acts as a store of biological information; this information is stored in a series of chromosomes, each of which are divided into a number of genes. Each gene in turn contains one ‘snippet’ of biological information. But how are these genes actually used? How is the information stored in these genes actually extracted to do something useful (if ‘useful’ isn’t too flippant a term for something that the very continuation of life depends upon).

Many (but not all) genes act as recipes for a family of biological molecules called proteins: they literally tell the cell what the ingredients for a particular protein are, and how they should be combined to create the protein itself. (Proteins have a range of essential roles in the human body. Some act as building materials for different components of the body, such as the keratin we find in our hair and nails. Others act as molecular transporters: haemoglobin, which is found in our red blood cells, carries oxygen from our lungs to other parts of the body. A family of proteins called the enzymes are arguably the most important, however. Enzymes cajole different chemicals in our body into reacting with one another. Without enzymes, our bodies would be unable to generate energy from the food we eat (and you’d not be reading this blog post).)

So, somehow, the information stored in a DNA molecule is deciphered by the cell and used as the recipe for a protein. But how?

To answer this question, let’s take a journey inside the cell. We can imagine a cell to be like a factory, but one that has been divided into a series of physically separated compartments. Unlike a factory filled with air, a cell is filled with a jelly-like fluid called the cytoplasm, which surrounds the various compartments enclosed within it. In an earlier post I likened a genome to a biological library. And, inside the cell, this library is stored within a particular compartment called the nucleus. 

I mentioned earlier that genes often act as recipes for proteins. But here comes a bit of a quandary: chromosomes – and the genes they contain – are locked away inside the cell’s nucleus. By contrast, proteins are manufactured by the cell in the cytoplasm, outside of the nucleus. So, for the genetic information to be used, it has to get out of nucleus and into the cytoplasm. How does this happen? Well, if we’re in a library with a book that contains information we really need, but we’re unable to take the book out of the library, we might make a photocopy of the page that holds the information we’re after. To get the information it needs out of the nucleus and into the cytoplasm the cell does something remarkably similar. The chromosome containing the gene of interest has to stay inside the nucleus, so the cell makes a copy of the gene – and that copy is then transported to where it is to be used: out of the nucleus and into the cytoplasm. 

The copy of the gene generated during this cellular photocopying is made not of DNA but of a close cousin called RNA. RNA is made of three of the same building blocks as DNA – A, C and G. Instead of the T found in DNA, however, RNA uses a different block represented by the letter U (for ‘uracil’). Despite this difference in building material, RNA stores biological information in the same way as DNA – by joining the building blocks together in a long chain (whereby the information is ‘coded’ in the ordering of the building blocks along the chain).

Let’s return to our cell, where a cellular photocopy has been made of a particular portion of a chromosome, a portion containing one particular gene, and it has been transported to the cytoplasm. We then face our next quandary: how is the information contained in our photocopy deciphered and used as a recipe for a protein?  

To answer this question we need to know what a protein is made of – whereupon we stumble upon a not-coincidental similarity with DNA and RNA. Proteins are also made of a series of building blocks joined together to form a long chain. However, the building blocks themselves are quite different. Unlike the four ‘letters’ of DNA and RNA – A, T, C and G for DNA, and A, U, C and G for RNA – proteins are made from twenty different building blocks, called amino acids. The role of the information stored in our gene (which has now been transferred to our cellular photocopy, RNA) is to determine the identity of each amino acid along the protein chain.

So, how does this happen? Enter a special biological machine called a ‘ribosome’. The ribosome is a protein assembly line: it constructs a protein by attaching one amino acid to another to form a chain-like structure. But it doesn’t just pick amino acids at random – picking any one of the twenty amino acids available to it, and bolting it to the end of the chain that it’s currently building. Instead, it ‘reads’ the RNA molecule – the ‘photocopy’ of our gene – and uses the information it contains to work out which amino acid should be added next. (I like to think of the ribosome as a modern-day Pac Man: just as the Pac Man of the 1980s computer game chomped its way along a string of blobs on the screen, the ribosome physically chomps its way along the RNA molecule, ‘reading’ the letters in sequence as it goes.)

But it’s not as straightforward as the ribosome reading the sequence of the RNA molecule one at a time,  and each letter of the RNA molecule referring neatly to one particular amino acid (after all, there are only four ‘letters’ in the RNA recipe, but twenty amino acids).  Instead, the ribosome scans the sequence of letters in the RNA molecule in groups of three; this three-letter barcode tells the ribosome which amino acid it needs to add next. So, a ribosome would encounter the sequence GCCUCAUGC and would read it three letters at a time to add the following three amino acids in sequence:
Alanine – Serine – Cysteine
[GCC]      [UCA]    [UGC]

This journey – from gene to protein – has brought us face-to-face with what some might consider the Holy Grail of biology. The ‘cipher’ used by the ribosome to decode a three-letter RNA sequence and translate it into the identity of an amino acid is called the ‘genetic code’ – a code so universal to life that it is used by every living organism on the planet, from the daffodils flowering in the garden as I write, to the cow whose milk is in the tea I’m currently drinking. It is a code that takes the simplicity of just four building blocks and opens it up into the complexity of life that each one of us represents.
 

Monday, 4 April 2011

What are genes and genomes?


I described in my last blog post how DNA acts as a store of biological information – information that serves as a set of instructions that direct our growth and function. Indeed, we could consider DNA to be the biological equivalent of a library – another repository of information with which we’re all probably much more familiar. The information we find in a library isn’t present in one huge tome, however. Rather, it is divided into discrete packages of information – namely books. And so it is with DNA: the biological information it stores isn’t captured in a single, huge molecule, but is divided into separate entities called chromosomes – the biological equivalent of individual books in a library.

I commented previously that DNA is composed of a long chain of four building blocks, A, C, G, and T. Rather than existing as an extended chain (like a stretched out length of rope), the DNA in a chromosome is tightly packaged. In fact, if stretched out (like our piece of rope), the DNA in a single chromosome would be around 2-8 cm long. Yet a typical chromosome is just 0.00002–0.002 cm long: that’s between 1000 and 100,000 times shorter than the unpackaged DNA would be. This packaging is quite the feat of space-saving efficiency.

Let’s return to our imaginary library of books. The information in a book isn’t presented as one long uninterrupted sequence of words. Rather, the information is divided into chapters. When we want to find out something from a book – to extract some specific information from it – we don’t read the whole thing cover-to-cover. Instead, we may just read a single chapter. In a fortuitous extension of our analogy, the same is true of information retrieval from chromosomes. The information captured in a single chromosome is stored in discrete ‘chunks’ (just as a book is divided into chapters), and these chunks can be read separately from one another. These ‘chunks’ – these discrete units of information – are what we call ‘genes’. In essence, one gene contains one snippet of biological information.

I’ve just likened chromosomes to books in a library. But is there a biological equivalent of the library itself? Well, yes, there is. Virtually every cell in the human body (with specific exceptions) contains 46 chromosomes – 23 from each of its parents. All of the genes found in this ‘library’ of chromosomes are collectively termed the ‘genome’. Put another way, a genome is a collection of all the genes found in a particular organism.

Different organisms have different-sized genomes. For example, the human genome comprises around 20,000-25,000 genes; the mouse genome, with 40 chromosomes, comprises a similar number of individual genes. However, the bacterium H. influenzae has just a single chromosome, containing around 1700 genes.

It is not just the number of genes (and chromosomes) in the genome that varies between organisms: the long stretches of DNA making up the genomes of different organisms have different sequences (and so store different information). These differences make sense, particularly if we imagine the genome of an organism to represent the ‘recipe’ for that organism: a human is quite a different organism from a mouse, so we would expect the instructions that direct the growth and function of the two organisms to differ. 

But we shouldn’t be fooled into thinking that individuals that look quite different to the observer must have vastly different genomes. Imagine walking along a busy shopping street on a Saturday afternoon. If you’re anything like me, you’ll find the simple observation of passers-by to be an endlessly fascinating pastime – an ever-changing mix of physical appearance and behaviour. Yet, if we scratch the surface, this seemingly endless variety belies a resoundingly similar genome. I mentioned in my previous post how the human genome is made up of around 3 billion (3,000,000,000) building blocks. However, between individual humans, our DNA differs by just 0.2%: around 2,994,000,000 of our building blocks (or 99.8%) are the same, and just 6,000,000 are different. And that handful of differences – that 0.2% – is enough to generate the huge variety we see in the population around us. It’s pretty amazing (to my mind, at least). 

The completion of the draft of the human genome (a result of the human genome sequencing project) was announced jointly by the then President of the US Bill Clinton and Prime Minister of the UK Tony Blair ten years ago, and the intervening years have witnessed the publication of a range of other genome sequences – from mouse, to dog, to chicken. As more genome sequences become available so we begin to learn more and more about just how similar (or different) the variety of life on Earth is at the level of our genes – all of which provide further clues to our evolution, and the evolution of the other species with which we share this planet. But perhaps that’s a topic for another post.