This blog has been created partly as a companion to Chemistry for the Biosciences, the textbook that I co-author with Tony Bradshaw, and to act as an archive of posts I write for other sites (particularly the OUPblog). Like the book itself, it explores how life on the scale of atoms and molecules has an impact on biology - at the scale of cells, tissues, and organisms - and seeks to demystify a range of biological and chemical concepts.

The blog's name takes as its inspiration the cover of the first edition of Chemistry for the Biosciences, which depicts a gecko seemingly clinging to its surface. To find out what links geckos to chemistry, read this.



Sunday, 1 May 2011

How is the information in a gene used by a cell?

In my last two posts I’ve introduced the notion that DNA acts as a store of biological information; this information is stored in a series of chromosomes, each of which are divided into a number of genes. Each gene in turn contains one ‘snippet’ of biological information. But how are these genes actually used? How is the information stored in these genes actually extracted to do something useful (if ‘useful’ isn’t too flippant a term for something that the very continuation of life depends upon).

Many (but not all) genes act as recipes for a family of biological molecules called proteins: they literally tell the cell what the ingredients for a particular protein are, and how they should be combined to create the protein itself. (Proteins have a range of essential roles in the human body. Some act as building materials for different components of the body, such as the keratin we find in our hair and nails. Others act as molecular transporters: haemoglobin, which is found in our red blood cells, carries oxygen from our lungs to other parts of the body. A family of proteins called the enzymes are arguably the most important, however. Enzymes cajole different chemicals in our body into reacting with one another. Without enzymes, our bodies would be unable to generate energy from the food we eat (and you’d not be reading this blog post).)

So, somehow, the information stored in a DNA molecule is deciphered by the cell and used as the recipe for a protein. But how?

To answer this question, let’s take a journey inside the cell. We can imagine a cell to be like a factory, but one that has been divided into a series of physically separated compartments. Unlike a factory filled with air, a cell is filled with a jelly-like fluid called the cytoplasm, which surrounds the various compartments enclosed within it. In an earlier post I likened a genome to a biological library. And, inside the cell, this library is stored within a particular compartment called the nucleus. 

I mentioned earlier that genes often act as recipes for proteins. But here comes a bit of a quandary: chromosomes – and the genes they contain – are locked away inside the cell’s nucleus. By contrast, proteins are manufactured by the cell in the cytoplasm, outside of the nucleus. So, for the genetic information to be used, it has to get out of nucleus and into the cytoplasm. How does this happen? Well, if we’re in a library with a book that contains information we really need, but we’re unable to take the book out of the library, we might make a photocopy of the page that holds the information we’re after. To get the information it needs out of the nucleus and into the cytoplasm the cell does something remarkably similar. The chromosome containing the gene of interest has to stay inside the nucleus, so the cell makes a copy of the gene – and that copy is then transported to where it is to be used: out of the nucleus and into the cytoplasm. 

The copy of the gene generated during this cellular photocopying is made not of DNA but of a close cousin called RNA. RNA is made of three of the same building blocks as DNA – A, C and G. Instead of the T found in DNA, however, RNA uses a different block represented by the letter U (for ‘uracil’). Despite this difference in building material, RNA stores biological information in the same way as DNA – by joining the building blocks together in a long chain (whereby the information is ‘coded’ in the ordering of the building blocks along the chain).

Let’s return to our cell, where a cellular photocopy has been made of a particular portion of a chromosome, a portion containing one particular gene, and it has been transported to the cytoplasm. We then face our next quandary: how is the information contained in our photocopy deciphered and used as a recipe for a protein?  

To answer this question we need to know what a protein is made of – whereupon we stumble upon a not-coincidental similarity with DNA and RNA. Proteins are also made of a series of building blocks joined together to form a long chain. However, the building blocks themselves are quite different. Unlike the four ‘letters’ of DNA and RNA – A, T, C and G for DNA, and A, U, C and G for RNA – proteins are made from twenty different building blocks, called amino acids. The role of the information stored in our gene (which has now been transferred to our cellular photocopy, RNA) is to determine the identity of each amino acid along the protein chain.

So, how does this happen? Enter a special biological machine called a ‘ribosome’. The ribosome is a protein assembly line: it constructs a protein by attaching one amino acid to another to form a chain-like structure. But it doesn’t just pick amino acids at random – picking any one of the twenty amino acids available to it, and bolting it to the end of the chain that it’s currently building. Instead, it ‘reads’ the RNA molecule – the ‘photocopy’ of our gene – and uses the information it contains to work out which amino acid should be added next. (I like to think of the ribosome as a modern-day Pac Man: just as the Pac Man of the 1980s computer game chomped its way along a string of blobs on the screen, the ribosome physically chomps its way along the RNA molecule, ‘reading’ the letters in sequence as it goes.)

But it’s not as straightforward as the ribosome reading the sequence of the RNA molecule one at a time,  and each letter of the RNA molecule referring neatly to one particular amino acid (after all, there are only four ‘letters’ in the RNA recipe, but twenty amino acids).  Instead, the ribosome scans the sequence of letters in the RNA molecule in groups of three; this three-letter barcode tells the ribosome which amino acid it needs to add next. So, a ribosome would encounter the sequence GCCUCAUGC and would read it three letters at a time to add the following three amino acids in sequence:
Alanine – Serine – Cysteine
[GCC]      [UCA]    [UGC]

This journey – from gene to protein – has brought us face-to-face with what some might consider the Holy Grail of biology. The ‘cipher’ used by the ribosome to decode a three-letter RNA sequence and translate it into the identity of an amino acid is called the ‘genetic code’ – a code so universal to life that it is used by every living organism on the planet, from the daffodils flowering in the garden as I write, to the cow whose milk is in the tea I’m currently drinking. It is a code that takes the simplicity of just four building blocks and opens it up into the complexity of life that each one of us represents.
 

Monday, 4 April 2011

What are genes and genomes?


I described in my last blog post how DNA acts as a store of biological information – information that serves as a set of instructions that direct our growth and function. Indeed, we could consider DNA to be the biological equivalent of a library – another repository of information with which we’re all probably much more familiar. The information we find in a library isn’t present in one huge tome, however. Rather, it is divided into discrete packages of information – namely books. And so it is with DNA: the biological information it stores isn’t captured in a single, huge molecule, but is divided into separate entities called chromosomes – the biological equivalent of individual books in a library.

I commented previously that DNA is composed of a long chain of four building blocks, A, C, G, and T. Rather than existing as an extended chain (like a stretched out length of rope), the DNA in a chromosome is tightly packaged. In fact, if stretched out (like our piece of rope), the DNA in a single chromosome would be around 2-8 cm long. Yet a typical chromosome is just 0.00002–0.002 cm long: that’s between 1000 and 100,000 times shorter than the unpackaged DNA would be. This packaging is quite the feat of space-saving efficiency.

Let’s return to our imaginary library of books. The information in a book isn’t presented as one long uninterrupted sequence of words. Rather, the information is divided into chapters. When we want to find out something from a book – to extract some specific information from it – we don’t read the whole thing cover-to-cover. Instead, we may just read a single chapter. In a fortuitous extension of our analogy, the same is true of information retrieval from chromosomes. The information captured in a single chromosome is stored in discrete ‘chunks’ (just as a book is divided into chapters), and these chunks can be read separately from one another. These ‘chunks’ – these discrete units of information – are what we call ‘genes’. In essence, one gene contains one snippet of biological information.

I’ve just likened chromosomes to books in a library. But is there a biological equivalent of the library itself? Well, yes, there is. Virtually every cell in the human body (with specific exceptions) contains 46 chromosomes – 23 from each of its parents. All of the genes found in this ‘library’ of chromosomes are collectively termed the ‘genome’. Put another way, a genome is a collection of all the genes found in a particular organism.

Different organisms have different-sized genomes. For example, the human genome comprises around 20,000-25,000 genes; the mouse genome, with 40 chromosomes, comprises a similar number of individual genes. However, the bacterium H. influenzae has just a single chromosome, containing around 1700 genes.

It is not just the number of genes (and chromosomes) in the genome that varies between organisms: the long stretches of DNA making up the genomes of different organisms have different sequences (and so store different information). These differences make sense, particularly if we imagine the genome of an organism to represent the ‘recipe’ for that organism: a human is quite a different organism from a mouse, so we would expect the instructions that direct the growth and function of the two organisms to differ. 

But we shouldn’t be fooled into thinking that individuals that look quite different to the observer must have vastly different genomes. Imagine walking along a busy shopping street on a Saturday afternoon. If you’re anything like me, you’ll find the simple observation of passers-by to be an endlessly fascinating pastime – an ever-changing mix of physical appearance and behaviour. Yet, if we scratch the surface, this seemingly endless variety belies a resoundingly similar genome. I mentioned in my previous post how the human genome is made up of around 3 billion (3,000,000,000) building blocks. However, between individual humans, our DNA differs by just 0.2%: around 2,994,000,000 of our building blocks (or 99.8%) are the same, and just 6,000,000 are different. And that handful of differences – that 0.2% – is enough to generate the huge variety we see in the population around us. It’s pretty amazing (to my mind, at least). 

The completion of the draft of the human genome (a result of the human genome sequencing project) was announced jointly by the then President of the US Bill Clinton and Prime Minister of the UK Tony Blair ten years ago, and the intervening years have witnessed the publication of a range of other genome sequences – from mouse, to dog, to chicken. As more genome sequences become available so we begin to learn more and more about just how similar (or different) the variety of life on Earth is at the level of our genes – all of which provide further clues to our evolution, and the evolution of the other species with which we share this planet. But perhaps that’s a topic for another post.

Tuesday, 1 March 2011

What is DNA, and what does it do?


We’ve all heard of DNA, and probably know that it’s ‘something to do with our genes’. But what actually is DNA, and what does it do? At the level of chemistry, DNA - or deoxyribonucleic acid, to give it its full name – is a collection of carbon, hydrogen, oxygen, nitrogen and phosphorus atoms, joined together to form a large molecule. There is nothing that special about the atoms found in a molecule of DNA: they are no different from the atoms found in the thousands of other molecules from which the human body is made. What makes DNA special, though, is its biological role: DNA stores information – specifically, the information needed by a living organism to direct its correct growth and function. 

But how does DNA, simply a collection of just a few different types of atom, actually store information? To answer this question, we need to consider the structure of DNA in a little more detail. DNA is like a long, thin chain – a chain that is constructed from a series of building blocks joined end-to-end. (In fact, a molecule of DNA features two chains, which line up side-by-side. But we only need to focus on one of these chains to be able to understand how DNA stores its information.) 

There are only four different building blocks; these are represented by the letters A, C, G and T. (Each building block has three component parts; one of these parts is made up of one of four molecules: adenine, cytosine, guanine or thymine. It is these names that give rise to letters used to represent the four complete building blocks themselves.)  A single DNA molecule is composed of a mixture of these four building blocks, joined together one by one  to form a long chain – and it is the order in which the four building blocks are joined together along the DNA chain that lies at the heart of DNA’s information-storing capability.

The order in which the four building blocks appear along a DNA molecule determines what we call its ‘sequence’; this sequence is represented using the single-letter shorthand mentioned above. If we imagine that we had a very small DNA molecule that is composed of just eight building blocks, and these blocks were joined together in the order cytosine-adenine-cytosine-guanine-guanine-thymine-adenine-cytosine, the sequence of this DNA molecule would be CACGGTAC.

The biological information stored in a DNA molecule depends upon the order of its building blocks – that is, its sequence. If a DNA sequence changes, so too does the information it contains. On reflection, this concept – that the order in which a selection of items appears in a linear sequence affects the information stored in that sequence – may not be as alien to us as it might first seem. Indeed, it is the concept on which written communication is based: each sentence in this blog post is composed of a selection of items – the letters of the alphabet – appearing in different sequences. These different sequences of letters spell out different words, which convey different information to the reader. And so it is with the sequence of DNA: as the sequence of the four building blocks of DNA varies, so too does the information being conveyed. (You may well be asking how the information stored in DNA is actually interpreted – how it actually determines how an organism develops and functions – but that’s a topic for a different blog post.)

You may be wondering how on earth just four different building blocks can come together with such variety to capture all the information needed to direct the growth of a living organism. Well, let’s pause for a moment to look back at our eight-letter DNA sequence: CACGGTAC. Notice that there isn’t an equal mix of the four different building blocks: cytosine appears three times, adenine and guanine twice, and thymine just once. There isn’t a rule that says a C must always be followed by an A, or that a T must always be followed by a G. Instead, any of the four building blocks can appear as the next link in a DNA chain (or, put another way, each of the four building blocks of DNA can appear at each position along a DNA chain). 

Imagine we had a DNA molecule just two blocks long. Even with this tiny molecule, being able to draw upon four different building blocks at each of the two positions along the two-block chain makes 16 different molecules possible:
AA          AC          AT           AG
CA          CC          CT           CG
TA          TC          TT           TG
GA          GC         GT           GG
If we scaled this up to our eight-block molecule, we’d find that a whopping 65536 different combinations would be possible. (I won’t write them all out.) 

When we consider that the DNA in a human cell is made up of around 3 billion of the four building blocks joined in sequence (that’s 3,000,000,000), we can begin to imagine just how much variety is actually possible – and all from just four starting ingredients. This is a trend we see throughout the natural world: of relative simplicity giving rise to quite remarkable complexity – as complex and sophisticated as the living organisms we each represent. 

This post first appeared on the OUPblog