This is the answered version of the Translational Biology practical. Every ❓ question is followed by a worked model answer (green box). All the interactive steps still work, so you can keep looking genes up and aligning proteins while you read, and the self-test quiz at the end is unchanged.
1 Why you would go looking for the same gene twice
In the previous practical you ended up with a group of Arabidopsis thaliana
genes that respond to a bacterial infection, and a suggestion of what those genes are
doing. That is a real result. It is also a result about a small roadside weed that
nobody eats and nobody grows.
Almost everything we know about how plants defend themselves was worked out in
Arabidopsis, and there is a good reason for that. Compare it with bread wheat,
one of the crops the answer is eventually meant for.
Arabidopsis thaliana
Bread wheat
Genome size
135 million bases
16,000 million bases
Sets of chromosomes
2
6
Protein-coding genes
about 27,000
about 107,000
Seed to seed
6 to 8 weeks
5 to 8 months
Space for 1,000 plants
a bench
a field
Adding or removing a gene
routine
slow and often variety-specific
That is why the experiments happen in Arabidopsis. A PhD student can run six
generations in the time a wheat breeder runs one, in a growth cabinet instead of a field.
The catch is obvious: the finding is now in the wrong plant. Getting it into the right
one means finding, in wheat, the gene that corresponds to the Arabidopsis gene
you care about. That gene is called an orthologue, and this practical is
about what that word means, how you find one, and how far you can trust it.
The 500 genes from the previous two practicals, looked up in Ensembl Plants:
❓ Questions
Why is finding orthologues of A. thaliana genes important for crop
resilience research? Try to phrase it as what you would not be able to do
without them.
Look at the numbers above. Which of the differences between the two plants do you
think matters most for how quickly research gets done, and why?
✅ Answer
Why orthologues matter. Turn it around and ask what you could do without them: almost nothing. Every result you produced this week is about Arabidopsis, and nobody grows Arabidopsis. Without a way of saying "this gene here corresponds to that gene there", a finding in the model plant stays in the model plant. The orthologue is the bridge. It tells a wheat breeder which of their 107,000 genes to look at, which one to screen their varieties for, which one to knock out or to select on. It also works in the other direction: if a gene is present and similar in many species, that is evidence it does something worth keeping, and a gene that only exists in Arabidopsis is a much weaker bet for a crop programme.
Which difference matters most. There is no single right answer, but the two that usually come first are generation time and how easy it is to add or remove a gene. Six to eight weeks from seed to seed against five to eight months means you can run six experiments in the time a wheat researcher runs one, and genetics is a subject where you often need several rounds of crossing before you learn anything. Being able to change a gene routinely is what lets you go from "this gene is correlated with resistance" to "this gene causes resistance", which is the step that actually settles the question. The genome size and the six sets of chromosomes matter too, but mostly because they make everything else harder: with six copies of every chromosome, knocking out one copy of a gene often changes nothing at all, which is a problem you will meet again in step 6.
2 Orthologue or paralogue: it depends on the history
Two genes in two different genomes can look almost identical and still not be the pair
you want. Whether they are depends entirely on how the two copies came to exist,
and there are only two ways that happens.
A species splits in two. One ancestral population becomes two separate
species. Every gene the ancestor had is now sitting in both of them, one copy each. The
two copies start out identical and slowly drift apart as the two species go their own
way. A pair of genes related like this is called a pair of orthologues.
A gene gets copied inside one genome. Copying accidents are common in
plants: a stretch of DNA gets duplicated, or a chromosome, or occasionally the entire
genome at once. Afterwards that one genome carries two copies of the same gene. A pair
of genes related like this is called a pair of paralogues.
Both kinds of pair descend from one ancestral gene, so both will look similar and both
will turn up if you simply search for lookalike sequences. Only one of the two is the
pair you want when you are moving a result from Arabidopsis into a crop. The
diagram below lets you build the history yourself and see what comes out.
History:
species split gene copied orthologues paralogues
❓ Questions
In your own words, what is the difference between an orthologue and a paralogue?
What has to have happened for a pair to be one rather than the other?
You want to know which wheat gene does the job that your Arabidopsis gene
does. Which of the two kinds of pair are you after, and why does the other kind not
answer your question?
Compare the last two histories. The four genes at the bottom are the same in both,
and only the order of events changed. Why does that change which pairs are
orthologues?
💡 Hint if the last one will not come out
Pick two genes, put a finger on each, and trace both fingers upwards until they
meet. The question is only ever what kind of event you land on when they do.
✅ Answer
The difference. Both are pairs of genes that descend from one single gene in some ancestor. What separates them is the event that made the two copies. If the two copies were separated because the species split in two, they are orthologues. If the two copies were made because the gene was copied inside one genome, they are paralogues. That is the entire definition, and notice that it says nothing about how similar the two sequences are. Two paralogues can be nearly identical and two orthologues can be quite different; similarity is a clue to the history, not the history itself.
Which one you want. You want the orthologue. The reasoning is about what the two kinds of pair have been doing since they separated. Orthologues have been sitting in two different species, each doing the same job the ancestral gene did, because there is only one copy in each species and nothing else to take that job over. Paralogues sit side by side in one genome, where one copy is enough to keep the plant running, so the second copy is free to drift, pick up a new role, or stop working altogether. So the orthologue is the pair with a reason to have kept the same function, and the paralogue is the pair with a reason not to. That is a tendency rather than a law, which is what step 6 is about.
Why the order of events changes everything. In "copied in both", the split happened first. Every gene on the Arabidopsis side and every gene on the rice side meet at that split when you trace them upwards, so all four cross-species pairs are orthologues, and because there are two copies on each side that is a many-to-many relationship. In "copied before the split", the copying happened first. Trace A1 and R1 upwards and you meet a species split, so they are orthologues, and the same is true of A2 and R2. But trace A1 and R2 upwards and you go past both splits before the two paths meet, at the copying event, so those two are paralogues even though they are in different species. The four genes at the bottom are identical in both pictures. What differs is only which event you land on when you trace two of them back, and that is what the words are defined by. This is also exactly why the tree-based method in step 3 beats simply looking for the most similar sequence: A1 and R2 might well be each other's best match, and they would still be the wrong answer.
3 How an orthologue is actually found
Nobody has written down a list of which wheat gene corresponds to which
Arabidopsis gene. There are hundreds of thousands of genes involved and the two
plants last shared an ancestor around 150 million years ago. The correspondence has to be
worked out from the sequences, and there are two ways of doing it.
Best match in both directions. Take the protein your
Arabidopsis gene codes for and search all wheat proteins for the one that looks
most like it. Then turn around: take that wheat protein and search all
Arabidopsis proteins. If you land back on the gene you started from, the two are
each other's best match and you call them orthologues. This is quick and it is what
"reciprocal best hit" means. It also assumes there is exactly one best answer on each
side, which step 2 should make you suspicious of.
Build the family tree. Collect every similar sequence from every species
you have, build a tree of how they are related, and then go through the tree branch point
by branch point deciding whether each one was a species split or a copying event. That is
the same exercise you just did by hand in step 2, done automatically for tens of
thousands of gene families. It is slower and it is what Ensembl Compara
does, and it is where the one-to-one, one-to-many and many-to-many labels come from.
Everything below this point is that second method's real output. Pick one of the 500
genes from the previous practicals and see what Ensembl has for it.
Pick a gene
Steps 4 and 5 follow this choice, so you can come back and change it at any time. If you
picked a gene you liked in the GO enrichment practical, use that one.
Identity is the percentage of the Arabidopsis protein that is
exactly the same in the other species. High confidence is Ensembl's own
flag for the orthologues its tree supports most strongly.
❓ Questions
Does the gene you picked have orthologues? In which species?
Repeat this for a few more genes, including some you pick at random rather than
choose. Why do you think some genes have far more orthologues than others?
Both methods above rely only on how similar two sequences are. What could go wrong
with that, given what you saw in step 2?
✅ Answer
Does your gene have orthologues? Almost certainly yes: 479 of the 500 genes have at least one in these ten species, so if you picked at random you probably got one. Which species you got depends on the gene. Cabbage is the most reliable, missing for only 38 of the 500, which makes sense because it is in the same plant family as Arabidopsis and separated from it most recently. Barley is the least reliable, missing for 240 of them. The 21 genes with nothing anywhere are worth a look on their own.
Why some genes have far more than others. Try AT1G79040 (PSBR) and then AT3G53260 (PAL2) to see the extremes: 7 orthologues against 149. Several things drive the difference and they stack up. Some genes belong to large families that have been copied over and over in every lineage, so one Arabidopsis gene meets dozens of relatives in each species. Some genomes have been doubled or tripled more recently than others, which multiplies every gene in them at once. Some genes are so central to staying alive that every species has kept one, while others are specific to a way of life and get lost in species that do not need them. And some of the difference is not biology at all: a well-studied genome is annotated better than a poorly studied one, so it is easier to find orthologues in it.
What can go wrong with similarity alone. Step 2 gave the answer: the most similar sequence is not always the orthologue. In the "copied before the split" history, A1 and R2 are paralogues, yet nothing about their sequences announces that, and if the copying event was recent they may well look more alike than the true orthologue pair. Similarity also gets misread in the other direction: two orthologues that have both changed a lot no longer look like each other's best match, so a real relationship gets missed. Similarity is evidence about history, and treating evidence as if it were the conclusion is the mistake.
4 The same gene, mapped onto the tree of plants
The table in step 3 is a list, and a list hides the thing that explains it. These ten
species are not ten independent samples: they are the tips of a tree, some of them close
cousins and some of them separated by hundreds of millions of years. Putting the counts
back on the tree usually makes the pattern obvious.
Each branch point below is a species split. The number beside each species is how many
orthologues of your selected gene Ensembl finds there, and the colour is the relationship
it reports. A species with nothing at all is worth as much of your attention as one with
six.
one-to-one one-to-many many-to-many none found
Two genes from the same dataset that behave very differently. Load one, look at the tree,
then load the other:
❓ Questions
What does this tree actually show? Be precise: what is a branch point, and what is
a number beside a species?
Why are orthologous relations sometimes one-to-one, sometimes one-to-many and
sometimes many-to-many? Use the two contrasting genes above to make the point.
Wheat almost always has more copies than the other species. Look back at the table
in step 1 for a reason why.
If a species shows no orthologue at all, name two quite different things that could
be true. Which would you check first?
✅ Answer
What the tree shows. The tree is a picture of how the ten species are related to each other. Each branch point is a moment in the past when one ancestral species split into two, and the further left a branch point is, the longer ago that happened. It is not a picture of your gene: the same tree is drawn no matter which gene you select. What changes with your gene is the annotation on the right. The number beside a species is how many genes in that species Ensembl calls orthologues of your one Arabidopsis gene, and the colour is the relationship it reports for them.
Why the relationship differs. It comes straight out of step 2, applied to the real tree. One-to-one means one copy on each side of the species split, so nothing was copied afterwards in either lineage. One-to-many means one copy stayed single in Arabidopsis while the other lineage copied it. Many-to-many means both lineages copied it. Compare the two contrasting genes: AT1G79040 (PSBR) has at most two copies anywhere and is missing entirely from the grasses and moss, while AT3G53260 (PAL2) is a member of a large family that has been copied repeatedly in every lineage, giving 54 copies in wheat alone. Same tree, same method, completely different picture, because the two genes have had completely different histories.
Why wheat has more of everything. Look at the table in step 1: bread wheat has six sets of chromosomes where Arabidopsis has two. Bread wheat formed when three related grass species combined their whole genomes, so it carries three near-complete copies of a grass genome at once. Every gene is therefore present roughly three times before you even start counting gene families, and in this dataset wheat averages 8.8 orthologues per gene where every other species averages between 1.5 and 3.6. This is the single most common reason a crop gives you a one-to-many result, and it is not a quirk of wheat: potato, soybean, maize and cabbage have all been through genome doublings of their own.
Nothing found. Two quite different things could be true, and they call for different responses. Either the gene genuinely is not there, because it was lost in that lineage or never existed before it, or the gene is there and we failed to find it, because the genome is incompletely sequenced, badly annotated, or the sequence has changed so much that the search no longer recognises it. Check the second one first, because it is much more common and much cheaper to check. A good test is to look at the neighbours on the tree: CAB1 has 54 orthologues in wheat and 3 in rice but none in barley, and it is not credible that barley alone among the grasses lost a photosynthesis gene. That is an annotation gap, not biology. When a whole branch of the tree comes back empty, that is when a real loss becomes the better explanation.
5 Look at the sequences yourself
Ensembl says these two genes are orthologues. That claim rests entirely on their
sequences, so it is worth putting them side by side and forming your own opinion instead
of taking the label on trust.
An alignment writes the two proteins one above the other and slides them
along until as many positions as possible line up, inserting gaps where one of the two
has gained or lost a stretch. Each letter is one amino acid, the building blocks a
protein is made of. The alignment below is computed in your browser when you press the
button, using the same kind of algorithm the databases use.
This uses the gene you picked in step 3. For each species the page carries the
closest of its orthologues, which is the one worth looking at; the rest are
still listed with their identities in the table above. Sequences are available for six of
the ten species.
Press the button to build the alignment.
The identity figure counts identical positions as a share of the Arabidopsis protein, which is how Ensembl counts it too, so it matches the number in the table in step 3.
identical different, but a chemically similar amino acid different gap, present in one protein only
❓ Questions
Looking at this alignment, do you agree that these two genes are orthologues? What
in the picture makes you say so?
Why is the alignment not a perfect match? Give a reason for the single differing
positions and a reason for the gaps.
Align the same gene against a close relative and against a distant one. What
changes, and does the amount of change fit the tree in step 4?
Some regions match almost perfectly while others are a mess. What might that tell
you about those parts of the protein?
✅ Answer
Do you agree? For most pairs, yes, and the reason is the pattern rather than the percentage. What convinces you is a long stretch of the protein matching almost letter for letter, running the full length of both sequences, with the mismatches scattered as single positions rather than piled up in one place. Two unrelated proteins do not do that. Around 20 to 25 percent of positions match by chance alone between any two protein sequences, so a figure near that means nothing, while 87 percent running end to end between two species that separated 150 million years ago is not something chance produces.
Why it is not perfect. The two lineages have been evolving separately ever since the species split, and two different kinds of change accumulate. Single differing positions come from point mutations, where one letter of the DNA changed and the amino acid at that position changed with it. Most of these are harmless, which is why they survive; notice how many of them are the "similar" colour, meaning the replacement amino acid has much the same chemical character and the protein carries on working. Gaps come from a different kind of event: a piece of DNA was inserted into one lineage or deleted from the other, so one protein has a stretch the other does not. Gaps in the middle of a protein are relatively rare because losing a chunk of a working protein usually breaks it, and you will notice that most gaps sit near the ends.
Close against distant. The closer the species, the higher the identity and the fewer the gaps. For CAB1 the cabbage orthologue comes in around 97 percent while the rice one is around 87 percent, and that ordering follows the tree in step 4 exactly: less time since the split means less time to accumulate change. This is worth noticing because it is the same logic running in reverse. We use trees to interpret sequences, and sequence differences are how the trees were built in the first place.
Conserved and variable regions. The regions that match almost perfectly are almost always the parts that have to be right: the active site of an enzyme, the surface where the protein grips another molecule, the core that holds it folded. A mutation there breaks the protein and the plant carrying it does worse, so those changes never spread. The messy regions are the parts where the exact sequence matters less, such as flexible linkers between the working parts, or the signal at the start of the protein that says which compartment of the cell to send it to, where the general character matters but the individual letters do not. So an alignment is not only evidence about ancestry, it is also a rough map of which parts of a protein matter, which is genuinely useful if you are deciding where to aim a mutation.
6 Do it for real, and know what it does not tell you
Everything so far ran on data packaged into this page. Ensembl Plants is where that data
came from and where you would go in practice, so open the real pages for the gene you
picked. These three links are built from your selection in step 3 and go straight there.
Summary is the gene itself: where it sits in the genome, what it is
called, and what is known about it.
Orthologues is the table from step 3, for all of Ensembl's species
rather than ten. Each row has a View Sequence Alignment link, which is
the real version of step 5. Try it on a species this page does not carry.
Gene tree is the evolutionary history the orthologues were derived
from, and the real version of the diagram you built in step 2. Blue nodes are species
splits and red nodes are gene copies, the same colours used here.
One last thing, and it is the thing that most often trips people up.
Orthologue means shared ancestry, not shared job. The definition is
entirely about history: these two genes descend from one gene in one ancestor. That is
all it claims. It is a good bet that they still do something similar, because that is
usually how it works out, and it is only a bet.
Genes drift into new roles. A gene may be switched on in the root in one species and in
the leaf in another, or matter enormously to a plant that lives in the cold and not at
all to one that does not. When a gene sits in a family of near-identical copies, as most
of the ones you just looked at do, the copies often divide the original job between them,
so no single one of them is the equivalent of your Arabidopsis gene. The
orthologue is where you start looking in the crop. Whether it does the same thing is a
question you answer with an experiment.
❓ Questions
Take one of your orthologues and see whether you can find any published work on it.
Does what people report about it match what the gene does in A. thaliana?
Suppose the orthologue in wheat turns out to have three near-identical copies, and
you can only afford to knock out one of them. What would you expect to happen, and
what would that do to your experiment?
Pulling the week together: you have a cluster of co-expressed genes, a GO term that
describes them, and now orthologues in a crop. Write down, in three or four
sentences, the experiment you would propose next.
✅ Answer
The literature check. What you find depends entirely on the gene, and the honest outcome for most of them is that nothing has been published on the crop orthologue at all. That is itself the answer to why this practical exists. Where you do find something, the usual pattern is a broad match with a specific mismatch: the crop gene turns out to be involved in the same general process, but switched on at a different time, in a different tissue, or in response to a different stress. Treat a match as encouragement rather than proof, and treat a mismatch as interesting rather than as an error, because a gene that has been repurposed in a crop is worth understanding. Also check what kind of evidence you have found: a paper that measured the gene is worth far more than a database entry that inferred its function from the Arabidopsis gene, which would just be your own assumption handed back to you.
Knocking out one of three copies. Most likely nothing visible happens, and your experiment tells you nothing. The other two copies are still there and still doing the job, so the plant carries on as before. This is called redundancy and it is the standard frustration of working in wheat: the six sets of chromosomes that make the genome big also mean that single knockouts are usually silent. The consequences are practical. You may need to knock out all three copies together to see any effect at all, which is far more work. A negative result from a single knockout is close to uninformative, so you should not publish one as evidence that the gene does not matter. And if you have to choose, it can be smarter to work in a crop with fewer genome copies first, such as barley or rice, and move to wheat once you know what you are looking for.
The experiment. A reasonable answer chains the whole week together and stays concrete. For example: the clustering put a group of genes together that rise in the AVR treatment at 6 and 12 hours but not in MOCK, GO enrichment says that group is enriched for programmed cell death, and Ensembl gives orthologues of the three strongest of those genes in barley, which is a real crop with only two sets of chromosomes. So: obtain or generate knockout lines for those three barley orthologues, infect them and wild-type barley with a comparable pathogen, and measure both the disease itself and the process the GO term pointed at, for instance by staining for cell death, at the same time points the original experiment used. Include a mock-inoculated control at every time point, for the reason step 4 of the GO practical made painfully clear. The point to get across is the shape of the argument: clustering says which genes move together, GO says what they are probably doing, the orthologue says where to test it, and only the experiment says whether it is true.