Fundamentholfundamenthol

Genetic Engineering

BiologyMolecular Basis of InheritanceFor NEET aspirants

Genetic engineering tools made it possible to read the whole human genome and to compare the DNA of any two people. This page covers the Human Genome Project: why it was a mega project, its goals, its methods (ESTs, sequence annotation, BAC and YAC vectors, Sanger sequencing) and the salient features of the human genome. It then explains DNA fingerprinting: repetitive and satellite DNA, polymorphism, VNTRs, the six steps of the technique and its uses. NEET often asks HGP numbers and fingerprinting steps from this part of genetic engineering.

On this page1Why sequence2Mega project3Goals4Methods5Salient features6Future7Fingerprinting8Polymorphism9Technique10Exam essentials11Quick revision12Solved examples13Practice
Key Points at a Glance
  1. ★ Must learn The Human Genome Project (HGP) began in 1990 and was completed in 2003, a 13-year project.
  2. Mega project: about bp at 3 US dollars per bp 9 billion US dollars; 3300 books of 1000 pages to store one sequence.
  3. The HGP drove the rise of a new field, bioinformatics.
  4. ★ Must learn Two approaches: Expressed Sequence Tags (ESTs) for expressed genes, and sequence annotation of the whole genome.
  5. Fragments were cloned in BAC and YAC vectors and read by automated sequencers based on Sanger's method.
  6. ★ Must learn Human genome: 3164.7 million bp, about 30,000 genes, 99.9% identical in all humans, under 2% coding for proteins.
  7. Chromosome 1 has the most genes (2968) and Y the fewest (231); about 1.4 million SNPs.
  8. ★ Must learn Satellite DNA: repetitive DNA seen as small peaks in density gradient centrifugation; highly polymorphic.
  9. DNA polymorphism: more than one allele at a locus, at a frequency greater than 0.01 in the population.
  10. ★ Must learn Alec Jeffreys: DNA fingerprinting with a VNTR (mini-satellite) probe by Southern blot hybridisation.

1. The Human Genome Project

  • The sequence of bases in DNA carries the genetic information of an organism; an individual's genetic make-up lies in its DNA sequence.
  • If two individuals differ, their DNA sequences must also differ, at least at some places.
  • These ideas led to the quest to find the complete DNA sequence of the human genome.
  • Genetic engineering techniques made it possible to isolate and clone any piece of DNA, and simple, fast methods of DNA sequencing became available.
  • ★ Exam imp So a very ambitious project to sequence the human genome was launched in 1990.
  • Determining the complete nucleotide sequence of the human genome began a new era of genomics.

1.1 Why the HGP was a mega project

  • The human genome has about bp.
  • ★ Exam imp At the cost estimated at the start, 3 US dollars per bp, the project would cost about 9 billion US dollars.
  • Suppose the sequence were typed in books, with 1000 letters on each page and 1000 pages in each book.
  • Then 3300 such books would be needed to store the DNA sequence of a single human cell.
  • The huge amount of data needed high-speed computers for storage, retrieval and analysis.
  • So the HGP was closely tied to the rapid growth of a new area of biology, bioinformatics.
Tips and Tricks Both numbers come from simple arithmetic. Cost: bp 3 US dollars = US dollars. Books: one book holds = letters, so letters (the haploid content) need 3300 books.

1.2 Goals of the HGP

  1. Identify all the approximately 20,000-25,000 genes in human DNA.
  2. Determine the sequences of the 3 billion chemical base pairs that make up human DNA.
  3. Store this information in databases.
  4. Improve tools for data analysis.
  5. Transfer related technologies to other sectors, such as industries.
  6. Address the ethical, legal and social issues (ELSI) that may arise from the project.
Memory Trick I Did Study In The Afternoon: Identify, Determine, Store, Improve, Transfer, Address (ELSI).

1.3 Who ran it

  • ★ Exam imp The HGP was a 13-year project coordinated by the U.S. Department of Energy and the National Institutes of Health.
  • In its early years the Wellcome Trust (U.K.) became a major partner.
  • More contributions came from Japan, France, Germany, China and others.
  • The project was completed in 2003.
Memory Trick 1990 to 2003 is thirteen years; chromosome 1, the biggest, trailed in last, in May 2006.

1.4 Why it matters

  • Knowing how DNA variations among individuals act could lead to new ways to diagnose, treat and someday prevent the thousands of disorders that affect humans.
  • Besides explaining human biology, the DNA sequences of non-human organisms reveal their natural capabilities.
  • These can be applied to challenges in health care, agriculture, energy production and environmental remediation.
  • Many non-human model organisms have also been sequenced.

Examples to Remember

Model organism sequencedWhat it is
BacteriaProkaryotes
YeastA fungus
Caenorhabditis elegansA free-living, non-pathogenic nematode
DrosophilaThe fruit fly
Rice and ArabidopsisPlants
Key idea
The HGP (1990-2003) was a mega project in cost and data, and it gave rise to bioinformatics.

2. Methods of the HGP

  • The methods followed two major approaches.
Expressed Sequence Tags (ESTs)Focused on identifying all the genes that are expressed as RNA.
Sequence annotationThe blind approach: sequence the whole genome, coding and non-coding, and later assign functions to its regions.
  1. The total DNA of a cell is isolated and broken into random fragments of relatively small size, since very long DNA cannot be sequenced in one piece.
  2. The fragments are cloned in a suitable host using specialised vectors. Cloning amplifies each fragment so that it can be sequenced easily.
  3. The common hosts were bacteria and yeast, and the vectors were BAC (bacterial artificial chromosomes) and YAC (yeast artificial chromosomes).
  4. The fragments were sequenced with automated DNA sequencers working on the principle of a method developed by Frederick Sanger.
  5. The sequences were arranged using the overlapping regions in them, which required overlapping fragments to be sequenced.
  6. Aligning these sequences by hand was impossible, so specialised computer-based programs were developed.
  7. The sequences were then annotated and assigned to each chromosome.
  • Sanger is also credited with developing the method for determining the amino acid sequences of proteins.
  • ★ Exam imp The sequence of chromosome 1 was completed only in May 2006. It was the last of the 24 human chromosomes (22 autosomes, X and Y) to be sequenced.
  • Another challenge was assigning the genetic and physical maps on the genome.
  • These maps used information on polymorphism of restriction endonuclease recognition sites and on repetitive DNA sequences called microsatellites.
A representative diagram of the Human Genome Project A flow from left to right: a human, one of their cells, a chromosome from the nucleus, the DNA double helix of that chromosome, and a computer screen showing base sequences, because the sequenced fragments were stored and aligned by computer. ATGCC GTTAC
Figure 1: A representative diagram of the Human Genome Project. DNA from a person's cells and chromosomes was cut, cloned, sequenced and aligned by computer.
Key idea
ESTs found the expressed genes; whole-genome sequencing with BAC and YAC clones and Sanger-based machines, aligned by computer, read the rest.

3. Salient Features of the Human Genome

  1. The human genome contains 3164.7 million bp.
  2. The average gene has 3000 bases, but sizes vary greatly. The largest known human gene is dystrophin, at 2.4 million bases.
  3. The total number of genes is estimated at 30,000, much lower than the earlier estimates of 80,000 to 1,40,000 genes.
  4. Almost all (99.9 per cent) nucleotide bases are exactly the same in all humans.
  5. The functions are unknown for over 50 per cent of the discovered genes.
  6. Less than 2 per cent of the genome codes for proteins.
  7. Repeated sequences make up a very large portion of the human genome.
  8. Repetitive sequences are stretches of DNA repeated many times, sometimes a hundred to a thousand times. They have no direct coding function, but they shed light on chromosome structure, dynamics and evolution.
  9. Chromosome 1 has the most genes (2968), and the Y has the fewest (231).
  10. About 1.4 million locations of single-base DNA differences have been found in humans: SNPs (single nucleotide polymorphism, pronounced 'snips').
  11. SNP information promises to change how we find the chromosomal locations of disease-associated sequences and trace human history.
★ Very important 3164.7 million bp, about 30,000 genes, 99.9% identical, less than 2% coding; dystrophin is the largest gene; chromosome 1 has the most genes and Y the fewest.
NEET Focus Three gene counts appear in this topic; do not mix them. The goal was to identify about 20,000-25,000 genes; the salient features give an estimate of about 30,000; and earlier estimates were 80,000 to 1,40,000. Also pair 2968 with chromosome 1 and 231 with Y.
Memory Trick Chromosome 1 is number one twice: it has the most genes, and it was the last to be finished (May 2006). Y has the fewest genes (231).
Quick Recall: tap to check
Which chromosome has the fewest genes?
The Y chromosome, with 231 genes.
Which vectors were used to clone DNA fragments in the HGP?
BAC (bacterial artificial chromosomes) and YAC (yeast artificial chromosomes).
Whose method do automated DNA sequencers use?
Frederick Sanger's.

3.1 Applications and future challenges

  • Drawing meaningful knowledge from DNA sequences will shape research for decades to come.
  • This needs the expertise and creativity of tens of thousands of scientists from many disciplines, in public and private sectors across the world.
  • One of the greatest impacts of the human genome sequence may be a radically new approach to biological research.
  • In the past, researchers studied one or a few genes at a time.
  • With whole-genome sequences and new high-throughput technologies, questions can be studied systematically and on a much broader scale.
  • For example, scientists can study all the genes in a genome, all the transcripts in a tissue, organ or tumour, or how tens of thousands of genes and proteins work together in networks.
Key idea
Most of the genome is repeated, non-coding DNA; less than 2% codes for proteins, and humans are 99.9% identical.

4. DNA Fingerprinting

  • 99.9 per cent of the base sequence is the same in all humans.
  • With a genome of bp, the remaining 0.1 per cent means differences at about bp.
  • These differences in DNA sequence make every individual unique in phenotypic appearance.
  • Sequencing DNA every time to find genetic differences between individuals would be a daunting and expensive task, since whole genomes of bp would have to be compared.
  • ★ Exam imp DNA fingerprinting is a very quick way to compare the DNA sequences of any two individuals.

4.1 Repetitive DNA and satellite DNA

  • DNA fingerprinting identifies differences in specific regions called repetitive DNA, in which a small stretch of DNA is repeated many times.
  • In density gradient centrifugation, repetitive DNA separates from the bulk genomic DNA as different peaks.
  • ★ Exam imp The bulk DNA forms a major peak, and the other, small peaks are called satellite DNA.
  • Satellite DNA is classified by base composition (A : T rich or G : C rich), length of segment and number of repetitive units, into micro-satellites, mini-satellites, etc.
  • These sequences normally do not code for any proteins, but they form a large portion of the human genome.
  • They show a high degree of polymorphism and form the basis of DNA fingerprinting.
  • DNA from every tissue of an individual (blood, hair follicle, skin, bone, saliva, sperm, etc.) shows the same polymorphism, so it is a very useful identification tool in forensic work.
  • The polymorphisms are inherited from parents to children, so DNA fingerprinting is the basis of paternity testing in disputes.
Repetitive DNAAny DNA in which a small stretch is repeated many times.
The region compared in DNA fingerprinting.
Satellite DNARepetitive DNA that separates as small peaks from bulk DNA in density gradient centrifugation.
Classed as micro-satellites, mini-satellites, etc.

4.2 DNA polymorphism

  • Polymorphism in DNA sequence is the basis of genetic mapping of the human genome as well as of DNA fingerprinting.
  • Polymorphism (variation at the genetic level) arises due to mutations.
  • New mutations may arise in an individual either in somatic cells or in germ cells (cells that generate gametes in sexually reproducing organisms).
  • A germ cell mutation that does not seriously impair the ability to have offspring can spread to other members of the population through sexual reproduction.
  • ★ Exam imp Allelic sequence variation is called a DNA polymorphism if more than one variant (allele) at a locus occurs in the human population with a frequency greater than 0.01.
  • In simple terms, an inheritable mutation seen in a population at high frequency is a DNA polymorphism.
  • Such variation is more likely in non-coding DNA, because mutations there may have no immediate effect on reproductive ability.
  • These mutations accumulate generation after generation and form one basis of variability.
  • Polymorphisms range from a single nucleotide change to very large-scale changes, and they play an important role in evolution and speciation.
★ Very important DNA polymorphism: an inheritable variant (allele) at a locus with a frequency greater than 0.01 in the population. It arises by mutation and is most common in non-coding DNA.

4.3 The technique

  • ★ Exam imp The technique was first developed by Alec Jeffreys.
  • He used as a probe a satellite DNA with a very high degree of polymorphism, called Variable Number of Tandem Repeats (VNTR).
  • The technique, as first used, involved Southern blot hybridisation with radiolabelled VNTR as the probe.
  1. Isolation of DNA.
  2. Digestion of DNA by restriction endonucleases.
  3. Separation of the DNA fragments by electrophoresis.
  4. Transferring (blotting) the separated fragments to synthetic membranes, such as nitrocellulose or nylon.
  5. Hybridisation using a labelled VNTR probe.
  6. Detection of the hybridised DNA fragments by autoradiography.
Memory Trick Maps used Micro, Jeffreys used Mini: microsatellites helped build the genome maps, while the VNTR probe of DNA fingerprinting is a mini-satellite.
Memory Trick I Delight in Sending Blotted Hybrid Data: Isolation, Digestion, Separation, Blotting, Hybridisation, Detection.
Schematic representation of DNA fingerprinting Three DNA samples: individual A, individual B and a crime scene sample C. Each shows the paternal and maternal copies of chromosomes 7, 2 and 16, with different numbers of short tandem repeats (VNTR copies). Below, each sample runs in its own lane of a gel, where each copy gives one band at the height set by its number of repeats. The band pattern of the crime scene lane matches individual B, not individual A. Chromosome 7 Chromosome 2 Chromosome 16 DNA from individual A Chromosome 7 Chromosome 2 Chromosome 16 DNA from individual B Chromosome 7 Chromosome 2 Chromosome 16 DNA from crime scene (C) Paternal chromosome Maternal chromosome 0 12 Number of short tandem repeats A B C 1 2 3 4 5 6 7 8 9 10 11 12 Number of short tandem repeats Amplified repeats, separated by size on a gel, give a DNA fingerprint
Figure 2: Schematic representation of DNA fingerprinting. Paternal and maternal copies of each chromosome carry different numbers of tandem repeats, and the crime scene band pattern matches individual B, not individual A.
  • In the figure, a few representative chromosomes carry different copy numbers of VNTR, and colours trace the origin of each band in the gel.
  • VNTR belongs to the class of satellite DNA called mini-satellites.
  • A small DNA sequence is arranged tandemly in many copies, and the copy number varies from chromosome to chromosome in an individual.
  • The number of repeats is highly polymorphic, so the size of a VNTR varies from 0.1 to 20 kb.
  • After hybridisation with the VNTR probe, the autoradiogram shows many bands of differing sizes, a characteristic pattern for each individual's DNA.
  • ★ Exam imp The pattern differs from individual to individual in a population, except in monozygotic (identical) twins.
  • The sensitivity of the technique has been increased by the polymerase chain reaction (PCR), so DNA from a single cell is enough.
  • Besides forensic science, it is used to determine population and genetic diversities. Many different probes are now used to make DNA fingerprints.
  • In short, DNA fingerprinting finds variation among individuals at the DNA level, and has wide uses in forensic science, genetic biodiversity and evolutionary biology.
NEET Focus Order the six steps exactly, and note the details examiners swap: the probe is radiolabelled VNTR, a mini-satellite; blotting is onto nitrocellulose or nylon; detection is by autoradiography; only monozygotic twins share a pattern; and PCR lets a single cell suffice.
Quick Recall: tap to check
With 99.9% identical bases and bp, at how many base positions do two humans differ?
About 0.1% of , that is bp.
In Figure 2, which lane matches the crime scene sample?
Lane B: every band of the crime scene DNA lines up with a band of individual B.
What is the size range of a VNTR?
0.1 to 20 kb.
Key idea
DNA fingerprinting compares highly polymorphic VNTRs by Southern blotting; the band pattern is unique except in identical twins.

5. Exam Essentials

Pairs to Match

List IList II
HGP launched1990
HGP completed2003
Chromosome 1 sequence completedMay 2006
Expressed Sequence TagsGenes expressed as RNA
Sequence annotationAssigning functions to sequenced regions
BAC and YACVectors for cloning DNA fragments
Frederick SangerSequencing method behind automated sequencers
DystrophinLargest human gene, 2.4 million bases
Chromosome 1Most genes (2968)
Y chromosomeFewest genes (231)
SNPsAbout 1.4 million single-base differences
Alec JeffreysDeveloped DNA fingerprinting
VNTRMini-satellite used as the probe
Southern blot hybridisationProbing blotted DNA with labelled VNTR
AutoradiographyDetection of hybridised fragments
Exceptions
  • Only monozygotic (identical) twins share the same DNA fingerprint.
  • Repetitive sequences make up a very large part of the genome, yet have no direct coding function.
  • Less than 2 per cent of the human genome codes for proteins.
  • Satellite DNA normally does not code for any protein, yet it is the basis of fingerprinting.
  • Chromosome 1, with the most genes, was the last of the 24 chromosomes to be sequenced.
  • Polymorphism is more likely in non-coding DNA than in coding DNA.

Numbers to Remember

  • HGP: launched 1990, completed 2003, 13 years; chromosome 1 finished May 2006.
  • About bp; 3 US dollars per bp; about 9 billion US dollars; 3300 books of 1000 pages, 1000 letters per page.
  • Goal: 20,000-25,000 genes and 3 billion base pairs. Estimate from the project: about 30,000 genes (earlier 80,000 to 1,40,000).
  • 3164.7 million bp; average gene 3000 bases; dystrophin 2.4 million bases.
  • 99.9% bases identical; over 50% of genes of unknown function; under 2% coding.
  • Chromosome 1: 2968 genes; Y: 231 genes; 24 human chromosomes (22 autosomes, X, Y).
  • 1.4 million SNPs; polymorphism frequency greater than 0.01; VNTR size 0.1 to 20 kb; differences at about bp.

6. Quick Revision

  • HGP launched in 1990 after cloning and fast sequencing became possible; completed in 2003.
  • Mega project: bp, 3 US dollars per bp, about 9 billion dollars, 3300 books; gave rise to bioinformatics.
  • Six goals: identify genes, determine sequence, store, improve tools, transfer technology, address ELSI.
  • Run by the U.S. Department of Energy and National Institutes of Health; Wellcome Trust a major partner.
  • Model organisms sequenced: bacteria, yeast, C. elegans, Drosophila, rice, Arabidopsis.
  • ESTs (expressed genes) and sequence annotation (whole genome, then functions).
  • Random fragments cloned in BAC and YAC, read by Sanger-based automated sequencers, aligned by computer.
  • Chromosome 1 completed last, in May 2006; maps from restriction-site polymorphism and microsatellites.
  • 3164.7 million bp; about 30,000 genes; dystrophin largest; 99.9% identical; under 2% coding.
  • Chromosome 1 has 2968 genes, Y has 231; about 1.4 million SNPs.
  • DNA fingerprinting compares repetitive DNA quickly; satellite DNA forms small peaks.
  • DNA polymorphism: allele frequency greater than 0.01; arises by mutation; common in non-coding DNA.
  • Alec Jeffreys; VNTR (mini-satellite) probe; Southern blot hybridisation.
  • Six steps: isolation, digestion, electrophoresis, blotting, hybridisation, autoradiography.
  • Unique pattern except monozygotic twins; PCR allows a single cell; forensics, paternity, diversity studies.

7. Solved Examples

Solved Example 1
Match List I with List II.
List I: A. Dystrophin, B. Chromosome 1, C. Y chromosome, D. SNPs
List II: I. 231 genes, II. About 1.4 million locations, III. 2.4 million bases, IV. 2968 genes
Choose the correct answer:
(A) A-III, B-IV, C-I, D-II
(B) A-IV, B-III, C-I, D-II
(C) A-III, B-I, C-IV, D-II
(D) A-II, B-IV, C-I, D-III
Solution:

Answer: (A). Dystrophin is 2.4 million bases (III), chromosome 1 has 2968 genes (IV), Y has 231 (I), and there are about 1.4 million SNPs (II).

Solved Example 2
Read the statements about the human genome.
A. It contains 3164.7 million bp.
B. More than half of it codes for proteins.
C. 99.9 per cent of the bases are the same in all humans.
D. The functions of over 50 per cent of the discovered genes are unknown.
E. Repetitive sequences code directly for most proteins.
Choose the correct answer:
(A) A, C and D only
(B) A, B and C only
(C) B, D and E only
(D) A, C, D and E only
Solution:

Answer: (A). B is wrong: less than 2 per cent codes for proteins. E is wrong: repetitive sequences have no direct coding function.

Solved Example 3
Arrange the steps of DNA fingerprinting in the correct order.
A. Blotting onto a nylon membrane
B. Digestion by restriction endonucleases
C. Autoradiography
D. Electrophoresis
E. Hybridisation with labelled VNTR probe
Choose the correct answer:
(A) B, D, A, E, C
(B) B, A, D, E, C
(C) D, B, A, E, C
(D) B, D, E, A, C
Solution:

Answer: (A). After isolation, DNA is digested (B), separated by electrophoresis (D), blotted (A), hybridised (E) and detected by autoradiography (C).

Solved Example 4
If 99.9 per cent of bases are identical among humans and the genome has bp, at about how many base positions do two people differ?
(A)
(B)
(C)
(D)
Solution:

Answer: (B). 0.1 per cent of is .

Solved Example 5
Which of the following was NOT a goal of the Human Genome Project?
(A) Identify all the approximately 20,000-25,000 human genes
(B) Store the information in databases
(C) Address ethical, legal and social issues
(D) Replace defective genes in every patient
Solution:

Answer: (D). The HGP aimed to identify and sequence genes, store data, improve tools, transfer technology and address ELSI, not to treat patients.

Solved Example 6
Statement I: VNTR belongs to the class of satellite DNA called mini-satellites.
Statement II: The size of a VNTR varies from 0.1 to 20 kb.
(A) Both Statement I and Statement II are correct
(B) Statement I is correct, Statement II is incorrect
(C) Statement I is incorrect, Statement II is correct
(D) Both are incorrect
Solution:

Answer: (A). VNTRs are mini-satellites, and their highly variable copy number gives sizes from 0.1 to 20 kb.

8. Practice Questions

Practice Questions
  1. Match List I with List II.
    List I: A. BAC, B. YAC, C. ESTs, D. SNPs
    List II: I. Single-base differences, II. Vector based on a yeast chromosome, III. Genes expressed as RNA, IV. Vector based on a bacterial chromosome
    (A) A-IV, B-II, C-III, D-I (B) A-II, B-IV, C-III, D-I (C) A-IV, B-II, C-I, D-III (D) A-III, B-II, C-IV, D-IAnswer: (A). BAC and YAC are bacterial and yeast artificial chromosomes, ESTs tag expressed genes and SNPs are single-base differences.
  2. Which statement is NOT correct? (A) DNA from blood and saliva of one person shows the same polymorphism (B) Monozygotic twins have different DNA fingerprints (C) Polymorphisms are inherited from parents to children (D) PCR allows fingerprinting from a single cellAnswer: (B). Monozygotic twins have identical DNA fingerprints.
  3. Arrange the HGP sequencing steps in order.
    A. Cloning fragments in BAC or YAC
    B. Breaking total DNA into random fragments
    C. Aligning overlapping sequences by computer
    D. Sequencing in automated sequencers
    (A) B, A, D, C (B) A, B, D, C (C) B, D, A, C (D) B, A, C, DAnswer: (A). Fragment, clone to amplify, sequence, then align the overlaps.
  4. Statement I: A DNA polymorphism is an allelic variant with a frequency greater than 0.01 in the population.
    Statement II: Polymorphism is less likely in non-coding DNA.
    (A) Both Statement I and Statement II are correct (B) Statement I is correct, Statement II is incorrect (C) Statement I is incorrect, Statement II is correct (D) Both are incorrectAnswer: (B). Polymorphism is more likely in non-coding DNA, where mutations may not affect reproductive ability.
  5. Read the statements.
    A. The HGP was coordinated by the U.S. Department of Energy and the National Institutes of Health.
    B. Chromosome 1 was the first chromosome to be sequenced.
    C. The Wellcome Trust (U.K.) was a major partner.
    D. The project was completed in 2003.
    Choose the correct answer: (A) A, C and D only (B) A, B and C only (C) B and D only (D) A, B, C and DAnswer: (A). B is wrong: chromosome 1 was the last, completed in May 2006.
  6. Differentiate between repetitive DNA and satellite DNA.Answer: Repetitive DNA is any DNA in which a small stretch is repeated many times. Satellite DNA is repetitive DNA that separates from bulk DNA as small peaks in density gradient centrifugation; it is classed as micro-satellites, mini-satellites, etc., and is highly polymorphic.
  7. Why is the Human Genome Project called a mega project?Answer: It aimed to sequence about bp at an estimated cost of about 9 billion US dollars, and its data would fill 3300 books of 1000 pages. Storing and analysing this needed high-speed computers and gave rise to bioinformatics.
  8. What is DNA fingerprinting? Mention its applications.Answer: It is a quick way to compare the DNA of individuals by identifying differences in highly polymorphic repetitive DNA (VNTRs). It is used in forensic identification, paternity testing, and studies of population and genetic diversity and evolution.
  9. Briefly describe (a) polymorphism, (b) bioinformatics.Answer: (a) Polymorphism is variation at the genetic level that arises from mutation; an inheritable variant at a locus with a frequency above 0.01 is a DNA polymorphism. (b) Bioinformatics is the field that uses high-speed computers to store, retrieve and analyse biological data such as genome sequences; it grew with the HGP.

Common Mistakes to Avoid

Watch out
  • Saying the HGP was completed in 1990. Correct: it was launched in 1990 and completed in 2003.
  • Giving the Y chromosome the most genes. Correct: chromosome 1 has the most (2968) and Y the fewest (231).
  • Quoting 80,000-1,40,000 as the human gene count. Correct: those were earlier estimates; the project estimates about 30,000.
  • Saying most of the genome codes for proteins. Correct: less than 2 per cent does.
  • Calling VNTR a micro-satellite. Correct: VNTR is a mini-satellite.
  • Saying identical twins have different fingerprints. Correct: monozygotic twins share the same pattern.
  • Leaving out the label on the probe. Correct: the VNTR probe is radiolabelled, and bands are detected by autoradiography.
  • Setting the polymorphism cut-off below 0.01. Correct: a variant must occur at a frequency greater than 0.01.

Frequently Asked Questions

Why was the Human Genome Project called a mega project?

The human genome has about 3 billion base pairs. At an estimated 3 US dollars per base pair, sequencing would cost about 9 billion dollars, and printing the sequence would fill 3300 books of 1000 pages each. Handling this data needed high-speed computers and gave rise to bioinformatics.

What were the goals of the Human Genome Project?

The goals were to identify all the approximately 20,000 to 25,000 human genes, determine the sequence of the 3 billion base pairs, store the data in databases, improve tools for analysis, transfer the technology to other sectors such as industry, and address ethical, legal and social issues.

What are Expressed Sequence Tags and sequence annotation?

They were the two approaches of the Human Genome Project. Expressed Sequence Tags identify all the genes that are expressed as RNA. Sequence annotation is the blind approach: the whole genome, coding and non-coding, is sequenced first, and functions are assigned to its regions later.

What are the salient features of the human genome?

The genome has 3164.7 million base pairs and about 30,000 genes. The average gene has 3000 bases, and dystrophin, at 2.4 million bases, is the largest. 99.9 percent of bases are the same in all humans, less than 2 percent codes for proteins, chromosome 1 has 2968 genes and Y has 231.

What is satellite DNA?

Satellite DNA is repetitive DNA, in which a short stretch is repeated many times. In density gradient centrifugation it forms small peaks separate from the major peak of bulk DNA. It is classed as micro-satellites, mini-satellites and so on, does not code for proteins, and is highly polymorphic.

What is DNA polymorphism?

DNA polymorphism is variation at the genetic level that arises from mutation. An allelic variant is called a polymorphism when more than one allele at a locus occurs in the population at a frequency greater than 0.01. It is more common in non-coding DNA and is the basis of DNA fingerprinting.

What are the steps of DNA fingerprinting?

DNA is isolated and digested with restriction endonucleases. The fragments are separated by electrophoresis and blotted onto a nitrocellulose or nylon membrane. They are then hybridised with a labelled VNTR probe, and the hybridised fragments are detected by autoradiography as a pattern of bands.

What are the applications of DNA fingerprinting?

Since every tissue of a person shows the same polymorphism, DNA fingerprinting identifies individuals in forensic science. Because polymorphisms are inherited, it is the basis of paternity testing. It is also used to study population and genetic diversity and in evolutionary biology.

Ready to master Molecular Basis of Inheritance?

Take a full mock test, practice concept-by-concept, and get an AI-powered rank prediction — all on Fundamenthol.