What the Shine-Dalgarno sequence is
The Shine-Dalgarno sequence is a short, purine-rich stretch a few bases before the start codon of many bacterial mRNAs. It pairs with the 3′ end of the 16S ribosomal RNA of the small ribosomal subunit, which helps the ribosome find the start codon. John Shine and Lynn Dalgarno sequenced the 3′ end of E. coli 16S rRNA, …ACCUCCUUA-3′, and saw that the ribosome binding sites of several phage mRNAs all contained part of its complement, GGAGGU (Shine and Dalgarno, 1974).
The complete consensus is UAAGGAGGU (TAAGGAGGT in DNA), the exact complement of that 3′ end, and most genes carry only part of it (Chen and colleagues, 1994). The finder's example, the lacZ gene of E. coli, has AGGA seven bases before its ATG.
Paste the bases before a bacterial start codon together with the codon, or a longer stretch of DNA or mRNA, and select Find. The finder lists every start codon, and for each one the best match to the consensus in the 20 bases before it. It starts on the start codon of the longest open reading frame; select Show on any other.
How it pairs with 16S rRNA
The mRNA runs 5′ to 3′ and the end of the 16S rRNA lies antiparallel beneath it, so the consensus UAAGGAGGU faces 3′-AUUCCUCCA-5′, base for base. The finder's diagram draws the rRNA under the matched bases and shades the pairs. A longer match can form more base pairs, and a 3-base match is found before almost any start codon by chance (see the table below).
Spacing and aligned spacing
The distance between the Shine-Dalgarno sequence and the start codon affects translation, but studies of the best distance reported optimal spacings from 5 to 13 bases. Chen and colleagues (1994) traced much of the disagreement to how spacing is counted, and defined two measures:
- Spacing: the number of bases between the end of the Shine-Dalgarno match and the A of the start codon.
- Aligned spacing: the number of bases between the base that lines up with the last U of the complete consensus and the A of the start codon. When the match ends with that U, the two are equal; a match such as UAAGG, which stops four bases earlier, has an aligned spacing four smaller than its spacing.
In their experiments, synthetic ribosome binding sites before a chloramphenicol acetyltransferase gene gave the most protein at an aligned spacing of 5 bases, with GAGGU at a spacing of 5 and with UAAGG at a spacing of 9. Spacings of 4 or 6 bases kept about 80% of the maximum, and spacings of 2 or 17 still about 15%. Natural mRNAs vary, with an average spacing of about 7 bases.
Shine-Dalgarno sequences in E. coli genes
Codon Chart ran the finder on the 20 bases before the start codon of all 4,294 complete protein-coding genes of E. coli K-12 MG1655 (RefSeq NC_000913.3). As a chance baseline, it shuffled the bases of each 20-base region ten times and ran the finder again.
| Longest match | E. coli genes | Shuffled regions |
|---|---|---|
| 3 bases or more | 97.6% | 89.6% |
| 4 bases or more | 73.3% | 38.8% |
| 5 bases or more | 38.6% | 11.0% |
| 6 bases or more | 14.0% | 2.6% |
A match of 4 bases or more is where the genes clearly differ from chance. Among the 3,148 genes with one, the aligned spacing is 3 to 6 bases for 74.8%, against 26.7% of the shuffled regions with such a match, and the median spacing is 7 bases.
Most of these genes start with ATG: 3,870 (90.1%), against 338 with GTG (7.9%), 80 with TTG (1.9%) and 6 with another codon. Under Start codons, choose ATG, GTG and TTG to list all three.
What the finder does and does not model
- It looks for exact matches of 3 bases or more to a stretch of TAAGGAGGT in the 20 bases before each start codon. When two matches are equally long, it keeps the one whose aligned spacing is closest to 5, then the one nearest the start codon.
- It reads DNA or RNA, one sequence, up to 30,000 bases. FASTA headers, spaces and numbers are ignored.
- It does not count G·U pairs or base-pairing energy, and it does not model mRNA structure. Structure that involves the Shine-Dalgarno sequence or the start codon strongly affects initiation (Chen and colleagues, 1994), and the same ribosome binding site can give different amounts of protein in different genetic contexts (Salis and colleagues, 2009).
A good match is evidence of a ribosome binding site, not a measure of expression. To check the protein a start codon begins, translate from it with DNA to protein.
Shine-Dalgarno and Kozak sequences
Eukaryotic ribosomes do not rely on a Shine-Dalgarno sequence: they bind at the mRNA’s 5′ end and scan to a start codon, whose surrounding bases are summarized by the Kozak consensus GCCRCCATGG. For eukaryotic genes, use the Kozak sequence checker.
Sources: Shine and Dalgarno (1974), Chen and colleagues (1994), their PubMed Central scans read in full; the E. coli counts are Codon Chart's, from NCBI's RefSeq record of the K-12 MG1655 genome. See sources and methods.