What the Kozak sequence is
The Kozak sequence is the set of bases around the start codon of a eukaryotic mRNA that helps ribosomes recognize it. Marilyn Kozak compiled the bases before the start codons of 699 vertebrate mRNAs and found the consensus (GCC)GCCRCCATGG, where R is A or G (Kozak, 1987). Positions count from the start codon: the A of ATG is +1, the base just before it is −1, and the base right after the ATG is +4.
Two positions matter most. A purine at −3, usually A, was the most conserved position: 97% of the 699 mRNAs had one. G was the preferred base at +4. When single bases around the ATG of a preproinsulin gene were changed, the amount of protein varied over a 20-fold range, and the purine at −3 had the dominant effect (Kozak, 1986).
Paste an mRNA or cDNA, from a few bases before the start codon or the whole transcript, and select Check. The checker lists every ATG, shows the chosen one's bases from −6 to +4 against the consensus, and says what each earlier ATG's reading frame does. It starts on the ATG of the longest open reading frame; select Show on any other ATG to check that one instead.
Strong and weak contexts
Because −3 and +4 have the strongest influence, Kozak (1987) grouped start codons as strong or weak by those two positions alone. The checker labels every ATG by them:
| Label | Position −3 | Position +4 | Kozak's 699 mRNAs | Human transcripts |
|---|---|---|---|---|
| Strong | A or G | G | 305 (43.6%) | 8,212 (43.6%) |
| One of two | A or G | A, C or T | 371 (53.1%) | 8,026 (42.6%) |
| One of two | C or T | G | 17 (2.4%) | 1,466 (7.8%) |
| Weak | C or T | A, C or T | 6 (0.9%) | 1,151 (6.1%) |
Kozak expected the six vertebrate mRNAs with neither preferred base to be translated inefficiently; four of them encoded hormones or lymphokines. The human column counts 18,855 MANE Select transcripts with at least three bases before an ATG start codon (see below).
The consensus, position by position
In both collections, the consensus base is the most common base at every position from −6 to +4:
| Position | Consensus | Kozak's 699 mRNAs | Human transcripts |
|---|---|---|---|
| −6 | G | 44% | 40.7% |
| −5 | C | 39% | 33.4% |
| −4 | C | 53% | 40.0% |
| −3 | A (or G) | 61% A, 36% G | 47.3% A, 38.8% G |
| −2 | C | 49% | 40.8% |
| −1 | C | 55% | 47.7% |
| +4 | G | 46% | 51.3% |
Kozak also noted G repeating at −3, −6 and −9; in the human transcripts, G is the most common base at −9 too (36.5%). Site-directed mutations had confirmed every base from −6 to −1 and the G at +4, while the role of the GCC at −9 to −7 was still open in 1987. The checker compares −6 to −1 and +4, seven bases, and shows the share of human transcripts with your base at each position.
Human transcripts today
The human figures come from MANE Select v1.5, one reference transcript for nearly every human protein-coding gene (Morales and colleagues, 2022). Of the 19,248 transcripts with a complete coding sequence and no translation exception, 19,221 start with ATG; Codon Chart counted the bases around each start codon and every ATG in each annotated 5′ untranslated region (5′ UTR).
- −3 and +4: 86.1% have A or G at −3, against 97% of Kozak's mRNAs, and 51.3% have G at +4. Weak contexts are rarer than one in ten, but more common than in 1987.
- Leader length: the median 5′ UTR is 133 bases long (the middle half from 65 to 255), and 60.3% are longer than 100 bases. Among Kozak's mRNAs with a mapped transcription start, only about a quarter were.
- Upstream ATGs: 41.9% of the transcripts have at least one ATG before the start codon. Kozak counted upstream ATGs in 9% of vertebrate mRNAs, leaving out proto-oncogenes and four mRNAs with many.
The two collections were built differently: Kozak's from cDNA and gene sequences published by 1987, MANE from the full-length transcripts NCBI and EMBL-EBI annotate today.
Upstream ATGs and scanning
Eukaryotic ribosomes bind at the mRNA's 5′ end and scan toward the start codon, so position as well as context decides which ATG is used: ribosomes start at the first ATG when its context is favorable (Kozak, 1987). An earlier ATG does not always block the one after it: some ribosomes pass an ATG in a weak context (leaky scanning), and ribosomes that finish a short upstream reading frame appear to start again at the next ATG (reinitiation). Kozak concluded that an upstream ATG is a problem when its context is favorable and no stop codon ends its reading frame before the main start codon; 5 of the 699 mRNAs had that arrangement.
The checker describes each ATG before the chosen start codon by its reading frame:
- Its frame stops before the start: an upstream open reading frame. 19,424 of the 22,909 upstream ATGs in the human transcripts are of this kind.
- Another frame, runs past the start: the upstream reading frame overlaps the start codon (3,077).
- Same frame, no stop before the start: translation from it would add amino acids to the start of the protein (408).
361 human transcripts (1.9%) have an upstream ATG in a strong context whose reading frame runs past the main start codon.
Adding a Kozak sequence to a construct
To give a coding sequence Kozak's consensus context, place GCCACC (or GCCGCC) directly before its ATG. The G at +4 is the first base of the second codon, so adding one changes the second amino acid unless its codon already starts with G: codons beginning with G encode valine, alanine, aspartate, glutamate and glycine (see the codon chart). Check the result here, and translate it with DNA to protein to confirm the protein. A consensus context does not guarantee expression, which also depends on the promoter, the mRNA's leader and the cells.
The consensus is from vertebrates. Surveys Kozak cited also found A most often at −3 in plants (53% A and 23% G in 47 mRNAs), Drosophila (82% A and 13% G in 77 mRNAs) and yeast (81% A in 96 mRNAs), but the effect of context had not yet been tested outside vertebrates.
Kozak and Shine-Dalgarno sequences
Bacteria find their start codons another way: a short sequence a few bases before the start codon pairs with the 3′ end of the 16S ribosomal RNA. For bacterial genes, use the Shine-Dalgarno sequence finder.
Sources: Kozak (1986) and (1987), with Table 1 and Table 2 of the 1987 paper checked number by number against its PubMed Central scan; human counts from NCBI's MANE v1.5 Select RefSeq file. See sources and methods.