GENETIC CODE REFERENCE

Kozak sequence checker

Start codon context, upstream ATGs and the consensus.

Example: human β-globin (HBB) mRNA, NCBI NM_000518.5.

ContextStrongA or G at −3 and G at +4
Position −3Aa purine, the position that matters most
Position +4GG, as in the consensus
Consensus bases6 of 7flanking bases matching GCCRCC·ATG·G

ATG at position 51 · 20 codons, no stop before the end

Bases −6 to +4 around the chosen ATG, the consensus, and the share of human transcripts with the same base
Position−6−5−4−3−2−1+1+2+3+4
YoursGACACCATGG
ConsensusGCCRCCATGG
Human %41174047414851
  • No ATG comes before this start codon. Ribosomes scan from the mRNA’s 5′ end and usually start at the first ATG in a good context.
  • Human shares come from 19,221 MANE Select transcripts that start with ATG; the consensus is Kozak’s (1987), from 699 vertebrate mRNAs.

Every ATG in your sequence (1)

PositionContext, −6 to +4StrengthReading frameAgainst the chosen startShow
51GACACCATGGStrong20 codons, no stop before the endChosen start codon

What the Kozak sequence is

The Kozak sequence is the set of bases around the start codon of a eukaryotic mRNA that helps ribosomes recognize it. Marilyn Kozak compiled the bases before the start codons of 699 vertebrate mRNAs and found the consensus (GCC)GCCRCCATGG, where R is A or G (Kozak, 1987). Positions count from the start codon: the A of ATG is +1, the base just before it is −1, and the base right after the ATG is +4.

Two positions matter most. A purine at −3, usually A, was the most conserved position: 97% of the 699 mRNAs had one. G was the preferred base at +4. When single bases around the ATG of a preproinsulin gene were changed, the amount of protein varied over a 20-fold range, and the purine at −3 had the dominant effect (Kozak, 1986).

Paste an mRNA or cDNA, from a few bases before the start codon or the whole transcript, and select Check. The checker lists every ATG, shows the chosen one's bases from −6 to +4 against the consensus, and says what each earlier ATG's reading frame does. It starts on the ATG of the longest open reading frame; select Show on any other ATG to check that one instead.

Strong and weak contexts

Because −3 and +4 have the strongest influence, Kozak (1987) grouped start codons as strong or weak by those two positions alone. The checker labels every ATG by them:

Label Position −3 Position +4 Kozak's 699 mRNAs Human transcripts
Strong A or G G 305 (43.6%) 8,212 (43.6%)
One of two A or G A, C or T 371 (53.1%) 8,026 (42.6%)
One of two C or T G 17 (2.4%) 1,466 (7.8%)
Weak C or T A, C or T 6 (0.9%) 1,151 (6.1%)

Kozak expected the six vertebrate mRNAs with neither preferred base to be translated inefficiently; four of them encoded hormones or lymphokines. The human column counts 18,855 MANE Select transcripts with at least three bases before an ATG start codon (see below).

The consensus, position by position

In both collections, the consensus base is the most common base at every position from −6 to +4:

Position Consensus Kozak's 699 mRNAs Human transcripts
−6 G 44% 40.7%
−5 C 39% 33.4%
−4 C 53% 40.0%
−3 A (or G) 61% A, 36% G 47.3% A, 38.8% G
−2 C 49% 40.8%
−1 C 55% 47.7%
+4 G 46% 51.3%

Kozak also noted G repeating at −3, −6 and −9; in the human transcripts, G is the most common base at −9 too (36.5%). Site-directed mutations had confirmed every base from −6 to −1 and the G at +4, while the role of the GCC at −9 to −7 was still open in 1987. The checker compares −6 to −1 and +4, seven bases, and shows the share of human transcripts with your base at each position.

Human transcripts today

The human figures come from MANE Select v1.5, one reference transcript for nearly every human protein-coding gene (Morales and colleagues, 2022). Of the 19,248 transcripts with a complete coding sequence and no translation exception, 19,221 start with ATG; Codon Chart counted the bases around each start codon and every ATG in each annotated 5′ untranslated region (5′ UTR).

The two collections were built differently: Kozak's from cDNA and gene sequences published by 1987, MANE from the full-length transcripts NCBI and EMBL-EBI annotate today.

Upstream ATGs and scanning

Eukaryotic ribosomes bind at the mRNA's 5′ end and scan toward the start codon, so position as well as context decides which ATG is used: ribosomes start at the first ATG when its context is favorable (Kozak, 1987). An earlier ATG does not always block the one after it: some ribosomes pass an ATG in a weak context (leaky scanning), and ribosomes that finish a short upstream reading frame appear to start again at the next ATG (reinitiation). Kozak concluded that an upstream ATG is a problem when its context is favorable and no stop codon ends its reading frame before the main start codon; 5 of the 699 mRNAs had that arrangement.

The checker describes each ATG before the chosen start codon by its reading frame:

361 human transcripts (1.9%) have an upstream ATG in a strong context whose reading frame runs past the main start codon.

Adding a Kozak sequence to a construct

To give a coding sequence Kozak's consensus context, place GCCACC (or GCCGCC) directly before its ATG. The G at +4 is the first base of the second codon, so adding one changes the second amino acid unless its codon already starts with G: codons beginning with G encode valine, alanine, aspartate, glutamate and glycine (see the codon chart). Check the result here, and translate it with DNA to protein to confirm the protein. A consensus context does not guarantee expression, which also depends on the promoter, the mRNA's leader and the cells.

The consensus is from vertebrates. Surveys Kozak cited also found A most often at −3 in plants (53% A and 23% G in 47 mRNAs), Drosophila (82% A and 13% G in 77 mRNAs) and yeast (81% A in 96 mRNAs), but the effect of context had not yet been tested outside vertebrates.

Kozak and Shine-Dalgarno sequences

Bacteria find their start codons another way: a short sequence a few bases before the start codon pairs with the 3′ end of the 16S ribosomal RNA. For bacterial genes, use the Shine-Dalgarno sequence finder.

Sources: Kozak (1986) and (1987), with Table 1 and Table 2 of the 1987 paper checked number by number against its PubMed Central scan; human counts from NCBI's MANE v1.5 Select RefSeq file. See sources and methods.