GENETIC CODE REFERENCE

Codon optimization tool

Host codon usage, CAI and optimized DNA.

Example: EGFP (NCBI U55762.1) for E. coli.

CAI0.696 → 1.000for E. coli K-12 MG1655, 1 is every codon the host’s most used
GC content61.5% → 48.9%of the coding sequence
Codons changed126of 240; the protein is the same
Rare codons in yours0under 0.2 of the most used codon for their amino acid

OPTIMIZED DNA · 720 BP

ATGGTGAGCAAAGGCGAAGAACTGTTTACC GGCGTGGTGCCGATTCTGGTGGAACTGGAT GGCGATGTGAACGGCCATAAATTTAGCGTG AGCGGCGAAGGCGAAGGCGATGCGACCTAT GGCAAACTGACCCTGAAATTTATTTGCACC ACCGGCAAACTGCCGGTGCCGTGGCCGACC CTGGTGACCACCCTGACCTATGGCGTGCAG TGCTTTAGCCGCTATCCGGATCATATGAAA CAGCATGATTTTTTTAAAAGCGCGATGCCG GAAGGCTATGTGCAGGAACGCACCATTTTT TTTAAAGATGATGGCAACTATAAAACCCGC GCGGAAGTGAAATTTGAAGGCGATACCCTG GTGAACCGCATTGAACTGAAAGGCATTGAT TTTAAAGAAGATGGCAACATTCTGGGCCAT AAACTGGAATATAACTATAACAGCCATAAC GTGTATATTATGGCGGATAAACAGAAAAAC GGCATTAAAGTGAACTTTAAAATTCGCCAT AACATTGAAGATGGCAGCGTGCAGCTGGCG GATCATTATCAGCAGAACACCCCGATTGGC GATGGCCCGGTGCTGCTGCCGGATAACCAT TATCTGAGCACCCAGAGCGCGCTGAGCAAA GATCCGAACGAAAAACGCGATCATATGGTG CTGCTGGAATTTGTGACCGCGGCGGGCATT ACCCTGGGCATGGATGAACTGTATAAATAA

E. coli K-12 MG1655 codon usage

Per thousand codons · share of its amino acid · count in your sequence

First baseTCAGThird base
TTTT Phe22.3 ‰ · 0.57 · 0TCT Ser8.4 ‰ · 0.15 · 0TAT Tyr16.0 ‰ · 0.57 · 1TGT Cys5.1 ‰ · 0.44 · 0T
TTC Phe16.6 ‰ · 0.43 · 12TCC Ser8.6 ‰ · 0.15 · 3TAC Tyr12.2 ‰ · 0.43 · 10TGC Cys6.5 ‰ · 0.56 · 2C
TTA Leu13.8 ‰ · 0.13 · 0TCA Ser7.0 ‰ · 0.12 · 0TAA Stop2.1 ‰ · 0.64 · 1TGA Stop0.9 ‰ · 0.29 · 0A
TTG Leu13.7 ‰ · 0.13 · 0TCG Ser8.9 ‰ · 0.15 · 0TAG Stop0.2 ‰ · 0.07 · 0TGG Trp15.2 ‰ · 1.00 · 1G
CCTT Leu11.0 ‰ · 0.10 · 0CCT Pro7.0 ‰ · 0.16 · 0CAT His12.9 ‰ · 0.57 · 0CGT Arg21.2 ‰ · 0.38 · 0T
CTC Leu11.2 ‰ · 0.10 · 3CCC Pro5.5 ‰ · 0.12 · 10CAC His9.7 ‰ · 0.43 · 9CGC Arg22.2 ‰ · 0.40 · 6C
CTA Leu3.9 ‰ · 0.04 · 0CCA Pro8.4 ‰ · 0.19 · 0CAA Gln15.4 ‰ · 0.35 · 0CGA Arg3.5 ‰ · 0.06 · 0A
CTG Leu53.3 ‰ · 0.50 · 18CCG Pro23.4 ‰ · 0.53 · 0CAG Gln29.0 ‰ · 0.65 · 8CGG Arg5.3 ‰ · 0.10 · 0G
AATT Ile30.5 ‰ · 0.51 · 0ACT Thr8.8 ‰ · 0.16 · 1AAT Asn17.4 ‰ · 0.45 · 0AGT Ser8.6 ‰ · 0.15 · 0T
ATC Ile25.3 ‰ · 0.42 · 12ACC Thr23.5 ‰ · 0.44 · 15AAC Asn21.5 ‰ · 0.55 · 13AGC Ser16.0 ‰ · 0.28 · 7C
ATA Ile4.1 ‰ · 0.07 · 0ACA Thr6.9 ‰ · 0.13 · 0AAA Lys33.7 ‰ · 0.77 · 1AGA Arg1.9 ‰ · 0.04 · 0A
ATG Met27.9 ‰ · 1.00 · 6ACG Thr14.4 ‰ · 0.27 · 0AAG Lys10.2 ‰ · 0.23 · 19AGG Arg1.1 ‰ · 0.02 · 0G
GGTT Val18.3 ‰ · 0.26 · 0GCT Ala15.2 ‰ · 0.16 · 0GAT Asp32.0 ‰ · 0.63 · 2GGT Gly24.7 ‰ · 0.34 · 0T
GTC Val15.3 ‰ · 0.22 · 4GCC Ala25.7 ‰ · 0.27 · 8GAC Asp19.1 ‰ · 0.37 · 16GGC Gly29.8 ‰ · 0.41 · 19C
GTA Val10.9 ‰ · 0.15 · 1GCA Ala20.2 ‰ · 0.21 · 0GAA Glu39.8 ‰ · 0.69 · 1GGA Gly7.8 ‰ · 0.11 · 0A
GTG Val26.4 ‰ · 0.37 · 13GCG Ala34.0 ‰ · 0.36 · 0GAG Glu17.8 ‰ · 0.31 · 15GGG Gly11.0 ‰ · 0.15 · 3G

Most used codon for its amino acidRare in E. coli K-12 MG1655: under 0.2 of the most used codon’s count

  • Each amino acid gets the codon E. coli K-12 MG1655 uses most for it; the protein is unchanged.
  • Check the result before ordering a gene: this method can raise or lower the GC content, repeat one codon many times and create restriction sites.

What codon optimization changes

Most amino acids have more than one codon, and each organism uses its synonymous codons unevenly. A gene moved into a new host, such as a human gene expressed in E. coli, can carry codons that the host rarely uses. Clusters of rare codons can slow translation and lower the yield of the protein (Kane, 1995). Codon optimization rewrites the DNA with codons the host uses more, without changing a single amino acid.

Paste a coding sequence from its start codon, or a protein in one-letter code, choose the host and select Optimize. The result is the new DNA, its codon adaptation index (CAI) and GC content next to your sequence's, and the host's codon usage table with your codon counts. Everything runs in your browser; the sequence is not uploaded.

How the tool optimizes

Each amino acid gets the codon the host uses most for it: the "one amino acid, one codon" method. It is what Biopython 1.85's CodonAdaptationIndex.optimize does, and every result here is tested against Biopython, trained on the same host genes. The protein stays the same, and a stop at the end becomes the host's most used stop codon.

The method is simple and predictable, but not the only one. It uses a single codon for every copy of an amino acid, so it can change the GC content a lot and repeat the same codon many times. Other methods pick codons in proportion to the host's usage, or also balance GC content and avoid repeats, restriction sites and RNA structure near the start codon (Gustafsson and colleagues, 2004).

Codon adaptation index (CAI)

The CAI of Sharp and Li (1987) scores how closely a gene follows a host's codon preferences. Each codon's relative adaptiveness, w, is its count in the host's genes divided by the count of the most used codon for the same amino acid. A codon the host never uses counts as 0.5. The CAI is the geometric mean of w over the gene's codons, from 0 to 1, where 1 means every codon is the host's most used one.

ATG and TGG are left out, because methionine and tryptophan have only one codon each, so their w is always 1. As in Biopython, the stop codon stays in the mean, scored against the host's three stops. Sharp and Li built their reference from highly expressed genes. The tables here count every gene, so this CAI measures how common a sequence's codons are in the host's whole genome.

Host codon usage tables

Each table counts the codons of every complete protein-coding gene in the host's reference annotation from NCBI, the start and stop codons included. Pseudogenes, partial genes, genes with a translation exception (such as selenocysteine) and any sequence with an internal stop are left out.

Host Source Genes Codons
E. coli K-12 MG1655 RefSeq assembly GCF_000005845.2 4,294 1,330,718
S. cerevisiae S288C RefSeq assembly GCF_000146045.2, mitochondrial genes left out 5,955 2,853,391
Human MANE Select v1.5 transcripts, one per gene (Morales and colleagues, 2022) 19,248 11,193,713

In the table, each codon shows its uses per thousand codons and its share of its amino acid's codons. The most used codon for each amino acid is outlined. A codon used less than 0.2 times as often as the most used one for its amino acid is striped as rare. In E. coli, these are the arginine codons AGA (1.9 per thousand), AGG (1.1) and CGA (3.5), the isoleucine codon ATA (4.1), the leucine codon CTA (3.9) and the TAG stop (0.2). Host table (CSV) downloads the chosen host's full table.

Worked example: EGFP for E. coli

The example is the coding sequence of EGFP, the enhanced green fluorescent protein, from NCBI record U55762.1 (the pEGFP-N1 cloning vector). Its codons follow human genes closely: its CAI is 0.957 for human, 0.696 for E. coli and 0.516 for yeast. Optimizing it for E. coli changes 126 of its 240 codons, raises the CAI to 1.000 and lowers the GC content from 61.5% to 48.9%.

Before you order a gene

Read the optimized sequence before using it. Check that your cloning sites are not created or lost with the restriction digest calculator on Plasmid Map, and that the GC content and any repeats suit your synthesis provider. Translate it back with DNA to protein to confirm the protein. A high CAI does not guarantee expression, which also depends on the promoter, the start of the mRNA (check a Kozak context or a Shine-Dalgarno sequence), the protein's folding and toxicity, and the growth conditions.

Sources: host annotations from NCBI RefSeq and MANE (see the table); Biopython 1.85 CodonAdaptationIndex. Automated tests compare every weight, CAI value and optimized sequence with Biopython 1.85; see sources and methods.