What codon optimization changes
Most amino acids have more than one codon, and each organism uses its synonymous codons unevenly. A gene moved into a new host, such as a human gene expressed in E. coli, can carry codons that the host rarely uses. Clusters of rare codons can slow translation and lower the yield of the protein (Kane, 1995). Codon optimization rewrites the DNA with codons the host uses more, without changing a single amino acid.
Paste a coding sequence from its start codon, or a protein in one-letter code, choose the host and select Optimize. The result is the new DNA, its codon adaptation index (CAI) and GC content next to your sequence's, and the host's codon usage table with your codon counts. Everything runs in your browser; the sequence is not uploaded.
How the tool optimizes
Each amino acid gets the codon the host uses most for it: the "one amino acid, one codon" method. It is what Biopython 1.85's CodonAdaptationIndex.optimize does, and every result here is tested against Biopython, trained on the same host genes. The protein stays the same, and a stop at the end becomes the host's most used stop codon.
The method is simple and predictable, but not the only one. It uses a single codon for every copy of an amino acid, so it can change the GC content a lot and repeat the same codon many times. Other methods pick codons in proportion to the host's usage, or also balance GC content and avoid repeats, restriction sites and RNA structure near the start codon (Gustafsson and colleagues, 2004).
Codon adaptation index (CAI)
The CAI of Sharp and Li (1987) scores how closely a gene follows a host's codon preferences. Each codon's relative adaptiveness, w, is its count in the host's genes divided by the count of the most used codon for the same amino acid. A codon the host never uses counts as 0.5. The CAI is the geometric mean of w over the gene's codons, from 0 to 1, where 1 means every codon is the host's most used one.
ATG and TGG are left out, because methionine and tryptophan have only one codon each, so their w is always 1. As in Biopython, the stop codon stays in the mean, scored against the host's three stops. Sharp and Li built their reference from highly expressed genes. The tables here count every gene, so this CAI measures how common a sequence's codons are in the host's whole genome.
Host codon usage tables
Each table counts the codons of every complete protein-coding gene in the host's reference annotation from NCBI, the start and stop codons included. Pseudogenes, partial genes, genes with a translation exception (such as selenocysteine) and any sequence with an internal stop are left out.
| Host | Source | Genes | Codons |
|---|---|---|---|
| E. coli K-12 MG1655 | RefSeq assembly GCF_000005845.2 | 4,294 | 1,330,718 |
| S. cerevisiae S288C | RefSeq assembly GCF_000146045.2, mitochondrial genes left out | 5,955 | 2,853,391 |
| Human | MANE Select v1.5 transcripts, one per gene (Morales and colleagues, 2022) | 19,248 | 11,193,713 |
In the table, each codon shows its uses per thousand codons and its share of its amino acid's codons. The most used codon for each amino acid is outlined. A codon used less than 0.2 times as often as the most used one for its amino acid is striped as rare. In E. coli, these are the arginine codons AGA (1.9 per thousand), AGG (1.1) and CGA (3.5), the isoleucine codon ATA (4.1), the leucine codon CTA (3.9) and the TAG stop (0.2). Host table (CSV) downloads the chosen host's full table.
Worked example: EGFP for E. coli
The example is the coding sequence of EGFP, the enhanced green fluorescent protein, from NCBI record U55762.1 (the pEGFP-N1 cloning vector). Its codons follow human genes closely: its CAI is 0.957 for human, 0.696 for E. coli and 0.516 for yeast. Optimizing it for E. coli changes 126 of its 240 codons, raises the CAI to 1.000 and lowers the GC content from 61.5% to 48.9%.
Before you order a gene
Read the optimized sequence before using it. Check that your cloning sites are not created or lost with the restriction digest calculator on Plasmid Map, and that the GC content and any repeats suit your synthesis provider. Translate it back with DNA to protein to confirm the protein. A high CAI does not guarantee expression, which also depends on the promoter, the start of the mRNA (check a Kozak context or a Shine-Dalgarno sequence), the protein's folding and toxicity, and the growth conditions.
Sources: host annotations from NCBI RefSeq and MANE (see the table); Biopython 1.85 CodonAdaptationIndex. Automated tests compare every weight, CAI value and optimized sequence with Biopython 1.85; see sources and methods.