GENETIC CODE REFERENCE

GC content calculator

Count bases, compare sequences and see where GC content changes.

Up to 100 records, 30,000 bases total and 100 KB. IUPAC letters accepted; whitespace and position numbers ignored. Your sequences stay in this browser.

Illustrative example · 600 bases.

Record 1 of 1

Sequence

GC among known bases45.83%

275 G or C / 600 A, C, G, T

Length
600 nt
Ambiguous
0 excluded
A
175
C
150
G
125
T
150

GC along the sequence

11 windows; GC among known bases ranges from 12.00% to 76.00%. 0 windows have no known bases.

Base position · points mark window centers. Gaps mean no known bases. Read exact windows.

100 nt window · 50 nt step · 11 windows. The final full window is anchored at the sequence end.

Inspect exact windows

Window 1: bases 1–10012.00% GC · 12 G or C / 100 known bases100 nt total · 0 ambiguous excluded

Exports include all records. Missing GC percentages stay blank in CSV.

How to calculate GC content

GC content is the share of guanine (G) and cytosine (C) bases in a sequence. For DNA with known bases:

GC % = 100 × (G + C) / (A + C + G + T)

For RNA, U replaces T in the denominator. For example, ATGCGC contains four G or C bases out of six, giving 66.67% GC. Its reverse complement, GCGCAT, has the same percentage. You can check both strands with the reverse complement tool.

Paste plain DNA or RNA, or open a FASTA file, and select Calculate GC content. Lowercase letters, whitespace and position numbers are accepted. Each FASTA record must have a header and a nonempty sequence. Mixed T and U within one record, alignment gaps and unsupported characters reject the input; records are never silently skipped. The limit is 100 records, 30,000 bases total and 100 KB.

How ambiguous bases are counted

The main percentage counts only known A, C, G and T/U letters. All IUPAC ambiguity letters are excluded from that denominator, including S and W. The result always gives the numerator, denominator and number excluded. If no known bases remain, the percentage is unavailable.

The separate possible GC range includes every position. S can only be G or C, so it always counts toward GC; W can only be A or T, so it never does. Each other ambiguity letter permits both a GC and an AT/U base. The lower bound picks a non-GC possibility wherever one exists; the upper bound picks a GC possibility wherever one exists. These are limits permitted by the letters, not probabilities or a confidence interval.

Sequence GC among known bases Possible GC over all positions
ATGC 50.00% (2 / 4) 50.00%
ATGCNN 50.00% (2 / 4) 33.33%–66.67%
SSWN Unavailable (0 / 0) 50.00%–75.00%
NNNN Unavailable (0 / 0) 0.00%–100.00%

Different calculators may distribute an N equally across bases or retain ambiguous positions in the denominator. Those policies answer different questions. This tool does neither for its main percentage. The reverse-complement reference lists each IUPAC symbol.

Read the sliding-window profile

Window is the number of sequence positions in each measurement. Step is how far the window moves each time. A 100-base window and a 50-base step first measure positions 1–100, then 51–150. Each point uses the same known-base denominator as the main result.

The last full window is anchored at the sequence end when the regular steps do not land there. A record shorter than the requested window uses one window containing the whole record. Step must be between 1 and the window size, so the profile covers the sequence without skipping stretches. Profiles are limited to 2,000 windows across the batch.

Points appear at window centers. Inspect exact windows shows the start, inclusive end, base counts and percentage for each window. A window with only ambiguous letters leaves a gap in the graph; it is not plotted as zero. Coordinates follow the supplied sequence from position 1, with headers, spaces and position numbers removed. Windows do not wrap around a circular sequence or cross between FASTA records.

GC percentage alone does not predict melting temperature, secondary structure or expression. After codon optimization, a profile can show where composition changed; it cannot establish that the optimized sequence will work better.

Compare records and save results

Use Inspect record to switch between sequences and Compare all records for the table. Duplicate FASTA names remain separate by record number. The pooled percentage divides the total G-plus-C count by the total known-base count; it is not the average of the displayed record percentages.

Records CSV contains every record's counts, known-base GC percentage and possible full-sequence range. Windows CSV contains every window, its record number, coordinates, actual length, requested size and step. An unavailable percentage is an empty CSV cell. Copy report and Report TXT include all record summaries and calculation assumptions. Changing the input or settings clears old results and disables their exports until you calculate again.

Sources and method

Nucleotide letters follow NCBI's FASTA guidance and the IUPAC-IUB recommendations on incompletely specified bases. The possible ranges are calculated directly from those allowed bases. Biopython's GC-fraction documentation describes several ambiguity policies; this tool's main result corresponds to calculating on A/C/G/T/U only, not Biopython's default treatment of S and W.

The opening example is an illustrative 600-base sequence with three composition blocks, not an identified gene. Its 275 G or C bases give 45.83% GC. All calculations run in your browser.