Analyze DNA & RNA sequences.
Understand every result.

A browser-based toolkit for sequence manipulation and analysis, with the algorithms behind every result explained.

  • Runs locally in your browser
  • Nothing uploaded
  • FASTA compatible
SEQUENCE PREVIEW
A T G C C A T G G A T C A
GC Content53.8%
Sequence Length13

Sequence Utilities

Quick sequence transformations and calculations.

DNA Nucleotide Count

Count A, C, G, and T bases in a DNA sequence.

Open tool →

Transcribe DNA to RNA

Convert DNA sequences to RNA by replacing thymine with uracil.

Open tool →

Complement DNA Strand

Generate the complementary DNA strand using standard base pairing.

Open tool →

Translate RNA to Protein

Translate RNA codons into amino acid sequences.

Open tool →

Sequence Analysis

Go beyond basic transformations with deeper sequence analysis.

Compute GC Content

Calculate GC percentage across one or more FASTA sequences.

Open tool →
🔍

DNA Motif Finder

Find sequence motifs with configurable matching and mismatch options.

Open tool →

ORF Analysis

Identify candidate open reading frames and translate coding regions.

Includes: Basic Finder · Advanced Translation

Why this platform

Runs locally

Sequence calculations happen directly in your browser.

Privacy first

Sequences are not uploaded to a remote server.

How it works

01
Paste your sequence
02
Run the analysis
03
Understand the result

Every tool explains its method alongside the result, so you can verify the calculation.

Ready to analyze a sequence?

Choose a tool and start analyzing directly in your browser.

Explore all tools →

DNA Nucleotide Count

This tool counts the individual frequency of Adenine (A), Cytosine (C), Guanine (G), and Thymine (T) bases across an input DNA strand. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.

Learning Panel: DNA Nucleotide Count Guide

What is nucleotide counting?

DNA is made of four nucleotide bases:

  • A — Adenine
  • T — Thymine
  • G — Guanine
  • C — Cytosine

Nucleotide counting tells you how many times each base occurs in a DNA sequence.

Why is it useful?

Counting bases is one of the simplest ways to analyze a DNA sequence. It provides basic information that can be used to calculate GC content, compare sequences, and perform quality checks.

How does it work?

DNA sequence ↓ Read each base ↓ Count A, T, G and C ↓ Calculate totals and percentages

Example:

DNA: ATGCCAT A = 2 T = 2 G = 1 C = 2 Total = 7
Remember: Every position in a DNA sequence contains one of four standard bases: A, T, G, or C.
Technical Details: Algorithm, Pseudocode & Implementation

Algorithm

  1. Clean the input: strip the FASTA header line (if any), remove line breaks, and uppercase the sequence.
  2. Count each base independently: scan the cleaned sequence once per base type (A, C, G, T) and tally occurrences.
  3. Build the outputs: populate the dictionary-style summary and the results table from the same four counts.

Pseudocode

sequence = clean(sequence) counts = { A: count_of("A", sequence), C: count_of("C", sequence), G: count_of("G", sequence), T: count_of("T", sequence) } display(counts)

Implementation Details

  • Four independent scans: the running tool counts each base with its own regular-expression match rather than a single pass that increments a shared tally, since sequences here are short enough that the simpler four-scan approach is fast and easy to read.
  • Character validation: cleanInput() strips anything outside A, T, C, G, N before counting, so unexpected characters are silently removed rather than causing an error.
  • Ambiguous base N: an N base still counts toward the total sequence length shown elsewhere in the app, but it is not tallied under A, C, G, or T, so the four displayed counts can sum to less than the total length when N bases are present.

Enter DNA sequence


INPUT (DNA Query)

OUTPUT (DNA Nucleotide Count)

1. Print dictionary

2. Display DataFrame

nucleotidecount

Interactive Sequence Visualization

Click any base below, or a nucleotide in the legend, to highlight every occurrence of that base and see its count and percentage. Click it again to clear the selection.

Transcribing DNA into RNA

This tool transcribes a coding (sense) DNA strand into its corresponding single-stranded RNA. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.

Learning Panel: RNA Transcription Guide

What is transcription?

Transcription is the process of making an RNA sequence from DNA. In RNA, uracil (U) takes the place of thymine (T); the other three bases (A, C, G) stay the same.

What do 5′ and 3′ mean?

A DNA or RNA strand has two different ends. This means the molecule has a direction, like a one-way street.

The direction is labeled 5′ (pronounced "five prime") at one end and 3′ ("three prime") at the other.

Each nucleotide is built around a sugar ring with numbered carbons. At the 5′ end, the first sugar's phosphate group sticks out. That phosphate sits on the sugar's 5th carbon, so this end is called the 5′ end. At the 3′ end, the last sugar added has an exposed hydroxyl group. That hydroxyl sits on the sugar's 3rd carbon, so this end is called the 3′ end.

Direction matters in biology. DNA replication and transcription only run one way along the strand, from 5′ to 3′, never backwards. That's why you'll find strands in this app labeled with their direction. It's not decoration; it shows which way the underlying process reads or builds the molecule.

DNA has two strands — which one does the cell actually read?

DNA is double-stranded. The strand you paste into this tool is the coding strand (also called the sense strand) — by convention, written 5′→3′, and it reads almost identically to the RNA that gets produced. But that is not the strand RNA polymerase actually reads.

RNA polymerase reads the other strand, the template strand (antisense strand), running in the opposite direction (3′→5′), and builds RNA by pairing a complementary base against each template base, growing the new RNA 5′→3′.

Coding strand (you provide this): 5′ — A T G C C A T — 3′ | | | | | | | Template strand (cell reads this): 3′ — T A C G G T A — 5′ | | | | | | | RNA (cell builds this): 5′ — A U G C C A U — 3′

Read down each column: the template base is the DNA complement of the coding base above it (A↔T, C↔G). The RNA base is then built by pairing against the template using DNA→RNA rules (template A pairs with RNA U, template T pairs with RNA A, template C pairs with RNA G, template G pairs with RNA C).

Why the result always matches "just swap T for U"

Follow one column all the way through: complementing a coding base to get the template base, then complementing the template base again to get the RNA base, is two complements in a row — and complementing twice cancels out. The only base where that cancellation doesn't return the original letter is T, because DNA has no U: T's DNA complement is A, and A's DNA→RNA pairing partner is U. So the net effect, column by column, is always: T becomes U, and A, C, and G pass straight through unchanged.

Coding A → template T → RNA A (unchanged) Coding T → template A → RNA U (the only real change) Coding G → template C → RNA G (unchanged) Coding C → template G → RNA C (unchanged)

The interactive visualization below shows all three strands explicitly, so you can see the full mechanism in action.

Why is it useful?

Cells use RNA as an intermediate step between DNA and protein production. Understanding transcription is essential for understanding the central dogma of molecular biology:

DNA → RNA → Protein
Remember: The strand you paste is the coding strand, not the strand the cell actually reads — the template strand is its complement, read in the opposite direction. Don't confuse transcription with translation (Transcription: DNA → RNA, Translation: RNA → Protein).
Technical Details: Algorithm, Pseudocode & Implementation

Algorithm

  1. Clean the input: strip the FASTA header line (if any), remove line breaks, and uppercase the sequence. This is treated as the coding (sense) strand, 5′→3′, matching every other tool in this suite.
  2. Derive the template strand (shown in the visualization): complement each coding base (A↔T, C↔G) to get the antisense strand the cell actually reads, conventionally labeled 3′→5′.
  3. Derive the RNA strand from the template: complement each template base again, using DNA→RNA pairing (template A→U, T→A, C→G, G→C), producing RNA 5′→3′.
  4. Simplify: since steps 2 and 3 are two complements in a row, they algebraically reduce to a single substitution: replace every T with U, leave A/C/G unchanged. The implementation computes this directly rather than building the template strand as an intermediate step, since the two approaches are mathematically guaranteed to agree.
  5. Report statistics: compare original and transcribed lengths and count how many substitutions were made.

Pseudocode

Full mechanism (what actually happens biologically, and what the visualization below renders):

coding = clean(sequence) # 5' -> 3' template = map(coding, base -> DNA_COMPLEMENT[base]) # 3' -> 5' rna = map(template, base -> TEMPLATE_TO_RNA[base]) # 5' -> 3'

Equivalent shortcut (what the tool actually computes, proven identical to the full mechanism above for every base):

coding = clean(sequence) rna = replace_all(coding, "T", "U") t_count = count_of("T", coding) report(rna, length(coding), length(rna), t_count)

Implementation Details

  • Single substitution pass: the tool performs one replaceAll('T', 'U') call on the coding strand rather than materializing an intermediate template strand, since the two are proven equivalent above.
  • Length is always preserved: transcription is a one-to-one base substitution, so the RNA output is always exactly the same length as the cleaned DNA input.
  • Coding strand convention: like every other tool in this suite (Complement, ORF Analysis, Motif Finder), the pasted sequence is treated as the coding strand, 5′→3′. The template strand is derived for display purposes only, on demand, in the interactive visualization below — it is not what the input is assumed to already be.

Enter DNA sequence


INPUT (DNA Sequence)

OUTPUT (Transcribed RNA Sequence)

1. Transcribed RNA Sequence

2. Transcription Details

Interactive Transcription Visualization

All three strands, shown biologically: the coding strand you entered, the template strand the cell actually reads (its complement, 3′→5′), and the RNA built from the template (5′→3′). Click any column to inspect it, or press Play to watch transcription happen once, one base at a time.

A T (DNA) / U (RNA) G C

Complementing a Strand of DNA

This tool generates both the direct complement and standard biological reverse complement (5' to 3') strands for a DNA sequence. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.

Learning Panel: DNA Complement Guide

What is a complementary DNA strand?

DNA has two strands whose bases pair according to specific rules:

A ↔ T C ↔ G

The complement of a DNA sequence replaces every base with its partner.

Why is it useful?

Complementary base pairing is fundamental to DNA replication, sequencing, molecular biology experiments, and many bioinformatics algorithms.

How does it work?

Original: A T G C C A Complement: T A C G G T

The algorithm is simply:

A → T T → A C → G G → C

What do 5′ and 3′ mean?

DNA and RNA strands have a direction, like a one-way street. One end is called 5′ ("five prime") and the other 3′ ("three prime"), named after which carbon of the sugar ring is exposed at that end.

The two paired strands of a DNA double helix always run in opposite directions. If one strand reads 5′→3′ left to right, its partner reads 3′→5′ left to right. Biologists call this antiparallel.

This is exactly why complement and reverse complement are different operations below. A straight complement keeps the original left-to-right order. But that doesn't match the other strand's own natural reading direction. Reversing it does.

Important distinction: Complement ≠ Reverse Complement

Complement:

ATGC ↓ TACG

Reverse complement:

ATGC ↓ GCAT

A reverse complement first finds the complement and then reverses the result.

Remember: Complementing changes the bases. Reversing changes their order.
Technical Details: Algorithm, Pseudocode & Implementation

Algorithm

  1. Clean the input: strip the FASTA header, remove whitespace, and uppercase the sequence.
  2. Compute the direct complement: map every base through the A↔T, C↔G pairing table, in the original order.
  3. Compute the reverse complement: reverse the direct-complement string.
  4. Report pairing statistics: count A-T pairs and G-C pairs across the original sequence.

Pseudocode

sequence = clean(sequence) complement = map(sequence, base -> COMPLEMENT[base]) reverse_complement = reverse(complement) report(complement, reverse_complement)

Implementation Details

  • Complement-then-reverse order: the tool always builds the direct complement first and reverses that result to get the reverse complement, rather than reversing first. Both orders are mathematically equivalent for this operation.
  • Shared helper: the same BASE_COMPLEMENT map and the equivalent of a getReverseComplement() pattern used here are reused by the Motif Finder and ORF Analysis tools elsewhere in this app.
  • Ambiguous base N: an N base maps to itself in the complement, since its identity is unknown.

Enter DNA sequence


INPUT (DNA Sequence)

OUTPUT (Complementary Sequences)

1. Reverse Complement Sequence (5' to 3')

2. Direct Complement Sequence

3. Sequence Details

Interactive Complement Visualization

Click any base to see exactly why the direct complement and the reverse complement look different — same base pairs, read in two different directions. Or press Play to watch it happen once, one base at a time.

A T G C
Reverse complement — same bases as the row above, read backwards (5′→3′):

Translating RNA into Protein

This tool translates an RNA sequence into its corresponding amino acid chain by mapping nucleotide triplets (codons) against the standard genetic code. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.

Learning Panel: Protein Translation Guide

What is translation?

Translation is the process of converting an RNA sequence into an amino-acid sequence. The ribosome reads RNA in groups of three nucleotides called codons.

RNA ↓ Codons ↓ Amino acids ↓ Protein

Why is it useful?

Proteins perform most of the structural and functional work inside cells. Translation lets us predict the protein sequence encoded by an RNA sequence.

How does it work?

Example:

RNA: AUG GCU AAA UGG ↓ ↓ ↓ ↓ M A K W

Each codon corresponds to an amino acid.

Some codons signal stop rather than an amino acid:

UAA UAG UGA

Start codon

The standard start codon is:

AUG → Methionine (M)

It can signal where translation begins.

Do all organisms use every codon equally?

Most amino acids can be written by more than one codon. For example, leucine has six. But real organisms don't pick between them at random. Each species has favorite codons, and the favorites differ from one species to the next. This is called codon usage bias.

The tool on this page always translates with the same standard genetic code, so the protein it gives you doesn't depend on the organism. The table below is here as a reference, so you can see how differently real organisms use the same code.

Open the codon usage table

Codon Usage Table

How to read this table (start here if this is new to you)

Every cell in your body builds proteins by reading its DNA like a recipe. The recipe is read three letters at a time. Each three-letter group is called a codon. DNA only has four possible letters — A, T, C, and G — so there are 4 × 4 × 4 = 64 possible codons in total. That's why this table has 64 boxes.

Each codon is an instruction for one amino acid — amino acids are the small building-block molecules that get strung together, one after another, to form a protein. Some codons don't code for an amino acid at all; instead they say "stop building" — these are the stop codons, shown in red.

There are only 20 amino acids but 64 codons, so most amino acids are spelled out by more than one codon. For example, Leucine can be written six different ways. Codons that code for the same amino acid are called synonymous codons — and organisms don't use their synonymous codons equally. That uneven preference is called codon usage bias, and it's the whole reason this table exists.

How to look up a codon, for example GAC:

  1. Take the 1st letter (G) — find that row on the far left.
  2. Take the 2nd letter (A) — find that column, headed "2nd Base: A".
  3. Take the 3rd letter (C) — within that row group, find the sub-row labelled C on the far right.
  4. The box where they all meet is your answer: GAC codes for Aspartate (Asp, D), and the number tells you how often the selected organism actually uses this codon versus its synonymous partners.
What's the difference between "Intuitional" and "Proportional"?

Intuitional answers: "Out of all the codons for this one amino acid, how often is this particular one picked?" The numbers for each amino acid add up to 1. This is the number most people reach for when they're choosing which codon to use in a gene they're designing.

Proportional answers a slightly different question: "Out of every 1000 codons anywhere in this organism's genes, how many are this exact codon?" This blends two things at once — how common the amino acid itself is, and how much the organism prefers this codon over its synonyms — so the numbers are comparable across the whole table, not just within one amino acid's group.

Most-used codon for that amino acid Stop codon
Remember: Three RNA bases = one codon. One codon specifies one amino acid or a stop signal.
Flow: DNA → RNA → Protein
Technical Details: Algorithm, Pseudocode & Implementation

Algorithm

  1. Clean and normalize the input: strip the FASTA header and whitespace, uppercase the sequence, and convert U back to T so the codon lookup can reuse the DNA-based codon table.
  2. Read codons from position 0: take consecutive non-overlapping 3-base windows starting at the very first base of the input, with no search for a start codon.
  3. Translate each codon: look up the corresponding amino acid in the standard codon table.
  4. Stop at the first stop codon: once a codon translates to "Stop," translation ends immediately and that codon is excluded from the protein.
  5. Report statistics: sequence length, codons processed, and final protein length.

Pseudocode

sequence = clean(sequence) sequence = replace_all(sequence, "U", "T") protein = [] for i from 0 to length(sequence) - 3 step 3: codon = sequence[i : i+3] aa = CODON_TABLE[codon] if aa == "Stop": break protein.append(aa) report(protein)

Implementation Details

  • No start-codon search: unlike the ORF Analysis tools, this translator does not look for AUG/ATG before translating. It reads from position 0 of whatever sequence you enter, so the input is expected to already begin at the intended reading frame.
  • Shared codon table: the RNA input is converted to its DNA-letter equivalent internally and looked up against the same STANDARD_CODON_TABLE used by the ORF Analysis tools, rather than maintaining a separate RNA-specific table.
  • Trailing partial codon: the loop condition only processes complete 3-base groups, so 1–2 leftover bases at the end of the sequence (if the length isn't a multiple of 3) are silently ignored rather than flagged.
  • No stop codon means no truncation: if the input never reaches an in-frame stop codon, translation simply continues until it runs out of complete codons.

Enter RNA sequence


INPUT (RNA Sequence)

OUTPUT (Translated Protein Sequence)

1. Amino Acid Sequence

2. Translation Details

Interactive Translation Visualization

Click any codon to see which amino acid it becomes — and every other codon that would have produced the exact same amino acid. The genetic code is redundant on purpose; this is where that becomes visible. Or press Play to translate once, one codon at a time.

A U G C

Computing GC Content

This tool evaluates multiple FASTA DNA records to compute and identify the entry with the highest GC content ratio. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.

Learning Panel: GC Content Guide

What is GC content?

GC content is the percentage of bases in a DNA sequence that are either Guanine (G) or Cytosine (C). The formula is:

GC Content = ((G + C) / Total) × 100

Why is it useful?

GC content provides information about the composition of a DNA sequence. It is useful when comparing sequences and can affect properties such as DNA stability and experimental behavior.

How does it work?

Example:

DNA: ATGCCGTA

Count:

G = 2 C = 2 G + C = 4 Total = 8

Therefore:

GC% = 4 / 8 × 100 = 50%

Visual explanation

A T G C C G T A ------- GC bases
Remember: GC content measures composition, not whether a sequence contains a gene or motif. A high GC percentage does not automatically mean a sequence is a gene.
Technical Details: Algorithm, Pseudocode & Implementation

Algorithm

  1. Split the FASTA dataset: break the input into individual records wherever a line starts with >.
  2. Parse each record: the first line of a record is treated as its ID; the remaining lines are joined and cleaned into the sequence.
  3. Count G and C bases: tally G and C occurrences in the cleaned sequence for each record.
  4. Compute the ratio: GC% = (G + C) / sequence length × 100 for each record.
  5. Find every tied maximum: after all records are scored, find the highest displayed GC percentage, then collect every record that shares it, rather than stopping at the first one found.
  6. Report results: display the highest-GC record (or the full set of tied records) plus a table of every record's stats, with tied rows visibly flagged.

Pseudocode

records = split(dataset, on_lines_starting_with = ">") for each record in records: id = first_line(record) seq = clean(remaining_lines(record)) gc_count = count_of("G" or "C", seq) gc_ratio = round((gc_count / length(seq)) * 100, 6) record.gc_ratio = gc_ratio add_row(id, length(seq), gc_count, gc_ratio) max_gc = max(record.gc_ratio for record in records) tied = [record for record in records if record.gc_ratio == max_gc] flag_rows(tied) if length(tied) > 1: report(count(tied), max_gc, ids_of(tied)) else: report(tied[0].id, max_gc)

Implementation Details

  • Record splitting: the dataset is split with a regular expression that matches > only at the start of a line, so a stray > character inside a sequence body would not incorrectly start a new record.
  • Empty-sequence guard: if a record's sequence has zero length after cleaning, its GC ratio is reported as 0% rather than causing a division-by-zero error.
  • Per-record cleaning: each record's sequence is independently uppercased and stripped of any character outside A/T/C/G before counting, so a malformed header or stray whitespace in one record does not affect others.
  • Tie-breaking: ties are detected at the same 6-decimal precision shown in the table. If two or more records share the highest displayed GC ratio, every tied record is flagged with a 🏆 marker and highlighted in the table, and the summary lists all tied sequence IDs rather than silently picking one.
  • Reported summary: the top result box always reflects the current tie state — a single ID and percentage when there's a unique highest record, or a count plus the full list of tied IDs when there isn't.

Enter dataset

Output (GC Content)

Seq_0808 60.919540
Sequence IDSequence LengthG+C CountGC Content (%)

Interactive GC Content Visualization

Click any row above to load that sequence here, or click any base below to change it and watch GC content update live. This is a sandbox — it never changes your real input or the table above. Use Reset to undo your edits.

—

Click a base to cycle it A → T → G → C → A. G and C bases are underlined throughout.

G count0
C count0
GC bases0
Total bases0
GC Content0.0%

DNA Motif Finder

This tool locates sequence motifs in DNA with support for overlapping matches, IUPAC ambiguity codes, mismatch tolerance, and reverse complement search.

Learning Panel: DNA Motif Finder Guide

What is a DNA motif?

A DNA motif is a short sequence pattern that can have biological significance. Motifs may occur in regions involved in processes such as:

  • transcription regulation
  • protein-DNA binding
  • replication
  • RNA processing
  • other molecular interactions

For example, a researcher may be interested in the pattern TATA and want to determine where that pattern occurs in a DNA sequence.

Important: a motif is a pattern, not automatically a gene or protein-coding region. Finding TATA does not by itself prove that a functional regulatory element exists.

Why do scientists search for motifs?

DNA contains enormous amounts of sequence information. Researchers may already know that a particular protein tends to bind near a certain sequence pattern. Instead of manually examining thousands of bases, a computer can search the sequence and report every location where the pattern occurs.

DNA: GGCATATACCGGTATAGC Motif: TATA Found: GGCA[TATA]CCGGTATAGC ↑ another region

This allows researchers to quickly investigate potential sites of biological interest.

Motif vs DNA sequence

Think of it this way: the DNA sequence is the complete string being examined. The motif is the smaller pattern being searched for inside it.

Entire book → DNA sequence Word/pattern → motif Finding every hit → motif search

How does the computer search for a motif?

Suppose Sequence = GATATAC and Motif = TATA (length 4). The computer examines every possible 4-base window and compares it with the motif. Only the window that equals TATA exactly counts as a match.

Sequence: G A T A T A C Window 1: [G A T A] T A C Window 2: G [A T A T] A C Window 3: G A [T A T A] C ✓ match Window 4: G A T [A T A C]

The window moves one nucleotide at a time. This one-nucleotide movement is important because it allows overlapping matches to be found.

Why overlapping matches matter

An overlapping match occurs when two motif occurrences share some nucleotides. Consider Sequence = ATATAT, Motif = ATAT:

[ATAT]AT → Position 1 AT[ATAT] → Position 3

The second match starts two bases after the first, and the two matches overlap. If the algorithm jumped forward by the motif's entire length after finding a match, it would miss the second occurrence. Overlapping matches should normally be reported, because they represent real occurrences of the requested pattern.

Positions: 0-based vs 1-based

Computers commonly count positions starting at 0, but biology users often prefer counting from 1, since it is easier for beginners to interpret. For Sequence ATATAT, Motif ATAT:

Sequence: A T A T A T 0-based: 0 1 2 3 4 5 1-based: 1 2 3 4 5 6 Reported as: Position 1, Position 3 (1-based)

This tool displays positions as 1-based for readability, and shows the 0-based index in the technical details below where relevant.

Lowercase input

DNA sequences may appear as atgcatata or ATGCATATA. Biologically these represent the same nucleotide sequence, so this tool normalizes all input to uppercase before searching. Matching is case-insensitive.

🔎 What does the tool report?

The result depends on the Matching Algorithm you choose.

1. Exact Matching

Finds positions where the DNA sequence matches the motif exactly, base for base.

DNA: GATATATGCATATACTT Motif: ATAT Matches: Position 3 Position 4 Position 10

Use when: you know the exact sequence you're looking for.

2. IUPAC Ambiguous Nucleotides

Real biological motifs are often not perfectly identical at every position. Instead of writing several exact motifs separately, an IUPAC ambiguity code represents a set of allowed bases at that position. For example, R = A or G, so a motif written TARA means the third position can be A or G.

CodeMeansMemory Trick
AA—
CC—
GG—
TT—
RA or GpuRine
YC or TpYrimidine
SG or CStrong (3 H bonds)
WA or TWeak (2 H bonds)
KG or TKeto
MA or CaMino
BC, G, or Tnot A (B comes right after A)
DA, G, or Tnot C (D comes right after C)
HA, C, or Tnot G (H comes right after G)
VA, C, or Gnot T/U (V comes right after U)
NA, C, G, or TaNy base

For B, D, H, and V, the pattern is: each letter means "everything except the base whose letter comes right before it in the alphabet." B excludes A, D excludes C, H excludes G, and V excludes T (thinking of it as U, the RNA version of T).

Motif: TARA TAAA ✓ match TAGA ✓ match TACA ✗ no match

The tool reports the position, matched sequence, and the motif pattern used. Use when: the biological pattern is known, but some positions can vary.

Where do motifs like this come from? Scientists don't invent them randomly — they study many DNA sequences where a particular protein is known to bind, notice that some positions are highly conserved while others vary, and represent that pattern with IUPAC notation (or, for more advanced analysis, a position weight matrix). This tool works from a motif you already supply, not automated motif discovery.

3. Mismatch Tolerance — Hamming Distance

Exact matching is not always enough. This mode finds sequences similar to the motif even when they contain a limited number of substitutions. Hamming distance simply counts how many positions differ between two equal-length sequences.

Motif: T A T A Sequence: T A G A Compare: T=T ✓ A=A ✓ T≠G ✗ A=A ✓ Hamming distance = 1

If Maximum Mismatches = 0, only exact matches count. If Maximum Mismatches = 1, this example is reported as a match. The tool reports the position, matched sequence, and the number of mismatches. Important: Hamming distance allows substitutions only, not insertions or deletions.

4. Search Reverse Complement Strand (+ and -)

DNA is double-stranded, so a motif may occur on either strand.

What do 5′ and 3′ mean? A DNA strand has a direction, like a one-way street. One end is labeled 5′ ("five prime"), the other 3′ ("three prime"), named after which carbon of the sugar ring is exposed there. The two strands of a DNA double helix always run in opposite directions: if one strand reads 5′→3′ left to right, its partner reads 3′→5′ left to right.

To search the opposite strand in the standard 5′→3′ orientation, the tool calculates the reverse complement — not to be confused with the plain complement.

For: ATGC Complement: TACG Reverse complement: GCAT

The tool searches the forward strand and, when this mode is selected, the reverse complement strand, then reports which strand produced each match:

Position: 42 Strand: + Sequence: ATGC Position: 87 Strand: - Sequence: GCAT

Use when: you don't want to miss a motif simply because it occurs on the opposite DNA strand.

🧠 Which search should I use?

I want to...Choose
Find an exact sequenceExact Matching
Search a pattern containing variable basesIUPAC Ambiguous Nucleotides
Find similar sequences with a few substitutionsMismatch Tolerance
Search both DNA strandsReverse Complement
Remember: Motif search starts with a pattern you already know and asks: "Where does it occur?" That is different from motif discovery, where the pattern itself is unknown.
Technical Details: Algorithm, Pseudocode & Implementation

Algorithm

Exact matching

Sequence ↓ Normalize to uppercase ↓ Validate DNA characters ↓ Read motif ↓ Move a window across sequence ↓ Compare window with motif ↓ Match? ── YES → Record position └─ NO → Continue

If reverse-complement searching is enabled

Original sequence ↓ Reverse complement ↓ Search both ↓ Combine results

Worked example: DNA = ATATCGTAT, Motif = TAT, motif length = 3.

Possible windows: ATA TAT ✓ ATC TCG CGT GTA TAT ✓ Positions (1-based): 2, 7

Complexity: for a sequence of length n and a motif of length m, exact and IUPAC matching run in O((n − m + 1) × m) time, since every window position performs a base-by-base comparison. Mismatch tolerance mode has the same bound, with an early exit once the mismatch count exceeds the allowed maximum. Reverse-complement mode runs the same scan twice, once per strand.

Pseudocode — Exact Matching

sequence = uppercase(sequence) motif = uppercase(motif) matches = [] for i from 0 to length(sequence) - length(motif): window = sequence[i : i + length(motif)] if window == motif: add i to matches # 1-based display position display_position = i + 1

Pseudocode — IUPAC Matching

Instead of asking window == motif, the program asks whether each nucleotide is allowed by the corresponding motif symbol.

for each candidate position i: is_match = true for each position j in motif: allowed_bases = IUPAC[motif[j]] if sequence[i + j] not in allowed_bases: is_match = false break if is_match: record match at position i

Only if every position is valid do we report a match.

Pseudocode — Mismatch Tolerance

for each possible position i: mismatches = 0 for each position j in motif: if sequence[i + j] != motif[j]: mismatches += 1 if mismatches > maximum_allowed: break if mismatches <= maximum_allowed: record match at position i, mismatches

Maximum mismatches = 0 behaves as exact matching. Maximum mismatches = 1 allows one substitution, and so on.

Pseudocode — Reverse Complement

forward_matches = find_matches(sequence, motif) rc_motif = reverse_complement(motif) reverse_matches = find_matches(sequence, rc_motif) # each match tagged strand = "-" all_matches = forward_matches + reverse_matches sort all_matches by position

Implementation Details

  • Overlapping matches are kept: the window advances one base at a time regardless of match outcome, so two matches that share nucleotides (e.g. both ATAT occurrences in ATATAT) are both reported rather than the second being skipped.
  • One motif field, multiple patterns: the motif input accepts several patterns separated by commas or newlines; each is searched independently and all results are merged into a single sorted-by-position table.
  • Reverse complement is computed once per motif, not per position: when Reverse Complement mode is on, the tool derives the reverse complement of the motif itself and runs the normal forward-scan logic against it, rather than reverse-complementing the DNA sequence.


Interactive Search Visualization

Click any position in the sequence to test the sliding window there, or press Play Search to watch the algorithm scan the whole sequence, one window at a time — matching whichever algorithm and settings are selected above. Uses the first motif if you entered several.

Scanning forward strand (+)

ORF Analysis

This tool identifies candidate Open Reading Frames (ORFs) and translates nucleotide sequences across reading frames. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.

Learning Panel: ORF Analysis Guide

What is an ORF?

An Open Reading Frame (ORF) is a stretch of DNA that can potentially encode a protein.

START STOP ↓ ↓ ATG --- codons --- TAA

In DNA, the standard start codon is ATG.

The standard stop codons are TAA, TAG, and TGA.

An ORF begins at a start codon and continues in groups of three nucleotides (codons) until an in-frame stop codon is reached.

Important: An ORF is a candidate protein-coding region. Finding an ORF does not prove that the DNA is actually a gene.

Why do we look for ORFs?

Proteins are built from amino acids, and the information for making a protein is encoded in DNA.

The ribosome ultimately reads the information as groups of three nucleotides:

DNA -> RNA -> Codons -> Amino acids -> Protein

Example:

DNA: ATG AAA TGC TAA ↓ ↓ ↓ ↓ M K C STOP ↓ MKC

An ORF finder helps us identify regions that could potentially be translated into proteins.

How does an ORF finder work?

Your tool essentially performs this process:

DNA sequence ↓ Clean / validate sequence ↓ Create reverse complement ↓ Generate 6 reading frames ↓ Search each frame for ATG ↓ Move forward 3 bases at a time ↓ Look for TAA / TAG / TGA ↓ Record complete ORFs ↓ Translate ORFs into amino acids ↓ Report results

Step 1: Read the DNA

Suppose we have:

CCATGAAACCC TAGG

Ignore spaces:

CCATGAAACCCTAGG

The tool examines the sequence in different reading frames.

Step 2: Find a start codon

The standard DNA start codon is ATG.

CC ATG AAA CCC TAG G ↑ START

Finding ATG alone does not automatically mean we have a complete ORF. We need to continue reading in the same frame.

Step 3: Read codons in groups of three

Starting at ATG:

ATG | AAA | CCC | TAG

Each group of three is one codon. We continue until we encounter a stop codon.

Step 4: Find an in-frame stop codon

The stop codons are TAA, TAG, and TGA.

ATG | AAA | CCC | TAG ↑ ↑ START STOP

Therefore, ATGAAACCCTAG is a complete candidate ORF.

In-frame sequence alignment is crucial. Consider:

ATG | AAA | CCC | TAG

TAG is in the same reading frame. A stop codon elsewhere in the DNA sequence is not automatically relevant. The tool must move +3 bases at a time from the start codon.

Step 5: Translate the ORF

Once a complete ORF is identified, its codons can be translated.

ATG | AAA | CCC | TAG ↓ ↓ ↓ ↓ M K P STOP

Protein: MKP

The stop codon signals termination and does not become an amino acid.

What happens if another stop codon appears?

Suppose we have:

ATG | AAA | TGA | CCC | TAA

The first stop codon is TGA. The ORF ends there: ATG | AAA | TGA. The later TAA is not part of that ORF.

What about multiple start codons?

Consider:

ATG | AAA | ATG | CCC | TAA

There are two possible starts:

  • ORF 1: ATG | AAA | ATG | CCC | TAA (Protein: MKMP)
  • ORF 2: ATG | CCC | TAA (Protein: MP)

The second ORF is nested inside the first. The tool reports both ORFs and flags the nested relationship rather than discarding the shorter sequence.

What is a reading frame?

DNA is read in groups of three nucleotides. Consider: ATGAAACCC

There are three possible ways to divide this sequence into codons on one strand:

  • +1: ATG | AAA | CCC
  • +2: TGA | AAC
  • +3: GAA | ACC

A shift of just one nucleotide changes every codon after that point.

Why are there six reading frames?

DNA has two strands, and either strand can serve as the coding strand. Each strand has three reading frames:

DNA / \ Forward Reverse complement / | \ / | \ +1 +2 +3 -1 -2 -3
3 forward frames + 3 reverse-complement frames = 6 reading frames
FORWARD +1 ATG | AAA | CCC | TAA +2 TGA | AAC | CCT +3 GAA | ACC | CTA REVERSE COMPLEMENT -1 ... -2 ... -3 ...

What do 5′ and 3′ mean?

A DNA or RNA strand has two different ends. This means the molecule has a direction, like a one-way street.

The direction is labeled 5′ (pronounced "five prime") at one end and 3′ ("three prime") at the other, named after which carbon of the sugar ring is exposed at that end.

The two strands of a DNA double helix always run in opposite directions. If one strand reads 5′→3′ left to right, its partner reads 3′→5′ left to right. That's why the reverse complement (below) needs reversing, not just complementing, to correctly represent the opposite strand's own reading direction.

What is the reverse complement?

DNA has two complementary strands. If you have 5' - ATGC - 3', the complementary strand is 3' - TACG - 5'.

Sequence analysis conventionally reads sequences in the 5' to 3' direction, so we reverse it to 5' - CGTA - 3':

ATGC ↓ TACG (complement) ↓ CGTA (reverse complement)

An ORF Finder needs the reverse complement when examining the opposite strand in the standard 5' to 3' orientation.

Why length matters

ORFs can be short. For example, ATG TAA is technically a start codon followed immediately by a stop codon, but that does not mean it is biologically meaningful.

Applying a minimum ORF length setting (such as 30 nucleotides) lets the tool distinguish candidate ORFs passing a length threshold from random noise.

Candidate ORF versus Gene

Warning: Finding an ORF does not prove that a functional gene or protein exists. An ORF is simply a sequence that has the characteristics of a possible protein-coding region.

Real gene identification involves additional evidence:

  • Biological context
  • Sequence conservation
  • Expression evidence
  • Regulatory signals
  • Experimental validation
  • Organism-specific features
Technical Details: Algorithm, Pseudocode & Implementation

Algorithm

  1. Prepare the sequence: Remove FASTA headers, spaces, and line breaks and validate the nucleotide sequence.
  2. Generate the reverse complement: This allows the tool to examine both DNA strands.
  3. Create six reading frames: Three frames are generated from the original sequence and three from its reverse complement.
  4. Search each frame: Starting from each possible position, scan codons three nucleotides at a time.
  5. Identify candidate ORFs: When an ATG is encountered, continue scanning until an in-frame TAA, TAG, or TGA is found.
  6. Record the ORF: Store its strand, reading frame, start position, end position, nucleotide sequence, and length.
  7. Translate: Convert the ORF's codons into amino acids, excluding the terminal stop codon from the protein sequence.
  8. Analyze results: Identify nested ORFs, apply optional length filters, and remove duplicate protein sequences when appropriate.

Pseudocode

sequence = clean(sequence) reverse = reverse_complement(sequence) for each strand in [sequence, reverse]: for frame in [0, 1, 2]: i = frame while i + 2 < length(strand): codon = strand[i : i+3] if codon == "ATG": start = i j = i + 3 while j + 2 < length(strand): codon = strand[j : j+3] if codon is a stop codon: record ORF(start, j+3) break j = j + 3 i = i + 3

Implementation Details

  • Coordinate Translation: Convert positions on the negative strand back to original forward sequence coordinates:
    • Forward Start = Sequence Length - Reverse End + 1
    • Forward End = Sequence Length - Reverse Start
  • Memory Optimization: Slice sequence strings efficiently during frame extraction to avoid unnecessary memory overhead on large inputs.
  • Nested ORF Handling: Track active start positions within reading frames to capture and flag nested ORFs without dropping valid candidate sequences.

Advanced Codon Settings

Interactive Reading Frame Explorer

This is the concept that makes 6-frame analysis confusing at first: the same DNA reads as completely different codons depending on where you start counting. Switch frames below and watch every codon boundary shift. Click any codon to check whether it's a start or stop signal, based on the start/stop codon boxes you've checked above.