Analyze DNA & RNA sequences.
Understand every result.
A browser-based toolkit for sequence manipulation and analysis, with the algorithms behind every result explained.
- Runs locally in your browser
- Nothing uploaded
- FASTA compatible
Sequence Utilities
Quick sequence transformations and calculations.
DNA Nucleotide Count
Count A, C, G, and T bases in a DNA sequence.
Open tool →Transcribe DNA to RNA
Convert DNA sequences to RNA by replacing thymine with uracil.
Open tool →Complement DNA Strand
Generate the complementary DNA strand using standard base pairing.
Open tool →Translate RNA to Protein
Translate RNA codons into amino acid sequences.
Open tool →Sequence Analysis
Go beyond basic transformations with deeper sequence analysis.
Compute GC Content
Calculate GC percentage across one or more FASTA sequences.
Open tool →DNA Motif Finder
Find sequence motifs with configurable matching and mismatch options.
Open tool →ORF Analysis
Identify candidate open reading frames and translate coding regions.
Includes: Basic Finder · Advanced TranslationWhy this platform
Runs locally
Sequence calculations happen directly in your browser.
Privacy first
Sequences are not uploaded to a remote server.
Explained algorithms
See how each result is calculated instead of receiving a black-box answer. Every tool includes the exact algorithm and pseudocode behind its output, not just the number.
How it works
Every tool explains its method alongside the result, so you can verify the calculation.
Ready to analyze a sequence?
Choose a tool and start analyzing directly in your browser.
Explore all tools →DNA Nucleotide Count
This tool counts the individual frequency of Adenine (A), Cytosine (C), Guanine (G), and Thymine (T) bases across an input DNA strand. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.
Learning Panel: DNA Nucleotide Count Guide
What is nucleotide counting?
DNA is made of four nucleotide bases:
- A — Adenine
- T — Thymine
- G — Guanine
- C — Cytosine
Nucleotide counting tells you how many times each base occurs in a DNA sequence.
Why is it useful?
Counting bases is one of the simplest ways to analyze a DNA sequence. It provides basic information that can be used to calculate GC content, compare sequences, and perform quality checks.
How does it work?
Example:
Technical Details: Algorithm, Pseudocode & Implementation
Algorithm
- Clean the input: strip the FASTA header line (if any), remove line breaks, and uppercase the sequence.
- Count each base independently: scan the cleaned sequence once per base type (A, C, G, T) and tally occurrences.
- Build the outputs: populate the dictionary-style summary and the results table from the same four counts.
Pseudocode
Implementation Details
- Four independent scans: the running tool counts each base with its own regular-expression match rather than a single pass that increments a shared tally, since sequences here are short enough that the simpler four-scan approach is fast and easy to read.
- Character validation:
cleanInput()strips anything outsideA,T,C,G,Nbefore counting, so unexpected characters are silently removed rather than causing an error. - Ambiguous base N: an
Nbase still counts toward the total sequence length shown elsewhere in the app, but it is not tallied under A, C, G, or T, so the four displayed counts can sum to less than the total length when N bases are present.
Enter DNA sequence
INPUT (DNA Query)
OUTPUT (DNA Nucleotide Count)
1. Print dictionary
2. Display DataFrame
| nucleotide | count |
|---|
Interactive Sequence Visualization
Click any base below, or a nucleotide in the legend, to highlight every occurrence of that base and see its count and percentage. Click it again to clear the selection.
Transcribing DNA into RNA
This tool transcribes a coding (sense) DNA strand into its corresponding single-stranded RNA. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.
Learning Panel: RNA Transcription Guide
What is transcription?
Transcription is the process of making an RNA sequence from DNA. In RNA, uracil (U) takes the place of thymine (T); the other three bases (A, C, G) stay the same.
What do 5′ and 3′ mean?
A DNA or RNA strand has two different ends. This means the molecule has a direction, like a one-way street.
The direction is labeled 5′ (pronounced "five prime") at one end and 3′ ("three prime") at the other.
Each nucleotide is built around a sugar ring with numbered carbons. At the 5′ end, the first sugar's phosphate group sticks out. That phosphate sits on the sugar's 5th carbon, so this end is called the 5′ end. At the 3′ end, the last sugar added has an exposed hydroxyl group. That hydroxyl sits on the sugar's 3rd carbon, so this end is called the 3′ end.
Direction matters in biology. DNA replication and transcription only run one way along the strand, from 5′ to 3′, never backwards. That's why you'll find strands in this app labeled with their direction. It's not decoration; it shows which way the underlying process reads or builds the molecule.
DNA has two strands — which one does the cell actually read?
DNA is double-stranded. The strand you paste into this tool is the coding strand (also called the sense strand) — by convention, written 5′→3′, and it reads almost identically to the RNA that gets produced. But that is not the strand RNA polymerase actually reads.
RNA polymerase reads the other strand, the template strand (antisense strand), running in the opposite direction (3′→5′), and builds RNA by pairing a complementary base against each template base, growing the new RNA 5′→3′.
Read down each column: the template base is the DNA complement of the coding base above it (A↔T, C↔G). The RNA base is then built by pairing against the template using DNA→RNA rules (template A pairs with RNA U, template T pairs with RNA A, template C pairs with RNA G, template G pairs with RNA C).
Why the result always matches "just swap T for U"
Follow one column all the way through: complementing a coding base to get the template base, then complementing the template base again to get the RNA base, is two complements in a row — and complementing twice cancels out. The only base where that cancellation doesn't return the original letter is T, because DNA has no U: T's DNA complement is A, and A's DNA→RNA pairing partner is U. So the net effect, column by column, is always: T becomes U, and A, C, and G pass straight through unchanged.
The interactive visualization below shows all three strands explicitly, so you can see the full mechanism in action.
Why is it useful?
Cells use RNA as an intermediate step between DNA and protein production. Understanding transcription is essential for understanding the central dogma of molecular biology:
Technical Details: Algorithm, Pseudocode & Implementation
Algorithm
- Clean the input: strip the FASTA header line (if any), remove line breaks, and uppercase the sequence. This is treated as the coding (sense) strand, 5′→3′, matching every other tool in this suite.
- Derive the template strand (shown in the visualization): complement each coding base (A↔T, C↔G) to get the antisense strand the cell actually reads, conventionally labeled 3′→5′.
- Derive the RNA strand from the template: complement each template base again, using DNA→RNA pairing (template A→U, T→A, C→G, G→C), producing RNA 5′→3′.
- Simplify: since steps 2 and 3 are two complements in a row, they algebraically reduce to a single substitution: replace every T with U, leave A/C/G unchanged. The implementation computes this directly rather than building the template strand as an intermediate step, since the two approaches are mathematically guaranteed to agree.
- Report statistics: compare original and transcribed lengths and count how many substitutions were made.
Pseudocode
Full mechanism (what actually happens biologically, and what the visualization below renders):
Equivalent shortcut (what the tool actually computes, proven identical to the full mechanism above for every base):
Implementation Details
- Single substitution pass: the tool performs one
replaceAll('T', 'U')call on the coding strand rather than materializing an intermediate template strand, since the two are proven equivalent above. - Length is always preserved: transcription is a one-to-one base substitution, so the RNA output is always exactly the same length as the cleaned DNA input.
- Coding strand convention: like every other tool in this suite (Complement, ORF Analysis, Motif Finder), the pasted sequence is treated as the coding strand, 5′→3′. The template strand is derived for display purposes only, on demand, in the interactive visualization below — it is not what the input is assumed to already be.
Enter DNA sequence
INPUT (DNA Sequence)
OUTPUT (Transcribed RNA Sequence)
1. Transcribed RNA Sequence
2. Transcription Details
Interactive Transcription Visualization
All three strands, shown biologically: the coding strand you entered, the template strand the cell actually reads (its complement, 3′→5′), and the RNA built from the template (5′→3′). Click any column to inspect it, or press Play to watch transcription happen once, one base at a time.
Complementing a Strand of DNA
This tool generates both the direct complement and standard biological reverse complement (5' to 3') strands for a DNA sequence. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.
Learning Panel: DNA Complement Guide
What is a complementary DNA strand?
DNA has two strands whose bases pair according to specific rules:
The complement of a DNA sequence replaces every base with its partner.
Why is it useful?
Complementary base pairing is fundamental to DNA replication, sequencing, molecular biology experiments, and many bioinformatics algorithms.
How does it work?
The algorithm is simply:
What do 5′ and 3′ mean?
DNA and RNA strands have a direction, like a one-way street. One end is called 5′ ("five prime") and the other 3′ ("three prime"), named after which carbon of the sugar ring is exposed at that end.
The two paired strands of a DNA double helix always run in opposite directions. If one strand reads 5′→3′ left to right, its partner reads 3′→5′ left to right. Biologists call this antiparallel.
This is exactly why complement and reverse complement are different operations below. A straight complement keeps the original left-to-right order. But that doesn't match the other strand's own natural reading direction. Reversing it does.
Important distinction: Complement ≠ Reverse Complement
Complement:
Reverse complement:
A reverse complement first finds the complement and then reverses the result.
Technical Details: Algorithm, Pseudocode & Implementation
Algorithm
- Clean the input: strip the FASTA header, remove whitespace, and uppercase the sequence.
- Compute the direct complement: map every base through the A↔T, C↔G pairing table, in the original order.
- Compute the reverse complement: reverse the direct-complement string.
- Report pairing statistics: count A-T pairs and G-C pairs across the original sequence.
Pseudocode
Implementation Details
- Complement-then-reverse order: the tool always builds the direct complement first and reverses that result to get the reverse complement, rather than reversing first. Both orders are mathematically equivalent for this operation.
- Shared helper: the same
BASE_COMPLEMENTmap and the equivalent of agetReverseComplement()pattern used here are reused by the Motif Finder and ORF Analysis tools elsewhere in this app. - Ambiguous base N: an
Nbase maps to itself in the complement, since its identity is unknown.
Enter DNA sequence
INPUT (DNA Sequence)
OUTPUT (Complementary Sequences)
1. Reverse Complement Sequence (5' to 3')
2. Direct Complement Sequence
3. Sequence Details
Interactive Complement Visualization
Click any base to see exactly why the direct complement and the reverse complement look different — same base pairs, read in two different directions. Or press Play to watch it happen once, one base at a time.
Translating RNA into Protein
This tool translates an RNA sequence into its corresponding amino acid chain by mapping nucleotide triplets (codons) against the standard genetic code. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.
Learning Panel: Protein Translation Guide
What is translation?
Translation is the process of converting an RNA sequence into an amino-acid sequence. The ribosome reads RNA in groups of three nucleotides called codons.
Why is it useful?
Proteins perform most of the structural and functional work inside cells. Translation lets us predict the protein sequence encoded by an RNA sequence.
How does it work?
Example:
Each codon corresponds to an amino acid.
Some codons signal stop rather than an amino acid:
Start codon
The standard start codon is:
It can signal where translation begins.
Do all organisms use every codon equally?
Most amino acids can be written by more than one codon. For example, leucine has six. But real organisms don't pick between them at random. Each species has favorite codons, and the favorites differ from one species to the next. This is called codon usage bias.
The tool on this page always translates with the same standard genetic code, so the protein it gives you doesn't depend on the organism. The table below is here as a reference, so you can see how differently real organisms use the same code.
Open the codon usage table
Codon Usage Table
How to read this table (start here if this is new to you)
Every cell in your body builds proteins by reading its DNA like a recipe. The recipe is read three letters at a time. Each three-letter group is called a codon. DNA only has four possible letters — A, T, C, and G — so there are 4 × 4 × 4 = 64 possible codons in total. That's why this table has 64 boxes.
Each codon is an instruction for one amino acid — amino acids are the small building-block molecules that get strung together, one after another, to form a protein. Some codons don't code for an amino acid at all; instead they say "stop building" — these are the stop codons, shown in red.
There are only 20 amino acids but 64 codons, so most amino acids are spelled out by more than one codon. For example, Leucine can be written six different ways. Codons that code for the same amino acid are called synonymous codons — and organisms don't use their synonymous codons equally. That uneven preference is called codon usage bias, and it's the whole reason this table exists.
How to look up a codon, for example GAC:
- Take the 1st letter (G) — find that row on the far left.
- Take the 2nd letter (A) — find that column, headed "2nd Base: A".
- Take the 3rd letter (C) — within that row group, find the sub-row labelled C on the far right.
- The box where they all meet is your answer: GAC codes for Aspartate (Asp, D), and the number tells you how often the selected organism actually uses this codon versus its synonymous partners.
What's the difference between "Intuitional" and "Proportional"?
Intuitional answers: "Out of all the codons for this one amino acid, how often is this particular one picked?" The numbers for each amino acid add up to 1. This is the number most people reach for when they're choosing which codon to use in a gene they're designing.
Proportional answers a slightly different question: "Out of every 1000 codons anywhere in this organism's genes, how many are this exact codon?" This blends two things at once — how common the amino acid itself is, and how much the organism prefers this codon over its synonyms — so the numbers are comparable across the whole table, not just within one amino acid's group.
Flow: DNA → RNA → Protein
Technical Details: Algorithm, Pseudocode & Implementation
Algorithm
- Clean and normalize the input: strip the FASTA header and whitespace, uppercase the sequence, and convert
Uback toTso the codon lookup can reuse the DNA-based codon table. - Read codons from position 0: take consecutive non-overlapping 3-base windows starting at the very first base of the input, with no search for a start codon.
- Translate each codon: look up the corresponding amino acid in the standard codon table.
- Stop at the first stop codon: once a codon translates to "Stop," translation ends immediately and that codon is excluded from the protein.
- Report statistics: sequence length, codons processed, and final protein length.
Pseudocode
Implementation Details
- No start-codon search: unlike the ORF Analysis tools, this translator does not look for
AUG/ATGbefore translating. It reads from position 0 of whatever sequence you enter, so the input is expected to already begin at the intended reading frame. - Shared codon table: the RNA input is converted to its DNA-letter equivalent internally and looked up against the same
STANDARD_CODON_TABLEused by the ORF Analysis tools, rather than maintaining a separate RNA-specific table. - Trailing partial codon: the loop condition only processes complete 3-base groups, so 1–2 leftover bases at the end of the sequence (if the length isn't a multiple of 3) are silently ignored rather than flagged.
- No stop codon means no truncation: if the input never reaches an in-frame stop codon, translation simply continues until it runs out of complete codons.
Enter RNA sequence
INPUT (RNA Sequence)
OUTPUT (Translated Protein Sequence)
1. Amino Acid Sequence
2. Translation Details
Interactive Translation Visualization
Click any codon to see which amino acid it becomes — and every other codon that would have produced the exact same amino acid. The genetic code is redundant on purpose; this is where that becomes visible. Or press Play to translate once, one codon at a time.
Computing GC Content
This tool evaluates multiple FASTA DNA records to compute and identify the entry with the highest GC content ratio. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.
Learning Panel: GC Content Guide
What is GC content?
GC content is the percentage of bases in a DNA sequence that are either Guanine (G) or Cytosine (C). The formula is:
Why is it useful?
GC content provides information about the composition of a DNA sequence. It is useful when comparing sequences and can affect properties such as DNA stability and experimental behavior.
How does it work?
Example:
Count:
Therefore:
Visual explanation
Technical Details: Algorithm, Pseudocode & Implementation
Algorithm
- Split the FASTA dataset: break the input into individual records wherever a line starts with
>. - Parse each record: the first line of a record is treated as its ID; the remaining lines are joined and cleaned into the sequence.
- Count G and C bases: tally G and C occurrences in the cleaned sequence for each record.
- Compute the ratio: GC% = (G + C) / sequence length × 100 for each record.
- Find every tied maximum: after all records are scored, find the highest displayed GC percentage, then collect every record that shares it, rather than stopping at the first one found.
- Report results: display the highest-GC record (or the full set of tied records) plus a table of every record's stats, with tied rows visibly flagged.
Pseudocode
Implementation Details
- Record splitting: the dataset is split with a regular expression that matches
>only at the start of a line, so a stray>character inside a sequence body would not incorrectly start a new record. - Empty-sequence guard: if a record's sequence has zero length after cleaning, its GC ratio is reported as 0% rather than causing a division-by-zero error.
- Per-record cleaning: each record's sequence is independently uppercased and stripped of any character outside A/T/C/G before counting, so a malformed header or stray whitespace in one record does not affect others.
- Tie-breaking: ties are detected at the same 6-decimal precision shown in the table. If two or more records share the highest displayed GC ratio, every tied record is flagged with a 🏆 marker and highlighted in the table, and the summary lists all tied sequence IDs rather than silently picking one.
- Reported summary: the top result box always reflects the current tie state — a single ID and percentage when there's a unique highest record, or a count plus the full list of tied IDs when there isn't.
Enter dataset
Output (GC Content)
| Sequence ID | Sequence Length | G+C Count | GC Content (%) |
|---|
Interactive GC Content Visualization
Click any row above to load that sequence here, or click any base below to change it and watch GC content update live. This is a sandbox — it never changes your real input or the table above. Use Reset to undo your edits.
Click a base to cycle it A → T → G → C → A. G and C bases are underlined throughout.
DNA Motif Finder
This tool locates sequence motifs in DNA with support for overlapping matches, IUPAC ambiguity codes, mismatch tolerance, and reverse complement search.
Learning Panel: DNA Motif Finder Guide
What is a DNA motif?
A DNA motif is a short sequence pattern that can have biological significance. Motifs may occur in regions involved in processes such as:
- transcription regulation
- protein-DNA binding
- replication
- RNA processing
- other molecular interactions
For example, a researcher may be interested in the pattern TATA and want to determine where that pattern occurs in a DNA sequence.
Important: a motif is a pattern, not automatically a gene or protein-coding region. Finding TATA does not by itself prove that a functional regulatory element exists.
Why do scientists search for motifs?
DNA contains enormous amounts of sequence information. Researchers may already know that a particular protein tends to bind near a certain sequence pattern. Instead of manually examining thousands of bases, a computer can search the sequence and report every location where the pattern occurs.
This allows researchers to quickly investigate potential sites of biological interest.
Motif vs DNA sequence
Think of it this way: the DNA sequence is the complete string being examined. The motif is the smaller pattern being searched for inside it.
How does the computer search for a motif?
Suppose Sequence = GATATAC and Motif = TATA (length 4). The computer examines every possible 4-base window and compares it with the motif. Only the window that equals TATA exactly counts as a match.
The window moves one nucleotide at a time. This one-nucleotide movement is important because it allows overlapping matches to be found.
Why overlapping matches matter
An overlapping match occurs when two motif occurrences share some nucleotides. Consider Sequence = ATATAT, Motif = ATAT:
The second match starts two bases after the first, and the two matches overlap. If the algorithm jumped forward by the motif's entire length after finding a match, it would miss the second occurrence. Overlapping matches should normally be reported, because they represent real occurrences of the requested pattern.
Positions: 0-based vs 1-based
Computers commonly count positions starting at 0, but biology users often prefer counting from 1, since it is easier for beginners to interpret. For Sequence ATATAT, Motif ATAT:
This tool displays positions as 1-based for readability, and shows the 0-based index in the technical details below where relevant.
Lowercase input
DNA sequences may appear as atgcatata or ATGCATATA. Biologically these represent the same nucleotide sequence, so this tool normalizes all input to uppercase before searching. Matching is case-insensitive.
🔎 What does the tool report?
The result depends on the Matching Algorithm you choose.
1. Exact Matching
Finds positions where the DNA sequence matches the motif exactly, base for base.
Use when: you know the exact sequence you're looking for.
2. IUPAC Ambiguous Nucleotides
Real biological motifs are often not perfectly identical at every position. Instead of writing several exact motifs separately, an IUPAC ambiguity code represents a set of allowed bases at that position. For example, R = A or G, so a motif written TARA means the third position can be A or G.
| Code | Means | Memory Trick |
|---|---|---|
| A | A | — |
| C | C | — |
| G | G | — |
| T | T | — |
| R | A or G | puRine |
| Y | C or T | pYrimidine |
| S | G or C | Strong (3 H bonds) |
| W | A or T | Weak (2 H bonds) |
| K | G or T | Keto |
| M | A or C | aMino |
| B | C, G, or T | not A (B comes right after A) |
| D | A, G, or T | not C (D comes right after C) |
| H | A, C, or T | not G (H comes right after G) |
| V | A, C, or G | not T/U (V comes right after U) |
| N | A, C, G, or T | aNy base |
For B, D, H, and V, the pattern is: each letter means "everything except the base whose letter comes right before it in the alphabet." B excludes A, D excludes C, H excludes G, and V excludes T (thinking of it as U, the RNA version of T).
The tool reports the position, matched sequence, and the motif pattern used. Use when: the biological pattern is known, but some positions can vary.
Where do motifs like this come from? Scientists don't invent them randomly — they study many DNA sequences where a particular protein is known to bind, notice that some positions are highly conserved while others vary, and represent that pattern with IUPAC notation (or, for more advanced analysis, a position weight matrix). This tool works from a motif you already supply, not automated motif discovery.
3. Mismatch Tolerance — Hamming Distance
Exact matching is not always enough. This mode finds sequences similar to the motif even when they contain a limited number of substitutions. Hamming distance simply counts how many positions differ between two equal-length sequences.
If Maximum Mismatches = 0, only exact matches count. If Maximum Mismatches = 1, this example is reported as a match. The tool reports the position, matched sequence, and the number of mismatches. Important: Hamming distance allows substitutions only, not insertions or deletions.
4. Search Reverse Complement Strand (+ and -)
DNA is double-stranded, so a motif may occur on either strand.
What do 5′ and 3′ mean? A DNA strand has a direction, like a one-way street. One end is labeled 5′ ("five prime"), the other 3′ ("three prime"), named after which carbon of the sugar ring is exposed there. The two strands of a DNA double helix always run in opposite directions: if one strand reads 5′→3′ left to right, its partner reads 3′→5′ left to right.
To search the opposite strand in the standard 5′→3′ orientation, the tool calculates the reverse complement — not to be confused with the plain complement.
The tool searches the forward strand and, when this mode is selected, the reverse complement strand, then reports which strand produced each match:
Use when: you don't want to miss a motif simply because it occurs on the opposite DNA strand.
🧠 Which search should I use?
| I want to... | Choose |
|---|---|
| Find an exact sequence | Exact Matching |
| Search a pattern containing variable bases | IUPAC Ambiguous Nucleotides |
| Find similar sequences with a few substitutions | Mismatch Tolerance |
| Search both DNA strands | Reverse Complement |
Technical Details: Algorithm, Pseudocode & Implementation
Algorithm
Exact matching
If reverse-complement searching is enabled
Worked example: DNA = ATATCGTAT, Motif = TAT, motif length = 3.
Complexity: for a sequence of length n and a motif of length m, exact and IUPAC matching run in O((n − m + 1) × m) time, since every window position performs a base-by-base comparison. Mismatch tolerance mode has the same bound, with an early exit once the mismatch count exceeds the allowed maximum. Reverse-complement mode runs the same scan twice, once per strand.
Pseudocode — Exact Matching
Pseudocode — IUPAC Matching
Instead of asking window == motif, the program asks whether each nucleotide is allowed by the corresponding motif symbol.
Only if every position is valid do we report a match.
Pseudocode — Mismatch Tolerance
Maximum mismatches = 0 behaves as exact matching. Maximum mismatches = 1 allows one substitution, and so on.
Pseudocode — Reverse Complement
Implementation Details
- Overlapping matches are kept: the window advances one base at a time regardless of match outcome, so two matches that share nucleotides (e.g. both
ATAToccurrences inATATAT) are both reported rather than the second being skipped. - One motif field, multiple patterns: the motif input accepts several patterns separated by commas or newlines; each is searched independently and all results are merged into a single sorted-by-position table.
- Reverse complement is computed once per motif, not per position: when Reverse Complement mode is on, the tool derives the reverse complement of the motif itself and runs the normal forward-scan logic against it, rather than reverse-complementing the DNA sequence.
Interactive Search Visualization
Click any position in the sequence to test the sliding window there, or press Play Search to watch the algorithm scan the whole sequence, one window at a time — matching whichever algorithm and settings are selected above. Uses the first motif if you entered several.
ORF Analysis
This tool identifies candidate Open Reading Frames (ORFs) and translates nucleotide sequences across reading frames. Sequence datasets are provided in FASTA format, where each entry begins with a > symbol followed by a header identifier.
Learning Panel: ORF Analysis Guide
What is an ORF?
An Open Reading Frame (ORF) is a stretch of DNA that can potentially encode a protein.
In DNA, the standard start codon is ATG.
The standard stop codons are TAA, TAG, and TGA.
An ORF begins at a start codon and continues in groups of three nucleotides (codons) until an in-frame stop codon is reached.
Why do we look for ORFs?
Proteins are built from amino acids, and the information for making a protein is encoded in DNA.
The ribosome ultimately reads the information as groups of three nucleotides:
Example:
An ORF finder helps us identify regions that could potentially be translated into proteins.
How does an ORF finder work?
Your tool essentially performs this process:
Step 1: Read the DNA
Suppose we have:
Ignore spaces:
The tool examines the sequence in different reading frames.
Step 2: Find a start codon
The standard DNA start codon is ATG.
Finding ATG alone does not automatically mean we have a complete ORF. We need to continue reading in the same frame.
Step 3: Read codons in groups of three
Starting at ATG:
Each group of three is one codon. We continue until we encounter a stop codon.
Step 4: Find an in-frame stop codon
The stop codons are TAA, TAG, and TGA.
Therefore, ATGAAACCCTAG is a complete candidate ORF.
In-frame sequence alignment is crucial. Consider:
TAG is in the same reading frame. A stop codon elsewhere in the DNA sequence is not automatically relevant. The tool must move +3 bases at a time from the start codon.
Step 5: Translate the ORF
Once a complete ORF is identified, its codons can be translated.
Protein: MKP
The stop codon signals termination and does not become an amino acid.
What happens if another stop codon appears?
Suppose we have:
The first stop codon is TGA. The ORF ends there: ATG | AAA | TGA. The later TAA is not part of that ORF.
What about multiple start codons?
Consider:
There are two possible starts:
- ORF 1:
ATG | AAA | ATG | CCC | TAA(Protein:MKMP) - ORF 2:
ATG | CCC | TAA(Protein:MP)
The second ORF is nested inside the first. The tool reports both ORFs and flags the nested relationship rather than discarding the shorter sequence.
What is a reading frame?
DNA is read in groups of three nucleotides. Consider: ATGAAACCC
There are three possible ways to divide this sequence into codons on one strand:
- +1:
ATG | AAA | CCC - +2:
TGA | AAC - +3:
GAA | ACC
A shift of just one nucleotide changes every codon after that point.
Why are there six reading frames?
DNA has two strands, and either strand can serve as the coding strand. Each strand has three reading frames:
What do 5′ and 3′ mean?
A DNA or RNA strand has two different ends. This means the molecule has a direction, like a one-way street.
The direction is labeled 5′ (pronounced "five prime") at one end and 3′ ("three prime") at the other, named after which carbon of the sugar ring is exposed at that end.
The two strands of a DNA double helix always run in opposite directions. If one strand reads 5′→3′ left to right, its partner reads 3′→5′ left to right. That's why the reverse complement (below) needs reversing, not just complementing, to correctly represent the opposite strand's own reading direction.
What is the reverse complement?
DNA has two complementary strands. If you have 5' - ATGC - 3', the complementary strand is 3' - TACG - 5'.
Sequence analysis conventionally reads sequences in the 5' to 3' direction, so we reverse it to 5' - CGTA - 3':
An ORF Finder needs the reverse complement when examining the opposite strand in the standard 5' to 3' orientation.
Why length matters
ORFs can be short. For example, ATG TAA is technically a start codon followed immediately by a stop codon, but that does not mean it is biologically meaningful.
Applying a minimum ORF length setting (such as 30 nucleotides) lets the tool distinguish candidate ORFs passing a length threshold from random noise.
Candidate ORF versus Gene
Real gene identification involves additional evidence:
- Biological context
- Sequence conservation
- Expression evidence
- Regulatory signals
- Experimental validation
- Organism-specific features
Technical Details: Algorithm, Pseudocode & Implementation
Algorithm
- Prepare the sequence: Remove FASTA headers, spaces, and line breaks and validate the nucleotide sequence.
- Generate the reverse complement: This allows the tool to examine both DNA strands.
- Create six reading frames: Three frames are generated from the original sequence and three from its reverse complement.
- Search each frame: Starting from each possible position, scan codons three nucleotides at a time.
- Identify candidate ORFs: When an ATG is encountered, continue scanning until an in-frame TAA, TAG, or TGA is found.
- Record the ORF: Store its strand, reading frame, start position, end position, nucleotide sequence, and length.
- Translate: Convert the ORF's codons into amino acids, excluding the terminal stop codon from the protein sequence.
- Analyze results: Identify nested ORFs, apply optional length filters, and remove duplicate protein sequences when appropriate.
Pseudocode
Implementation Details
- Coordinate Translation: Convert positions on the negative strand back to original forward sequence coordinates:
• Forward Start = Sequence Length - Reverse End + 1
• Forward End = Sequence Length - Reverse Start - Memory Optimization: Slice sequence strings efficiently during frame extraction to avoid unnecessary memory overhead on large inputs.
- Nested ORF Handling: Track active start positions within reading frames to capture and flag nested ORFs without dropping valid candidate sequences.
Advanced Codon Settings
Interactive Reading Frame Explorer
This is the concept that makes 6-frame analysis confusing at first: the same DNA reads as completely different codons depending on where you start counting. Switch frames below and watch every codon boundary shift. Click any codon to check whether it's a start or stop signal, based on the start/stop codon boxes you've checked above.