Proto is not affiliated with EMBL-EBI. This toolkit is open source and builds on the implementation produced by this organization. Product names, logos, and trademarks are the property of their respective owners.
Background
Rfam is developed at EMBL-EBI. An RNA family is a group of sequences believed to be evolutionarily related through similarity in sequence or secondary structure. Related families may be grouped into clans. The database and its release 15.0 updates are described in Rfam 15: RNA families database in 2025 by Ontiveros-Palacios et al., published in Nucleic Acids Research. Each family has a manually curated seed alignment, a representative set of sequences annotated with a consensus secondary structure. Rfam uses this alignment to build a covariance model, a statistical model that scores both sequence and secondary structure similarity. Infernal searches these models against the Rfamseq sequence database to identify additional candidate homologues. A curator-defined gathering cutoff specifies the bit-score threshold for inclusion in the family. The family-building documentation describes this process and the sources of structural annotations. The toolkit retrieves these existing records through the Rfam API.rfam-family reads the family description as JSON and extracts the consensus structure (#=GC SS_cons) and reference annotation (#=GC RF) from the Stockholm seed alignment. rfam-regions parses the family’s region table and separates strand orientation from the start and end coordinates. Both outputs report the Rfam release.
Learning Resources
- How Rfam families are built (Rfam) - seed alignments, structural annotations, and covariance-model searches.
- Rfam glossary (Rfam) - definitions of families, clans, gathering cutoffs, and alignment formats.
- Rfam API (Rfam) - reference for family records, sequence regions, and alignments.
- Infernal documentation (Eddy lab) - the software used to build and search RNA covariance models.
Tools
Rfam Family (rfam-family)
Retrieves a family by accession or family ID and returns its description, RNA type, curation information, sequence and species counts, clan membership when available, and the gathering cutoff for family membership. The output also contains the consensus secondary structure, reference annotation, and database release information. The complete Stockholm seed alignment can be included through configuration.API Reference
Input: RfamFamilyInput
Input: RfamFamilyInput
Config: RfamFamilyConfig
Config: RfamFamilyConfig
True is coerced to 1 and False to 0.None waits indefinitely.BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.Output: RfamFamilyOutput
Output: RfamFamilyOutput
Applications
Family records provide context for interpreting RNA annotations. The consensus structure describes conserved pairing patterns across the alignment, while the curation fields identify the sources of the alignment and structure. These records support comparisons of representative family sequences and interpretation of model scores alongside the reported cutoffs. The seed alignment can also be used for further alignment or structural analysis.Usage Tips
- Families can be identified by accession or ID. For example,
RF01731andTwoAYGGAYidentify the same Rfam family. The output reports both identifiers. - Consensus annotations use alignment coordinates.
consensus_structurecontains the StockholmSS_consannotation in WUSS notation;consensus_sequencecontains theRFreference annotation. These strings include alignment columns and should not be interpreted as an unaligned nucleotide sequence or genomic coordinates. - Structural annotations have different sources. Rfam includes both experimentally supported and computationally predicted structures. The
structure_sourcefield records provenance when available; the Rfam documentation explains why an underlying publication may be needed to establish the type of evidence. - The seed alignment is optional in the output. Set
include_seed_alignment=Trueto retain it and enablestoexport. The tool downloads the alignment to extract the consensus annotations even when this option is disabled. Family records can also be exported as JSON.
Rfam Regions (rfam-regions)
Retrieves the annotated sequence regions for a family, with optional filters for NCBI taxonomy ID, species name, or sequence accession. Each region contains a versioned sequence accession, Infernal bit score, start and end coordinates, strand, sequence description, species name, and taxonomy ID. The output includes the family identifiers, total and filtered region counts, and a flag indicating whether the returned list was truncated.API Reference
Input: RfamRegionsInput
Input: RfamRegionsInput
Config: RfamRegionsConfig
Config: RfamRegionsConfig
True is coerced to 1 and False to 0.None waits indefinitely.BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.Output: RfamRegionsOutput
Output: RfamRegionsOutput
Applications
Region records locate candidate family members in the sequences represented by Rfam. Filtering by species or sequence accession supports examination of annotated loci in a particular organism or genome record. The accession, coordinates, and strand can be passed toncbi-efetch to retrieve the corresponding nucleotide subsequence for comparative analysis. Taxonomic fields also support examination of a family’s distribution within the Rfam dataset.Usage Tips
- Coordinates are 1-indexed and inclusive. The tool normalizes each region to
start <= endand reports orientation separately as+or-. These values correspond toncbi-efetch’sseq_start,seq_stop, andstrandinputs. - Filters are applied after download.
taxidmatches an exact taxonomy ID,speciesmatches a case-insensitive substring, andsequence_accessionaccepts either a versioned or an unversioned accession. When several filters are provided, a region must satisfy all of them. - The return limit applies after filtering.
max_regionsdefaults to 500.total_regionsreports the family-wide count,matched_regionscounts all regions satisfying the filters, andtruncatedindicates that some matching regions were omitted. The limit does not reduce the size of the download. - Some families are too large for the endpoint. The Rfam API documentation states that the server can refuse region downloads for very large families. Local filters cannot bypass this restriction.
- Region tables can be exported. JSON preserves the full output, including counts and release information; TSV and CSV contain the returned region rows.
Toolkit Notes
These apply to every Rfam tool in this toolkit (rfam-family, rfam-regions).
- Requires network access. Both tools retrieve data from the Rfam website using HTTPS requests and execute in the current Python process.
- Results depend on the Rfam release. The tools query the live website rather than selecting a fixed database release. Retain the reported release information with exported results for provenance.

