Skip to main content
License: Rfam retrieves data from the Rfam database, distributed under CC0-1.0 (public domain; no attribution required). The client wrapper code is MIT-licensed. Please refer to the data terms for full terms.

Proto is not affiliated with EMBL-EBI. This toolkit is open source and builds on the implementation produced by this organization. Product names, logos, and trademarks are the property of their respective owners.


/
/
View repo
rfam.org
Visit website
Rfam 15: RNA families database in 2025
Nancy Ontiveros-Palacios, Emma Cooke, … Blake Sweeney
Nucleic Acids Research (2025)
Read paper
Copy citation
evo-design/proto-tools/proto_tools/tools/database_retrieval/rfam
View source
Open Notebook
Open notebook
proto-tools on GitHub
Run locally with proto-tools
Toolkit contributors

Background

Rfam is developed at EMBL-EBI. An RNA family is a group of sequences believed to be evolutionarily related through similarity in sequence or secondary structure. Related families may be grouped into clans. The database and its release 15.0 updates are described in Rfam 15: RNA families database in 2025 by Ontiveros-Palacios et al., published in Nucleic Acids Research. Each family has a manually curated seed alignment, a representative set of sequences annotated with a consensus secondary structure. Rfam uses this alignment to build a covariance model, a statistical model that scores both sequence and secondary structure similarity. Infernal searches these models against the Rfamseq sequence database to identify additional candidate homologues. A curator-defined gathering cutoff specifies the bit-score threshold for inclusion in the family. The family-building documentation describes this process and the sources of structural annotations. The toolkit retrieves these existing records through the Rfam API. rfam-family reads the family description as JSON and extracts the consensus structure (#=GC SS_cons) and reference annotation (#=GC RF) from the Stockholm seed alignment. rfam-regions parses the family’s region table and separates strand orientation from the start and end coordinates. Both outputs report the Rfam release.

Learning Resources

  • How Rfam families are built (Rfam) - seed alignments, structural annotations, and covariance-model searches.
  • Rfam glossary (Rfam) - definitions of families, clans, gathering cutoffs, and alignment formats.
  • Rfam API (Rfam) - reference for family records, sequence regions, and alignments.
  • Infernal documentation (Eddy lab) - the software used to build and search RNA covariance models.

Tools

Rfam Family (rfam-family)

Retrieves a family by accession or family ID and returns its description, RNA type, curation information, sequence and species counts, clan membership when available, and the gathering cutoff for family membership. The output also contains the consensus secondary structure, reference annotation, and database release information. The complete Stockholm seed alignment can be included through configuration.

API Reference

Source
string
required
Rfam accession (e.g. ‘RF01731’) or family ID (e.g. ‘TwoAYGGAY’).
Source
boolean
default:"False"
Return the full Stockholm seed alignment, not only its consensus lines.
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cpu"
Device to run the tool on.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
Source
string
required
Rfam accession.
string
required
Rfam family ID.
string
required
One-line family description.
string
required
Rfam type annotation (e.g. ‘Cis-reg’, ‘Gene; snRNA’).
string
Curator comment, when the family has one.
string
required
Family authors.
string
required
Where the seed alignment came from.
string
Where the consensus structure came from.
integer
required
Sequences in the seed alignment.
integer
required
Annotated regions across all sequences.
integer
required
Species with at least one annotated region.
string
Accession of the clan the family belongs to.
string
ID of the clan the family belongs to.
number
required
Bit-score threshold for family membership.
string
required
Rfam release the record comes from.
string
required
Consensus secondary structure (#=GC SS_cons), WUSS notation.
string
required
Reference consensus sequence (#=GC RF), aligned with it.
string
Full Stockholm seed alignment, when requested.

Applications

Family records provide context for interpreting RNA annotations. The consensus structure describes conserved pairing patterns across the alignment, while the curation fields identify the sources of the alignment and structure. These records support comparisons of representative family sequences and interpretation of model scores alongside the reported cutoffs. The seed alignment can also be used for further alignment or structural analysis.

Usage Tips

  • Families can be identified by accession or ID. For example, RF01731 and TwoAYGGAY identify the same Rfam family. The output reports both identifiers.
  • Consensus annotations use alignment coordinates. consensus_structure contains the Stockholm SS_cons annotation in WUSS notation; consensus_sequence contains the RF reference annotation. These strings include alignment columns and should not be interpreted as an unaligned nucleotide sequence or genomic coordinates.
  • Structural annotations have different sources. Rfam includes both experimentally supported and computationally predicted structures. The structure_source field records provenance when available; the Rfam documentation explains why an underlying publication may be needed to establish the type of evidence.
  • The seed alignment is optional in the output. Set include_seed_alignment=True to retain it and enable sto export. The tool downloads the alignment to extract the consensus annotations even when this option is disabled. Family records can also be exported as JSON.

Rfam Regions (rfam-regions)

Retrieves the annotated sequence regions for a family, with optional filters for NCBI taxonomy ID, species name, or sequence accession. Each region contains a versioned sequence accession, Infernal bit score, start and end coordinates, strand, sequence description, species name, and taxonomy ID. The output includes the family identifiers, total and filtered region counts, and a flag indicating whether the returned list was truncated.

API Reference

Source
integer
Keep only hits in this NCBI taxonomy ID.
string
Keep only hits whose species contains this text (case-insensitive).
string
Keep only hits on this sequence accession. The version suffix is optional (‘AM181176’ matches ‘AM181176.4’).
string
required
Rfam accession (e.g. ‘RF01731’) or family ID (e.g. ‘TwoAYGGAY’).
Source
integer
default:"500"
Most regions returned after filtering; the rest are counted, not listed.
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cpu"
Device to run the tool on.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
Source
string
required
Rfam accession of the family.
string
required
Rfam family ID.
string
Rfam release the regions were built from.
integer
required
Regions the family has across all sequences.
integer
required
Regions left after the input filters.
boolean
required
Whether matched regions exceeded max_regions.
List[RfamRegion]
Matched regions, at most max_regions of them.

Applications

Region records locate candidate family members in the sequences represented by Rfam. Filtering by species or sequence accession supports examination of annotated loci in a particular organism or genome record. The accession, coordinates, and strand can be passed to ncbi-efetch to retrieve the corresponding nucleotide subsequence for comparative analysis. Taxonomic fields also support examination of a family’s distribution within the Rfam dataset.

Usage Tips

  • Coordinates are 1-indexed and inclusive. The tool normalizes each region to start <= end and reports orientation separately as + or -. These values correspond to ncbi-efetch’s seq_start, seq_stop, and strand inputs.
  • Filters are applied after download. taxid matches an exact taxonomy ID, species matches a case-insensitive substring, and sequence_accession accepts either a versioned or an unversioned accession. When several filters are provided, a region must satisfy all of them.
  • The return limit applies after filtering. max_regions defaults to 500. total_regions reports the family-wide count, matched_regions counts all regions satisfying the filters, and truncated indicates that some matching regions were omitted. The limit does not reduce the size of the download.
  • Some families are too large for the endpoint. The Rfam API documentation states that the server can refuse region downloads for very large families. Local filters cannot bypass this restriction.
  • Region tables can be exported. JSON preserves the full output, including counts and release information; TSV and CSV contain the returned region rows.

Toolkit Notes

These apply to every Rfam tool in this toolkit (rfam-family, rfam-regions).
  • Requires network access. Both tools retrieve data from the Rfam website using HTTPS requests and execute in the current Python process.
  • Results depend on the Rfam release. The tools query the live website rather than selecting a fixed database release. Retain the reported release information with exported results for provenance.
Example notebook: See the full working example for a copy-paste-ready walkthrough.

Infrastructure Guides

The following guides cover how to run tools efficiently and at scale.

Tool Persistence

Keep a tool’s model warm across calls instead of reloading it every invocation.

Device Management

How GPUs are allocated to tools and how to target specific devices.

Parallel Execution

Fan a batch of inputs out across multiple GPUs.