Skip to main content
Minerva
License: Minerva is open source and free for academic and commercial use under an Apache-2.0 license. Please refer to the license for full terms.

This toolkit is open source. Any third-party models, product names, or trademarks referenced are the property of their respective owners, and Proto is not affiliated with them.


garykbrixi/minerva
garykbrixi/minerva
Coevolutionary discovery using genome language models
24 stars
View repo
Coevolutionary mining of prokaryotic non-coding elements with a genome language model
David B. Li, Garyk Brixi, … Brian L. Hie
bioRxiv (2026)
Read preprint
Copy citation
evo-design/proto-tools/proto_tools/tools/masked_models/minerva
View source
Open Notebook
Open notebook
proto-tools on GitHub
Run locally with proto-tools
Toolkit contributors

Background

Many non-coding elements retain their structure or repeated organization despite substantial sequence divergence. Minerva reads out the coevolutionary patterns learned by a genome language model as separate maps of base pairing, protein contacts, and repeats. Each head combines attention features to predict its interaction type in a single forward pass, enabling rapid scans of genomic regions without constructing multiple sequence alignments. Speed makes genome-scale mining practical. In the Minerva study, the authors scanned 150 bacterial genomes in 100 minutes on one NVIDIA H100 GPU, plus 31 minutes to write the maps. That upstream workflow used overlapping 4096-token windows and the last-two-layer interaction heads; this tool exposes the inference step on prepared loci. The study used these maps to identify unannotated structured regions, extend the known TwoAYGGAY RNA architecture, and discover arrays of structurally conserved but sequence-diverse ncRNAs beside prophage UG27 reverse transcriptases. Experimental follow-up showed that the UG27 RNAs template short cDNA hairpins. These examples illustrate how interaction patterns can guide candidate selection, comparative analysis, and experimental discovery. See Li et al. (2026). Minerva-MLM was initialized from gLM2 and retains its mixed protein/DNA representation, with amino acids for coding regions, nucleotides for intergenic regions, and strand markers.

Tools

Minerva Interactions (minerva-interactions)

Predicts base-pairing, protein-contact, and repetitive-motif maps in a single model pass. Every selected head returns a labeled (L, L) probability matrix, with one result bundle per input locus.

API Reference

Source
List[string]
required
Uppercase proteins and lowercase DNA, optionally separated by +/- strand markers (<+>/<-> are stored as +/-). A single string is normalized to a list.
Source
List[string]
Selected named interaction heads.
enum
default:"2"
Number of final transformer layers used by the heads.Available options: 2, 6
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cuda"
Device used for model inference.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
enum
default:"gbrixi/minerva-mlm"
Public model checkpoint.Available options: gbrixi/minerva-mlm, gbrixi/minerva-mlm-8k
integer
default:"1"
Equal-token-length sequences per forward pass; one limits matrix memory.
Source
List[SequenceInteractions]
required
Interaction bundles in input order.

Applications

Screen intergenic regions for structured ncRNAs and repeat arrays, inspect their organization around nearby genes, and prioritize unannotated loci for comparative analysis or experimental follow-up. Combining the base-pairing and repeat maps can highlight arrays of structured elements, as in the study’s UG27 discovery. The protein map adds monomeric contact predictions for the coding parts of the same locus.

Usage Tips

  • heads selects the returned channels. All three are returned by default. Select fewer to reduce output size; this does not change the meaning of an individual channel.
  • interaction_layers=2 is the fast screening default. It uses the final two transformer layers, as in the paper’s genome scans. The released six-layer variant (interaction_layers=6) offers slightly higher accuracy at greater memory and compute cost.
  • The axes include strand markers and ambiguous context. Padding is removed. Use axis_labels (one per input character) to select biological subregions; matrix indices are not genome coordinates. The tool returns the upstream probabilities without symmetrizing, thresholding, removing the diagonal, or combining strands.
  • plot() draws the maps. result.plot() overlays base pairing, repeats, and protein contacts in the Minerva publication colors, with a track marking genes, intergenic DNA, and strand markers; plot(channel="base_pairing") shows one map with a colorbar, and window=(start, end) zooms to 1-indexed positions. It returns a matplotlib figure.
  • Dense output grows quadratically with token count. Doubling the locus length quadruples each matrix’s element count. The default export is compressed NPZ, with keys such as 0_protein and 0_protein_axis_labels; arrays load with numpy.load(..., allow_pickle=False). JSON is also supported, and the public Python values remain nested lists.

Minerva Embeddings (minerva-embedding)

Returns one mean-pooled embedding per prepared locus, with optional logits over the 24 canonical protein/DNA tokens.

API Reference

Source
List[string]
required
Uppercase proteins and lowercase DNA, optionally separated by +/- strand markers (<+>/<-> are stored as +/-). A single string is normalized to a list.
Source
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cuda"
Device used for model inference.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
enum
default:"gbrixi/minerva-mlm"
Public model checkpoint.Available options: gbrixi/minerva-mlm, gbrixi/minerva-mlm-8k
integer
default:"1"
Sequences per forward pass.
boolean
default:"False"
Include token-aligned biological logits.
integer
default:"-1"
0=embedding table, 1..N=transformer output, -1=last layer.
Source
List[MixedSequenceEmbedding]
required
Per-locus embedding bundles in input order.

Applications

Cluster or retrieve candidate loci after an interaction-map screen, or use the pooled vectors as features in a downstream classifier. Keep the checkpoint and representation layer fixed when comparing embeddings.

Usage Tips

  • repr_layer=-1 selects the last transformer output. Layer 0 selects the input embedding table; positive indices select transformer outputs. Pooling includes every unpadded token, including strand markers.
  • return_logits=True adds an (L, 24) matrix. Columns follow ACDEFGHIKLMNPQRSTVWYacgt; these are raw logits, not normalized probabilities. The matrix includes strand-marker rows, one row per input character.
  • CSV, NPY, and PT exports contain pooled vectors. JSON also preserves vocab and optional logits for position-level analysis; logits rows follow the input sequence.

Minerva Scoring (minerva-score)

Computes masked pseudo-log-likelihood by masking each canonical protein or DNA position in turn and predicting its original token from the remaining context. Returns summed and mean log-likelihood, perplexity, and the positions included in the score.

API Reference

Source
List[string]
required
Uppercase proteins and lowercase DNA, optionally separated by +/- strand markers (<+>/<-> are stored as +/-). A single string is normalized to a list.
Source
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cuda"
Device used for model inference.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
enum
default:"gbrixi/minerva-mlm"
Public model checkpoint.Available options: gbrixi/minerva-mlm, gbrixi/minerva-mlm-8k
integer
default:"1"
Individually masked variants per forward pass.
boolean
default:"False"
Include masked biological logits; unscored context rows are zero.
Source
List[MixedScoringMetrics]
required
Per-locus scores in input order.
Metrics (one set per scores item)

Applications

Rank related sequence variants with a contextual sequence prior. Scores reflect compatibility with the learned distribution and do not establish biological function or binding.

Usage Tips

  • scored_positions uses 1-indexed model-token positions. Strand markers and accepted ambiguous protein symbols provide context but are excluded from score targets and the mean’s denominator.
  • batch_size controls masked variants per forward pass. Scoring requires work proportional to the number of target positions; start with the default 1 when memory is limited.
  • Optional scoring logits come from masked passes. Their (L, 24) rows align with the input sequence; unscored marker and ambiguous-context rows are zero. These differ from the unmasked logits returned by embeddings.
  • PLL uses the full upstream vocabulary of 37 tokens. Exported logits contain only the 24 biological tokens, so applying softmax to those columns does not reproduce PLL.
  • Compare similar contexts. Higher mean log-likelihood and lower perplexity mean the model predicts the sequence more readily. Summed log-likelihood also depends on the number of scored positions.

Minerva Sampling (minerva-sample)

Refills selected protein and DNA positions while preserving their original modality, sequence length, strand markers, and noneditable ambiguous protein symbols. Each input produces one sampled locus.

API Reference

Source
array
One mapping per locus, assigning every mask’s 1-indexed token position to protein or DNA. Positions index the +/- form. Omit for fully specified sequences.
List[string]
required
Prepared loci, optionally carrying _ masks.
Source
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cuda"
Device used for model inference.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
enum
default:"gbrixi/minerva-mlm"
Public model checkpoint.Available options: gbrixi/minerva-mlm, gbrixi/minerva-mlm-8k
integer
default:"1"
Sequences per forward pass.
MaskingStrategy
Select editable token positions when no explicit masks are supplied.
enum
default:"single_pass"
Sampling algorithm.Available options: single_pass, iterative_refinement
number
default:"1.0"
Temperature within each editable position’s modality.
number
default:"1.0"
Nucleus threshold for iterative refinement.
integer
default:"20"
Number of iterative refinement rounds.
enum
default:"cosine"
Iterative unmask schedule.Available options: cosine, linear
enum
default:"random"
Random or confidence-based commitment selection.Available options: random, entropy
boolean
default:"True"
Cool temperature during iterative refinement.
boolean
default:"False"
Include biological logits from the completed sequence.
Source
List[MixedSequenceSample]
required
Per-locus sampling bundles in input order.

Applications

Propose local sequence edits for a design or variant-screening loop while keeping the prepared locus layout intact.

Usage Tips

  • Automatic masking uses MaskingStrategy. By default it selects about 30% of eligible canonical biological positions, with at least one when any are eligible. Set either num_mutations or mask_fraction; fixed_positions counts all model tokens, including markers. Entropy and max-logit selection use the same checkpoint as sampling.
  • Premasked inputs require mask_modalities. Supply one mapping per sequence that labels every _ position as "protein" or "dna". For example, +M_+a_ needs mask_modalities=[{3: "protein", 6: "dna"}]. Explicit masks bypass automatic selection for the batch.
  • The allowed alphabet is preserved per site. Protein positions sample among the 20 canonical amino acids; DNA positions sample among lowercase acgt. A selected position can draw its original token again, so the number of masks is not a guaranteed number of substitutions.
  • sampling_method selects the fill algorithm. single_pass fills all masks together. iterative_refinement progressively commits predictions and uses num_steps, schedule, strategy, top_p, and temperature_annealing. temperature controls sampling diversity and seed makes proposals reproducible.
  • Optional sampling logits describe the completed locus. They come from a final unmasked forward pass. Export prepared mixed strings as JSON or text; they are not nucleotide or protein FASTA records.

Minerva Gradient (minerva-gradient)

Returns the gradient of mean masked negative log-likelihood with respect to a relaxed (L, 24) sequence state. A fully specified sequence template fixes each position’s modality and the locations of strand markers and ambiguous protein context.

API Reference

Source
string
required
Fully specified template fixing modality and +/- strand markers.
List[array]
required
One 24-column row per model token in ACDEFGHIKLMNPQRSTVWYacgt order. Fixed context rows must be zero.
number
default:"1.0"
Softmax temperature; None requires probability distributions over the permitted modality and zero elsewhere.
Source
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cuda"
Device used for model inference.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
enum
default:"gbrixi/minerva-mlm"
Public model checkpoint.Available options: gbrixi/minerva-mlm, gbrixi/minerva-mlm-8k
integer
default:"1"
Individually masked variants per forward and backward pass.
boolean
default:"False"
Hard forward tokens with soft-probability gradients.
boolean
default:"True"
Compute the gradient, or evaluate only the masked PLL loss.
Source
array | array
Derivative of masked NLL with respect to the input state.
number
required
Mean masked negative log-likelihood.
Dict[string, any]
Masked PLL metrics and checkpoint metadata.
List[string]
required
Protein/DNA column order shared by input and gradient.

Applications

Use the masked-language-model objective as a differentiable prior in sequence optimization while retaining the original protein/DNA layout.

Usage Tips

  • The matrix axis includes every model token. Columns are ACDEFGHIKLMNPQRSTVWYacgt. Strand-marker and ambiguous-context rows must be zero; their returned gradients are also zero. Opposite-modality columns do not influence a position.
  • temperature=1.0 treats the input as logits. The worker applies a softmax within the position’s modality. With temperature=None, provide a nonnegative probability distribution within that modality and zeros elsewhere. one_hot_mixed_logits() builds a valid starting state from a template.
  • The objective masks canonical sites one at a time. The target at each site is its current discrete argmax token. use_ste=True uses hard forward tokens with soft derivatives; compute_gradient=False returns the objective with gradient=None.

Toolkit Notes

These apply to every Minerva tool in this toolkit.
  • Inputs are prepared, case-sensitive strings. Uppercase letters denote protein; lowercase acgt denotes DNA. The uppercase protein symbols X, B, U, Z, and O are accepted as fixed context. Lowercase ambiguous nucleotides, whitespace, literal <mask>, and other control tokens are rejected. Sampling accepts _ only with explicit modality metadata.
  • Strand markers are + (forward) and - (reverse). The upstream <+>/<-> spelling is also accepted and converted to +/-. Token positions are 1-indexed string indices of the +/- form, not genomic nucleotide coordinates. No case conversion, strand insertion, reverse complementation, translation, or gene calling is performed.
  • Prepare biologically appropriate orientation yourself. Upstream examples use + for intergenic DNA and strand markers for translated coding regions. The tool preserves the submitted orientation and does not combine strands.
  • Model setup and weights are managed automatically. The isolated environment is built on first use and public Hugging Face checkpoints are downloaded into the shared model cache. A Hugging Face token is not required for these public checkpoints.
  • Execution defaults to CUDA. Reduce batch_size when memory is limited. Repeated calls with one checkpoint can reuse a persistent worker through the standard ToolInstance.persist() API.
  • The default checkpoint is gbrixi/minerva-mlm. It accepts up to 4096 model tokens; gbrixi/minerva-mlm-8k accepts up to 8192. These counts include strand markers. Overlength inputs raise an error rather than truncating.
  • RNA base-pairing input uses the DNA alphabet. Submit lowercase acgt as in upstream examples; the tool does not convert u to t or call an RNA structure from the map.
  • Scope is prepared-string inference. Upstream gene calling, GenBank ingestion, fine-tuning, categorical Jacobians, and strand-ensemble workflows are not exposed by these tools.
Example notebook: See the full working example for a copy-paste-ready walkthrough.

Infrastructure Guides

The following guides cover how to run tools efficiently and at scale.

Tool Persistence

Keep a tool’s model warm across calls instead of reloading it every invocation.

Device Management

How GPUs are allocated to tools and how to target specific devices.

Parallel Execution

Fan a batch of inputs out across multiple GPUs.