
This toolkit is open source. Any third-party models, product names, or trademarks referenced are the property of their respective owners, and Proto is not affiliated with them.
Background
Many non-coding elements retain their structure or repeated organization despite substantial sequence divergence. Minerva reads out the coevolutionary patterns learned by a genome language model as separate maps of base pairing, protein contacts, and repeats. Each head combines attention features to predict its interaction type in a single forward pass, enabling rapid scans of genomic regions without constructing multiple sequence alignments.Tools
Minerva Interactions (minerva-interactions)
Predicts base-pairing, protein-contact, and repetitive-motif maps in a single model pass. Every selected head returns a labeled (L, L) probability matrix, with one result bundle per input locus.API Reference
Input: MixedSequenceInput
Input: MixedSequenceInput
+/- strand markers (<+>/<-> are stored as +/-). A single string is normalized to a list.Config: MinervaInteractionsConfig
Config: MinervaInteractionsConfig
2, 6True is coerced to 1 and False to 0.None waits indefinitely.BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.gbrixi/minerva-mlm, gbrixi/minerva-mlm-8kOutput: SequenceInteractionsOutput
Output: SequenceInteractionsOutput
Applications
Screen intergenic regions for structured ncRNAs and repeat arrays, inspect their organization around nearby genes, and prioritize unannotated loci for comparative analysis or experimental follow-up. Combining the base-pairing and repeat maps can highlight arrays of structured elements, as in the study’s UG27 discovery. The protein map adds monomeric contact predictions for the coding parts of the same locus.Usage Tips
headsselects the returned channels. All three are returned by default. Select fewer to reduce output size; this does not change the meaning of an individual channel.interaction_layers=2is the fast screening default. It uses the final two transformer layers, as in the paper’s genome scans. The released six-layer variant (interaction_layers=6) offers slightly higher accuracy at greater memory and compute cost.- The axes include strand markers and ambiguous context. Padding is removed. Use
axis_labels(one per input character) to select biological subregions; matrix indices are not genome coordinates. The tool returns the upstream probabilities without symmetrizing, thresholding, removing the diagonal, or combining strands. plot()draws the maps.result.plot()overlays base pairing, repeats, and protein contacts in the Minerva publication colors, with a track marking genes, intergenic DNA, and strand markers;plot(channel="base_pairing")shows one map with a colorbar, andwindow=(start, end)zooms to 1-indexed positions. It returns a matplotlib figure.- Dense output grows quadratically with token count. Doubling the locus length quadruples each matrix’s element count. The default export is compressed NPZ, with keys such as
0_proteinand0_protein_axis_labels; arrays load withnumpy.load(..., allow_pickle=False). JSON is also supported, and the public Python values remain nested lists.
Minerva Scoring (minerva-score)
Computes masked pseudo-log-likelihood by masking each canonical protein or DNA position in turn and predicting its original token from the remaining context. Returns summed and mean log-likelihood, perplexity, and the positions included in the score.API Reference
Input: MixedSequenceInput
Input: MixedSequenceInput
+/- strand markers (<+>/<-> are stored as +/-). A single string is normalized to a list.Config: MinervaScoringConfig
Config: MinervaScoringConfig
True is coerced to 1 and False to 0.None waits indefinitely.BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.gbrixi/minerva-mlm, gbrixi/minerva-mlm-8kOutput: MixedScoringOutput
Output: MixedScoringOutput
Applications
Rank related sequence variants with a contextual sequence prior. Scores reflect compatibility with the learned distribution and do not establish biological function or binding.Usage Tips
scored_positionsuses 1-indexed model-token positions. Strand markers and accepted ambiguous protein symbols provide context but are excluded from score targets and the mean’s denominator.batch_sizecontrols masked variants per forward pass. Scoring requires work proportional to the number of target positions; start with the default 1 when memory is limited.- Optional scoring logits come from masked passes. Their
(L, 24)rows align with the input sequence; unscored marker and ambiguous-context rows are zero. These differ from the unmasked logits returned by embeddings. - PLL uses the full upstream vocabulary of 37 tokens. Exported logits contain only the 24 biological tokens, so applying softmax to those columns does not reproduce PLL.
- Compare similar contexts. Higher mean log-likelihood and lower perplexity mean the model predicts the sequence more readily. Summed log-likelihood also depends on the number of scored positions.
Minerva Sampling (minerva-sample)
Refills selected protein and DNA positions while preserving their original modality, sequence length, strand markers, and noneditable ambiguous protein symbols. Each input produces one sampled locus.API Reference
Config: MinervaSampleConfig
Config: MinervaSampleConfig
True is coerced to 1 and False to 0.None waits indefinitely.BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.gbrixi/minerva-mlm, gbrixi/minerva-mlm-8ksingle_pass, iterative_refinementcosine, linearrandom, entropyOutput: MixedSampleOutput
Output: MixedSampleOutput
Applications
Propose local sequence edits for a design or variant-screening loop while keeping the prepared locus layout intact.Usage Tips
- Automatic masking uses
MaskingStrategy. By default it selects about 30% of eligible canonical biological positions, with at least one when any are eligible. Set eithernum_mutationsormask_fraction;fixed_positionscounts all model tokens, including markers. Entropy and max-logit selection use the same checkpoint as sampling. - Premasked inputs require
mask_modalities. Supply one mapping per sequence that labels every_position as"protein"or"dna". For example,+M_+a_needsmask_modalities=[{3: "protein", 6: "dna"}]. Explicit masks bypass automatic selection for the batch. - The allowed alphabet is preserved per site. Protein positions sample among the 20 canonical amino acids; DNA positions sample among lowercase
acgt. A selected position can draw its original token again, so the number of masks is not a guaranteed number of substitutions. sampling_methodselects the fill algorithm.single_passfills all masks together.iterative_refinementprogressively commits predictions and usesnum_steps,schedule,strategy,top_p, andtemperature_annealing.temperaturecontrols sampling diversity andseedmakes proposals reproducible.- Optional sampling logits describe the completed locus. They come from a final unmasked forward pass. Export prepared mixed strings as JSON or text; they are not nucleotide or protein FASTA records.
Minerva Gradient (minerva-gradient)
Returns the gradient of mean masked negative log-likelihood with respect to a relaxed (L, 24) sequence state. A fully specified sequence template fixes each position’s modality and the locations of strand markers and ambiguous protein context.API Reference
Input: MixedSequenceGradientInput
Input: MixedSequenceGradientInput
+/- strand markers.ACDEFGHIKLMNPQRSTVWYacgt order. Fixed context rows must be zero.None requires probability distributions over the permitted modality and zero elsewhere.Config: MinervaGradientConfig
Config: MinervaGradientConfig
True is coerced to 1 and False to 0.None waits indefinitely.BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.gbrixi/minerva-mlm, gbrixi/minerva-mlm-8kApplications
Use the masked-language-model objective as a differentiable prior in sequence optimization while retaining the original protein/DNA layout.Usage Tips
- The matrix axis includes every model token. Columns are
ACDEFGHIKLMNPQRSTVWYacgt. Strand-marker and ambiguous-context rows must be zero; their returned gradients are also zero. Opposite-modality columns do not influence a position. temperature=1.0treats the input as logits. The worker applies a softmax within the position’s modality. Withtemperature=None, provide a nonnegative probability distribution within that modality and zeros elsewhere.one_hot_mixed_logits()builds a valid starting state from a template.- The objective masks canonical sites one at a time. The target at each site is its current discrete argmax token.
use_ste=Trueuses hard forward tokens with soft derivatives;compute_gradient=Falsereturns the objective withgradient=None.
Toolkit Notes
These apply to every Minerva tool in this toolkit.- Inputs are prepared, case-sensitive strings. Uppercase letters denote protein; lowercase
acgtdenotes DNA. The uppercase protein symbolsX,B,U,Z, andOare accepted as fixed context. Lowercase ambiguous nucleotides, whitespace, literal<mask>, and other control tokens are rejected. Sampling accepts_only with explicit modality metadata. - Strand markers are
+(forward) and-(reverse). The upstream<+>/<->spelling is also accepted and converted to+/-. Token positions are 1-indexed string indices of the+/-form, not genomic nucleotide coordinates. No case conversion, strand insertion, reverse complementation, translation, or gene calling is performed. - Prepare biologically appropriate orientation yourself. Upstream examples use
+for intergenic DNA and strand markers for translated coding regions. The tool preserves the submitted orientation and does not combine strands. - Model setup and weights are managed automatically. The isolated environment is built on first use and public Hugging Face checkpoints are downloaded into the shared model cache. A Hugging Face token is not required for these public checkpoints.
- Execution defaults to CUDA. Reduce
batch_sizewhen memory is limited. Repeated calls with one checkpoint can reuse a persistent worker through the standardToolInstance.persist()API. - The default checkpoint is
gbrixi/minerva-mlm. It accepts up to 4096 model tokens;gbrixi/minerva-mlm-8kaccepts up to 8192. These counts include strand markers. Overlength inputs raise an error rather than truncating. - RNA base-pairing input uses the DNA alphabet. Submit lowercase
acgtas in upstream examples; the tool does not convertutotor call an RNA structure from the map. - Scope is prepared-string inference. Upstream gene calling, GenBank ingestion, fine-tuning, categorical Jacobians, and strand-ensemble workflows are not exposed by these tools.