Skip to main content
License: gLM2 is open source and free for academic and commercial use under an Apache-2.0 license. Please refer to the license for full terms.

Proto is not affiliated with Tatta Bio. This toolkit is open source and builds on the implementation produced by this organization. Product names, logos, and trademarks are the property of their respective owners.


TattaBio/gLM2
TattaBio/gLM2
97 stars
View repo
The OMG dataset: An Open MetaGenomic corpus for mixed-modality genomic language modeling
Andre Cornman, Jacob West-Roberts, … Yunha Hwang
bioRxiv (2024)
Read preprint
Copy citation
evo-design/proto-tools/proto_tools/tools/masked_models/glm2
View source
Open Notebook
Open notebook
proto-tools on GitHub
Run locally with proto-tools
Toolkit contributors

Background

gLM2 was introduced with the Open MetaGenomic (OMG) corpus (Cornman et al., 2024). Its bidirectional transformer predicts masked tokens from genomic context: coding regions are represented by amino acids and intergenic regions by individual nucleotides. Uppercase protein and lowercase DNA alphabets avoid token collisions; + and - represent strand orientation. The public 150M and 650M checkpoints both support 4096 tokens. This representation lets protein tokens condition on neighboring DNA and proteins without treating a DNA base and an identically named amino acid as the same symbol. The tool operates directly on that prepared representation.

Tools

gLM2 Embeddings (glm2-embedding)

Returns one mean-pooled embedding per prepared locus, with optional logits over the 24 canonical protein/DNA tokens.

API Reference

Source
List[string]
required
Uppercase proteins and lowercase DNA, optionally separated by +/- strand markers (<+>/<-> are stored as +/-). A single string is normalized to a list.
Source
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cuda"
Device used for model inference.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
enum
default:"tattabio/gLM2_650M"
Public model checkpoint.Available options: tattabio/gLM2_150M, tattabio/gLM2_650M
integer
default:"1"
Sequences per forward pass.
boolean
default:"False"
Include token-aligned biological logits.
integer
default:"-1"
0=embedding table, 1..N=transformer output, -1=last layer.
Source
List[MixedSequenceEmbedding]
required
Per-locus embedding bundles in input order.

Applications

Use the pooled vector as a feature for clustering, retrieval, or a downstream supervised model. Keep the checkpoint and representation layer fixed when comparing embeddings.

Usage Tips

  • repr_layer=-1 selects the last transformer output. Layer 0 selects the input embedding table; positive indices select transformer outputs. Pooling includes every unpadded token, including strand markers.
  • return_logits=True adds an (L, 24) matrix. Columns follow ACDEFGHIKLMNPQRSTVWYacgt; these are raw logits, not normalized probabilities. The matrix includes strand-marker rows, one row per input character.
  • CSV, NPY, and PT exports contain pooled vectors. JSON also preserves vocab and optional logits for position-level analysis; logits rows follow the input sequence.

gLM2 Scoring (glm2-score)

Computes masked pseudo-log-likelihood by masking each canonical protein or DNA position in turn and predicting its original token from the remaining context. Returns summed and mean log-likelihood, perplexity, and the positions included in the score.

API Reference

Source
List[string]
required
Uppercase proteins and lowercase DNA, optionally separated by +/- strand markers (<+>/<-> are stored as +/-). A single string is normalized to a list.
Source
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cuda"
Device used for model inference.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
enum
default:"tattabio/gLM2_650M"
Public model checkpoint.Available options: tattabio/gLM2_150M, tattabio/gLM2_650M
integer
default:"1"
Individually masked variants per forward pass.
boolean
default:"False"
Include masked biological logits; unscored context rows are zero.
Source
List[MixedScoringMetrics]
required
Per-locus scores in input order.
Metrics (one set per scores item)

Applications

Rank related sequence variants with a contextual sequence prior. Scores reflect compatibility with the learned distribution and do not establish biological function or binding.

Usage Tips

  • scored_positions uses 1-indexed model-token positions. Strand markers and accepted ambiguous protein symbols provide context but are excluded from score targets and the mean’s denominator.
  • batch_size controls masked variants per forward pass. Scoring requires work proportional to the number of target positions; start with the default 1 when memory is limited.
  • Optional scoring logits come from masked passes. Their (L, 24) rows align with the input sequence; unscored marker and ambiguous-context rows are zero. These differ from the unmasked logits returned by embeddings.
  • PLL uses the full upstream vocabulary of 37 tokens. Exported logits contain only the 24 biological tokens, so applying softmax to those columns does not reproduce PLL.
  • Compare similar contexts. Higher mean log-likelihood and lower perplexity mean the model predicts the sequence more readily. Summed log-likelihood also depends on the number of scored positions.

gLM2 Sampling (glm2-sample)

Refills selected protein and DNA positions while preserving their original modality, sequence length, strand markers, and noneditable ambiguous protein symbols. Each input produces one sampled locus.

API Reference

Source
array
One mapping per locus, assigning every mask’s 1-indexed token position to protein or DNA. Positions index the +/- form. Omit for fully specified sequences.
List[string]
required
Prepared loci, optionally carrying _ masks.
Source
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cuda"
Device used for model inference.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
enum
default:"tattabio/gLM2_650M"
Public model checkpoint.Available options: tattabio/gLM2_150M, tattabio/gLM2_650M
integer
default:"1"
Sequences per forward pass.
MaskingStrategy
Select editable token positions when no explicit masks are supplied.
enum
default:"single_pass"
Sampling algorithm.Available options: single_pass, iterative_refinement
number
default:"1.0"
Temperature within each editable position’s modality.
number
default:"1.0"
Nucleus threshold for iterative refinement.
integer
default:"20"
Number of iterative refinement rounds.
enum
default:"cosine"
Iterative unmask schedule.Available options: cosine, linear
enum
default:"random"
Random or confidence-based commitment selection.Available options: random, entropy
boolean
default:"True"
Cool temperature during iterative refinement.
boolean
default:"False"
Include biological logits from the completed sequence.
Source
List[MixedSequenceSample]
required
Per-locus sampling bundles in input order.

Applications

Propose local sequence edits for a design or variant-screening loop while keeping the prepared locus layout intact.

Usage Tips

  • Automatic masking uses MaskingStrategy. By default it selects about 30% of eligible canonical biological positions, with at least one when any are eligible. Set either num_mutations or mask_fraction; fixed_positions counts all model tokens, including markers. Entropy and max-logit selection use the same checkpoint as sampling.
  • Premasked inputs require mask_modalities. Supply one mapping per sequence that labels every _ position as "protein" or "dna". For example, +M_+a_ needs mask_modalities=[{3: "protein", 6: "dna"}]. Explicit masks bypass automatic selection for the batch.
  • The allowed alphabet is preserved per site. Protein positions sample among the 20 canonical amino acids; DNA positions sample among lowercase acgt. A selected position can draw its original token again, so the number of masks is not a guaranteed number of substitutions.
  • sampling_method selects the fill algorithm. single_pass fills all masks together. iterative_refinement progressively commits predictions and uses num_steps, schedule, strategy, top_p, and temperature_annealing. temperature controls sampling diversity and seed makes proposals reproducible.
  • Optional sampling logits describe the completed locus. They come from a final unmasked forward pass. Export prepared mixed strings as JSON or text; they are not nucleotide or protein FASTA records.

gLM2 Gradient (glm2-gradient)

Returns the gradient of mean masked negative log-likelihood with respect to a relaxed (L, 24) sequence state. A fully specified sequence template fixes each position’s modality and the locations of strand markers and ambiguous protein context.

API Reference

Source
string
required
Fully specified template fixing modality and +/- strand markers.
List[array]
required
One 24-column row per model token in ACDEFGHIKLMNPQRSTVWYacgt order. Fixed context rows must be zero.
number
default:"1.0"
Softmax temperature; None requires probability distributions over the permitted modality and zero elsewhere.
Source
integer
default:"0"
Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). True is coerced to 1 and False to 0.
string
default:"cuda"
Device used for model inference.
integer
default:"3600"
Maximum execution time in seconds. None waits indefinitely.
integer
Random seed. When set, tools run reproducibly up to small GPU float noise (see BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
enum
default:"tattabio/gLM2_650M"
Public model checkpoint.Available options: tattabio/gLM2_150M, tattabio/gLM2_650M
integer
default:"1"
Individually masked variants per forward and backward pass.
boolean
default:"False"
Hard forward tokens with soft-probability gradients.
boolean
default:"True"
Compute the gradient, or evaluate only the masked PLL loss.
Source
array | array
Derivative of masked NLL with respect to the input state.
number
required
Mean masked negative log-likelihood.
Dict[string, any]
Masked PLL metrics and checkpoint metadata.
List[string]
required
Protein/DNA column order shared by input and gradient.

Applications

Use the masked-language-model objective as a differentiable prior in sequence optimization while retaining the original protein/DNA layout.

Usage Tips

  • The matrix axis includes every model token. Columns are ACDEFGHIKLMNPQRSTVWYacgt. Strand-marker and ambiguous-context rows must be zero; their returned gradients are also zero. Opposite-modality columns do not influence a position.
  • temperature=1.0 treats the input as logits. The worker applies a softmax within the position’s modality. With temperature=None, provide a nonnegative probability distribution within that modality and zeros elsewhere. one_hot_mixed_logits() builds a valid starting state from a template.
  • The objective masks canonical sites one at a time. The target at each site is its current discrete argmax token. use_ste=True uses hard forward tokens with soft derivatives; compute_gradient=False returns the objective with gradient=None.

Toolkit Notes

These apply to every gLM2 tool in this toolkit.
  • Inputs are prepared, case-sensitive strings. Uppercase letters denote protein; lowercase acgt denotes DNA. The uppercase protein symbols X, B, U, Z, and O are accepted as fixed context. Lowercase ambiguous nucleotides, whitespace, literal <mask>, and other control tokens are rejected. Sampling accepts _ only with explicit modality metadata.
  • Strand markers are + (forward) and - (reverse). The upstream <+>/<-> spelling is also accepted and converted to +/-. Token positions are 1-indexed string indices of the +/- form, not genomic nucleotide coordinates. No case conversion, strand insertion, reverse complementation, translation, or gene calling is performed.
  • Prepare biologically appropriate orientation yourself. Upstream examples use + for intergenic DNA and strand markers for translated coding regions. The tool preserves the submitted orientation and does not combine strands.
  • Model setup and weights are managed automatically. The isolated environment is built on first use and public Hugging Face checkpoints are downloaded into the shared model cache. A Hugging Face token is not required for these public checkpoints.
  • Execution defaults to CUDA. Reduce batch_size when memory is limited. Repeated calls with one checkpoint can reuse a persistent worker through the standard ToolInstance.persist() API.
  • The default checkpoint is tattabio/gLM2_650M. Select tattabio/gLM2_150M for the smaller model. Both accept at most 4096 model tokens, including strand markers; longer inputs raise an error rather than truncating.
  • This toolkit exposes MLM operations. Upstream categorical-Jacobian interaction analysis is not part of the tool.
Example notebook: See the full working example for a copy-paste-ready walkthrough.

Infrastructure Guides

The following guides cover how to run tools efficiently and at scale.

Tool Persistence

Keep a tool’s model warm across calls instead of reloading it every invocation.

Device Management

How GPUs are allocated to tools and how to target specific devices.

Parallel Execution

Fan a batch of inputs out across multiple GPUs.