Proto is not affiliated with Tatta Bio. This toolkit is open source and builds on the implementation produced by this organization. Product names, logos, and trademarks are the property of their respective owners.
Background
gLM2 was introduced with the Open MetaGenomic (OMG) corpus (Cornman et al., 2024). Its bidirectional transformer predicts masked tokens from genomic context: coding regions are represented by amino acids and intergenic regions by individual nucleotides. Uppercase protein and lowercase DNA alphabets avoid token collisions;+ and - represent strand orientation. The public 150M and 650M checkpoints both support 4096 tokens.
This representation lets protein tokens condition on neighboring DNA and proteins without treating a DNA base and an identically named amino acid as the same symbol. The tool operates directly on that prepared representation.
Tools
gLM2 Scoring (glm2-score)
Computes masked pseudo-log-likelihood by masking each canonical protein or DNA position in turn and predicting its original token from the remaining context. Returns summed and mean log-likelihood, perplexity, and the positions included in the score.API Reference
Input: MixedSequenceInput
Input: MixedSequenceInput
+/- strand markers (<+>/<-> are stored as +/-). A single string is normalized to a list.Config: GLM2ScoringConfig
Config: GLM2ScoringConfig
True is coerced to 1 and False to 0.None waits indefinitely.BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.tattabio/gLM2_150M, tattabio/gLM2_650MOutput: MixedScoringOutput
Output: MixedScoringOutput
Applications
Rank related sequence variants with a contextual sequence prior. Scores reflect compatibility with the learned distribution and do not establish biological function or binding.Usage Tips
scored_positionsuses 1-indexed model-token positions. Strand markers and accepted ambiguous protein symbols provide context but are excluded from score targets and the mean’s denominator.batch_sizecontrols masked variants per forward pass. Scoring requires work proportional to the number of target positions; start with the default 1 when memory is limited.- Optional scoring logits come from masked passes. Their
(L, 24)rows align with the input sequence; unscored marker and ambiguous-context rows are zero. These differ from the unmasked logits returned by embeddings. - PLL uses the full upstream vocabulary of 37 tokens. Exported logits contain only the 24 biological tokens, so applying softmax to those columns does not reproduce PLL.
- Compare similar contexts. Higher mean log-likelihood and lower perplexity mean the model predicts the sequence more readily. Summed log-likelihood also depends on the number of scored positions.
gLM2 Sampling (glm2-sample)
Refills selected protein and DNA positions while preserving their original modality, sequence length, strand markers, and noneditable ambiguous protein symbols. Each input produces one sampled locus.API Reference
Config: GLM2SampleConfig
Config: GLM2SampleConfig
True is coerced to 1 and False to 0.None waits indefinitely.BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.tattabio/gLM2_150M, tattabio/gLM2_650Msingle_pass, iterative_refinementcosine, linearrandom, entropyOutput: MixedSampleOutput
Output: MixedSampleOutput
Applications
Propose local sequence edits for a design or variant-screening loop while keeping the prepared locus layout intact.Usage Tips
- Automatic masking uses
MaskingStrategy. By default it selects about 30% of eligible canonical biological positions, with at least one when any are eligible. Set eithernum_mutationsormask_fraction;fixed_positionscounts all model tokens, including markers. Entropy and max-logit selection use the same checkpoint as sampling. - Premasked inputs require
mask_modalities. Supply one mapping per sequence that labels every_position as"protein"or"dna". For example,+M_+a_needsmask_modalities=[{3: "protein", 6: "dna"}]. Explicit masks bypass automatic selection for the batch. - The allowed alphabet is preserved per site. Protein positions sample among the 20 canonical amino acids; DNA positions sample among lowercase
acgt. A selected position can draw its original token again, so the number of masks is not a guaranteed number of substitutions. sampling_methodselects the fill algorithm.single_passfills all masks together.iterative_refinementprogressively commits predictions and usesnum_steps,schedule,strategy,top_p, andtemperature_annealing.temperaturecontrols sampling diversity andseedmakes proposals reproducible.- Optional sampling logits describe the completed locus. They come from a final unmasked forward pass. Export prepared mixed strings as JSON or text; they are not nucleotide or protein FASTA records.
gLM2 Gradient (glm2-gradient)
Returns the gradient of mean masked negative log-likelihood with respect to a relaxed (L, 24) sequence state. A fully specified sequence template fixes each position’s modality and the locations of strand markers and ambiguous protein context.API Reference
Input: MixedSequenceGradientInput
Input: MixedSequenceGradientInput
+/- strand markers.ACDEFGHIKLMNPQRSTVWYacgt order. Fixed context rows must be zero.None requires probability distributions over the permitted modality and zero elsewhere.Config: GLM2GradientConfig
Config: GLM2GradientConfig
True is coerced to 1 and False to 0.None waits indefinitely.BaseToolOutput.approx_equal), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.tattabio/gLM2_150M, tattabio/gLM2_650MApplications
Use the masked-language-model objective as a differentiable prior in sequence optimization while retaining the original protein/DNA layout.Usage Tips
- The matrix axis includes every model token. Columns are
ACDEFGHIKLMNPQRSTVWYacgt. Strand-marker and ambiguous-context rows must be zero; their returned gradients are also zero. Opposite-modality columns do not influence a position. temperature=1.0treats the input as logits. The worker applies a softmax within the position’s modality. Withtemperature=None, provide a nonnegative probability distribution within that modality and zeros elsewhere.one_hot_mixed_logits()builds a valid starting state from a template.- The objective masks canonical sites one at a time. The target at each site is its current discrete argmax token.
use_ste=Trueuses hard forward tokens with soft derivatives;compute_gradient=Falsereturns the objective withgradient=None.
Toolkit Notes
These apply to every gLM2 tool in this toolkit.- Inputs are prepared, case-sensitive strings. Uppercase letters denote protein; lowercase
acgtdenotes DNA. The uppercase protein symbolsX,B,U,Z, andOare accepted as fixed context. Lowercase ambiguous nucleotides, whitespace, literal<mask>, and other control tokens are rejected. Sampling accepts_only with explicit modality metadata. - Strand markers are
+(forward) and-(reverse). The upstream<+>/<->spelling is also accepted and converted to+/-. Token positions are 1-indexed string indices of the+/-form, not genomic nucleotide coordinates. No case conversion, strand insertion, reverse complementation, translation, or gene calling is performed. - Prepare biologically appropriate orientation yourself. Upstream examples use
+for intergenic DNA and strand markers for translated coding regions. The tool preserves the submitted orientation and does not combine strands. - Model setup and weights are managed automatically. The isolated environment is built on first use and public Hugging Face checkpoints are downloaded into the shared model cache. A Hugging Face token is not required for these public checkpoints.
- Execution defaults to CUDA. Reduce
batch_sizewhen memory is limited. Repeated calls with one checkpoint can reuse a persistent worker through the standardToolInstance.persist()API. - The default checkpoint is
tattabio/gLM2_650M. Selecttattabio/gLM2_150Mfor the smaller model. Both accept at most 4096 model tokens, including strand markers; longer inputs raise an error rather than truncating. - This toolkit exposes MLM operations. Upstream categorical-Jacobian interaction analysis is not part of the tool.

Tatta Bio