> ## Documentation Index
> Fetch the complete documentation index at: https://proto.evodesign.org/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Gene/Protein Similarity

> Score percent identity via MMseqs2 (DNA is ORF-predicted first; proteins search directly).

<div class="page-hero">
  <img class="page-hero-banner" src="https://proto-bio.github.io/proto-assets/images/constraint/mmseqs-gene-similarity/hero.png" alt="Gene/Protein Similarity" />
</div>

<Note>
  **License:** This constraint can use multiple tools, each under its own license. See the **Tools Used** tab and each tool's page for license details.
</Note>

<p class="entity-disclaimer">This constraint is open source. Any third-party models, product names, or trademarks referenced are the property of their respective owners, and Proto is not affiliated with them.</p>

<hr class="entity-rule" />

<input type="radio" name="tab-constraint-mmseqs-gene-similarity" id="none-constraint-mmseqs-gene-similarity" class="tab-radio-input" />

<input type="radio" name="tab-constraint-mmseqs-gene-similarity" id="tools-constraint-mmseqs-gene-similarity" class="tab-radio-input" defaultChecked />

<input type="radio" name="tab-constraint-mmseqs-gene-similarity" id="source-constraint-mmseqs-gene-similarity" class="tab-radio-input" />

<div class="tool-tab-bar"><span class="tool-tab-wrap"><label for="tools-constraint-mmseqs-gene-similarity" class="tool-tab tab-open badge-tools"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect width="7" height="7" x="3" y="3" rx="1" /><rect width="7" height="7" x="14" y="3" rx="1" /><rect width="7" height="7" x="14" y="14" rx="1" /><rect width="7" height="7" x="3" y="14" rx="1" /></svg> Tools Used</label><label for="none-constraint-mmseqs-gene-similarity" class="tool-tab tab-close badge-tools"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect width="7" height="7" x="3" y="3" rx="1" /><rect width="7" height="7" x="14" y="3" rx="1" /><rect width="7" height="7" x="14" y="14" rx="1" /><rect width="7" height="7" x="3" y="14" rx="1" /></svg> Tools Used</label></span> <span class="tool-tab-wrap"><label for="source-constraint-mmseqs-gene-similarity" class="tool-tab tab-open badge-source"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</label><label for="none-constraint-mmseqs-gene-similarity" class="tool-tab tab-close badge-source"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</label></span></div>

<div class="tab-panel tools-panel" data-tab="tools-constraint-mmseqs-gene-similarity">
  <div class="tools-used-grid"><a href="/docs/tools/sequence-alignment/mmseqs2" class="tools-used-tile">  <img noZoom src="https://proto-bio.github.io/proto-assets/images/tool/mmseqs2/social.png" alt="" loading="lazy" /></a><a href="/docs/tools/orf-prediction/overview" class="tools-used-tile tools-used-cat-collage"><div class="tool-catalog-tile-collage"><div class="tool-catalog-tile-cell">  <img noZoom src="https://proto-bio.github.io/proto-assets/images/tool/prodigal/carousel.png" alt="" loading="lazy" /></div><div class="tool-catalog-tile-cell">  <img noZoom src="https://proto-bio.github.io/proto-assets/images/tool/orfipy/carousel.png" alt="" loading="lazy" /></div><div class="tool-catalog-tile-cell tools-cat-cell-blank" /><div class="tool-catalog-tile-cell tools-cat-cell-blank" /></div><div class="tool-catalog-tile-label"><span>ORF Prediction · 2 tools</span></div></a></div>
</div>

<a href="https://github.com/evo-design/proto-language/blob/d3b7822f74ea64747cc751a3b2ab1aa6b799ac47/proto_language/constraint/sequence_annotation/mmseqs_similarity_constraint.py#L150" target="_blank" class="tab-panel source-panel" data-tab="source-constraint-mmseqs-gene-similarity">
  <div class="source-info">
    <img noZoom src="https://github.com/evo-design.png?size=40" class="source-avatar" width="36" height="36" />

    <span class="source-path">evo-design/proto-language<span class="source-subpath">/proto\_language/constraint/sequence\_annotation/mmseqs\_similarity\_constraint.py</span></span>
  </div>

  <span class="panel-goto-btn source-goto-btn"><span><svg width="14" height="14" viewBox="0 0 24 24" fill="currentColor"><path d="M12 0C5.37 0 0 5.37 0 12c0 5.31 3.435 9.795 8.205 11.385.6.105.825-.255.825-.57 0-.285-.015-1.23-.015-2.235-3.015.555-3.795-.735-4.035-1.41-.135-.345-.72-1.41-1.23-1.695-.42-.225-1.02-.78-.015-.795.945-.015 1.62.87 1.845 1.23 1.08 1.815 2.805 1.305 3.495.99.105-.78.42-1.305.765-1.605-2.67-.3-5.46-1.335-5.46-5.925 0-1.305.465-2.385 1.23-3.225-.12-.3-.54-1.53.12-3.18 0 0 1.005-.315 3.3 1.23.96-.27 1.98-.405 3-.405s2.04.135 3 .405c2.295-1.56 3.3-1.23 3.3-1.23.66 1.65.24 2.88.12 3.18.765.84 1.23 1.905 1.23 3.225 0 4.605-2.805 5.625-5.475 5.925.435.375.81 1.095.81 2.22 0 1.605-.015 2.895-.015 3.3 0 .315.225.69.825.57A12.02 12.02 0 0024 12c0-6.63-5.37-12-12-12z" /></svg> View source</span></span>
</a>

<div class="entity-contributors"><span class="entity-contributors-label">Constraint contributors</span><span class="entity-contributors-people"><a class="entity-contributor" href="https://github.com/dguo8412" target="_blank" rel="noopener" title="dguo8412: 6 commits"><img noZoom class="entity-contributor-avatar" src="https://avatars.githubusercontent.com/u/46211285?v=4&s=64" alt="" loading="lazy" /><span class="entity-contributor-login">dguo8412</span></a><a class="entity-contributor" href="https://github.com/bviggiano" target="_blank" rel="noopener" title="bviggiano: 2 commits"><img noZoom class="entity-contributor-avatar" src="https://avatars.githubusercontent.com/u/21143637?v=4&s=64" alt="" loading="lazy" /><span class="entity-contributor-login">bviggiano</span></a></span></div>
Evaluate sequence similarity using MMseqs2 protein database search.

This constraint function evaluates whether protein sequences (or proteins
predicted from DNA sequences) have percent identity to known proteins within
an acceptable range. It uses MMseqs2, an ultra-fast sequence search tool,
to search against a reference protein database and calculates similarity scores.

For DNA sequences, the function first predicts open reading frames (ORFs)
using either Prodigal (for prokaryotes) or ORFipy (viral), then
searches the translated proteins. For protein sequences, the search is
performed directly. The constraint is satisfied when all database hits have
percent identity within the specified \[min\_similarity, max\_similarity] range.

## API Reference

<div class="api-model-section api-model-static api-config-section">
  <div class="api-model-header"><span class="api-model-badge api-config-badge">Config</span><span class="api-model-name">MMseqsSimilarityConfig</span><a href="https://github.com/evo-design/proto-language/blob/d3b7822f74ea64747cc751a3b2ab1aa6b799ac47/proto_language/constraint/sequence_annotation/mmseqs_similarity_constraint.py#L29" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a></div>

  Configuration for MMseqs gene similarity constraint.

  This class defines configuration parameters for evaluating sequence similarity
  (percent identity) to known proteins using MMseqs2, an ultra-fast sequence
  search tool. For DNA sequences, the constraint first predicts open reading
  frames (ORFs) using either Prodigal or ORFipy, then searches the translated
  proteins against a reference database. For protein sequences, the search is
  performed directly.

  <Note>
    For examples with tool configuration, see:

    > > > from proto\_tools import Mmseqs2SearchProteinsConfig
    > > > The similarity range \[min\_similarity, max\_similarity] defines acceptable percent
    > > > identity. Sequences with hits outside this range are penalized. For example:

    * \[40, 70]: Moderate similarity, useful for inferring functional similarity while
      avoiding identical sequences
    * \[0, 40]: Low similarity filter, for novelty/uniqueness constraints
    * \[80, 100]: High similarity filter, for functional conservation requirements
  </Note>

  <ParamField path="min_similarity" type="number" required>
    Minimum acceptable percent identity (0-100). Lower values are more permissive.
  </ParamField>

  <ParamField path="max_similarity" type="number" required>
    Maximum acceptable percent identity (0-100). Higher values allow more similar hits.
  </ParamField>

  <ParamField path="mmseqs_db" type="string" required>
    Path to MMseqs2 protein database for similarity search
  </ParamField>

  <ParamField path="mmseqs_config" type="Mmseqs2SearchProteinsConfig">
    MMseqs configuration (threads, sensitivity, etc.).
  </ParamField>

  <ParamField path="orf_predictor" type="enum" default="prodigal">
    ORF prediction tool (DNA only): 'orfipy' (viral) or 'prodigal' (prokaryotic).

    Options: `orfipy`, `prodigal`
  </ParamField>

  <ParamField path="orfipy_config" type="OrfipyConfig">
    ORFipy configuration (DNA only, used if orf\_predictor='orfipy').
  </ParamField>

  <ParamField path="prodigal_config" type="ProdigalConfig">
    Prodigal configuration (DNA only, used if orf\_predictor='prodigal').
  </ParamField>
</div>

<div class="api-model-section api-model-static api-output-section">
  <div class="api-model-header"><span class="api-model-badge api-output-badge">Returns</span><span class="api-model-name">ConstraintOutput</span></div>

  One result per sequence. Score 0.0 means all hits
  fall within \[min\_similarity, max\_similarity]; higher scores indicate
  greater deviation. Score 1.0 (MAX\_ENERGY) is returned if no ORFs are
  found (DNA) or no database hits are found. `metadata` carries:

  **For DNA sequences (with Prodigal):**

  * `prodigal_orfs`: List of dictionaries containing predicted ORF information
    (id, start, end, strand, protein\_sequence, etc.)
  * `mmseqs_results`: List of dictionaries with MMseqs2 hit information
    (target\_id, pident, evalue)
  * `unique_orfs_with_hits`: Integer count of distinct ORFs with at least
    one database match
  * `orfs_with_acceptable_similarity`: Integer count of ORFs with hits in
    acceptable range
  * `total_orfs_with_hits`: Integer total number of ORF-hit pairs
  * `similarity_compliance_rate`: Float fraction of hits within acceptable
    range (0.0-1.0)

  **For DNA sequences (with ORFipy):**

  * `orfipy_orfs`: List of dictionaries with ORFipy ORF predictions
  * Other fields same as Prodigal above

  **For protein sequences:**

  * `direct_protein`: Dictionary with protein information (id, sequence, length)
  * `mmseqs_results`: List of MMseqs2 hit dictionaries
  * `unique_orfs_with_hits`: Count of distinct ORFs with at least one hit
    (always 1 or 0 for a single protein, which has exactly one ORF)
  * `orfs_with_acceptable_similarity`: Count of acceptable hits
  * `total_orfs_with_hits`: Total hit count
  * `similarity_compliance_rate`: Fraction of hits in range
</div>

## Usage

Filtering for sequences with low similarity to existing proteins:

```python python icon="python" theme={null}
>>> from proto_language.core import Sequence, SequenceType
>>> protein_seq = Sequence("MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSF", "protein")
>>> config = MMseqsSimilarityConfig(
...     min_similarity=10.0, max_similarity=30.0, mmseqs_db="/data/databases/uniref90"
... )
>>> results = mmseqs_similarity_constraint([(protein_seq,)], config)
>>> print(results[0].score)  # 0.0 means no high-similarity hits
>>> print(results[0].metadata["similarity_compliance_rate"])
```

## Metadata

| Property        | Value                          |
| --------------- | ------------------------------ |
| Key             | `mmseqs-gene-similarity`       |
| Function        | `mmseqs_similarity_constraint` |
| Category        | `sequence_annotation`          |
| Mode            | `discrete`                     |
| Uses GPU        | `False`                        |
| Supported Types | `dna`, `protein`               |
