> ## Documentation Index
> Fetch the complete documentation index at: https://proto.evodesign.org/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# MMseqs2

> [MMseqs2](https://github.com/soedinglab/MMseqs2) (Many-against-Many sequence searching) is an ultra-fast protein and nucleotide sequence search and clustering toolkit from the [Steinegger](https://steineggerlab.com/) and [Söding](https://www.mpinat.mpg.de/soeding) labs. It searches very large databases at speeds substantially beyond BLAST with comparable sensitivity, supports GPU acceleration on recent NVIDIA hardware, and powers the homology-search step of the ColabFold structure-prediction pipeline. The toolkit exposes protein search, nucleotide-genome search, clustering, and ColabFold-style MSA generation as four registered tools.

<div class="page-hero"><img class="page-hero-banner" src="https://proto-bio.github.io/proto-assets/images/tool/mmseqs2/hero.png" alt="MMseqs2" /><div class="tool-org-badges page-hero-badges"><a href="/docs/tools/organizations/steinegger-lab" class="tool-org-badge" style={{background: "#2E86C1"}} title="Steinegger Lab"><img src="https://mintcdn.com/bio-pro/UeudeF7pW-Dj-pIN/assets/images/cached/5e3c631c9751.png?fit=max&auto=format&n=UeudeF7pW-Dj-pIN&q=85&s=5b9aab8b7ad1af8d3310a125bc9e540c" alt="" class="tool-org-badge-logo" width="200" height="200" data-path="assets/images/cached/5e3c631c9751.png" /> Steinegger Lab</a> <a href="/docs/tools/organizations/s-ding-lab" class="tool-org-badge" style={{background: "#116656"}} title="Söding Lab"><img src="https://mintcdn.com/bio-pro/_UGa2jUMKeVPCbLk/assets/images/cached/ff28b89e17f3.png?fit=max&auto=format&n=_UGa2jUMKeVPCbLk&q=85&s=4c430d789fb73ff3fee2ef1fb75dc90b" alt="" class="tool-org-badge-logo" width="200" height="200" data-path="assets/images/cached/ff28b89e17f3.png" /> Söding Lab</a></div></div>

<Note>
  **License:** MMseqs2 is open source and free for academic and commercial use under an MIT license. Please refer to [the license](https://github.com/soedinglab/MMseqs2/blob/master/LICENSE.md) for full terms.
</Note>

<p class="entity-disclaimer">Proto is not affiliated with the Steinegger Lab and the Söding Lab. This toolkit is open source and builds on the implementations produced by these organizations. Product names, logos, and trademarks are the property of their respective owners.</p>

<hr class="entity-rule" />

<input type="radio" name="tab-mmseqs2" id="none-mmseqs2" class="tab-radio-input" />

<input type="radio" name="tab-mmseqs2" id="github-mmseqs2" class="tab-radio-input" defaultChecked />

<input type="radio" name="tab-mmseqs2" id="paper-mmseqs2" class="tab-radio-input" />

<input type="radio" name="tab-mmseqs2" id="cite-mmseqs2" class="tab-radio-input" />

<input type="radio" name="tab-mmseqs2" id="source-mmseqs2" class="tab-radio-input" />

<input type="radio" name="tab-mmseqs2" id="proto-mmseqs2" class="tab-radio-input" />

<div class="tool-tab-bar">
  <span class="tool-tab-wrap"><label for="github-mmseqs2" class="tool-tab tab-open badge-github"><svg width="14" height="14" viewBox="0 0 24 24" fill="currentColor"><path d="M12 0C5.37 0 0 5.37 0 12c0 5.31 3.435 9.795 8.205 11.385.6.105.825-.255.825-.57 0-.285-.015-1.23-.015-2.235-3.015.555-3.795-.735-4.035-1.41-.135-.345-.72-1.41-1.23-1.695-.42-.225-1.02-.78-.015-.795.945-.015 1.62.87 1.845 1.23 1.08 1.815 2.805 1.305 3.495.99.105-.78.42-1.305.765-1.605-2.67-.3-5.46-1.335-5.46-5.925 0-1.305.465-2.385 1.23-3.225-.12-.3-.54-1.53.12-3.18 0 0 1.005-.315 3.3 1.23.96-.27 1.98-.405 3-.405s2.04.135 3 .405c2.295-1.56 3.3-1.23 3.3-1.23.66 1.65.24 2.88.12 3.18.765.84 1.23 1.905 1.23 3.225 0 4.605-2.805 5.625-5.475 5.925.435.375.81 1.095.81 2.22 0 1.605-.015 2.895-.015 3.3 0 .315.225.69.825.57A12.02 12.02 0 0024 12c0-6.63-5.37-12-12-12z" /></svg> GitHub</label><label for="none-mmseqs2" class="tool-tab tab-close badge-github"><svg width="14" height="14" viewBox="0 0 24 24" fill="currentColor"><path d="M12 0C5.37 0 0 5.37 0 12c0 5.31 3.435 9.795 8.205 11.385.6.105.825-.255.825-.57 0-.285-.015-1.23-.015-2.235-3.015.555-3.795-.735-4.035-1.41-.135-.345-.72-1.41-1.23-1.695-.42-.225-1.02-.78-.015-.795.945-.015 1.62.87 1.845 1.23 1.08 1.815 2.805 1.305 3.495.99.105-.78.42-1.305.765-1.605-2.67-.3-5.46-1.335-5.46-5.925 0-1.305.465-2.385 1.23-3.225-.12-.3-.54-1.53.12-3.18 0 0 1.005-.315 3.3 1.23.96-.27 1.98-.405 3-.405s2.04.135 3 .405c2.295-1.56 3.3-1.23 3.3-1.23.66 1.65.24 2.88.12 3.18.765.84 1.23 1.905 1.23 3.225 0 4.605-2.805 5.625-5.475 5.925.435.375.81 1.095.81 2.22 0 1.605-.015 2.895-.015 3.3 0 .315.225.69.825.57A12.02 12.02 0 0024 12c0-6.63-5.37-12-12-12z" /></svg> GitHub</label></span> <span class="tool-tab-wrap"><label for="paper-mmseqs2" class="tool-tab tab-open badge-paper"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z" /><polyline points="14 2 14 8 20 8" /><line x1="16" y1="13" x2="8" y2="13" /><line x1="16" y1="17" x2="8" y2="17" /><polyline points="10 9 9 9 8 9" /></svg> Publication</label><label for="none-mmseqs2" class="tool-tab tab-close badge-paper"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z" /><polyline points="14 2 14 8 20 8" /><line x1="16" y1="13" x2="8" y2="13" /><line x1="16" y1="17" x2="8" y2="17" /><polyline points="10 9 9 9 8 9" /></svg> Publication</label></span> <span class="tool-tab-wrap"><label for="cite-mmseqs2" class="tool-tab tab-open badge-cite"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M3 21c3 0 7-1 7-8V5c0-1.25-.756-2.017-2-2H4c-1.25 0-2 .75-2 1.972V11c0 1.25.75 2 2 2 1 0 1 0 1 1v1c0 1-1 2-2 2s-1 .008-1 1.031V20c0 1 0 1 1 1z" /><path d="M15 21c3 0 7-1 7-8V5c0-1.25-.757-2.017-2-2h-4c-1.25 0-2 .75-2 1.972V11c0 1.25.75 2 2 2h.75c0 2.25.25 4-2.75 4v3c0 1 0 1 1 1z" /></svg> Cite</label><label for="none-mmseqs2" class="tool-tab tab-close badge-cite"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M3 21c3 0 7-1 7-8V5c0-1.25-.756-2.017-2-2H4c-1.25 0-2 .75-2 1.972V11c0 1.25.75 2 2 2 1 0 1 0 1 1v1c0 1-1 2-2 2s-1 .008-1 1.031V20c0 1 0 1 1 1z" /><path d="M15 21c3 0 7-1 7-8V5c0-1.25-.757-2.017-2-2h-4c-1.25 0-2 .75-2 1.972V11c0 1.25.75 2 2 2h.75c0 2.25.25 4-2.75 4v3c0 1 0 1 1 1z" /></svg> Cite</label></span> <span class="tool-tab-wrap"><label for="source-mmseqs2" class="tool-tab tab-open badge-source"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Tool Source</label><label for="none-mmseqs2" class="tool-tab tab-close badge-source"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Tool Source</label></span> <span class="tool-tab-wrap"><label for="proto-mmseqs2" class="tool-tab tab-open badge-proto"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M13 2L3 14h9l-1 8 10-12h-9l1-8z" /></svg> Open on Proto</label><label for="none-mmseqs2" class="tool-tab tab-close badge-proto"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M13 2L3 14h9l-1 8 10-12h-9l1-8z" /></svg> Open on Proto</label></span>
</div>

<a href="https://github.com/soedinglab/MMseqs2" target="_blank" class="tab-panel github-panel" data-tab="github-mmseqs2">
  <div class="gh-card-wrap">
    <img src="https://opengraph.githubassets.com/1/soedinglab/MMseqs2" class="gh-card-img img-fallback" alt="soedinglab/MMseqs2" />

    <div class="gh-card-fallback">
      <div class="gh-fallback-org"><svg width="14" height="14" viewBox="0 0 24 24" fill="currentColor"><path d="M12 0C5.37 0 0 5.37 0 12c0 5.31 3.435 9.795 8.205 11.385.6.105.825-.255.825-.57 0-.285-.015-1.23-.015-2.235-3.015.555-3.795-.735-4.035-1.41-.135-.345-.72-1.41-1.23-1.695-.42-.225-1.02-.78-.015-.795.945-.015 1.62.87 1.845 1.23 1.08 1.815 2.805 1.305 3.495.99.105-.78.42-1.305.765-1.605-2.67-.3-5.46-1.335-5.46-5.925 0-1.305.465-2.385 1.23-3.225-.12-.3-.54-1.53.12-3.18 0 0 1.005-.315 3.3 1.23.96-.27 1.98-.405 3-.405s2.04.135 3 .405c2.295-1.56 3.3-1.23 3.3-1.23.66 1.65.24 2.88.12 3.18.765.84 1.23 1.905 1.23 3.225 0 4.605-2.805 5.625-5.475 5.925.435.375.81 1.095.81 2.22 0 1.605-.015 2.895-.015 3.3 0 .315.225.69.825.57A12.02 12.02 0 0024 12c0-6.63-5.37-12-12-12z" /></svg> soedinglab/MMseqs2</div>
    </div>
  </div>

  <span class="panel-goto-btn gh-goto-btn"><span><svg width="14" height="14" viewBox="0 0 24 24" fill="currentColor"><path d="M12 0C5.37 0 0 5.37 0 12c0 5.31 3.435 9.795 8.205 11.385.6.105.825-.255.825-.57 0-.285-.015-1.23-.015-2.235-3.015.555-3.795-.735-4.035-1.41-.135-.345-.72-1.41-1.23-1.695-.42-.225-1.02-.78-.015-.795.945-.015 1.62.87 1.845 1.23 1.08 1.815 2.805 1.305 3.495.99.105-.78.42-1.305.765-1.605-2.67-.3-5.46-1.335-5.46-5.925 0-1.305.465-2.385 1.23-3.225-.12-.3-.54-1.53.12-3.18 0 0 1.005-.315 3.3 1.23.96-.27 1.98-.405 3-.405s2.04.135 3 .405c2.295-1.56 3.3-1.23 3.3-1.23.66 1.65.24 2.88.12 3.18.765.84 1.23 1.905 1.23 3.225 0 4.605-2.805 5.625-5.475 5.925.435.375.81 1.095.81 2.22 0 1.605-.015 2.895-.015 3.3 0 .315.225.69.825.57A12.02 12.02 0 0024 12c0-6.63-5.37-12-12-12z" /></svg> View repo</span></span>
</a>

<a href="https://doi.org/10.1038/nbt.3988" target="_blank" class="tab-panel paper-panel" data-tab="paper-mmseqs2">
  <div class="paper-info">
    <div class="paper-title">MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets</div>
    <div class="paper-meta">Martin Steinegger and Johannes Soding</div>
    <div class="paper-meta paper-venue">Nature Biotechnology (2017)</div>
  </div>

  <span class="panel-goto-btn pub-goto-btn"><span><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z" /><polyline points="14 2 14 8 20 8" /><line x1="16" y1="13" x2="8" y2="13" /><line x1="16" y1="17" x2="8" y2="17" /><polyline points="10 9 9 9 8 9" /></svg> Read paper</span></span>
</a>

<div class="tab-panel cite-panel" data-tab="cite-mmseqs2">
  <div class="cite-code-wrap">
    ```bibtex theme={null}
    @article{steinegger2017mmseqs2,
      title={MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets},
      author={Steinegger, Martin and S{\"o}ding, Johannes},
      journal={Nature Biotechnology},
      volume={35},
      number={11},
      pages={1026--1028},
      year={2017},
      publisher={Nature Publishing Group},
      doi={10.1038/nbt.3988}
    }

    @article{kallenborn2025mmseqs2gpu,
      title={GPU-accelerated homology search with MMseqs2},
      author={Kallenborn, Felix and Chacon, Alvaro and Hundt, Christian and Sirelkhatim, Hassan and Didi, Kieran and Cha, Sooyoung and Dallago, Christian and Mirdita, Milot and Schmidt, Bertil and Steinegger, Martin},
      journal={Nature Methods},
      year={2025},
      publisher={Nature Publishing Group},
      doi={10.1038/s41592-025-02819-8}
    }

    @article{mirdita2022colabfold,
      title={ColabFold: making protein folding accessible to all},
      author={Mirdita, Milot and Sch{\"u}tze, Konstantin and Moriwaki, Yoshitaka and Heo, Lim and Ovchinnikov, Sergey and Steinegger, Martin},
      journal={Nature Methods},
      volume={19},
      number={6},
      pages={679--682},
      year={2022},
      publisher={Nature Publishing Group},
      doi={10.1038/s41592-022-01488-1}
    }
    ```
  </div>

  <span class="panel-goto-btn cite-copy-btn"><span><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M3 21c3 0 7-1 7-8V5c0-1.25-.756-2.017-2-2H4c-1.25 0-2 .75-2 1.972V11c0 1.25.75 2 2 2 1 0 1 0 1 1v1c0 1-1 2-2 2s-1 .008-1 1.031V20c0 1 0 1 1 1z" /><path d="M15 21c3 0 7-1 7-8V5c0-1.25-.757-2.017-2-2h-4c-1.25 0-2 .75-2 1.972V11c0 1.25.75 2 2 2h.75c0 2.25.25 4-2.75 4v3c0 1 0 1 1 1z" /></svg> Copy citation</span></span>
</div>

<a href="https://github.com/evo-design/proto-tools/tree/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2" target="_blank" class="tab-panel source-panel" data-tab="source-mmseqs2">
  <div class="source-info">
    <img src="https://github.com/evo-design.png?size=40" class="source-avatar" width="36" height="36" />

    <span class="source-path">evo-design/proto-tools<span class="source-subpath">/proto\_tools/tools/sequence\_alignment/mmseqs2</span></span>
  </div>

  <span class="panel-goto-btn source-goto-btn"><span><svg width="14" height="14" viewBox="0 0 24 24" fill="currentColor"><path d="M12 0C5.37 0 0 5.37 0 12c0 5.31 3.435 9.795 8.205 11.385.6.105.825-.255.825-.57 0-.285-.015-1.23-.015-2.235-3.015.555-3.795-.735-4.035-1.41-.135-.345-.72-1.41-1.23-1.695-.42-.225-1.02-.78-.015-.795.945-.015 1.62.87 1.845 1.23 1.08 1.815 2.805 1.305 3.495.99.105-.78.42-1.305.765-1.605-2.67-.3-5.46-1.335-5.46-5.925 0-1.305.465-2.385 1.23-3.225-.12-.3-.54-1.53.12-3.18 0 0 1.005-.315 3.3 1.23.96-.27 1.98-.405 3-.405s2.04.135 3 .405c2.295-1.56 3.3-1.23 3.3-1.23.66 1.65.24 2.88.12 3.18.765.84 1.23 1.905 1.23 3.225 0 4.605-2.805 5.625-5.475 5.925.435.375.81 1.095.81 2.22 0 1.605-.015 2.895-.015 3.3 0 .315.225.69.825.57A12.02 12.02 0 0024 12c0-6.63-5.37-12-12-12z" /></svg> View source</span></span>
</a>

<div class="tab-panel proto-panel" data-tab="proto-mmseqs2">
  <div class="proto-info">
    <div class="proto-cloud">
      <svg class="proto-cloud-bg" viewBox="0 0 640 512" xmlns="http://www.w3.org/2000/svg">
        <path d="M0 336c0 79.5 64.5 144 144 144H512c70.7 0 128-57.3 128-128c0-61.9-44-113.6-102.4-125.4c4.1-10.7 6.4-22.4 6.4-34.6c0-53-43-96-96-96c-19.7 0-38.1 6-53.3 16.2C367 64.2 315.3 32 256 32C167.6 32 96 103.6 96 192c0 2.7 .1 5.4 .2 8.1C40.2 219.8 0 273.2 0 336z" />
      </svg>

      <img noZoom src="https://mintcdn.com/bio-pro/KVh0EKV-IKblvXR8/assets/logo/evo-logo-light.svg?fit=max&auto=format&n=KVh0EKV-IKblvXR8&q=85&s=0cb66034ba45618505501aee6ea5f5c1" class="proto-panel-logo block dark:hidden" alt="Proto" width="198" height="151" data-path="assets/logo/evo-logo-light.svg" />

      <img noZoom src="https://mintcdn.com/bio-pro/KVh0EKV-IKblvXR8/assets/logo/evo-logo-dark.svg?fit=max&auto=format&n=KVh0EKV-IKblvXR8&q=85&s=2c9e23a14635e60384a434e220788f54" class="proto-panel-logo hidden dark:block" alt="Proto" width="198" height="151" data-path="assets/logo/evo-logo-dark.svg" />
    </div>
  </div>

  <div class="proto-actions">
    <a href="https://proto.evodesign.org/tools/mmseqs2-clustering" target="_blank" class="proto-action-btn"><span>MMseqs2 Clustering</span><svg width="13" height="13" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><line x1="7" y1="17" x2="17" y2="7" /><polyline points="7 7 17 7 17 17" /></svg></a>
    <a href="https://proto.evodesign.org/tools/mmseqs2-homology-search" target="_blank" class="proto-action-btn"><span>MMseqs2 Homology Search</span><svg width="13" height="13" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><line x1="7" y1="17" x2="17" y2="7" /><polyline points="7 7 17 7 17 17" /></svg></a>
    <a href="https://proto.evodesign.org/tools/mmseqs2-search-genomes" target="_blank" class="proto-action-btn"><span>MMseqs2 Genome Search</span><svg width="13" height="13" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><line x1="7" y1="17" x2="17" y2="7" /><polyline points="7 7 17 7 17 17" /></svg></a>
    <a href="https://proto.evodesign.org/tools/mmseqs2-search-proteins" target="_blank" class="proto-action-btn"><span>MMseqs2 Protein Search</span><svg width="13" height="13" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><line x1="7" y1="17" x2="17" y2="7" /><polyline points="7 7 17 7 17 17" /></svg></a>
  </div>
</div>

<div class="entity-contributors"><span class="entity-contributors-label">Toolkit contributors</span><span class="entity-contributors-people"><a class="entity-contributor" href="https://github.com/bviggiano" target="_blank" rel="noopener" title="bviggiano: 26 commits"><img noZoom class="entity-contributor-avatar" src="https://avatars.githubusercontent.com/u/21143637?v=4&s=64" alt="" loading="lazy" /><span class="entity-contributor-login">bviggiano</span></a><a class="entity-contributor" href="https://github.com/dguo8412" target="_blank" rel="noopener" title="dguo8412: 19 commits"><img noZoom class="entity-contributor-avatar" src="https://avatars.githubusercontent.com/u/46211285?v=4&s=64" alt="" loading="lazy" /><span class="entity-contributor-login">dguo8412</span></a><a class="entity-contributor" href="https://github.com/leba01" target="_blank" rel="noopener" title="leba01: 2 commits"><img noZoom class="entity-contributor-avatar" src="https://avatars.githubusercontent.com/u/124846286?v=4&s=64" alt="" loading="lazy" /><span class="entity-contributor-login">leba01</span></a></span></div>

| Function                        | Description                                                                           |                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| ------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `run_mmseqs2_clustering()`      | Perform sequence clustering using MMseqs2 to reduce redundancy                        | <a href="#api-run-mmseqs2-clustering" class="func-table-btn func-api-btn"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M4 19.5v-15A2.5 2.5 0 0 1 6.5 2H19a1 1 0 0 1 1 1v18a1 1 0 0 1-1 1H6.5a1 1 0 0 1 0-5H20" /></svg> Docs</a> <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/clustering.py#L296" target="_blank" class="func-table-btn func-source-btn"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>           |
| `run_mmseqs2_homology_search()` | Generate MSAs by searching protein sequences against MMseqs2-indexed databases. (GPU) | <a href="#api-run-mmseqs2-homology-search" class="func-table-btn func-api-btn"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M4 19.5v-15A2.5 2.5 0 0 1 6.5 2H19a1 1 0 0 1 1 1v18a1 1 0 0 1-1 1H6.5a1 1 0 0 1 0-5H20" /></svg> Docs</a> <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/homology_search.py#L452" target="_blank" class="func-table-btn func-source-btn"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a> |
| `run_mmseqs2_search_genomes()`  | Execute nucleotide genome-to-genome search workflow                                   | <a href="#api-run-mmseqs2-search-genomes" class="func-table-btn func-api-btn"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M4 19.5v-15A2.5 2.5 0 0 1 6.5 2H19a1 1 0 0 1 1 1v18a1 1 0 0 1-1 1H6.5a1 1 0 0 1 0-5H20" /></svg> Docs</a> <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/search_genomes.py#L302" target="_blank" class="func-table-btn func-source-btn"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>   |
| `run_mmseqs2_search_proteins()` | Search protein sequences using MMseqs2 with per-sequence results (GPU)                | <a href="#api-run-mmseqs2-search-proteins" class="func-table-btn func-api-btn"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M4 19.5v-15A2.5 2.5 0 0 1 6.5 2H19a1 1 0 0 1 1 1v18a1 1 0 0 1-1 1H6.5a1 1 0 0 1 0-5H20" /></svg> Docs</a> <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/search_proteins.py#L404" target="_blank" class="func-table-btn func-source-btn"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a> |

## Background

MMseqs2 ([Steinegger and Söding, 2017](https://doi.org/10.1038/nbt.3988)) implements sequence-similarity search and clustering through a cascaded prefilter-align approach. Short k-mer matches between the query and the database are located first, surviving candidates are scored with an ungapped extension, and final hits are realigned with gapped [Smith-Waterman](https://en.wikipedia.org/wiki/Smith%E2%80%93Waterman_algorithm) alignment. Clustering uses greedy [set cover](https://en.wikipedia.org/wiki/Set_cover_problem) over the alignment graph. The cascade reduces the search space by several orders of magnitude while retaining sensitivity comparable to BLAST, making analyses over databases with billions of sequences tractable on a single workstation.

The GPU build ([Kallenborn et al., 2025](https://doi.org/10.1038/s41592-025-02819-8)) accelerates the prefilter and alignment stages on NVIDIA Turing-generation or newer hardware. On top of the search engine, the ColabFold homology-search pipeline ([Mirdita et al., 2022](https://doi.org/10.1038/s41592-022-01488-1)) iterates MMseqs2 searches against clustered reference databases such as UniRef30 to produce the multiple sequence alignments that AlphaFold-class structure predictors consume. This toolkit exposes that pipeline as `mmseqs2-homology-search` in addition to the more general search and clustering operations.

### Learning Resources

* [soedinglab/MMseqs2](https://github.com/soedinglab/MMseqs2) (Söding and Steinegger labs). Official repository and the source of the `mmseqs` command-line program that this toolkit invokes.
* [MMseqs2 wiki](https://github.com/soedinglab/MMseqs2/wiki) (Söding and Steinegger labs). The reference wiki for the command-line surface, including the workflow modules that the four registered tools wrap.
* [ColabFold homology-search documentation](https://github.com/sokrypton/ColabFold/wiki) (Steinegger and Ovchinnikov labs). Walks through the iterative MSA pipeline that `mmseqs2-homology-search` runs internally.

## Tools

<a name="api-run-mmseqs2-search-proteins" />

<div class="tool-section-card tool-section-card--search">
  ### MMseqs2 Protein Search (`mmseqs2-search-proteins`)

  Performs `mmseqs easy-search` of one or more protein query sequences against either a user-supplied target database or an inline list of target proteins. Returns the alignment hits per query, each with target identifier, percent identity, and E-value. The local execution mode runs on CPU by default and supports an opt-in GPU mode for searches against a prebuilt database with a GPU-padded index.

  #### API Reference

  <div class="api-model-section api-input-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/search_proteins.py#L117" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Input: Mmseqs2SearchProteinsInput">
      <ParamField path="query_sequences" type="List[string]" required>
        List of protein sequence strings (amino acid sequences) to search. Labeled positionally (`seq_0`, `seq_1`, ...) as `query_id` in the output; results are returned in input order.
      </ParamField>
    </Accordion>
  </div>

  <div class="api-model-section api-config-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/search_proteins.py#L228" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Config: Mmseqs2SearchProteinsConfig">
      <ParamField path="mmseqs_db" type="string">
        Target DB (path/slug/AssetRef). Mutually exclusive with `target_sequences`; required at dispatch.
      </ParamField>

      <ParamField path="target_sequences" type="array">
        Inline target sequences. Mutually exclusive with `mmseqs_db`; required at dispatch.
      </ParamField>

      <ParamField path="threads" type="integer" default="0">
        CPU threads; `0` auto-detects all cores (the wrapper omits `--threads` since `mmseqs` rejects `--threads 0`).
      </ParamField>

      <ParamField path="split" type="integer" default="0">
        Split into N chunks to bound memory; `0` = auto.
      </ParamField>

      <ParamField path="split_memory_limit" type="string">
        Max prefilter memory per split (e.g. `"90G"`); `None` uses all system memory.
      </ParamField>

      <ParamField path="sensitivity" type="number" default="5.7">
        Prefilter sensitivity (1.0-7.5); higher = slower but finds more remote homologs.
      </ParamField>

      <ParamField path="evalue" type="number" default="0.001">
        E-value threshold for reported hits.
      </ParamField>

      <ParamField path="min_seq_id" type="number" default="0.0">
        Minimum sequence identity (0.0-1.0) for reported hits.
      </ParamField>

      <ParamField path="coverage" type="number" default="0.0">
        Minimum aligned-residue fraction (0.0-1.0); semantics depend on `cov_mode`.
      </ParamField>

      <ParamField path="cov_mode" type="enum" default="0">
        0=query AND target, 1=target, 2=query, 3-5=length-ratio variants.

        Available options: `0`, `1`, `2`, `3`, `4`, `5`
      </ParamField>

      <ParamField path="max_seqs" type="integer" default="300">
        Max prefilter results per query.
      </ParamField>

      <ParamField path="only_top_hits" type="boolean" default="True">
        Wrapper filter; keep only the best hit (highest pident) per query sequence.
      </ParamField>

      <ParamField path="extra_args" type="List[string]" default="[]">
        Verbatim `mmseqs easy-search` CLI tokens for niche flags (e.g. `["--alignment-mode", "3"]`).
      </ParamField>

      <ParamField path="verbose" type="integer" default="0">
        Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). `True` is coerced to `1` and `False` to `0`.
      </ParamField>

      <ParamField path="device" type="string" default="cpu">
        A `cuda` device runs MMseqs2-GPU (`--gpu 1`); requires a `.idx_pad` sibling on the target DB (built via `mmseqs makepaddedseqdb`).
      </ParamField>

      <ParamField path="timeout" type="integer" default="3600">
        Maximum execution time in seconds. `None` waits indefinitely.
      </ParamField>

      <ParamField path="seed" type="integer">
        Random seed. When set, tools run reproducibly up to small GPU float noise (see `BaseToolOutput.approx_equal`), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
      </ParamField>
    </Accordion>
  </div>

  <div class="api-model-section api-output-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/search_proteins.py#L145" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Output: Mmseqs2SearchProteinsOutput">
      <ResponseField name="results" type="List[Mmseqs2SequenceSearchResult]" required>
        List of search results, one per input sequence. The order matches the input sequences order.

        <Expandable title="Mmseqs2SequenceSearchResult">
          <ResponseField name="query_id" type="string" required>
            Identifier of the query sequence.
          </ResponseField>

          <ResponseField name="query_sequence" type="string" required>
            The input query sequence.
          </ResponseField>

          <ResponseField name="hits" type="List[Mmseqs2Hit]">
            All hits found for this query, sorted by pident descending.
          </ResponseField>
        </Expandable>
      </ResponseField>
    </Accordion>
  </div>

  #### Applications

  This tool is appropriate for ad-hoc protein homology search at scale, ranking hits against a custom reference database, identifying functional homologs across a sequenced library, and any analysis in which the BLAST-style hit table is the deliverable. The MMseqs2 sensitivity model finds remote homologs that fall well below the sequence-identity range where standard BLAST searches lose signal.

  #### Usage Tips

  * **Targets are specified on `Mmseqs2SearchProteinsConfig` via either `mmseqs_db` or `target_sequences`, but not both.** Use `mmseqs_db` for a path to a FASTA file or a prebuilt MMseqs2 database when the target set is large or reused across calls. Use `target_sequences` for short inline lists. They live on Config because every query in a run searches against the same target collection.
  * **`mmseqs_db` is cached by path, not by contents.** The per-item cache key includes the path string but not the bytes at that path, so mutating the database in place (overwriting `/dbs/uniref90` with a new build at the same path) will silently return stale hits. If you swap a DB at a stable path, either clear the cache or use a versioned filename (`/dbs/uniref90_v2`) so the new path forces a fresh key.
  * **`sensitivity=5.7` is the wrapper default and matches upstream `easy-search`.** Higher values recover more distant homologs at the cost of additional runtime. The accepted range is 1.0 to 7.5.
  * **`only_top_hits=True` (the default) returns only the best hit per query by percent identity.** Set it to `False` to retain every hit that passes the configured thresholds.
  * **A `cuda` device requires a GPU-padded index alongside the target database.** Build the index once with `mmseqs makepaddedseqdb <db> <db>.idx_pad`. The configuration validator hard-errors when the `.idx_pad` companion is missing or when a `cuda` device is combined with inline `target_sequences` (the GPU path does not accept inline targets).
  * **`extra_args` accepts verbatim [`mmseqs easy-search`](https://github.com/soedinglab/MMseqs2/wiki#easy-search) CLI tokens.** Pass any flag not exposed as a typed field through this list (for example `["--alignment-mode", "3"]`). Tokens are appended after the typed flags.

  <a name="api-run-mmseqs2-search-genomes" />
</div>

<div class="tool-section-card tool-section-card--search">
  ### MMseqs2 Genome Search (`mmseqs2-search-genomes`)

  Performs the full MMseqs2 nucleotide search pipeline against either a user-supplied target database or an inline list of target genomes. The tool builds query and target databases with `createdb`, runs `search`, and converts the result to the BLAST-style tabular schema with `convertalis`. Runs on CPU only.

  #### API Reference

  <div class="api-model-section api-input-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/search_genomes.py#L44" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Input: Mmseqs2SearchGenomesInput">
      <ParamField path="query_genomes" type="List[string]" required>
        List of nucleotide sequence strings (DNA/RNA) to use as queries. Labeled positionally (`seq_0`, `seq_1`, ...) as `query_id` in the output; results are returned in query order.
      </ParamField>
    </Accordion>
  </div>

  <div class="api-model-section api-config-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/search_genomes.py#L151" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Config: Mmseqs2SearchGenomesConfig">
      <ParamField path="target_genomes" type="array">
        Inline target genomes. Mutually exclusive with `target_db`; required at dispatch.
      </ParamField>

      <ParamField path="target_db" type="string">
        Target FASTA or MMseqs2 DB stem (path/slug/AssetRef). Mutually exclusive with `target_genomes`; required at dispatch.
      </ParamField>

      <ParamField path="threads" type="integer" default="0">
        CPU threads; `0` auto-detects all cores (the wrapper omits `--threads` since `mmseqs` rejects `--threads 0`).
      </ParamField>

      <ParamField path="sensitivity" type="number" default="7.5">
        Prefilter sensitivity (1.0-7.5). Wrapper default 7.5 (upstream MMseqs2 = 5.7).
      </ParamField>

      <ParamField path="evalue" type="number" default="0.001">
        E-value threshold for reported hits.
      </ParamField>

      <ParamField path="min_seq_id" type="number" default="0.0">
        Minimum sequence identity (0.0-1.0) for reported hits.
      </ParamField>

      <ParamField path="coverage" type="number" default="0.0">
        Minimum aligned-residue fraction (0.0-1.0); semantics depend on `cov_mode`.
      </ParamField>

      <ParamField path="cov_mode" type="enum" default="0">
        0=query AND target, 1=target, 2=query, 3-5=length-ratio variants.

        Available options: `0`, `1`, `2`, `3`, `4`, `5`
      </ParamField>

      <ParamField path="max_seqs" type="integer" default="300">
        Max prefilter results per query.
      </ParamField>

      <ParamField path="strand" type="enum" default="2">
        0=reverse, 1=forward, 2=both. Wrapper default 2 (upstream MMseqs2 = 1).

        Available options: `0`, `1`, `2`
      </ParamField>

      <ParamField path="extra_args" type="List[string]" default="[]">
        Verbatim `mmseqs search` CLI tokens for niche flags (e.g. `["--alignment-mode", "2"]`).
      </ParamField>

      <ParamField path="verbose" type="integer" default="0">
        Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). `True` is coerced to `1` and `False` to `0`.
      </ParamField>

      <ParamField path="device" type="string" default="cpu">
        Device to run the tool on.
      </ParamField>

      <ParamField path="timeout" type="integer" default="3600">
        Maximum execution time in seconds. `None` waits indefinitely.
      </ParamField>

      <ParamField path="seed" type="integer">
        Random seed. When set, tools run reproducibly up to small GPU float noise (see `BaseToolOutput.approx_equal`), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
      </ParamField>
    </Accordion>
  </div>

  <div class="api-model-section api-output-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/search_genomes.py#L71" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Output: Mmseqs2SearchGenomesOutput">
      <ResponseField name="results" type="List[Mmseqs2SequenceSearchResult]" required>
        List of search results, one per input query genome. The order matches the input query genomes order.

        <Expandable title="Mmseqs2SequenceSearchResult">
          <ResponseField name="query_id" type="string" required>
            Identifier of the query sequence.
          </ResponseField>

          <ResponseField name="query_sequence" type="string" required>
            The input query sequence.
          </ResponseField>

          <ResponseField name="hits" type="List[Mmseqs2Hit]">
            All hits found for this query, sorted by pident descending.
          </ResponseField>
        </Expandable>
      </ResponseField>
    </Accordion>
  </div>

  #### Applications

  This tool is appropriate for genome-to-genome similarity analysis, locating homologous regions between assembled genomes, comparative genomics over closely related strains, and any nucleotide analog of the protein-search workflow.

  #### Usage Tips

  * **Targets are specified on `Mmseqs2SearchGenomesConfig` via either `target_db` or `target_genomes`, but not both.** Use `target_db` for a FASTA file or a prebuilt MMseqs2 database; use `target_genomes` for inline nucleotide sequences. They live on Config because every query in a run scans against the same target collection.
  * **`target_db` is cached by path, not by contents.** The per-item cache key includes the path string but not the bytes at that path, so mutating the database in place will silently return stale hits. If you swap a DB at a stable path, either clear the cache or use a versioned filename so the new path forces a fresh key.
  * **`sensitivity=7.5` is the wrapper default for nucleotide search.** This is a wrapper bias above the upstream MMseqs2 default of 5.7, chosen because nucleotide searches typically benefit from the higher sensitivity setting. The accepted range is 1.0 to 7.5.
  * **`strand=2` (both strands) is the wrapper default.** Upstream defaults to forward strand only. Set `strand=1` to restrict to the forward strand or `strand=0` for reverse only.
  * **`extra_args` accepts verbatim [`mmseqs search`](https://github.com/soedinglab/MMseqs2/wiki#search-workflow) CLI tokens.** Tokens are appended after the typed flags.

  <a name="api-run-mmseqs2-clustering" />
</div>

<div class="tool-section-card tool-section-card--cluster">
  ### MMseqs2 Clustering (`mmseqs2-clustering`)

  Performs `mmseqs cluster` over an inline list of sequences or a prebuilt MMseqs2 database and returns per-sequence cluster assignments. Each result records the cluster identifier and whether the sequence is the cluster representative. Runs on CPU only.

  #### API Reference

  <div class="api-model-section api-input-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/clustering.py#L88" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Input: Mmseqs2ClusteringInput">
      <ParamField path="input_sequences" type="array">
        Inline sequences to cluster. Mutually exclusive with `mmseqs_db`.
      </ParamField>

      <ParamField path="mmseqs_db" type="string">
        Pre-built MMseqs2 DB (path/slug/AssetRef). Mutually exclusive with `input_sequences`.
      </ParamField>

      <ParamField path="sequence_ids" type="array">
        Optional IDs for inline sequences (defaults to seq\_0, seq\_1, ...).
      </ParamField>
    </Accordion>
  </div>

  <div class="api-model-section api-config-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/clustering.py#L208" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Config: Mmseqs2ClusteringConfig">
      <ParamField path="min_seq_id" type="number" default="0.6">
        Min identity (0.0-1.0) to share a cluster. Wrapper default 0.6 (upstream MMseqs2 = 0.0).
      </ParamField>

      <ParamField path="coverage" type="number" default="0.8">
        Minimum aligned-residue fraction (0.0-1.0); semantics depend on `cov_mode`.
      </ParamField>

      <ParamField path="cov_mode" type="enum" default="0">
        0=query AND target, 1=target, 2=query, 3-5=length-ratio variants.

        Available options: `0`, `1`, `2`, `3`, `4`, `5`
      </ParamField>

      <ParamField path="evalue" type="number" default="0.001">
        E-value threshold for the prefilter step.
      </ParamField>

      <ParamField path="cluster_mode" type="enum" default="0">
        0=Set-Cover (greedy), 1=Connected component (BLASTclust), 2-3=Greedy by length (CD-HIT).

        Available options: `0`, `1`, `2`, `3`
      </ParamField>

      <ParamField path="max_seqs" type="integer" default="20">
        Max prefilter results per query.
      </ParamField>

      <ParamField path="sensitivity" type="number" default="4.0">
        Prefilter sensitivity (1.0-7.5).
      </ParamField>

      <ParamField path="extra_args" type="List[string]" default="[]">
        Verbatim `mmseqs cluster` CLI tokens for niche flags (e.g. `["--similarity-type", "2"]`).
      </ParamField>

      <ParamField path="verbose" type="integer" default="0">
        Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). `True` is coerced to `1` and `False` to `0`.
      </ParamField>

      <ParamField path="device" type="string" default="cpu">
        Device to run the tool on.
      </ParamField>

      <ParamField path="timeout" type="integer" default="3600">
        Maximum execution time in seconds. `None` waits indefinitely.
      </ParamField>

      <ParamField path="seed" type="integer">
        Random seed. When set, tools run reproducibly up to small GPU float noise (see `BaseToolOutput.approx_equal`), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
      </ParamField>
    </Accordion>
  </div>

  <div class="api-model-section api-output-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/clustering.py#L139" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Output: Mmseqs2ClusteringOutput">
      <ResponseField name="results" type="List[Mmseqs2ClusterResult]" required>
        List of clustering results, one per input sequence. The order matches the input sequences order.

        <Expandable title="Mmseqs2ClusterResult">
          <ResponseField name="sequence_id" type="string" required>
            Identifier of the input sequence.
          </ResponseField>

          <ResponseField name="input_sequence" type="string">
            Original input sequence; `None` when the caller used `mmseqs_db`.
          </ResponseField>

          <ResponseField name="cluster_id" type="string" required>
            Identifier of the cluster (usually the representative's ID).
          </ResponseField>

          <ResponseField name="is_representative" type="boolean">
            Whether this sequence is the cluster representative.
          </ResponseField>
        </Expandable>
      </ResponseField>
    </Accordion>
  </div>

  #### Applications

  This tool is appropriate for deduplicating a sequence set before downstream analysis, partitioning a protein library into functional families, selecting representative sequences from a redundant collection, and any analysis that benefits from a similarity-based grouping of sequences.

  #### Usage Tips

  * **Inputs are specified via either `input_sequences` or `mmseqs_db`, but not both.** Use `input_sequences` for inline sequences; use `mmseqs_db` for a prebuilt database that may be reused across calls.
  * **`mmseqs_db` is cached by path, not by contents.** The cache key includes the path string but not the bytes at that path, so mutating the database in place will silently return stale cluster assignments. If you swap a DB at a stable path, either clear the cache or use a versioned filename so the new path forces a fresh key.
  * **`min_seq_id=0.6` is the wrapper default.** This is a wrapper bias above the upstream MMseqs2 default of 0.0, chosen as a reasonable starting point for grouping proteins into functional families. Set it higher (for example `0.95`) to remove near-duplicates, or lower (for example `0.3`) to group remote homologs.
  * **`cluster_mode=0` (set-cover) is the default greedy algorithm.** Alternative modes are `1` (connected-component, BLASTclust-style) and `2` or `3` (greedy by length, CD-HIT-style).
  * **The cluster representative is the first sequence to cover the cluster during greedy set-cover.** It is not necessarily the longest or most central sequence. Choose an alternative `cluster_mode` if a different representative-selection policy is needed.
  * **`extra_args` accepts verbatim [`mmseqs cluster`](https://github.com/soedinglab/MMseqs2/wiki#clustering) CLI tokens.** Tokens are appended after the typed flags.

  <a name="api-run-mmseqs2-homology-search" />
</div>

<div class="tool-section-card tool-section-card--search">
  ### MMseqs2 Homology Search (`mmseqs2-homology-search`)

  Generates a multiple sequence alignment per query protein using the ColabFold homology-search pipeline. Returns one `MSA` object per query, suitable as the MSA input to AlphaFold-class structure predictors. By default (`search_mode="remote"`) it queries the hosted ColabFold MSA API over the network and needs no local database; set `search_mode="local"` to run MMseqs2 against a registry-provisioned reference database on disk (GPU-accelerated by default on supported hardware).

  #### API Reference

  <div class="api-model-section api-input-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/homology_search.py#L112" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Input: Mmseqs2HomologySearchInput">
      <ParamField path="queries" type="List[Mmseqs2HomologySearchQuery | List[Mmseqs2HomologySearchQuery]]" required>
        List of query groups, in input order.
      </ParamField>
    </Accordion>
  </div>

  <div class="api-model-section api-config-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/homology_search.py#L294" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Config: Mmseqs2HomologySearchConfig">
      <ParamField path="search_mode" type="enum" default="remote">
        `"local"` runs MMseqs2 against a registry-provisioned DB on disk; `"remote"` (the default) queries the ColabFold MSA API over the network and needs no local DB.

        Available options: `local`, `remote`
      </ParamField>

      <ParamField path="dataset" type="enum" default="uniref30-2302">
        Local-only (ignored when remote). Registered key of the searchable reference database; one ColabFold protein DB.

        Available options: `colabfold-envdb-202108`, `uniref30-2302`
      </ParamField>

      <ParamField path="use_metagenomic_db" type="boolean" default="False">
        Include the metagenomic/environmental DB (ColabFoldDB envdb) to deepen unpaired MSAs. Works in both modes; local mode requires the `colabfold-envdb-202108` dataset provisioned. Default `False`. Does not affect cross-chain pairing.
      </ParamField>

      <ParamField path="pairing_strategy" type="enum" default="greedy">
        Cross-chain pairing strategy for paired (multi-chain) groups. `"greedy"` pairs a species found in at least two chains; `"complete"` only pairs a species present in every chain. Ignored for singleton groups.

        Available options: `greedy`, `complete`
      </ParamField>

      <ParamField path="sensitivity" type="number">
        Local-only (ignored when remote). MMseqs2 `-s` override; ignored on a `cuda` device. `None` uses the dataset's registered default.
      </ParamField>

      <ParamField path="num_threads" type="integer">
        Local-only. CPU threads; `None` auto-detects all cores.
      </ParamField>

      <ParamField path="verbose" type="integer" default="0">
        Verbosity level (0=quiet, 1=info, 2=debug, 3=raw subprocess stderr). `True` is coerced to `1` and `False` to `0`.
      </ParamField>

      <ParamField path="device" type="string" default="cpu">
        Local-only (ignored when remote). A `cuda` device runs MMseqs2-GPU and requires a `.idx_pad` index, an NVIDIA GPU (Turing+), and a Linux host; `"cpu"` (the default) runs the CPU pipeline. A specific device such as `"cuda:1"` is honoured.
      </ParamField>

      <ParamField path="timeout" type="integer" default="21600">
        Local subprocess cap, or seconds to wait for a remote submission. Generous by default: the MSA server queues, and a submission carries many queries. `None` waits indefinitely.
      </ParamField>

      <ParamField path="seed" type="integer">
        Random seed. When set, tools run reproducibly up to small GPU float noise (see `BaseToolOutput.approx_equal`), and the seed participates in cache keys. When None, cacheable seed-sensitive tools skip cache until seeded.
      </ParamField>
    </Accordion>
  </div>

  <div class="api-model-section api-output-section">
    <a href="https://github.com/evo-design/proto-tools/blob/47e34afa5ea240a3b406e323dc38aa5dc85f223e/proto_tools/tools/sequence_alignment/mmseqs2/homology_search.py#L239" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a>

    <Accordion title="Output: Mmseqs2HomologySearchOutput">
      <ResponseField name="results" type="List[Mmseqs2HomologySearchResult]" required>
        One result per input group (matches the order of `Mmseqs2HomologySearchInput.queries`).

        <Expandable title="Mmseqs2HomologySearchResult">
          <ResponseField name="sequence_ids" type="List[string]" required>
            Identifiers for the chains in this group.
          </ResponseField>

          <ResponseField name="msas" type="List[MSA]" required>
            Per-chain unpaired MSAs. `None` when no homologs were found beyond the query itself.
          </ResponseField>

          <ResponseField name="paired_msas" type="List[MSA]" required>
            Per-chain taxonomy-paired MSAs, row-aligned across the group's chains. `[None]` for singleton groups (nothing to pair).
          </ResponseField>

          <ResponseField name="datasets_searched" type="List[string]" required>
            Registry keys of datasets hit for this group.
          </ResponseField>

          <ResponseField name="num_homologs_found" type="List[integer]" required>
            Number of homologs per chain (excludes the query itself).
          </ResponseField>
        </Expandable>
      </ResponseField>
    </Accordion>
  </div>

  #### Applications

  This tool is the proto-tools entry point for generating the MSA input to structure-prediction tools. It also drives coevolutionary analyses that identify covarying residue pairs as candidate spatial contacts, conservation analyses that highlight functionally important residues, and homolog mining for protein engineering and design pipelines.

  #### Usage Tips

  * **The `dataset` field selects one registered reference database.** The default is `uniref30-2302`. It is a scalar enum of the searchable ColabFold-style protein databases; non-searchable or non-protein datasets are rejected by validation.
  * **`search_mode="remote"` is the default.** It queries the hosted ColabFold MSA API over the network; `dataset`, `device`, and `sensitivity` are ignored, and no local database or GPU is required. Set `search_mode="local"` to run MMseqs2 against a provisioned on-disk database instead.
  * **Local mode (`search_mode="local"`) runs on CPU unless `device` names a CUDA device.** Set `device="cuda"` (or a specific `cuda:N`) to run MMseqs2-GPU; the configuration validator hard-errors for a CUDA device on a local search on macOS or Windows, since GPU search is Linux-only.
  * **Local mode requires the reference database to be provisioned once before the first call.** Run `python -m proto_tools.tools.sequence_alignment.mmseqs2.setup_databases <dataset>`, where the dataset key matches the value of `Mmseqs2HomologySearchConfig.dataset`. The wrapper does not auto-download databases at call time. (Remote mode skips this entirely.)
  * **Local search needs enough RAM to hold the dataset's sequence database.** When available memory (cgroup-aware) is below that file's size, the tool logs a warning and falls back to a disk-paged (mmap) search that completes but is much slower; for a responsive search, allocate more memory or use `search_mode="remote"`.
  * **Each query produces an `MSA` object or `None`.** Always check `result.msas[i] is not None` before accessing alignment properties. The `num_homologs_found` list returns `0` for queries that produced no homologs. `MSA` objects serialise to A3M or FASTA through the standard export interface.
</div>

## Toolkit Notes

These apply to every MMseqs2 tool in this toolkit (`mmseqs2-search-proteins`, `mmseqs2-search-genomes`, `mmseqs2-clustering`, `mmseqs2-homology-search`).

* **All four tools share a single MMseqs2 installation.** The local installation downloads the GPU-capable MMseqs2 build, which is a strict superset of the CPU-only build and runs CPU subcommands without enabling GPU code paths.

## Infrastructure Guides

The following guides cover how to run tools efficiently and at scale.

<CardGroup cols={2}>
  <Card title="Tool Persistence" icon="repeat" href="/docs/tools/guides/tool-persistence">Keep a tool's model warm across calls instead of reloading it every invocation.</Card>
  <Card title="Device Management" icon="cpu" href="/docs/tools/guides/device-management">How GPUs are allocated to tools and how to target specific devices.</Card>
  <Card title="Parallel Execution" icon="layers" href="/docs/tools/guides/parallel-execution">Fan a batch of inputs out across multiple GPUs.</Card>
</CardGroup>
