> ## Documentation Index
> Fetch the complete documentation index at: https://proto.evodesign.org/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Gradient Optimizer

> Continuous, differentiable optimization: represents the sequence as per-position residue logits and descends them with SGD or Adam using gradients backpropagated through differentiable constraints, annealing from a soft probabilistic sequence to a hard discrete one.

<div class="page-hero">
  <div className="block dark:hidden">
    <svg viewBox="0 0 720 432" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Gradient optimizer: loss contours descended by a gradient trajectory to the minimum" style={{width:"100%",height:"auto",display:"block"}}><defs><pattern id="gridgradL" width="22" height="22" patternUnits="userSpaceOnUse"><circle cx="2" cy="2" r="1.2" fill="#344649" fillOpacity="0.10" /></pattern><marker id="arrgradL" viewBox="0 0 10 10" refX="8.5" refY="5" markerWidth="6.5" markerHeight="6.5" orient="auto-start-reverse"><path d="M0,0 L10,5 L0,10 L3,5 z" fill="#768b8e" /></marker></defs><rect x="12" y="12" width="696" height="408" rx="16" fill="#f9fcfc" stroke="#dee9e8" strokeWidth="1.2" /><rect x="12" y="12" width="696" height="408" rx="16" fill="url(#gridgradL)" /><text x="360" y="56" fontFamily="Geist, ui-sans-serif, system-ui, -apple-system, sans-serif" fontSize="13.5" fontWeight="600" fill="#1d2c2f" textAnchor="middle">each step:  backprop ∇ → merge → update logits</text><ellipse cx="340" cy="248" rx="46" ry="26" fill="#046e7a" fillOpacity="0.06" stroke="#9eb4b2" strokeWidth="1.2" /><ellipse cx="340" cy="248" rx="94" ry="54" fill="none" stroke="#9eb4b2" strokeWidth="1.2" /><ellipse cx="340" cy="248" rx="144" ry="84" fill="none" stroke="#9eb4b2" strokeWidth="1.2" /><ellipse cx="340" cy="248" rx="196" ry="114" fill="none" stroke="#9eb4b2" strokeWidth="1.2" /><ellipse cx="340" cy="248" rx="248" ry="144" fill="none" stroke="#9eb4b2" strokeWidth="1.2" /><text x="340" y="120" fontFamily="Geist, ui-sans-serif, system-ui, -apple-system, sans-serif" fontSize="10.5" fontWeight="400" fill="#768b8e" textAnchor="middle">loss contours</text><path d="M120,250 L190,232" fill="none" stroke="#046e7a" strokeWidth="2.4" /><path d="M0,0 L-6.5,-3.575 L-6.5,3.575 Z" fill="#046e7a" transform="translate(190,232) rotate(-14.42)" /><path d="M190,232 L236,272" fill="none" stroke="#046e7a" strokeWidth="2.4" /><path d="M0,0 L-6.5,-3.575 L-6.5,3.575 Z" fill="#046e7a" transform="translate(236,272) rotate(41.01)" /><path d="M236,272 L282,238" fill="none" stroke="#046e7a" strokeWidth="2.4" /><path d="M0,0 L-6.5,-3.575 L-6.5,3.575 Z" fill="#046e7a" transform="translate(282,238) rotate(-36.47)" /><path d="M282,238 L316,262" fill="none" stroke="#046e7a" strokeWidth="2.4" /><path d="M0,0 L-6.5,-3.575 L-6.5,3.575 Z" fill="#046e7a" transform="translate(316,262) rotate(35.22)" /><path d="M316,262 L340,248" fill="none" stroke="#046e7a" strokeWidth="2.4" /><path d="M0,0 L-6.5,-3.575 L-6.5,3.575 Z" fill="#046e7a" transform="translate(340,248) rotate(-30.26)" /><circle cx="340" cy="248" r="7" fill="#2f8f6b" /><path d="M245,396 l20,0" fill="none" stroke="#046e7a" strokeWidth="2.6" /><text x="272" y="400" fontFamily="Geist, ui-sans-serif, system-ui, -apple-system, sans-serif" fontSize="11" fontWeight="400" fill="#768b8e" textAnchor="start">gradient step</text><circle cx="396" cy="396" r="5.5" fill="#2f8f6b" /><text x="412" y="400" fontFamily="Geist, ui-sans-serif, system-ui, -apple-system, sans-serif" fontSize="11" fontWeight="400" fill="#768b8e" textAnchor="start">minimum</text></svg>
  </div>

  <div className="hidden dark:block">
    <svg viewBox="0 0 720 432" xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Gradient optimizer: loss contours descended by a gradient trajectory to the minimum" style={{width:"100%",height:"auto",display:"block"}}><defs><pattern id="gridgradD" width="22" height="22" patternUnits="userSpaceOnUse"><circle cx="2" cy="2" r="1.2" fill="#9eb4b2" fillOpacity="0.10" /></pattern><marker id="arrgradD" viewBox="0 0 10 10" refX="8.5" refY="5" markerWidth="6.5" markerHeight="6.5" orient="auto-start-reverse"><path d="M0,0 L10,5 L0,10 L3,5 z" fill="#7e9498" /></marker></defs><rect x="12" y="12" width="696" height="408" rx="16" fill="#0e1718" stroke="#2b3c3e" strokeWidth="1.2" /><rect x="12" y="12" width="696" height="408" rx="16" fill="url(#gridgradD)" /><text x="360" y="56" fontFamily="Geist, ui-sans-serif, system-ui, -apple-system, sans-serif" fontSize="13.5" fontWeight="600" fill="#eef5f4" textAnchor="middle">each step:  backprop ∇ → merge → update logits</text><ellipse cx="340" cy="248" rx="46" ry="26" fill="#0a7e8c" fillOpacity="0.06" stroke="#566b6e" strokeWidth="1.2" /><ellipse cx="340" cy="248" rx="94" ry="54" fill="none" stroke="#566b6e" strokeWidth="1.2" /><ellipse cx="340" cy="248" rx="144" ry="84" fill="none" stroke="#566b6e" strokeWidth="1.2" /><ellipse cx="340" cy="248" rx="196" ry="114" fill="none" stroke="#566b6e" strokeWidth="1.2" /><ellipse cx="340" cy="248" rx="248" ry="144" fill="none" stroke="#566b6e" strokeWidth="1.2" /><text x="340" y="120" fontFamily="Geist, ui-sans-serif, system-ui, -apple-system, sans-serif" fontSize="10.5" fontWeight="400" fill="#9eb4b2" textAnchor="middle">loss contours</text><path d="M120,250 L190,232" fill="none" stroke="#0a7e8c" strokeWidth="2.4" /><path d="M0,0 L-6.5,-3.575 L-6.5,3.575 Z" fill="#0a7e8c" transform="translate(190,232) rotate(-14.42)" /><path d="M190,232 L236,272" fill="none" stroke="#0a7e8c" strokeWidth="2.4" /><path d="M0,0 L-6.5,-3.575 L-6.5,3.575 Z" fill="#0a7e8c" transform="translate(236,272) rotate(41.01)" /><path d="M236,272 L282,238" fill="none" stroke="#0a7e8c" strokeWidth="2.4" /><path d="M0,0 L-6.5,-3.575 L-6.5,3.575 Z" fill="#0a7e8c" transform="translate(282,238) rotate(-36.47)" /><path d="M282,238 L316,262" fill="none" stroke="#0a7e8c" strokeWidth="2.4" /><path d="M0,0 L-6.5,-3.575 L-6.5,3.575 Z" fill="#0a7e8c" transform="translate(316,262) rotate(35.22)" /><path d="M316,262 L340,248" fill="none" stroke="#0a7e8c" strokeWidth="2.4" /><path d="M0,0 L-6.5,-3.575 L-6.5,3.575 Z" fill="#0a7e8c" transform="translate(340,248) rotate(-30.26)" /><circle cx="340" cy="248" r="7" fill="#2f8f6b" /><path d="M245,396 l20,0" fill="none" stroke="#0a7e8c" strokeWidth="2.6" /><text x="272" y="400" fontFamily="Geist, ui-sans-serif, system-ui, -apple-system, sans-serif" fontSize="11" fontWeight="400" fill="#9eb4b2" textAnchor="start">gradient step</text><circle cx="396" cy="396" r="5.5" fill="#2f8f6b" /><text x="412" y="400" fontFamily="Geist, ui-sans-serif, system-ui, -apple-system, sans-serif" fontSize="11" fontWeight="400" fill="#9eb4b2" textAnchor="start">minimum</text></svg>
  </div>
</div>

<p class="entity-disclaimer">This optimizer is open source. Any third-party models, product names, or trademarks referenced are the property of their respective owners, and Proto is not affiliated with them.</p>

<hr class="entity-rule" />

<div class="tool-tab-bar entity-source-bar"><span class="tool-tab-wrap"><span class="tool-tab badge-source entity-source-tab"><svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</span></span></div>

<a href="https://github.com/evo-design/proto-language/blob/d3b7822f74ea64747cc751a3b2ab1aa6b799ac47/proto_language/optimizer/gradient_optimizer.py#L415" target="_blank" class="tab-panel source-panel entity-source-panel">
  <div class="source-info">
    <img noZoom src="https://github.com/evo-design.png?size=40" class="source-avatar" width="36" height="36" />

    <span class="source-path">evo-design/proto-language<span class="source-subpath">/proto\_language/optimizer/gradient\_optimizer.py</span></span>
  </div>

  <span class="panel-goto-btn source-goto-btn"><span><svg width="14" height="14" viewBox="0 0 24 24" fill="currentColor"><path d="M12 0C5.37 0 0 5.37 0 12c0 5.31 3.435 9.795 8.205 11.385.6.105.825-.255.825-.57 0-.285-.015-1.23-.015-2.235-3.015.555-3.795-.735-4.035-1.41-.135-.345-.72-1.41-1.23-1.695-.42-.225-1.02-.78-.015-.795.945-.015 1.62.87 1.845 1.23 1.08 1.815 2.805 1.305 3.495.99.105-.78.42-1.305.765-1.605-2.67-.3-5.46-1.335-5.46-5.925 0-1.305.465-2.385 1.23-3.225-.12-.3-.54-1.53.12-3.18 0 0 1.005-.315 3.3 1.23.96-.27 1.98-.405 3-.405s2.04.135 3 .405c2.295-1.56 3.3-1.23 3.3-1.23.66 1.65.24 2.88.12 3.18.765.84 1.23 1.905 1.23 3.225 0 4.605-2.805 5.625-5.475 5.925.435.375.81 1.095.81 2.22 0 1.605-.015 2.895-.015 3.3 0 .315.225.69.825.57A12.02 12.02 0 0024 12c0-6.63-5.37-12-12-12z" /></svg> View source</span></span>
</a>

<div class="entity-contributors"><span class="entity-contributors-label">Optimizer contributors</span><span class="entity-contributors-people"><a class="entity-contributor" href="https://github.com/dguo8412" target="_blank" rel="noopener" title="dguo8412: 6 commits"><img noZoom class="entity-contributor-avatar" src="https://avatars.githubusercontent.com/u/46211285?v=4&s=64" alt="" loading="lazy" /><span class="entity-contributor-login">dguo8412</span></a></span></div>
Continuous gradient descent on per-position logits of a single segment.

The optimization variable is `seq.logits`: an `(L, |vocab|)` matrix carried on
each of `num_results` parallel proposal Sequences (one independent trajectory per
result). The companion `PositionWeightGenerator` is the only thing that maps those
continuous logits back to a discrete sequence — it never proposes; it merely decodes
(argmax or categorical) at tracked steps so snapshots, `result_sequences`, and the
next-stage handoff carry a real sequence. Logits are seeded once (zeros, or
`initial_logits`/`sequence_bias`, optionally plus per-trajectory Gumbel noise so
parallel trajectories diverge) and then mutated in place every step; they are never
re-proposed.

Each step: (1) interpolate `soft` (relax↔hard blend), `hard` (straight-through),
softmax `temperature` and learning rate from their start→end configs/schedules with
`progress = step / num_steps`; (2) ask every compiled gradient provider to backpropagate
its differentiable (gradient-mode or compiler-backed) constraint into a per-trajectory logit
gradient and per-trajectory loss, given the current `temperature`/`soft`/`hard` and the
step's effective weight; non-finite gradients raise. Then per trajectory: (3) align per-constraint
gradient norms (`norm_alignment`), scale each by its effective weight, and merge them with the
configured `merger` (`weighted_sum`/`pcgrad`/`mgda`); (4) zero `fixed_positions` and
optionally normalize the merged gradient; (5) take one `ml_optimizer` step (SGD or Adam) at the
effective learning rate (`_effective_lr` optionally scales it by `(1 - soft) + soft * temp`,
floored at `min_lr_scale`). After updating, `energy_scores` is set to the summed weighted
constraint losses — the exact objective being minimized. At `tracking_interval` steps (and the
last step) the generator decodes logits, proposals sync to results, and a snapshot is saved.

Per-constraint weights can ramp over steps via `constraint_weight_schedules` (a
`ConstraintWeightSchedule` keyed by `Constraint.label`); unknown labels warn and are
ignored. With `save_best=True` (default) the lowest-loss logits per trajectory are restored
and re-decoded at the end instead of returning the final step. Constraints: single target
segment only; exactly one `PositionWeightGenerator`; every constraint must support gradient
evaluation. Chain stages in a `Program` for multi-phase pipelines (logit-relaxation phase
via `germinal_logit_preset` → softmax-annealing phase via `germinal_softmax_preset`).

## How It Works

The gradient optimizer relaxes the sequence into continuous logits and takes gradient steps that lower the constraint loss, sharpening the relaxation from soft to hard before decoding.

The discrete sequence is relaxed into a continuous logit matrix `L×|V|`, one per trajectory. Each step sharpens a softmax relaxation, backpropagates every differentiable constraint into a per-trajectory gradient, merges them, and applies an SGD or Adam update:

```
progress = step / num_steps
soft, hard, τ   interpolate  start → end  with progress
E_k = Σ_i  w_i(step) · L_i(logits_k)          (weighted constraint losses)
logits_k ← Optimizer(logits_k, ∇E_k, lr)      (SGD or Adam)
```

Gradients merge by `weighted_sum`, `pcgrad`, or `mgda`; `fixed_positions` stay frozen (gradient set to 0). With `save_best`, the lowest-loss logits seen across all steps are decoded at the end through the `PositionWeightGenerator` (argmax by default, or categorical sampling).

## API Reference

<div class="api-model-section api-model-static api-config-section">
  <div class="api-model-header"><span class="api-model-badge api-config-badge">Config</span><span class="api-model-name">GradientOptimizerConfig</span><a href="https://github.com/evo-design/proto-language/blob/d3b7822f74ea64747cc751a3b2ab1aa6b799ac47/proto_language/optimizer/gradient_optimizer.py#L112" target="_blank" class="func-table-btn func-source-btn api-model-source"><svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="16 18 22 12 16 6" /><polyline points="8 6 2 12 8 18" /></svg> Source</a></div>

  Configuration for gradient-based sequence optimization.

  Each GradientOptimizer runs one mode (fixed or ramping soft, with optional
  temperature annealing). Chain multiple in a `Program` for multi-phase
  pipelines (e.g. logit phase → softmax phase).

  <Note>
    Ramps use `progress = step / num_steps` with `step` starting at 1,
    so step 1 evaluates to `start + (end - start) / num_steps` (not exactly
    `start`); step `num_steps` evaluates exactly to `end`.
  </Note>

  <ParamField path="num_results" type="integer">
    Candidate designs for this optimizer. Overrides program-level count.
  </ParamField>

  <ParamField path="num_steps" type="integer" default="1">
    Number of gradient descent steps.
  </ParamField>

  <ParamField path="lr" type="number" default="0.05">
    Base learning rate for gradient updates.
  </ParamField>

  <ParamField path="sequence_bias" type="SequenceLogitBiasConfig">
    Per-position logit bias for the target vocabulary; added to initial logits to seed the search.
  </ParamField>

  <ParamField path="soft_start" type="number" default="1.0">
    Soft sampling weight at the first step. 0 uses hard logits; 1 uses the full softmax over logits.
  </ParamField>

  <ParamField path="soft_end" type="number" default="1.0">
    Soft sampling weight at the final step. 0 uses hard logits; 1 uses the full softmax.
  </ParamField>

  <ParamField path="hard_start" type="number" default="0.0">
    Straight-through blend at step 1. 0 is fully relaxed; 1 is argmax forward + relaxed gradient.
  </ParamField>

  <ParamField path="hard_end" type="number" default="0.0">
    Straight-through blend at the final step. 0 is fully relaxed; 1 = argmax forward + relaxed grad.
  </ParamField>

  <ParamField path="temperature_start" type="number" default="1.0">
    Softmax temperature at the first step. Lower values produce sharper distributions.
  </ParamField>

  <ParamField path="temperature_end" type="number" default="1.0">
    Softmax temperature at the final step. Lower values produce sharper distributions.
  </ParamField>

  <ParamField path="softmax_schedule" type="enum" default="constant">
    Curve interpolating the softmax temperature from start to end across optimization steps.

    Options: `constant`, `cosine`, `exponential`, `hinge`, `linear`, `quadratic`
  </ParamField>

  <ParamField path="lr_schedule" type="enum" default="constant">
    LR curve over the temperature endpoints; only active when scale\_lr\_by\_temperature=True.

    Options: `constant`, `cosine`, `exponential`, `hinge`, `linear`, `quadratic`
  </ParamField>

  <ParamField path="merger" type="enum" default="weighted_sum">
    Strategy for merging gradients from multiple constraints.

    Options: `weighted_sum`, `pcgrad`, `mgda`
  </ParamField>

  <ParamField path="ml_optimizer" type="enum" default="sgd">
    Gradient update rule applied each step. Currently 'sgd' or 'adam'.

    Options: `sgd`, `adam`
  </ParamField>

  <ParamField path="adam_config" type="AdamConfig">
    Beta and epsilon parameters used when the update algorithm is 'adam'.
  </ParamField>

  <ParamField path="norm_alignment" type="enum" default="none">
    How per-constraint gradients are rescaled before merging: as-is, unit-normalized, or match-first.

    Options: `none`, `unit`, `match_first`
  </ParamField>

  <ParamField path="zero_norm_eps" type="number" default="0.0">
    In match\_first mode, zero out gradients with norm below this threshold.
  </ParamField>

  <ParamField path="normalize_gradients" type="boolean" default="True">
    Normalize the merged gradient before each update.
  </ParamField>

  <ParamField path="normalize_mode" type="enum" default="unit">
    'unit' rescales the gradient to unit L2 norm; 'sqrt\_length' scales magnitude by sqrt(length).

    Options: `unit`, `sqrt_length`
  </ParamField>

  <ParamField path="fixed_positions" type="array">
    Zero-based positions to freeze during optimization. Pair with sequence\_bias to anchor each position.
  </ParamField>

  <ParamField path="scale_lr_by_temperature" type="boolean" default="False">
    Multiply LR by a blend of soft weight and softmax temperature; slows updates as sharpness rises.
  </ParamField>

  <ParamField path="min_lr_scale" type="number" default="0.0">
    Lower bound on the learning-rate scale factor when temperature scaling is enabled.
  </ParamField>

  <ParamField path="save_best" type="boolean" default="True">
    Return the lowest-loss result instead of the last iteration.
  </ParamField>

  <ParamField path="constraint_weight_schedules" type="array">
    Per-constraint weight schedules that override the constraint's static weight at each step.
  </ParamField>

  <ParamField path="gumbel_logit_init" type="boolean" default="False">
    Add Gumbel noise to default-init logits (frozen positions excluded) to diverge trajectories.
  </ParamField>

  <ParamField path="gumbel_init_alpha" type="number" default="1.0">
    Divisor for the default-path Gumbel init noise. 1.0 = unscaled; larger shrinks it.
  </ParamField>

  <ParamField path="initial_logits" type="array">
    Base logit matrix (rows=positions, cols=vocab) that replaces default initialization.
  </ParamField>

  <ParamField path="softmax_init_positions" type="array">
    Zero-based positions perturbed with Gumbel noise and passed through a softmax over initial logits.
  </ParamField>

  <ParamField path="seed" type="integer">
    Random seed for reproducible optimization, generator, and constraint tool streams.
  </ParamField>

  <ParamField path="tracking_interval" type="integer" default="1">
    Save history and log progress every N steps. Step 0 and final step always saved.
  </ParamField>

  <ParamField path="track_proposals" type="boolean" default="False">
    Save granular per-proposal results (accept/reject) in history snapshots.
  </ParamField>

  <ParamField path="verbose" type="boolean" default="False">
    Emit per-step debug information about proposals, scores, and acceptance through the logger.
  </ParamField>
</div>

## Usage

```python python icon="python" theme={null}
>>> from proto_language.constraint import MalinoisActivityConfig, malinois_activity_constraint
>>> from proto_language.core import Constraint, Construct, Program, Segment
>>> from proto_language.generator import PositionWeightGenerator, PositionWeightGeneratorConfig
>>> from proto_language.optimizer import GradientOptimizer, GradientOptimizerConfig
>>> seg = Segment(sequence="A" * 200, sequence_type="dna", label="enhancer")
>>> gen = PositionWeightGenerator(PositionWeightGeneratorConfig(sampling_mode="argmax"))
>>> gen.assign(seg)
>>> on_target = Constraint(  # differentiable path
...     inputs=[seg],
...     function=malinois_activity_constraint,
...     function_config=MalinoisActivityConfig(cell_type="K562", direction="max"),
...     label="malinois_k562_max",
...     weight=1.0,
... )
>>> off_target = Constraint(
...     inputs=[seg],
...     function=malinois_activity_constraint,
...     function_config=MalinoisActivityConfig(cell_type="HepG2", direction="min"),
...     label="malinois_hepg2_min",
...     weight=1.0,
... )
>>> optimizer = GradientOptimizer(
...     target_segment=seg,
...     constructs=[Construct([seg])],
...     generators=[gen],
...     constraints=[on_target, off_target],
...     config=GradientOptimizerConfig(
...         num_results=20,
...         num_steps=300,
...         lr=0.5,
...         ml_optimizer="adam",
...         merger="weighted_sum",
...         gumbel_logit_init=True,  # diverge the 20 trajectories
...     ),
... )
>>> program = Program([optimizer], num_results=20)
>>> # program.run()  # needs GPU
```

## Metadata

| Property                 | Value               |
| ------------------------ | ------------------- |
| Key                      | `gradient`          |
| Class                    | `GradientOptimizer` |
| Targets Single Segment   | `True`              |
| Uses GPU                 | `False`             |
| Required Constraint Mode | `gradient`          |
| Compatible Generators    | `position-weight`   |
