Table of Contents

Property LookupDraftEqualization

Namespace
LMKit.Global
Assembly
LM-Kit.NET.dll

LookupDraftEqualization

Chain-length rule for lookup-drafted passes that carry several slots. The default follows the loaded backend. Where the backend replays a captured graph per pass shape (the CUDA backends) it is Longest: every drafting slot verifies the whole depth, a decision that depends on the proposals alone, so a run's drafting is the same from one run to the next, and there the priced rule gains nothing while declining passes a shared device's moving timings mislead it into refusing. Everywhere else it is ExpectedValue: a quantized small-batch product there costs close to one more decode per verified row, so verifying the whole depth on every pass is slower than not drafting at all, and only a rule that prices the rows drafts where it pays. Read on every pass, so a change applies to the next pass without a rebuild. Reading answers with the inherited value while LookupDraftEqualizationOverride is null; assigning sets the override.

public SharedSlotPoolOptions.DraftEqualization LookupDraftEqualization { get; set; }

Property Value

SharedSlotPoolOptions.DraftEqualization

Remarks

On a model carrying recurrent state, every drafting slot in a pass must verify the same number of rows: the recurrent memory decodes unequal per-sequence lengths as separate micro-batches, one per sequence, which costs more than the drafts return. The rows that bring a short proposal up to the pass length are fillers: decoded, verified, rejected and rolled back, each costing a row of decode time for nothing. Attention-only models have no such constraint; on them only ExpectedValue changes anything, by capping the depth when rows are expensive.

Share