Property LookupDraftEqualization
LookupDraftEqualization
Chain-length rule for lookup-drafted passes that carry several slots. The default follows the loaded backend. Where the backend replays a captured graph per pass shape (the CUDA backends) it is Longest: every drafting slot verifies the whole depth, a decision that depends on the proposals alone, so a run's drafting is the same from one run to the next, and there the priced rule gains nothing while declining passes a shared device's moving timings mislead it into refusing. Everywhere else it is ExpectedValue: a quantized small-batch product there costs close to one more decode per verified row, so verifying the whole depth on every pass is slower than not drafting at all, and only a rule that prices the rows drafts where it pays. Read on every pass, so a change applies to the next pass without a rebuild. Reading answers with the inherited value while LookupDraftEqualizationOverride is null; assigning sets the override.
public SharedSlotPoolOptions.DraftEqualization LookupDraftEqualization { get; set; }
Property Value
Remarks
On a model carrying recurrent state, every drafting slot in a pass must verify the same number of rows: the recurrent memory decodes unequal per-sequence lengths as separate micro-batches, one per sequence, which costs more than the drafts return. The rows that bring a short proposal up to the pass length are fillers: decoded, verified, rejected and rolled back, each costing a row of decode time for nothing. Attention-only models have no such constraint; on them only ExpectedValue changes anything, by capping the depth when rows are expensive.