Field FavorDistributedInference
Determines whether models load spread across every eligible GPU instead of living on their main device alone.
public static bool FavorDistributedInference
Returns
- bool
trueto spread models across all eligible GPUs;falseto place each model on one device. Default isfalse.
Remarks
When true, a model with no explicit
LM.TensorDistribution proportions has its layers
distributed over all eligible GPUs in proportion to their free
memory (integrated GPUs join only when no discrete device exists).
This pools the devices' memory, so models larger than any single
card load fully on GPU, and lets large prompt batches overlap
across devices when every layer is offloaded.
Layer distribution does not make single-request token generation faster: the devices hold different layers, so a token's forward pass visits them one after another, and the slower card's share is paid on every token. The spread is therefore taken only when the weights do not fit the elected main device whole (with the margins the loader prices): a model that fits one GPU is placed there alone even with this setting on, and the load log says so. An explicit TensorDistribution spreads a model whatever its size.
The value is captured when a model loads; changing it affects subsequent loads, never models already resident.