Class LoraTrainingParameters
- Namespace
- LMKit.Finetuning
- Assembly
- LM-Kit.NET.dll
Hyperparameters for a LoRA (Low-Rank Adaptation) fine-tuning run.
public sealed class LoraTrainingParameters
- Inheritance
-
LoraTrainingParameters
- Inherited Members
Remarks
The defaults are a sensible starting point for a small instruction fine-tune: rank 8 attention adapters, a cosine-decayed learning rate of 1e-4, and three epochs. Raise Rank and widen TargetModules for tasks that must absorb new knowledge; keep them low for stylistic adaptation.
Properties
- Alpha
LoRA alpha: the adapter's scaling factor. The effective scale applied to the adapter is
alpha / rank. A common rule is alpha = 2 x rank. Default 16.
- AssistantLossOnly
When true, loss is computed only on assistant-role tokens; user and system tokens are masked out. This is the standard setup for instruction fine-tuning and is the default for chat datasets.
- Epochs
Number of passes over the training set. Default 3.
- FirstLayer
First transformer block that receives adapters, inclusive. 0 (the default) starts at the first block. Restricting the trainable range cuts adapter memory and backward compute proportionally; the last blocks carry most task-specific behavior, so a range covering the last third often keeps most of the quality.
- GradientAccumulation
Samples accumulated per optimizer step: gradients from this many samples combine before the weights update, multiplying the effective batch size without growing memory. 1 (the default) steps on every sample.
- InitStandardDeviation
Standard deviation of the Gaussian used to initialize the adapter's A matrix. The B matrix starts at zero so the adapter is a no-op at step zero. Default 0.02.
- LastLayer
Last transformer block that receives adapters, inclusive. 0 (the default) extends to the last block.
- LearningRate
Initial AdamW learning rate. Default 1e-4.
- LoraPlusRatio
LoRA+ learning-rate ratio: the zero-initialized B matrices train at this multiple of LearningRate while the A side keeps the base rate, which speeds convergence at unchanged memory. 16 is the commonly used value; 0 (the default) trains both sides at the base rate.
- MaxGradNorm
Gradient clipping by global norm: right before each optimizer step, every gradient is scaled by
MaxGradNorm / max(MaxGradNorm, norm), wherenormis the global L2 norm over all trainable gradients. Outlier steps are rescaled; ordinary steps pass through unchanged. Default 1.0. Set to 0 to disable clipping.
- MaxPackedSamples
Upper bound on how many samples may share one training window. Each packed sample trains on its own attention sequence with positions restarting at zero, so co-resident samples never attend to each other, and loss never crosses a sample boundary. 1 (the default) gives every sample its own padded window; 0 picks the largest value that divides GradientAccumulation (at most 8); an explicit value must divide GradientAccumulation, so the samples combined per optimizer step stay exactly that setting. Samples carrying images are never packed, and architectures whose batches split per sequence (hybrid and recurrent) always train unpacked.
- MinLearningRate
Floor learning rate for the cosine and linear schedules. Ignored for Constant. Default 0.
- NeftuneAlpha
NEFTune noise alpha. When positive, training adds uniform noise scaled by alpha / sqrt(tokens x embedding width) onto the input embeddings of every training window, a regularizer that counters overfitting on small datasets. Validation passes always run without noise. Typical values are 5 to 15; 0 (the default) disables it.
- Rank
LoRA rank: the inner dimension of each adapter's low-rank pair. Higher ranks add capacity and adapter size. Default 8.
- Schedule
Learning-rate curve. Default Cosine, decaying to MinLearningRate over the run.
- Seed
Seed for reproducible adapter initialization. 0 (the default) means a non-deterministic seed.
- SequencePacking
Packs consecutive short samples into shared training windows. Superseded by MaxPackedSamples, which additionally isolates the packed samples' attention; true maps to the automatic packing bound and false to unpacked training.
- TargetModules
Which base-model weight matrices receive adapters. Default Attention.
- UseRsLora
Rank-stabilized LoRA: scales the adapter by alpha divided by the square root of the rank instead of alpha divided by the rank, which keeps higher ranks from training with a vanishing effective scale. The scaling is folded into the saved adapter, so artifacts stay compatible with every consumer.
- ValidationSplit
Fraction of samples held out for per-epoch validation. Clamped to [0, 0.5]. Default 0.05.
- WarmupRatio
Fraction of total steps spent linearly warming the learning rate up from zero before the schedule begins. 0 (the default) disables warmup. Clamped to [0, 1].
- WeightDecay
AdamW weight decay. Default 1e-2.