Table of Contents

Class TrainingDataset

Namespace
LMKit.Finetuning
Assembly
LM-Kit.NET.dll

A supervised fine-tuning dataset: a collection of chat conversations (ChatTrainingSample) and optional raw-text samples, with loaders for the common industry formats and ZIP archives.

public class TrainingDataset
Inheritance
TrainingDataset
Derived
Inherited Members

Remarks

Load from files with Load(LM, string) (format auto-detected) or a specific loader, feed the whole set to a run with AddDataset(TrainingDataset), and export back to ShareGPT with ExportAsSharegpt(string, bool, string, CancellationToken).

Supported formats:

  • Chat JSONL: one object per line, {"messages":[{"role","content"}]} (the OpenAI fine-tuning shape).
  • ShareGPT: a JSON array of {"conversations":[{"from","value"}]}.
  • Alpaca: a JSON array of {"instruction","input","output"}.
  • Plain text: a corpus split into samples by a delimiter, for continued-pretraining-style data.
  • ZIP archive: any mix of the above bundled together; every dataset file inside is loaded and merged.

Constructors

TrainingDataset()

Initializes a new empty dataset.

Properties

Count

Total number of samples (chat plus raw text).

RawTextSamples

The raw-text samples in this dataset (continued-pretraining style).

ResolvedImageCount

Number of image references that resolved to image data.

Samples

The chat conversations in this dataset.

SkipSummary

One-sentence description of what the load skipped and why, or null when nothing was skipped. Names the supported shapes so a rejected dataset is fixable from the message alone.

SkippedEntryCount

Number of dataset entries the load skipped: lines that were not valid JSON, and JSON entries that matched no supported shape. A dataset that loads zero samples with a non-zero skip count is a format problem, not an empty file; consumers surface SkipSummary so the error names the cause.

UnresolvedImageCount

Number of image references that could not be resolved. Relative paths resolve beside the dataset file or inside its ZIP archive; data URIs and inline base64 resolve anywhere. A dataset whose images all fail to resolve would silently train text-only, so consumers surface this count instead.

UnresolvedImageRefs

A few examples of unresolved image references, for error reporting.

Methods

AddRawText(string)

Adds a raw-text sample (every token supervised).

AddSample(ChatTrainingSample)

Adds a chat training sample.

ExportAsSharegpt(string, bool, string, CancellationToken)

Exports the chat samples as a ShareGPT-compatible JSON file.

Load(LM, Stream, TrainingDatasetFormat, string)

Loads a dataset from a stream with an explicit format.

Load(LM, string)

Loads a dataset from a file, auto-detecting the format from its extension and content. Supports .jsonl, .json (ShareGPT, Alpaca, chat, or prompt/completion pairs), .txt, and .zip.

Share