DPO is currently available for selected models and configurations.
PREFERENCE DATA
Each training record must include a shared prompt, a preferred response inchosen, and a non-preferred response in rejected. The canonical Oumi format represents all three as message lists:
messages must contain a non-empty user message and end with a user or tool message. chosen and rejected must differ, and each must contain exactly one assistant message. An assistant response may contain text, tool calls, or both; tool calls also require top-level tools definitions, and their arguments must be a JSON-encoded object string.
SUPPORTED IMPORT FORMATS
Smart Upload also accepts common Hugging Face TRL preference formats and converts them to the canonical Oumi format. These include explicit string prompts and responses:prompt, chosen, and rejected fields may instead contain conversational message lists. Smart Upload also accepts implicit conversational pairs in which chosen and rejected contain the same shared message prefix and differ only in their final assistant message:
DPO HYPERPARAMETERS
Start with the model preset defaults and change one setting at a time. Evaluate the base and trained models on the same task-specific dataset to verify that the target metric improves without unacceptable regressions elsewhere.BETA
beta controls the strength of the penalty for moving away from the reference model. The default 0.1 strongly steers the model toward the chosen responses; mid-range values such as 0.3 to 0.5 are more conservative, and values of 1.0 or greater are increasingly conservative. beta must be greater than zero.
LOSS TYPE
loss_type selects the preference loss:
sigmoidis the standard DPO loss and the default.nca_pairis an alternative to consider if overall generation quality degrades after training with standard sigmoid loss.