Skip to main content
Direct preference optimization (DPO) trains a model to favor a preferred response over a rejected response for the same prompt. It is well suited to improving qualities for which multiple responses may be valid but one is more useful, such as style, instruction following, reasoning approach, or tool use. Use supervised fine-tuning instead when you have only one target response per prompt rather than preference pairs.
DPO is currently available for selected models and configurations.

PREFERENCE DATA

Each training record must include a shared prompt, a preferred response in chosen, and a non-preferred response in rejected. The canonical Oumi format represents all three as message lists:
The shared messages must contain a non-empty user message and end with a user or tool message. chosen and rejected must differ, and each must contain exactly one assistant message. An assistant response may contain text, tool calls, or both; tool calls also require top-level tools definitions, and their arguments must be a JSON-encoded object string.

SUPPORTED IMPORT FORMATS

Smart Upload also accepts common Hugging Face TRL preference formats and converts them to the canonical Oumi format. These include explicit string prompts and responses:
The prompt, chosen, and rejected fields may instead contain conversational message lists. Smart Upload also accepts implicit conversational pairs in which chosen and rejected contain the same shared message prefix and differ only in their final assistant message:

DPO HYPERPARAMETERS

Start with the model preset defaults and change one setting at a time. Evaluate the base and trained models on the same task-specific dataset to verify that the target metric improves without unacceptable regressions elsewhere.

BETA

beta controls the strength of the penalty for moving away from the reference model. The default 0.1 strongly steers the model toward the chosen responses; mid-range values such as 0.3 to 0.5 are more conservative, and values of 1.0 or greater are increasingly conservative. beta must be greater than zero.

LOSS TYPE

loss_type selects the preference loss:
  • sigmoid is the standard DPO loss and the default.
  • nca_pair is an alternative to consider if overall generation quality degrades after training with standard sigmoid loss.