> ## Documentation Index
> Fetch the complete documentation index at: https://docs.oumi.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# DIRECT PREFERENCE OPTIMIZATION

> Train models from preferred and rejected responses

Direct preference optimization (DPO) trains a model to favor a preferred response over a rejected response for the same prompt. It is well suited to improving qualities for which multiple responses may be valid but one is more useful, such as style, instruction following, reasoning approach, or tool use. Use supervised fine-tuning instead when you have only one target response per prompt rather than preference pairs.

<Note>DPO is currently available for selected models and configurations.</Note>

***

## PREFERENCE DATA

Each training record must include a shared prompt, a preferred response in `chosen`, and a non-preferred response in `rejected`. The canonical Oumi format represents all three as message lists:

```json theme={null}
{
  "messages": [
    {"role": "user", "content": "Name a noteworthy city in France."}
  ],
  "chosen": [
    {"role": "assistant", "content": "Paris."}
  ],
  "rejected": [
    {"role": "assistant", "content": "Lyon."}
  ]
}
```

The shared `messages` must contain a non-empty user message and end with a user or tool message. `chosen` and `rejected` must differ, and each must contain exactly one assistant message. An assistant response may contain text, tool calls, or both; tool calls also require top-level `tools` definitions, and their `arguments` must be a JSON-encoded object string.

### SUPPORTED IMPORT FORMATS

Smart Upload also accepts common [Hugging Face TRL preference formats](https://huggingface.co/docs/trl/main/en/dataset_formats#preference) and converts them to the canonical Oumi format. These include explicit string prompts and responses:

```json theme={null}
{
  "prompt": "Name a noteworthy city in France.",
  "chosen": "Paris.",
  "rejected": "Lyon."
}
```

The `prompt`, `chosen`, and `rejected` fields may instead contain conversational message lists. Smart Upload also accepts implicit conversational pairs in which `chosen` and `rejected` contain the same shared message prefix and differ only in their final assistant message:

```json theme={null}
{
  "chosen": [
    {"role": "user", "content": "Name a noteworthy city in France."},
    {"role": "assistant", "content": "Paris."}
  ],
  "rejected": [
    {"role": "user", "content": "Name a noteworthy city in France."},
    {"role": "assistant", "content": "Lyon."}
  ]
}
```

***

## DPO HYPERPARAMETERS

Start with the model preset defaults and change one setting at a time. Evaluate the base and trained models on the same task-specific dataset to verify that the target metric improves without unacceptable regressions elsewhere.

### BETA

`beta` controls the strength of the penalty for moving away from the reference model. The default `0.1` strongly steers the model toward the chosen responses; mid-range values such as `0.3` to `0.5` are more conservative, and values of `1.0` or greater are increasingly conservative. `beta` must be greater than zero.

### LOSS TYPE

`loss_type` selects the preference loss:

* `sigmoid` is the standard DPO loss and the default.
* `nca_pair` is an alternative to consider if overall generation quality degrades after training with standard sigmoid loss.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.