PORT: Preference Optimization via Robust Token-Level Reweighting
Published in The Conference on Neural Information Processing Systems (NeurIPS), 2026
Recommended citation: D. Zhu, X. Wei, T. Xie, Z. Zhu, X. Zhang, M. Khalili, "PORT: Preference Optimization via Robust Token-Level Reweighting," The Conference on Neural Information Processing Systems (NeurIPS), 2026.
Ding Zhu, Xiukun Wei, Tian Xie, Zhihui Zhu, Xueru Zhang, Mohammad Mahdi Khalili
Paper (coming soon) · Code · BibTeX
Overview of PORT. Token logits from a single forward pass are pooled across the batch into a global empirical CDF. The per-token gap between the rejected and chosen percentile ranks, Δ = Fl − Fw, is mapped to a weight w = γ exp(ρΔ), and these weights rescale the token-level DPO signal so that likely-corrupted tokens contribute less to the gradient.
Abstract
Preference optimization has become a central approach for aligning large language models with human values, but its effectiveness depends heavily on the quality of preference annotations. In practice, preference data is often noisy due to annotation errors and ambiguity. Existing robust preference optimization methods primarily operate at the sequence level, implicitly treating entire responses as uniformly correct or incorrect. However, rejected responses may still contain informative reasoning steps, while preferred responses can include subtle errors. In this work, we propose Preference Optimization via Robust Token-level reweighting (PORT), a fine-grained framework for robust alignment under noisy preferences. PORT performs token-level reweighting using the empirical cumulative distribution function (CDF) of token logits, yielding an efficient forward-pass-only proxy to gradient-norm penalization that selectively suppresses corrupted tokens without additional backward-pass computation. We provide theoretical analysis showing that PORT reduces the gradient bias from noisy preference labels by minimizing an upper bound on the bias term. Extensive experiments across diverse noise settings demonstrate that PORT consistently improves robustness and outperforms existing sequence-level baselines.
Why token level?
Consider a pair where the preferred response is “The capital of Australia is Canberra, which was purpose-built as a compromise between Sydney and Melbourne” and the rejected one is “The capital of Australia is Sydney, as it is the largest city.” If the label is flipped by noise, sequence-level methods penalize every token of the response equally, even though most tokens (e.g., “as it is the largest city”) are fluent and not responsible for the error. Label smoothing (cDPO) shrinks the whole objective uniformly, pair-level reweighting (Dr.DPO) gives one scalar per pair, and token-level importance sampling (TIS-DPO) needs an external reward model. PORT adapts at the token level with no external supervision.
Key contributions
- Token-level CDF reweighting. A robust alignment objective that needs no external reward model and adds zero backward-pass overhead.
- Theory. The weighting scheme minimizes an upper bound on the gradient bias under instance-dependent preference noise, with conditions under which it strictly reduces the bias for both mislabeled and uncertain noise.
- Experiments. Consistent gains over standard, sequence-level robust, and filtering-based baselines on AlpacaFarm, Anthropic HH, and PKU-Beaver with LLaMA2-7B and Qwen3-3B.
Results
AlpacaFarm, preference labels flipped with probability 0.3:
| Method | LLaMA2-7B Gold RM ↑ | LLaMA2-7B Win Rate (%) ↑ | Qwen3-3B Gold RM ↑ | Qwen3-3B Win Rate (%) ↑ |
|---|---|---|---|---|
| DPO | 2.66 | 56.8 | 2.18 | 46.3 |
| SamS | 3.68 | 64.8 | 2.84 | 53.8 |
| β-DPO | 3.42 | 63.2 | 2.73 | 50.2 |
| Dr.DPO | 3.39 | 61.6 | 2.69 | 52.5 |
| PORT | 3.82 | 67.6 | 2.98 | 53.6 |
Anthropic HH, safety rate (PKU-Beaver) as label-flip probability p grows:
| Method | p = 0.2 | p = 0.3 | p = 0.4 |
|---|---|---|---|
| DPO | 0.45 | 0.42 | 0.38 |
| SamS | 0.53 | 0.53 | 0.46 |
| β-DPO | 0.54 | 0.51 | 0.47 |
| PORT | 0.56 | 0.55 | 0.53 |
PORT’s advantage grows with noise, reaching +0.06 over the strongest baseline at p = 0.4.
BibTeX
@inproceedings{zhu2026port,
title = {{PORT}: Preference Optimization via Robust Token-Level Reweighting},
author = {Zhu, Ding and Wei, Xiukun and Xie, Tian and Zhu, Zhihui and Zhang, Xueru and Khalili, Mohammad Mahdi},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026}
}
