Papers
Topics
Authors
Recent
Search
2000 character limit reached

CURATRON: Complete and Robust Preference Data for Rigorous Alignment of Large Language Models

Published 5 Mar 2024 in cs.AI and cs.CL | (2403.02745v2)

Abstract: This paper addresses the challenges of aligning LLMs with human values via preference learning (PL), focusing on incomplete and corrupted data in preference datasets. We propose a novel method for robustly and completely recalibrating values within these datasets to enhance LLMs' resilience against the issues. In particular, we devise a guaranteed polynomial time ranking algorithm that robustifies several existing models, such as the classic Bradley-Terry-Luce (BTL) (Bradley and Terry, 1952) model and certain generalizations of it. To the best of our knowledge, our present work is the first to propose an algorithm that provably recovers an $\epsilon$-optimal ranking with high probability while allowing as large as $O(n)$ perturbed pairwise comparison results per model response. Furthermore, we show robust recovery results in the partially observed setting. Our experiments confirm that our algorithms handle adversarial noise and unobserved comparisons well in both general and LLM preference dataset settings. This work contributes to the development and scaling of more reliable and ethically aligned AI models by equipping the dataset curation pipeline with the ability to handle missing and maliciously manipulated inputs.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (75)
  1. 01.ai (2023). Yi-34b. https://www.01.ai. Accessed 03-03-2024.
  2. Towards a human-like open-domain chatbot. CoRR, abs/2001.09977.
  3. The falcon series of open language models.
  4. Anthropic (2023). Introducing Claude. https://www.anthropic.com/news/introducing-claude. Accessed 03-03-2024.
  5. A general theoretical paradigm to understand learning from human preferences.
  6. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  7. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324–345.
  8. Sparks of Artificial General Intelligence: Early experiments with GPT-4.
  9. Weak-to-strong generalization: Eliciting strong capabilities with weak supervision.
  10. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint arXiv:2307.15217.
  11. Self-play fine-tuning converts weak language models to strong language models.
  12. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. https://lmsys.org/blog/2023-03-30-vicuna/. Accessed 03-03-2024.
  13. Provably robust dpo: Aligning language models with noisy feedback.
  14. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30.
  15. Crowd control: Effectively utilizing unscreened crowd workers for biomedical data annotation. Journal of Biomedical Informatics, 69:86–92.
  16. Copeland, A. H. (1951). A reasonable social welfare function. In Mimeographed notes from a Seminar on Applications of Mathematics to the Social Sciences, University of Michigan.
  17. Ordering by weighted number of wins gives a good ranking for weighted tournaments. In Proceedings of the seventeenth annual ACM-SIAM symposium on Discrete algorithm, pages 776–782. Society for Industrial and Applied Mathematics.
  18. Cucuringu, M. (2016). Sync-rank: Robust ranking, constrained ranking and rank aggregation via eigenvector and sdp synchronization. IEEE Transactions on Network Science and Engineering, 3(1):58–79.
  19. Quality control in crowdsourcing: A survey of quality attributes, assessment techniques and assurance actions. CoRR, abs/1801.02546.
  20. A checklist to combat cognitive biases in crowdsourcing.
  21. GPTs are GPTs: An early look at the labor market impact potential of Large Language Models.
  22. Human-centered loss functions (HALOs). Technical report, Contextual AI. https://github.com/ContextualAI/HALOs/blob/main/assets/report.pdf.
  23. Fishburn, P. C. (1973). Binary choice probabilities: on the varieties of stochastic transitivity. Journal of Mathematical psychology, 10(4):327–352.
  24. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  25. Gordon, C. (2024). Google Pauses Gemini AI Model After Latest Debacle — forbes.com. https://www.forbes.com/sites/cindygordon/2024/02/29/google-latest-debacle-has-paused-gemini-ai-model/?sh=3f0a20b4536c. Accessed 03-03-2024.
  26. Hartford, E. (2023). Dolphin-2.2.1-mistral-7b. https://huggingface.co/cognitivecomputations/dolphin-2.2.1-mistral-7b. Accessed 03-03-2024.
  27. Active ranking from pairwise comparisons and when parametric assumptions don’t help. arXiv preprint arXiv:1606.08842.
  28. Horne, B. G. (1997). Lower bounds for the spectral radius of a matrix. Linear algebra and its applications, 263:261–273.
  29. Robust matrix decomposition with sparse corruptions. IEEE Transactions on Information Theory, 57(11):7221–7234.
  30. Hugging Face (2023). Huggingchat. https://huggingface.co/chat. Accessed 03-03-2024.
  31. Mistral 7b.
  32. Mixtral of experts.
  33. LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion.
  34. Statistical ranking and combinatorial hodge theory. Mathematical Programming, 127(1):203–244.
  35. Estimating the personality of white-box language models. 32rd Annual Workshop on Information Technologies and Systems.
  36. Matrix completion from noisy entries. Journal of Machine Learning Research, 11(Jul):2057–2078.
  37. Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling.
  38. Lambert, N. (2023). Llama 2 follow-up: too much rlhf, gpu sizing, technical details. https://www.interconnects.ai/p/llama-2-part-2. Accessed 03-03-2024.
  39. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. arXiv preprint arXiv:2309.00267.
  40. Lee, P. (2016). Learning from Tay’s introduction. https://blogs.microsoft.com/blog/2016/03/25/learning-tays-introduction/. Accessed 03-03-2024.
  41. AlpacaEval: An Automatic Evaluator of Instruction-following Models.
  42. Lipo: Listwise preference optimization through learning-to-rank.
  43. Statistical rejection sampling improves preference optimization.
  44. Mccarthy, L. (2023). A wellness chatbot is offline after its “harmful” focus on weight loss. https://www.nytimes.com/2023/06/08/us/ai-chatbot-tessa-eating-disorders-association.html. Accessed 03-03-2024.
  45. Mitchell, E. (2023). A note on DPO with noisy preferences & relationship to IPO. https://ericmitchell.ai/cdpo.pdf. Accessed 03-03-2024.
  46. Mok, A. (09-03-2023). Two Google engineers built a ChatGPT-like AI chatbot years ago, but execs reportedly shut it down due to safety concerns — businessinsider.com. https://www.businessinsider.com/google-ai-chatbot-chatgpt-years-ago-execs-shut-down-report-2023-3. Accessed 03-03-2024.
  47. PENTATRON: Personalized context-aware transformer for retrieval-based conversational understanding.
  48. Iterative ranking from pair-wise comparisons. In Advances in Neural Information Processing Systems, pages 2474–2482.
  49. Rank centrality: Ranking from pairwise comparisons. Operations Research, 65(1):266–287.
  50. Non-convex robust PCA. In Advances in Neural Information Processing Systems, pages 1107–1115.
  51. User friendly and adaptable discriminative AI: Using the lessons from the success of LLMs and image generation models.
  52. Rosa: Accurate parameter-efficient fine-tuning via robust adaptation.
  53. Inductive pairwise ranking: Going beyond the n log (n) barrier. In AAAI, pages 2436–2442.
  54. Provable inductive robust PCA via iterative hard thresholding. 33rd Conference on Uncertainty in Artificial Intelligence.
  55. GPT-4 technical report.
  56. Training language models to follow instructions with human feedback.
  57. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290.
  58. A statistical convergence perspective of algorithms for rank aggregation from pairwise data. In ICML, pages 118–126.
  59. When can we rank well from comparisons of o⁢(n⁢log⁡(n))𝑜𝑛𝑛o(n\log(n))italic_o ( italic_n roman_log ( italic_n ) ) non-actively chosen pairs? In 29th Annual Conference on Learning Theory, pages 1376–1401.
  60. Proximal policy optimization algorithms.
  61. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca. Accessed 03-03-2024.
  62. Teknium1 (2023). Openhermes 2.5 - Mistral 7B. https://huggingface.co/teknium/OpenHermes-2.5-Mistral-7B. Accessed 03-03-2024.
  63. Thurstone, L. L. (1927). A law of comparative judgment. Psychological review, 34(4):273.
  64. Llama: Open and efficient foundation language models.
  65. Llama 2: Open foundation and fine-tuned chat models.
  66. Robust subspace learning: Robust pca, robust subspace tracking, and robust subspace recovery. IEEE Signal Processing Magazine, 35(4):32–55.
  67. Openchat: Advancing open-source language models with mixed-quality data.
  68. Robust ranking models via risk-sensitive optimization. In Proceedings of the 35th international ACM SIGIR conference on Research and development in information retrieval, pages 761–770. ACM.
  69. Wittenstein, J. (2023). Bard AI chatbot just cost Google $100 Billion — time.com. https://time.com/6254226/alphabet-google-bard-100-billion-ai-error/. Accessed 03-03-2024.
  70. Some things are more CRINGE than others: Preference optimization with the pairwise cringe loss.
  71. Fast algorithms for Robust PCA via Gradient Descent. arXiv preprint arXiv:1605.07784.
  72. Self-rewarding language models.
  73. SLiC-HF: Sequence likelihood calibration with human feedback.
  74. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena.
  75. A robust ranking algorithm to spamming. EPL (Europhysics Letters), 94(4):48002.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 3 tweets with 0 likes about this paper.