Preference Optimization
The process of optimizing an AI system to produce outputs that better match preferred responses, behaviors, or outcomes.
Explore more about Feedback Loops
Related terms
A learning approach in which an AI system learns from preferences between possible outputs or actions to better align its behavior with desired outcomes.
Preference ModelingA method for representing and estimating preferences between possible actions or outcomes to guide an AI system's decision-making.
Preference DataData that records preferences between AI outputs or behaviors, typically used to train or optimize models toward preferred responses.
Reward ModelingThe process of training a mathematical model to score AI outputs based on human or automated preferences.
Reinforcement LearningRLA machine learning paradigm where an agent learns to make optimal decisions by taking actions in an environment to maximize cumulative rewards.