BubbleRank: Safe Online Learning to Re-Rank via Implicit Click Feedback

Li, C.
Kveton, B.
Lattimore, T.
Markov, I.
de Rijke, M.
Szepesvári, C.
Zoghi, M.
Globerson, A.
Silva, R.

Publication date

January 2019

Publisher

AUAI Press

Abstract

In this paper, we study the problem of safe online learning to re-rank, where user feedback is used to improve the quality of displayed lists. Learning to rank has traditionally been studied in two settings. In the offline setting, rankers are typically learned from relevance labels created by judges. This approach has generally become standard in industrial applications of ranking, such as search. However, this approach lacks exploration and thus is limited by the information content of the offline training data. In the online setting, an algorithm can experiment with lists and learn from feedback on them in a sequential fashion. Bandit algorithms are well-suited for this setting but they tend to learn user preferences from scratch, which ...

Extracted data

We use cookies to provide a better user experience.

Data Protection

BubbleRank: Safe Online Learning to Re-Rank via Implicit Click Feedback

Abstract

Extracted data

BubbleRank: Safe Online Learning to Re-Rank via Implicit Click Feedback

Abstract

Extracted data

Related items

Related items