跳到主要导航 跳到搜索 跳到主要内容

Open bandit processes with uncountable states and time-Ackward effects

  • Macquarie University

科研成果: 期刊稿件文章同行评审

摘要

Bandit processes and the Gittins index have provided powerful and elegant theory and tools for the optimization of allocating limited resources to competitive demands. In this paper we extend the Gittins theory to more general branching bandit processes, also referred to as open bandit processes, that allow uncountable states and backward times. We establish the optimality of the Gittins index policy with uncountably many states, which is useful in such problems as dynamic scheduling with continuous random processing times. We also allow negative time durations for discounting a reward to account for the present value of the reward that was received before the present time, which we refer to as time-backward effects. This could model the situation of offering bonus rewards for completing jobs above expectation. Moreover, we discover that a common belief on the optimality of the Gittins index in the generalized bandit problem is not always true without additional conditions, and provide a counterexample. We further apply our theory of open bandit processes with time-backward effects to prove the optimality of the Gittins index in the generalized bandit problem under a sufficient condition.

源语言英语
页(从-至)388-402
页数15
期刊Journal of Applied Probability
50
2
DOI
出版状态已出版 - 6月 2013

学术指纹

探究 'Open bandit processes with uncountable states and time-Ackward effects' 的科研主题。它们共同构成独一无二的学术指纹。

引用此