Trading Value and Information in MDPs

Rubin, Jonathan; Shamir, Ohad; Tishby, Naftali

doi:10.1007/978-3-642-24647-0_3

Jonathan Rubin⁶,
Ohad Shamir⁷ &
Naftali Tishby⁶

Part of the book series: Intelligent Systems Reference Library ((ISRL,volume 28))

1515 Accesses
23 Citations
1 Altmetric

Abstract

Interactions between an organism and its environment are commonly treated in the framework of Markov Decision Processes (MDP). While standard MDP is aimed solely at maximizing expected future rewards (value), the circular flow of information between the agent and its environment is generally ignored. In particular, the information gained from the environment by means of perception and the information involved in the process of action selection (i.e., control) are not treated in the standard MDP setting. In this paper, we focus on the control information and show how it can be combined with the reward measure in a unified way. Both of these measures satisfy the familiar Bellman recursive equations, and their linear combination (the free-energy) provides an interesting new optimization criterion. The tradeoff between value and information, explored using our info-rl algorithm, provides a principled justification for stochastic (soft) policies. We use computational learning theory to show that these optimal policies are also robust to uncertainties in settings with only partial knowledge of the MDP parameters.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 129.00; Price excludes VAT (USA)

Softcover Book: USD 169.99; Price excludes VAT (USA)

Hardcover Book: USD 169.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

References

Bertsekas, D.P.: Dynamic Programming and Optimal Control. Athena Scientific (1995)
Google Scholar
Braun, D.A., Ortega, P.A., Theodorou, E., Schaal, S.: Path integral control and bounded rationality. To appear in Approximate Dynamic Programming and Reinforcement Learnig (2011), http://www-clmc.usc.edu/publications//D/DanielADPRL2011.pdf
Cover, T.M., Thomas, J.A.: Elements of Information Theory. Wiley, New York (1991)
Book MATH Google Scholar
Friston, K.: The free-energy principle: a rough guide to the brain? Trends Cogn. Sci. 13(7), 293–301 (2009), doi:10.1016/j.tics.2009.04.005
Article Google Scholar
Fuster, J.M.: The prefrontal cortex — an update: Time is of the essence. Neuron 30, 319–333 (2001)
Article Google Scholar
Kappen, B., Gomez, V., Opper, M.: Optimal control as a graphical model inference problem. ArXiv e-prints (2009)
Google Scholar
Mcallester, D.: Simplified pac-bayesian margin bounds. In: Proc. 2007 IEEE International Symposium on Approximate Dynamic Programming and Reinforcement Learning, Hawaii of the 16th Annual Conference on Learning Theory, April 1-5 (2003)
Google Scholar
Shannon, C.: Coding theorems for a discrete source with a fidelity criterion. IRE NATO Conv. Rec. 4, 142–163 (1959)
Google Scholar
Sutton, R.S., Barto, A.G.: Reinforcement Learning. MIT Press, Cambridge (1998)
Google Scholar
Tishby, N., Pereira, F.C., Bialek, W.: The information bottleneck method. In: Proc. 37th Annual Allerton Conference on Communication, Control and Computing (1999)
Google Scholar
Tishby, N., Polani, D.: Information theory of decisions and actions. In: Vassilis, Hussain, Taylor (eds.) Perception-Reason-Action, Cognitive Neuroscience. Springer, Heidelberg (2010)
Google Scholar
Todorov, E.: Efficient computation of optimal actions. PNAS 106(28), 11,478–11,483 (2009)
Google Scholar

Download references

Author information

Authors and Affiliations

Hebrew University Jerusalem, Israel
Jonathan Rubin & Naftali Tishby
Microsoft Research Cambridge,
Ohad Shamir

Authors

Jonathan Rubin
View author publications
You can also search for this author in PubMed Google Scholar
Ohad Shamir
View author publications
You can also search for this author in PubMed Google Scholar
Naftali Tishby
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

Department of Adaptive Systems, Institute of Information Theory and Automation of the ASCR, PO Box 18, 18208, Praha 8, Czech Republic
Tatiana Valentine Guy
Institute of Information Theory and Automation of the ASCR, PO Box 18, 18208, Praha 8, Czech Republic
Miroslav Kárný
Intelligent Systems Division Ames Research Center, NASA, Mail Stop 269-1, 94035, Moffett Field, CA, USA
David H. Wolpert

Rights and permissions

Reprints and permissions

Copyright information

About this chapter

Cite this chapter

Rubin, J., Shamir, O., Tishby, N. (2012). Trading Value and Information in MDPs. In: Guy, T.V., Kárný, M., Wolpert, D.H. (eds) Decision Making with Imperfect Decision Makers. Intelligent Systems Reference Library, vol 28. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-24647-0_3

Download citation

DOI: https://doi.org/10.1007/978-3-642-24647-0_3
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-642-24646-3
Online ISBN: 978-3-642-24647-0
eBook Packages: EngineeringEngineering (R0)

Publish with us

Policies and ethics