> For the complete documentation index, see [llms.txt](https://json007.gitbook.io/deeplearning/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://json007.gitbook.io/deeplearning/dqn.md).

# DQN

[DQN 从入门到放弃4 动态规划与Q-Learning](https://zhuanlan.zhihu.com/p/21378532)

<https://zhuanlan.zhihu.com/sharerl>\
1\. 首先Bellman方程\
2\. 策略迭代Policy Iteration求解\
3\. Value Iteration 价值迭代求解\
4\. Q-Learning

[浅述：从 Minimax 到 AlphaZero，完全信息博弈之路（2）](https://zhuanlan.zhihu.com/p/32073374)

[GMIS 2017 | NIPS最佳论文作者之一吴翼：价值迭代网络](https://mp.weixin.qq.com/s/4XEwBDn0zK-PRy2QswUcGg)

[独家对话NIPS 2016最佳论文作者：如何打造新型强化学习观](http://www.jiqizhixin.com/article/1963)

[策略梯度下降过时了，OpenAI 拿出一种新的策略优化算法PPO](https://mp.weixin.qq.com/s?__biz=MzI5NTIxNTg0OA==\&mid=2247486306\&idx=3\&sn=6900a0177cb26df0fd4ecd7d1431500d)

[荐译一篇通俗易懂的策略梯度方法讲解](https://zhuanlan.zhihu.com/p/27699682)

[DeepMind ICML 2017论文： 超越传统强化学习的价值分布方法](https://mp.weixin.qq.com/s/cO1VlYGwdRBAbPs7IgvcAA)

[Neural Fictitious Self Play——从博弈论到深度强化学习](https://tigerneil.wordpress.com/2016/04/15/neural-fictitious-self-play-从博弈论到深度强化学习/)

[三十分钟理解博弈论“纳什均衡” -- Nash Equilibrium](http://blog.csdn.net/xbinworld/article/details/50932559)\
[深度强化学习初探](http://www.52cs.org/?p=764)\
[DQN 从入门到放弃 第一篇：DQN与增强学习](https://zhuanlan.zhihu.com/p/21262246?refer=intelligentunit)\
[强化学习系列之一:马尔科夫决策过程](http://www.gotoli.us/强化学习-马尔科夫决策过程)\
[强化学习系列之一:马尔科夫决策过程](http://www.algorithmdog.com/强化学习-马尔科夫决策过程)\
[Python code for Reinforcement Learning: An Introduction](https://github.com/ShangtongZhang/reinforcement-learning-an-introduction)\
[Reinforcement\_Learning\_Blog](https://github.com/algorithmdog/Reinforcement_Learning_Blog)\
[A Painless Q-learning Tutorial (一个 Q-learning 算法的简明教程)](http://blog.csdn.net/itplus/article/details/9361915)

[深度强化学习导引](http://www.jianshu.com/p/dfd987aa765a)\
[AlphaGo作者又一力作，攻克德州扑克](http://i.dataguru.cn/mportal.php?mod=view\&aid=9166)\
[【David Silver强化学习公开课之一】强化学习入门](http://chenrudan.github.io/blog/2016/06/06/reinforcementlearninglesssion1.html)

<https://zhuanlan.zhihu.com/p/24761972> 吴恩达对于增强学习的形象论述（上）

[吴恩达对于增强学习的形象论述（下）](https://zhuanlan.zhihu.com/p/24996278)

[增强学习的解释——学习基于长期回报的行为](https://mp.weixin.qq.com/s?__biz=MzIyODE5MTcxNw==∣=2650372430\&idx=1\&sn=763885f70a9f1f01cd25fb4b1d0d441f\&chksm=f05865c4c72fecd26d4fdc7b8c2a2e0caee936255e5c7fde0c4f93018836f7444b5f362b5865\&scene=25\&pass_ticket=TyrCO9ZWtuuBYeGfCLfObps9Mskg9kt6UT8PvM2Bj2o%3D#wechat_redirect)

[深度增强学习【1】走向通用人工智能之路](http://blog.greenwicher.com/2016/12/18/drl-general_ai-intro/)

[深度增强学习【2】从多臂赌博机问题到蒙特卡洛树搜索](http://blog.greenwicher.com/2016/12/24/drl-from_mab_to_mcts/)

[「模仿学习」的确强大，但你知道它和「强化学习」的关系吗？](https://mp.weixin.qq.com/s?__biz=MzA5MzQwMDk4Mg==\&mid=2651042957\&idx=1\&sn=c6e124bd4f18b2b59af1b5b7391ff1e9)

[从Q学习到DDPG，一文简述多种强化学习算法](https://www.jiqizhixin.com/articles/2018-01-22-5)

[DeepMind SR论文：非对称博弈的对称分解](https://zhuanlan.zhihu.com/p/33103175)

[Role of RL in Text Generation by GAN(强化学习在生成对抗网络文本生成中扮演的角色)](https://zhuanlan.zhihu.com/p/29168803)
