OpenAI: Derin RL'de Dönme
Özgün başlık: OpenAI: Spinning Up in Deep RL
latest User Documentation Introduction Installation Algorithms Running Experiments Experiment Outputs Plotting Results Introduction to RL Part 1: Key Concepts in RL Part 2: Kinds of RL Algorithms Part 3: Intro to Policy Optimization Resources Spinning Up as a Deep RL Researcher Key Papers in Deep RL Exercises Benchmarks for Spinning Up Implementations Algorithms Docs Vanilla Policy Gradient Trust Region Policy Optimization Proximal Policy Optimization Deep Deterministic Policy Gradient Twin Delayed DDPG Soft Actor-Critic Utilities Docs Logger Plotter MPI Tools Run Utils Etc. Acknowledgements About the Author Spinning Up Docs » Welcome to Spinning Up in Deep RL! Edit on GitHub Welcome to Spinning Up in Deep RL User Documentation Introduction What This Is Why We Built This How This Serves Our Mission Code Design Philosophy Long-Term Support and Support History Installation Installing Python Installing OpenMPI Installing Spinning Up Check Your Install Installing MuJoCo (Optional) Algorithms What’s Included Why These Algorithms? Code Format Running Experiments Launching from the Command Line Launching from Scripts Experiment Outputs Algorithm Outputs Save Directory Location Loading and Running Trained Policies Plotting Results Introduction to RL Part 1: Key Concepts in RL What Can RL Do? Key Concepts and Terminology (Optional) Formalism Part 2: Kinds of RL Algorithms A Taxonomy of RL Algorithms Links to Algorithms in Taxonomy Part 3: Intro to Policy Optimization Deriving the Simpl
est Policy Gradient Implementing the Simplest Policy Gradient Expected Grad-Log-Prob Lemma Don’t Let the Past Distract You Implementing Reward-to-Go Policy Gradient Baselines in Policy Gradients Other Forms of the Policy Gradient Recap Resources Spinning Up as a Deep RL Researcher The Right Background Learn by Doing Developing a Research Project Doing Rigorous Research in RL Closing Thoughts PS: Other Resources References Key Papers in Deep RL 1. Model-Free RL 2. Exploration 3. Transfer and Multitask RL 4. Hierarchy 5. Memory 6. Model-Based RL 7. Meta-RL 8. Scaling RL 9. RL in the Real World 10. Safety 11. Imitation Learning and Inverse Reinforcement Learning 12. Reproducibility, Analysis, and Critique 13.