Agent–Environment Interaction
智能体与环境EssentialThe basic RL framework: the decision-maker is the agent, and everything outside it that it can observe and affect is the environment.
This is the basic framework reinforcement learning uses to describe a problem. The agent is the decision-maker; the environment is everything outside it that the agent can observe and act on. The two interact in a loop: the agent sees the current observation and picks an action; the environment transitions to a new state accordingly and returns a new observation and a reward (a number measuring how good the outcome was); this repeats until the episode ends. In embodied AI, the agent is usually a policy model running on a robot, and the environment is the real world or a simulator, including the tabletop, objects, and any nearby people. This shared framework lets very different problems — navigation, grasping, walking — be described with the same vocabulary; simulation libraries such as Gymnasium build their APIs around it too, using reset to start a new episode and step to execute one action.
ExampleA robot arm folding a towel: the agent is the folding policy, the environment is the tabletop, the towel, and the camera feed. At each step the policy outputs a set of joint actions, and the environment returns the new image that results.
- Also called
- Agent-Environment Loop, Agent-Environment Interface
- Related
- Reinforcement Learning · Markov Decision Process · Observation · Action Space · Reward Function · Environment (Env; reset/step interface)
- Sources
- OpenAI Spinning Up: Key Concepts in RL
Gymnasium Documentation: Basic Usage