graph LR Agent[Agent] --> |"Action a_t"| Env[Environment] Env --> |Reward r_t| Agent Env --> |"New State s_{t+1}"| Agent style Agent fill:#2196f3 style Env fill:#4caf50
graph LR State[State s] --> Q[Q-function] Q -->|compute| Values[Q-values] Values -->|argmax| Action[Greedy Action a] Action -->|execute| Result[Reward + New State] Result -->|update| Q style State fill:#4caf50 style Q fill:#2196f3 style Values fill:#ff9800 style Action fill:#f44336 style Result fill:#9c27b0
graph TD Agent[RL Agent] --> Choice[Action Selection] Choice --> Explore[Exploration] Choice --> Exploit[Exploitation] Explore -->|random| New[Discover New Strategies] Exploit -->|greedy| Best[Use Known Best] New -->|improve| Future[Better Long-term Reward] Best -->|maximize| Now[Higher Short-term Reward] style Agent fill:#2196f3 style Choice fill:#ff9800 style Explore fill:#4caf50 style Exploit fill:#f44336 style New fill:#4caf50 style Best fill:#f44336 style Future fill:#9c27b0 style Now fill:#9c27b0
graph LR Policy[Policy pi] -->|sample| Traj[Trajectory] Traj -->|compute| Return[Cumulative Return] Return -->|gradient ascent| Update[Update Policy] Update --> Policy style Policy fill:#2196f3 style Traj fill:#ff9800 style Return fill:#4caf50 style Update fill:#9c27b0