TorchRL and PettingZoo MPE API Tutorial: Multi-Agent Cooperation
Goal of this tutorial¶
This API tutorial introduces a lightweight wrapper layer for training and evaluating cooperative multi-agent RL policies with communication on PettingZoo’s MPE tasks using TorchRL-friendly primitives.
Most MARL tutorials stop at moving losses or rising rewards, but in multi-agent RL those curves alone are not proof of cooperation. This notebook focuses on the training interface; the example notebook handles the deeper coordination checks.
The focus is on:
- configuring the training run using a single config object,
- running training end-to-end through one clean entrypoint,
- producing standardized plots and evaluation metrics,
- validating communication behavior with measurable signals.
This file explains the interface and design decisions, and demonstrates minimal usage.
Function imports¶
The notebook uses two public functions:
default_cfg()Returns a configuration object with sane defaults.train_wrapper(cfg)Runs the full training loop and produces:- training curves (loss, return, entropy),
- checkpointing (saved automatically after training).
from TorchRL_MAC_utils import train_wrapper, default_cfgConfiguration object (what you control)¶
The API notebook demonstrates editing these fields for stability and reproducibility:
Data per update (critical for stability)¶
cfg.rollout_len: number of steps collected per environment per updatecfg.num_envs: number of parallel environments
In the notebook we explicitly set:
rollout_len = 128
num_envs = 4This yields 128 * 4 = 512 samples per update, which is typically much more stable than very small on-policy batches.
PPO-style optimization knobs (as used in the notebook)¶
cfg.ppo_epochs: number of optimization epochs per collected batchcfg.clip_param: PPO clipping parameter for stable policy updatescfg.entropy_coef: exploration regularization
Training runtime knobs¶
cfg.num_iters: total training iterationscfg.lr: learning ratecfg.hidden_dim: network widthcfg.log_interval: print/log frequency
Evaluation knobs¶
cfg.eval_episodes: number of eval episodes per evaluation phase (used when callingevaluate()separately)cfg.eval_interval: evaluation interval (currently set in notebook but evaluation is done separately after training)
For this project, the most important post-training checks are:
- binary success on the task
- goal-distance debugging when success is low
- communication structure via
message_entropyandmessage_change_rate - observation-derived checks that verify messages are actually reflected in the observations
# Create configuration
cfg = default_cfg()
# --- CRITICAL FOR PPO STABILITY ---
# PPO needs more data per update.
# 128 steps * 4 envs = 512 total samples per update (Minimum for stable PPO)
cfg.rollout_len = 128
cfg.num_envs = 4
# PPO Hyperparameters
cfg.ppo_epochs = 4
cfg.clip_param = 0.2
cfg.entropy_coef = 0.01
# Training run
cfg.num_iters = 2000
cfg.lr = 3e-4
cfg.hidden_dim = 128
cfg.log_interval = 10
# Evaluation settings
cfg.eval_episodes = 20
cfg.eval_interval = 200
print("Configuration:")
print(f" Training iterations: {cfg.num_iters}")
print(f" Batch size (steps * envs): {cfg.rollout_len * cfg.num_envs}")
print(f" Learning rate: {cfg.lr}")Configuration:
Training iterations: 2000
Batch size (steps * envs): 512
Learning rate: 0.0003
Function call to training wrapper¶
This runs the training loop and plots training curves (training loss, episode returns, and policy entropy). Those curves are useful diagnostics, but they are not by themselves evidence that coordination emerged.
What train_wrapper(cfg) does¶
At a high level, train_wrapper(cfg) performs:
- Initialize env(s) + policy/value networks from config
- For each training iteration:
- Collect rollouts across
num_envs * rollout_lensteps - Update policy/value using PPO-style minibatch optimization for
ppo_epochs - Log training metrics (loss, return, entropy) at
log_interval
- Collect rollouts across
- Generate and display training curves (loss, return, entropy plots)
- Save checkpoint with trained models
Note: Evaluation (success metrics, communication statistics, and communication ablations) is performed separately using the
evaluate()function after training completes. SeeTorchRL_MAC.example.ipynbfor examples of post-training evaluation.Recommended checks after training:
- compare normal communication against
disable_commorrandom_comm- inspect goal distances if success remains near zero
- confirm action-derived and observation-derived communication metrics agree
Quick troubleshooting:
- if success is always
0.0, recheck the success threshold and environment version- if communication metrics stay random, train longer or retune
entropy_coef- if wrapper/spec errors appear, inspect the API notebook outputs and package versions first
train_wrapper(cfg)Starting training with config:
MacConfig(env_name='simple_reference', num_envs=4, max_cycles=25, continuous_actions=False, hidden_dim=128, actor_layers=2, critic_layers=2, num_iters=2000, rollout_len=128, lr=0.0003, gamma=0.99, entropy_coef=0.01, value_coef=0.5, max_grad_norm=0.5, eval_episodes=20, success_threshold=-150.0, seed=42, device='cpu', checkpoint_dir='./checkpoints', log_interval=10)
/opt/anaconda3/envs/hw-env/lib/python3.11/site-packages/pygame/pkgdata.py:25: UserWarning: pkg_resources is deprecated as an API. See https://setuptools.pypa.io/en/latest/pkg_resources.html. The pkg_resources package is slated for removal as early as 2025-11-30. Refrain from using this package or pin to Setuptools<81.
from pkg_resources import resource_stream, resource_exists
/opt/anaconda3/envs/hw-env/lib/python3.11/site-packages/torchrl/envs/libs/pettingzoo.py:1018: UserWarning: PettingZoo failed to load all modules with error message No module named 'multi_agent_ale_py', trying to load individual modules.
warnings.warn(
/opt/anaconda3/envs/hw-env/lib/python3.11/site-packages/torchrl/envs/libs/pettingzoo.py:51: UserWarning: SISL environments failed to load with error message No module named 'Box2D'.
warnings.warn(f"SISL environments failed to load with error message {err}.")
/opt/anaconda3/envs/hw-env/lib/python3.11/site-packages/torchrl/envs/libs/pettingzoo.py:57: UserWarning: Classic environments failed to load with error message No module named 'chess'.
warnings.warn(f"Classic environments failed to load with error message {err}.")
/opt/anaconda3/envs/hw-env/lib/python3.11/site-packages/torchrl/envs/libs/pettingzoo.py:63: UserWarning: Atari environments failed to load with error message No module named 'multi_agent_ale_py'.
warnings.warn(f"Atari environments failed to load with error message {err}.")
/opt/anaconda3/envs/hw-env/lib/python3.11/site-packages/torchrl/envs/libs/pettingzoo.py:69: UserWarning: Butterfly environments failed to load with error message No module named 'pymunk'.
warnings.warn(
/opt/anaconda3/envs/hw-env/lib/python3.11/site-packages/torchrl/envs/libs/pettingzoo.py:281: UserWarning: PettingZoo in TorchRL is tested using version == 1.24.3 , If you are using a different version and are experiencing compatibility issues,please raise an issue in the TorchRL github.
warnings.warn(
/Users/abhinavsingh/UMD_MSDS/Fall25/Data610/class_project/MSML610/Fall2025/Projects/UmdTask21_Fall2025_TorchRL_Multi-Agent_Cooperation/TorchRL_MAC_utils.py:704: UserWarning: Using a target size (torch.Size([512])) that is different to the input size (torch.Size([512, 1])). This will likely lead to incorrect results due to broadcasting. Please ensure they have the same size.
final_critic_loss = F.mse_loss(critic(states), returns)
Iter 10/2000 | Loss: 272.219 | Return: -20.727 | Entropy: 7.824
Iter 20/2000 | Loss: 283.786 | Return: -21.528 | Entropy: 7.824
Iter 30/2000 | Loss: 238.190 | Return: -19.437 | Entropy: 7.824
Iter 40/2000 | Loss: 240.562 | Return: -19.860 | Entropy: 7.824
Iter 50/2000 | Loss: 251.035 | Return: -20.100 | Entropy: 7.824
Iter 60/2000 | Loss: 217.081 | Return: -18.850 | Entropy: 7.824
Iter 70/2000 | Loss: 237.417 | Return: -20.624 | Entropy: 7.824
Iter 80/2000 | Loss: 241.440 | Return: -20.853 | Entropy: 7.824
Iter 90/2000 | Loss: 250.086 | Return: -21.271 | Entropy: 7.824
Iter 100/2000 | Loss: 268.014 | Return: -22.373 | Entropy: 7.824
Iter 110/2000 | Loss: 228.202 | Return: -20.840 | Entropy: 7.824
Iter 120/2000 | Loss: 251.697 | Return: -22.309 | Entropy: 7.824
Iter 130/2000 | Loss: 215.168 | Return: -20.777 | Entropy: 7.824
Iter 140/2000 | Loss: 244.807 | Return: -22.018 | Entropy: 7.824
Iter 150/2000 | Loss: 203.651 | Return: -20.399 | Entropy: 7.824
Iter 160/2000 | Loss: 209.817 | Return: -21.069 | Entropy: 7.824
Iter 170/2000 | Loss: 264.378 | Return: -23.563 | Entropy: 7.824
Iter 180/2000 | Loss: 195.709 | Return: -20.952 | Entropy: 7.824
Iter 190/2000 | Loss: 208.231 | Return: -21.361 | Entropy: 7.824
Iter 200/2000 | Loss: 204.849 | Return: -21.827 | Entropy: 7.824
Iter 210/2000 | Loss: 192.441 | Return: -21.417 | Entropy: 7.824
Iter 220/2000 | Loss: 192.897 | Return: -21.794 | Entropy: 7.824
Iter 230/2000 | Loss: 223.933 | Return: -23.453 | Entropy: 7.824
Iter 240/2000 | Loss: 227.054 | Return: -23.616 | Entropy: 7.824
Iter 250/2000 | Loss: 191.151 | Return: -22.462 | Entropy: 7.824
Iter 260/2000 | Loss: 234.027 | Return: -24.523 | Entropy: 7.824
Iter 270/2000 | Loss: 198.505 | Return: -23.140 | Entropy: 7.824
Iter 280/2000 | Loss: 186.768 | Return: -23.021 | Entropy: 7.824
Iter 290/2000 | Loss: 220.150 | Return: -24.428 | Entropy: 7.824
Iter 300/2000 | Loss: 198.833 | Return: -23.473 | Entropy: 7.824
Iter 310/2000 | Loss: 202.513 | Return: -24.599 | Entropy: 7.824
Iter 320/2000 | Loss: 158.454 | Return: -22.182 | Entropy: 7.824
Iter 330/2000 | Loss: 188.632 | Return: -24.372 | Entropy: 7.824
Iter 340/2000 | Loss: 184.924 | Return: -24.062 | Entropy: 7.824
Iter 350/2000 | Loss: 216.704 | Return: -25.952 | Entropy: 7.824
Iter 360/2000 | Loss: 161.028 | Return: -23.399 | Entropy: 7.824
Iter 370/2000 | Loss: 189.443 | Return: -25.307 | Entropy: 7.824
Iter 380/2000 | Loss: 212.249 | Return: -26.602 | Entropy: 7.824
Iter 390/2000 | Loss: 175.836 | Return: -24.346 | Entropy: 7.824
Iter 400/2000 | Loss: 154.737 | Return: -23.995 | Entropy: 7.824
Iter 410/2000 | Loss: 180.330 | Return: -25.461 | Entropy: 7.824
Iter 420/2000 | Loss: 161.149 | Return: -24.449 | Entropy: 7.824
Iter 430/2000 | Loss: 186.942 | Return: -26.297 | Entropy: 7.824
Iter 440/2000 | Loss: 131.880 | Return: -23.877 | Entropy: 7.824
Iter 450/2000 | Loss: 135.673 | Return: -23.542 | Entropy: 7.824
Iter 460/2000 | Loss: 169.219 | Return: -25.830 | Entropy: 7.824
Iter 470/2000 | Loss: 159.329 | Return: -25.831 | Entropy: 7.824
Iter 480/2000 | Loss: 166.124 | Return: -26.660 | Entropy: 7.824
Iter 490/2000 | Loss: 137.071 | Return: -25.266 | Entropy: 7.824
Iter 500/2000 | Loss: 146.829 | Return: -25.746 | Entropy: 7.824
Iter 510/2000 | Loss: 128.433 | Return: -25.264 | Entropy: 7.824
Iter 520/2000 | Loss: 153.761 | Return: -26.865 | Entropy: 7.824
Iter 530/2000 | Loss: 154.788 | Return: -27.449 | Entropy: 7.824
Iter 540/2000 | Loss: 144.498 | Return: -26.601 | Entropy: 7.824
Iter 550/2000 | Loss: 130.101 | Return: -26.399 | Entropy: 7.824
Iter 560/2000 | Loss: 129.165 | Return: -25.980 | Entropy: 7.824
Iter 570/2000 | Loss: 149.247 | Return: -27.918 | Entropy: 7.824
Iter 580/2000 | Loss: 155.469 | Return: -28.082 | Entropy: 7.824
Iter 590/2000 | Loss: 122.999 | Return: -26.408 | Entropy: 7.824
Iter 600/2000 | Loss: 128.385 | Return: -27.180 | Entropy: 7.824
Iter 610/2000 | Loss: 125.464 | Return: -26.817 | Entropy: 7.824
Iter 620/2000 | Loss: 137.167 | Return: -27.616 | Entropy: 7.824
Iter 630/2000 | Loss: 138.750 | Return: -29.508 | Entropy: 7.824
Iter 640/2000 | Loss: 130.827 | Return: -28.543 | Entropy: 7.824
Iter 650/2000 | Loss: 123.976 | Return: -28.539 | Entropy: 7.824
Iter 660/2000 | Loss: 129.067 | Return: -29.076 | Entropy: 7.824
Iter 670/2000 | Loss: 132.543 | Return: -29.004 | Entropy: 7.824
Iter 680/2000 | Loss: 143.127 | Return: -29.255 | Entropy: 7.824
Iter 690/2000 | Loss: 120.384 | Return: -29.251 | Entropy: 7.824
Iter 700/2000 | Loss: 135.363 | Return: -30.535 | Entropy: 7.824
Iter 710/2000 | Loss: 131.897 | Return: -29.639 | Entropy: 7.824
Iter 720/2000 | Loss: 119.251 | Return: -29.723 | Entropy: 7.824
Iter 730/2000 | Loss: 134.599 | Return: -30.807 | Entropy: 7.824
Iter 740/2000 | Loss: 111.218 | Return: -29.840 | Entropy: 7.824
Iter 750/2000 | Loss: 135.076 | Return: -30.522 | Entropy: 7.824
Iter 760/2000 | Loss: 123.852 | Return: -30.770 | Entropy: 7.824
Iter 770/2000 | Loss: 121.137 | Return: -30.798 | Entropy: 7.824
Iter 780/2000 | Loss: 127.408 | Return: -30.461 | Entropy: 7.824
Iter 790/2000 | Loss: 105.242 | Return: -29.312 | Entropy: 7.824
Iter 800/2000 | Loss: 124.324 | Return: -30.428 | Entropy: 7.824
Iter 810/2000 | Loss: 108.142 | Return: -31.155 | Entropy: 7.824
Iter 820/2000 | Loss: 123.512 | Return: -31.766 | Entropy: 7.824
Iter 830/2000 | Loss: 120.032 | Return: -30.098 | Entropy: 7.824
Iter 840/2000 | Loss: 121.658 | Return: -31.289 | Entropy: 7.824
Iter 850/2000 | Loss: 119.804 | Return: -31.195 | Entropy: 7.824
Iter 860/2000 | Loss: 114.878 | Return: -31.381 | Entropy: 7.824
Iter 870/2000 | Loss: 123.373 | Return: -33.549 | Entropy: 7.823
Iter 880/2000 | Loss: 136.105 | Return: -33.336 | Entropy: 7.823
Iter 890/2000 | Loss: 119.431 | Return: -32.809 | Entropy: 7.823
Iter 900/2000 | Loss: 125.352 | Return: -32.829 | Entropy: 7.823
Iter 910/2000 | Loss: 112.484 | Return: -31.248 | Entropy: 7.823
Iter 920/2000 | Loss: 132.778 | Return: -32.977 | Entropy: 7.823
Iter 930/2000 | Loss: 115.815 | Return: -32.372 | Entropy: 7.822
Iter 940/2000 | Loss: 129.861 | Return: -33.814 | Entropy: 7.823
Iter 950/2000 | Loss: 122.683 | Return: -33.432 | Entropy: 7.823
Iter 960/2000 | Loss: 132.228 | Return: -34.600 | Entropy: 7.822
Iter 970/2000 | Loss: 124.876 | Return: -32.571 | Entropy: 7.822
Iter 980/2000 | Loss: 127.689 | Return: -33.493 | Entropy: 7.822
Iter 990/2000 | Loss: 131.896 | Return: -35.043 | Entropy: 7.822
Iter 1000/2000 | Loss: 137.901 | Return: -33.531 | Entropy: 7.823
Iter 1010/2000 | Loss: 141.120 | Return: -34.071 | Entropy: 7.822
Iter 1020/2000 | Loss: 139.853 | Return: -34.858 | Entropy: 7.823
Iter 1030/2000 | Loss: 128.034 | Return: -33.193 | Entropy: 7.823
Iter 1040/2000 | Loss: 136.603 | Return: -33.402 | Entropy: 7.823
Iter 1050/2000 | Loss: 136.885 | Return: -35.759 | Entropy: 7.823
Iter 1060/2000 | Loss: 129.463 | Return: -34.831 | Entropy: 7.824
Iter 1070/2000 | Loss: 128.981 | Return: -34.252 | Entropy: 7.824
Iter 1080/2000 | Loss: 128.580 | Return: -34.055 | Entropy: 7.823
Iter 1090/2000 | Loss: 130.577 | Return: -33.744 | Entropy: 7.824
Iter 1100/2000 | Loss: 133.263 | Return: -34.781 | Entropy: 7.824
Iter 1110/2000 | Loss: 133.603 | Return: -34.109 | Entropy: 7.824
Iter 1120/2000 | Loss: 129.330 | Return: -34.194 | Entropy: 7.824
Iter 1130/2000 | Loss: 122.525 | Return: -33.402 | Entropy: 7.824
Iter 1140/2000 | Loss: 128.424 | Return: -35.014 | Entropy: 7.824
Iter 1150/2000 | Loss: 128.775 | Return: -33.834 | Entropy: 7.824
Iter 1160/2000 | Loss: 119.694 | Return: -33.134 | Entropy: 7.824
Iter 1170/2000 | Loss: 117.584 | Return: -33.442 | Entropy: 7.824
Iter 1180/2000 | Loss: 127.881 | Return: -34.535 | Entropy: 7.824
Iter 1190/2000 | Loss: 123.023 | Return: -34.018 | Entropy: 7.824
Iter 1200/2000 | Loss: 135.240 | Return: -35.580 | Entropy: 7.824
Iter 1210/2000 | Loss: 141.598 | Return: -34.583 | Entropy: 7.824
Iter 1220/2000 | Loss: 138.867 | Return: -34.330 | Entropy: 7.824
Iter 1230/2000 | Loss: 120.681 | Return: -33.525 | Entropy: 7.824
Iter 1240/2000 | Loss: 137.870 | Return: -35.492 | Entropy: 7.824
Iter 1250/2000 | Loss: 131.109 | Return: -34.736 | Entropy: 7.824
Iter 1260/2000 | Loss: 119.272 | Return: -33.574 | Entropy: 7.824
Iter 1270/2000 | Loss: 129.100 | Return: -34.314 | Entropy: 7.824
Iter 1280/2000 | Loss: 145.171 | Return: -35.822 | Entropy: 7.824
Iter 1290/2000 | Loss: 127.316 | Return: -32.960 | Entropy: 7.824
Iter 1300/2000 | Loss: 141.976 | Return: -35.044 | Entropy: 7.824
Iter 1310/2000 | Loss: 132.350 | Return: -33.530 | Entropy: 7.824
Iter 1320/2000 | Loss: 138.986 | Return: -35.041 | Entropy: 7.824
Iter 1330/2000 | Loss: 136.296 | Return: -35.078 | Entropy: 7.824
Iter 1340/2000 | Loss: 128.152 | Return: -34.780 | Entropy: 7.824
Iter 1350/2000 | Loss: 132.387 | Return: -35.271 | Entropy: 7.824
Iter 1360/2000 | Loss: 122.335 | Return: -33.618 | Entropy: 7.824
Iter 1370/2000 | Loss: 132.943 | Return: -33.000 | Entropy: 7.824
Iter 1380/2000 | Loss: 117.663 | Return: -33.066 | Entropy: 7.824
Iter 1390/2000 | Loss: 136.465 | Return: -35.461 | Entropy: 7.824
Iter 1400/2000 | Loss: 127.772 | Return: -33.854 | Entropy: 7.824
Iter 1410/2000 | Loss: 138.642 | Return: -35.739 | Entropy: 7.824
Iter 1420/2000 | Loss: 130.790 | Return: -34.733 | Entropy: 7.824
Iter 1430/2000 | Loss: 133.290 | Return: -34.068 | Entropy: 7.824
Iter 1440/2000 | Loss: 119.001 | Return: -33.137 | Entropy: 7.824
Iter 1450/2000 | Loss: 129.622 | Return: -34.483 | Entropy: 7.824
Iter 1460/2000 | Loss: 126.021 | Return: -33.947 | Entropy: 7.824
Iter 1470/2000 | Loss: 127.835 | Return: -33.531 | Entropy: 7.824
Iter 1480/2000 | Loss: 127.488 | Return: -32.304 | Entropy: 7.824
Iter 1490/2000 | Loss: 124.828 | Return: -33.300 | Entropy: 7.824
Iter 1500/2000 | Loss: 144.105 | Return: -34.026 | Entropy: 7.824
Iter 1510/2000 | Loss: 123.314 | Return: -32.939 | Entropy: 7.824
Iter 1520/2000 | Loss: 144.808 | Return: -35.060 | Entropy: 7.824
Iter 1530/2000 | Loss: 127.927 | Return: -33.978 | Entropy: 7.824
Iter 1540/2000 | Loss: 134.678 | Return: -34.356 | Entropy: 7.824
Iter 1550/2000 | Loss: 125.394 | Return: -34.657 | Entropy: 7.824
Iter 1560/2000 | Loss: 126.071 | Return: -34.040 | Entropy: 7.824
Iter 1570/2000 | Loss: 127.383 | Return: -34.143 | Entropy: 7.824
Iter 1580/2000 | Loss: 134.967 | Return: -34.283 | Entropy: 7.824
Iter 1590/2000 | Loss: 130.888 | Return: -33.791 | Entropy: 7.824
Iter 1600/2000 | Loss: 121.994 | Return: -33.131 | Entropy: 7.824
Iter 1610/2000 | Loss: 133.154 | Return: -34.052 | Entropy: 7.824
Iter 1620/2000 | Loss: 127.125 | Return: -32.983 | Entropy: 7.824
Iter 1630/2000 | Loss: 125.718 | Return: -33.795 | Entropy: 7.824
Iter 1640/2000 | Loss: 125.356 | Return: -31.941 | Entropy: 7.824
Iter 1650/2000 | Loss: 132.501 | Return: -34.211 | Entropy: 7.824
Iter 1660/2000 | Loss: 131.549 | Return: -33.690 | Entropy: 7.824
Iter 1670/2000 | Loss: 121.862 | Return: -34.058 | Entropy: 7.824
Iter 1680/2000 | Loss: 134.054 | Return: -33.779 | Entropy: 7.824
Iter 1690/2000 | Loss: 125.647 | Return: -34.280 | Entropy: 7.824
Iter 1700/2000 | Loss: 127.945 | Return: -33.293 | Entropy: 7.824
Iter 1710/2000 | Loss: 130.306 | Return: -34.077 | Entropy: 7.824
Iter 1720/2000 | Loss: 124.567 | Return: -33.640 | Entropy: 7.824
Iter 1730/2000 | Loss: 138.494 | Return: -34.346 | Entropy: 7.824
Iter 1740/2000 | Loss: 138.243 | Return: -34.821 | Entropy: 7.824
Iter 1750/2000 | Loss: 139.490 | Return: -35.332 | Entropy: 7.824
Iter 1760/2000 | Loss: 135.553 | Return: -32.960 | Entropy: 7.824
Iter 1770/2000 | Loss: 130.569 | Return: -33.463 | Entropy: 7.824
Iter 1780/2000 | Loss: 127.610 | Return: -33.195 | Entropy: 7.824
Iter 1790/2000 | Loss: 125.524 | Return: -34.318 | Entropy: 7.824
Iter 1800/2000 | Loss: 137.067 | Return: -33.671 | Entropy: 7.824
Iter 1810/2000 | Loss: 128.741 | Return: -33.161 | Entropy: 7.824
Iter 1820/2000 | Loss: 131.411 | Return: -33.000 | Entropy: 7.824
Iter 1830/2000 | Loss: 132.155 | Return: -33.598 | Entropy: 7.824
Iter 1840/2000 | Loss: 131.916 | Return: -32.702 | Entropy: 7.824
Iter 1850/2000 | Loss: 130.272 | Return: -33.885 | Entropy: 7.824
Iter 1860/2000 | Loss: 141.205 | Return: -34.522 | Entropy: 7.824
Iter 1870/2000 | Loss: 128.554 | Return: -34.314 | Entropy: 7.824
Iter 1880/2000 | Loss: 139.736 | Return: -33.479 | Entropy: 7.824
Iter 1890/2000 | Loss: 131.810 | Return: -33.502 | Entropy: 7.824
Iter 1900/2000 | Loss: 133.640 | Return: -32.967 | Entropy: 7.824
Iter 1910/2000 | Loss: 122.707 | Return: -33.815 | Entropy: 7.824
Iter 1920/2000 | Loss: 125.687 | Return: -34.250 | Entropy: 7.824
Iter 1930/2000 | Loss: 128.451 | Return: -34.683 | Entropy: 7.824
Iter 1940/2000 | Loss: 126.490 | Return: -33.862 | Entropy: 7.824
Iter 1950/2000 | Loss: 128.122 | Return: -34.110 | Entropy: 7.824
Iter 1960/2000 | Loss: 123.299 | Return: -33.446 | Entropy: 7.824
Iter 1970/2000 | Loss: 138.924 | Return: -33.922 | Entropy: 7.824
Iter 1980/2000 | Loss: 139.877 | Return: -35.000 | Entropy: 7.824
Iter 1990/2000 | Loss: 128.982 | Return: -34.637 | Entropy: 7.824
Iter 2000/2000 | Loss: 135.665 | Return: -34.567 | Entropy: 7.824
Training complete! Checkpoint saved to: ./checkpoints/mac_checkpoint_iter2000.pt
✓ Training complete!
Checkpoint: ./checkpoints/mac_checkpoint_iter2000.pt
Final loss: 135.665
Final success score: 34.567
Final entropy: 7.824

Training curves saved to: mac_training_curves.png