For each benchmark, we'd want: - static actions - random dynamic - tuned RL - any known solution policies (e.g. CSA) - optimal solution if it exists
For each benchmark, we'd want: