There has been a lot of excitement around "LLM agents", but how capable are they in open-ended multi-agent coordination problems?
To study this, we designed a long-horizon, open-ended multi-agent coordination environment and compared zero-shot LLM agents with trained MARL agents. We find that the two paradigms have distinct strengths and limitations, highlighting that coordination is a bottleneck separate from standard long-horizon task competence.
