https://www.mdu.se/

mdu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Improving representation learning in MARL using teammate-focused auxiliary task
Mälardalen University.
Mälardalen University.
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

In cooperative Multi-Agent Reinforcement Learning (MARL), agents jointly interact with theenvironment and receive a shared team reward. However, learning cooperative behaviour that im-proves both performance and generalisation remains a difficult problem, especially when rewardsare sparse and execution is decentralised. Although value-based methods such as QMIX addressparts of this challenge, effective team play and generalisation remain challenging. This thesis ex-plores team play and generalisation in a cooperative MARL setting and introduces Teammate Ac-tion Forecast-QMIX (TAF-QMIX), a QMIX-based architecture augmented with Teammate ActionForecast (TAF) as an auxiliary task. TAF takes the recurrent agents’ hidden states as input andpredicts the next greedy action selected by each teammate’s current policy. The aim is to increasesample efficiency and encourage the agents’ hidden representations to capture teammate-relevantinformation, thereby improving team play while reducing reliance on opponent-specific modellingthat may limit generalisation. The architecture was trained and evaluated in the Half Field Of-fence (HFO) robot-soccer environment across several scenarios. The results indicated more stableattention to teammate-related and environment-related information, as well as a higher ratio ofsuccessful passes relative to pass attempts. However, the results did not indicate any consistentimprovement in training or evaluation performance. This thesis therefore concludes that the be-havioural differences encouraged by TAF did not, in this case, translate into reliable gains in goalrate or generalisation to unseen opponents.

Place, publisher, year, edition, pages
2026. , p. 51
Keywords [en]
MARL, auxiliary task, HFO, representation learning
National Category
Robotics and automation
Identifiers
URN: urn:nbn:se:mdh:diva-77336OAI: oai:DiVA.org:mdh-77336DiVA, id: diva2:2067982
Subject / course
Miscellaneous
Supervisors
Examiners
Available from: 2026-06-22 Created: 2026-06-08 Last updated: 2026-06-22Bibliographically approved

Open Access in DiVA

fulltext(2139 kB)30 downloads
File information
File name FULLTEXT01.pdfFile size 2139 kBChecksum SHA-512
034a93150ad55964c342e692e21d7b239ad5ddb57f2f29440e17548d14fa18af0ee196d5b264d4e6eaa948d85450806947f7c98858aef6dcda1eea1de41094d2
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Tidström, MattiasPearson, Andreas
By organisation
Mälardalen University
Robotics and automation

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 101 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf