To Örebro University

oru.seÖrebro University Publications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Can context bridge the reality gap? Sim-to-real transfer of context-aware policies
Örebro University, School of Science and Technology. (AASS Research Centre)ORCID iD: 0000-0002-2142-6516
Örebro University, School of Science and Technology. (AASS Research Centre)ORCID iD: 0000-0003-1528-4301
Örebro University, School of Science and Technology. (AASS Research Centre)ORCID iD: 0000-0003-3958-6179
Technology Transfer Center Kitzingen, Technical University of Applied Sciences Würzburg-Schweinfurt, Kitzingen, Germany.
Show others and affiliations
2026 (English)In: Robotics and Autonomous Systems, ISSN 0921-8890, E-ISSN 1872-793X, Vol. 205, article id 105594Article in journal (Refereed) Published
Abstract [en]

Sim-to-real transfer remains a major challenge in reinforcement learning (RL) for robotics, as policies trained in simulation often fail to generalize to the real world due to discrepancies in environment dynamics. Domain Randomization (DR) mitigates this issue by exposing the policy to a wide range of randomized dynamics during training, yet leading to a reduction in performance. While standard approaches typically train policies agnostic to these variations, we investigate whether sim-to-real transfer can be improved by conditioning the policy on an estimate of the dynamics parameters - referred to as context. To this end, we integrate a context estimation module into a DR-based RL framework and systematically compare SOTA supervision strategies. We evaluate the resulting context-aware policies in both a canonical control benchmark and a real-world pushing task using a Franka Emika Panda robot. Results show that context-aware policies outperform the context-agnostic baseline across all settings, although the best supervision strategy depends on the task.

Place, publisher, year, edition, pages
Elsevier, 2026. Vol. 205, article id 105594
Keywords [en]
Robotics, Reinforcement learning, Sim-to-real
National Category
Computer Sciences Artificial Intelligence Robotics and automation
Identifiers
URN: urn:nbn:se:oru:diva-130287DOI: 10.1016/j.robot.2026.105594ISI: 001819233500001OAI: oai:DiVA.org:oru-130287DiVA, id: diva2:2088614
Funder
Knowledge Foundation, 20190128Knut and Alice Wallenberg FoundationWallenberg AI, Autonomous Systems and Software Program (WASP)
Note

This work was supported in part by Industrial Graduate Schoo lCollaborative AI & Robotics (CoAIRob), in part by the Swedish Knowledge Foundation under Grant Dnr:20190128, and the Knut and Alice Wallenberg Foundation through Wallenberg AI, Autonomous Systems and Software Program (WASP).

Available from: 2026-07-29 Created: 2026-07-29 Last updated: 2026-07-29Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full text

Authority records

Iannotta, MarcoYang, YuxuanStork, Johannes A.Stoyanov, Todor

Search in DiVA

By author/editor
Iannotta, MarcoYang, YuxuanStork, Johannes A.Stoyanov, Todor
By organisation
School of Science and Technology
In the same journal
Robotics and Autonomous Systems
Computer SciencesArtificial IntelligenceRobotics and automation

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 6 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf