To Örebro University

oru.seÖrebro University Publications
Change search
Link to record
Permanent link

Direct link
Publications (10 of 11) Show all publications
Toledo, E., Hambardzumyan, K., Josifoski, M., Hazra, R., Baldwin, N., Audran-Reiss, A., . . . Bachrach, Y. (2025). AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench. In: : . Paper presented at 39th Conference on Neural Information Processing Systems (NeurIPS 2025), San Diego Convention Center, San Diego, CA, United States, December 2-7, 2025. San Diego
Open this publication in new window or tab >>AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
Show others...
2025 (English)Conference paper, Published paper (Refereed)
Abstract [en]

AI research agents are demonstrating great potential to accelerate scientific progress by automating the design, implementation, and training of machine learning models. We focus on methods for improving agents' performance on MLE-bench, a challenging benchmark where agents compete in Kaggle competitions to solve real-world machine learning problems. We formalize AI research agents as search policies that navigate a space of candidate solutions, iteratively modifying them using operators. By designing and systematically varying different operator sets and search policies (Greedy, MCTS, Evolutionary), we show that their interplay is critical for achieving high performance. Our best pairing of search strategy and operator set achieves a state-of-the-art result on MLE-bench lite, increasing the success rate of achieving a Kaggle medal from 39.6% to 47.7%. Our investigation underscores the importance of jointly considering the search strategy, operator design, and evaluation methodology in advancing automated machine learning.

Place, publisher, year, edition, pages
San Diego: , 2025
Keywords
Deep learning (e.g., architectures, generative models, optimization for deep networks, foundation models, LLMs)
National Category
Computer Sciences
Identifiers
urn:nbn:se:oru:diva-125556 (URN)
Conference
39th Conference on Neural Information Processing Systems (NeurIPS 2025), San Diego Convention Center, San Diego, CA, United States, December 2-7, 2025
Note

NeurIPS 2025 spotlight

Available from: 2025-12-11 Created: 2025-12-11 Last updated: 2025-12-12Bibliographically approved
Hazra, R., Venturato, G., Zuidberg dos Martires, P. & De Raedt, L. (2025). Can Large Language Models Reason? A Characterization via 3-SAT. In: : . Paper presented at 13th International Conference on Learning Representations (ICLR 2025), Singapore, April 24-28, 2025.
Open this publication in new window or tab >>Can Large Language Models Reason? A Characterization via 3-SAT
2025 (English)Conference paper, Published paper (Refereed)
Abstract [en]

Large Language Models (LLMs) have been touted as AI models possessing advanced reasoning abilities. However, recent works have shown that LLMs often bypass true reasoning using shortcuts, sparking skepticism. To study the reasoning capabilities in a principled fashion, we adopt a computational theory perspective and propose an experimental protocol centered on 3-SAT – the prototypical NP-complete problem lying at the core of logical reasoning and constraint satisfaction tasks. Specifically, we examine the phase transitions in random 3-SAT and characterize the reasoning abilities of LLMs by varying the inherent hardness of the problem instances. Our experimental evidence shows that LLMs are incapable of performing true reasoning, as required for solving 3-SAT problems. Moreover, we observe significant performance variation based on the inherent hardness of the problems – performing poorly on harder instances and vice versa. Importantly ,we show that integrating external reasoners can considerably enhance LLM performance. By following a principled experimental protocol, our study draws concrete conclusions and moves beyond the anecdotal evidence often found in LLM reasoning research.

National Category
Computer Sciences
Identifiers
urn:nbn:se:oru:diva-123280 (URN)10.48550/arXiv.2408.07215 (DOI)
Conference
13th International Conference on Learning Representations (ICLR 2025), Singapore, April 24-28, 2025
Note

Published at ICLR 2025 Workshop on Reasoning and Planning for LLMs

Available from: 2025-09-01 Created: 2025-09-01 Last updated: 2025-09-01Bibliographically approved
Bachrach, Y., Toledo, E., Hambardzumyan, K., Magka, D., Josifoski, M., Jiang, M., . . . Hazra, R. (2025). Combining Code Generating Large Language Models and Self-Play to Iteratively Refine Strategies in Games. In: James Kwok (Ed.), Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI-25): . Paper presented at 34th International Joint Conference on Artificial Intelligence (IJCAI 2025), Montreal, Canada, August 16-22, 2025 (pp. 10999-11003).
Open this publication in new window or tab >>Combining Code Generating Large Language Models and Self-Play to Iteratively Refine Strategies in Games
Show others...
2025 (English)In: Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI-25) / [ed] James Kwok, 2025, p. 10999-11003Conference paper, Published paper (Refereed)
Abstract [en]

We propose a self-play approach to generating strategies for playing in multi-player games, where strategies are represented as computer code. We use large language models (LLMs) to generate pieces of code to play in the game, which we refer to as generated bots. We engage the LLM generated bots in competitions, designed to generate increasingly stronger strategies. We follow game theoretic principles in organizing these tournaments, and use a Policy Space Response Oracle (PSRO) approach. We start with an initial set of LLM generated bots, and continue in rounds for adding new bots into the population. Each round adds a bot to the population by asking the LLM to produce code for playing against a bot representing the Nash equilibrium mixture over the current population. Our analysis shows that even a few rounds are sufficient to produces strong bots for playing the game. Our demo shows the process for the game of Checkers. We allow users to select initial bots in the population, run the process, inspect how the bots evolve over time, and play against the generated bots.

Keywords
Agent-based and Multi-agent Systems, MAS, Applications Game Theory and Economic Paradigms, GTEP, Other Natural Language Processing, NLP, Language models
National Category
Computer Sciences
Identifiers
urn:nbn:se:oru:diva-125558 (URN)10.24963/ijcai.2025/1249 (DOI)001634925300605 ()
Conference
34th International Joint Conference on Artificial Intelligence (IJCAI 2025), Montreal, Canada, August 16-22, 2025
Note

Demo Track

Available from: 2025-12-11 Created: 2025-12-11 Last updated: 2026-01-22Bibliographically approved
Schreiter, T., Rüppel, J. V., Hazra, R., Rudenko, A., Magnusson, M. & Lilienthal, A. J. (2025). Evaluating Efficiency and Engagement in Scripted and LLM-Enhanced Human-Robot Interactions. In: 2025 20th ACM IEEE International Conference on Human Robot Interaction (HRI): . Paper presented at 20th International Conference on Human Robot Interaction (HRI 2025), Melbourne, Australia, March 4-6, 2025 (pp. 1608-1612). IEEE
Open this publication in new window or tab >>Evaluating Efficiency and Engagement in Scripted and LLM-Enhanced Human-Robot Interactions
Show others...
2025 (English)In: 2025 20th ACM IEEE International Conference on Human Robot Interaction (HRI), IEEE , 2025, p. 1608-1612Conference paper, Published paper (Refereed)
Abstract [en]

To achieve natural and intuitive interaction with people, HRI frameworks combine a wide array of methods for human perception, intention communication, human-aware navigation and collaborative action. In practice, when encountering unpredictable behavior of people or unexpected states of the environment, these frameworks may lack the ability to dynamically recognize such states, adapt and recover to resume the interaction. Large Language Models (LLMs), owing to their advanced reasoning capabilities and context retention, present a promising solution for enhancing robot adaptability. This potential, however, may not directly translate to improved interaction metrics. This paper considers a representative interaction with an industrial robot involving approach, instruction, and object manipulation, implemented in two conditions: (1) fully scripted and (2) including LLM-enhanced responses. We use gaze tracking and questionnaires to measure the participants' task efficiency, engagement, and robot perception. The results indicate higher SUbjective ratings for the LLM condition, but objective metrics show that the scripted condition performs comparably, particularly in efficiency and focus during simple tasks. We also note that the scripted condition may have an edge over LLM-enhanced responses in terms of response latency and energy consumption, especially for trivial and repetitive interactions.

Place, publisher, year, edition, pages
IEEE, 2025
Series
ACM/IEEE International Conference on Human-Robot Interaction (HRI), ISSN 2167-2121, E-ISSN 2167-2148
Keywords
Human-Robot Interaction, AI-Enabled Robotics
National Category
Human Computer Interaction Computer Sciences
Identifiers
urn:nbn:se:oru:diva-124901 (URN)10.1109/HRI61500.2025.10974124 (DOI)001492540600219 ()9798350378948 (ISBN)9798350378931 (ISBN)
Conference
20th International Conference on Human Robot Interaction (HRI 2025), Melbourne, Australia, March 4-6, 2025
Funder
EU, Horizon 2020, 101017274
Available from: 2025-11-11 Created: 2025-11-11 Last updated: 2025-11-11Bibliographically approved
Hazra, R., Venturato, G., Zuidberg dos Martires, P. & De Raedt, L. (2025). Have Large Language Models Learned to Reason? A Characterization via 3-SAT. In: : . Paper presented at Second Conference on Language Modeling (COLM 2025), Montreal, Canada, October 7-10, 2025.
Open this publication in new window or tab >>Have Large Language Models Learned to Reason? A Characterization via 3-SAT
2025 (English)Conference paper, Published paper (Refereed)
Abstract [en]

Large Language Models (LLMs) have been touted as AI models possessing advanced reasoning abilities. In theory, autoregressive LLMs with Chain-of-Thought (CoT) can perform more serial computations to solve complex reasoning tasks. However, recent studies suggest that, despite this capacity, LLMs do not truly learn to reason but instead fit on statistical features. To study the reasoning capabilities in a principled fashion, we adopt a computational theory perspective and propose an experimental protocol centered on 3-SAT -- the prototypical NP-complete problem lying at the core of logical reasoning and constraint satisfaction tasks. Specifically, we examine the phase transitions in random 3-SAT and characterize the reasoning abilities of state-of-the-art LLMs by varying the inherent hardness of the problem instances. By comparing DeepSeek R1 with other LLMs, our findings reveal two key insights (1) LLM accuracy drops significantly on harder instances, suggesting all current models struggle when statistical shortcuts are unavailable (2) Unlike other LLMs, R1 shows signs of having learned the underlying reasoning. Following a principled experimental protocol, our study moves beyond the benchmark-driven evidence often found in LLM reasoning research. Our findings highlight important gaps and suggest clear directions for future research. 

Keywords
Large Language Models, Reasoning, Computational Complexity, Logic, Satisfiability, Phase Transitions
National Category
Computer Sciences
Identifiers
urn:nbn:se:oru:diva-125557 (URN)
Conference
Second Conference on Language Modeling (COLM 2025), Montreal, Canada, October 7-10, 2025
Available from: 2025-12-11 Created: 2025-12-11 Last updated: 2025-12-12Bibliographically approved
Mantenoglou, P., Hazra, R., Zuidberg dos Martires, P. & De Raedt, L. (2025). LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language. In: : . Paper presented at 39th Conference on Neural Information Processing Systems (NeurIPS 2025), San Diego Convention Center, San Diego, CA, United States, December 2-7, 2025.
Open this publication in new window or tab >>LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
2025 (English)Conference paper, Published paper (Refereed)
Abstract [en]

Owing to their reasoning capabilities, large language models (LLMs) have been evaluated on planning tasks described in natural language. However, LLMs have largely been tested on planning domains without constraints. In order to deploy them in real-world settings where adherence to constraints, in particular safety constraints, is critical, we need to evaluate their performance on constrained planning tasks. We introduce LexiCon -- a natural language-based (Lexi) constrained (Con) planning benchmark, consisting of a suite of environments, that can be used to evaluate the planning capabilities of LLMs in a principled fashion. The core idea behind LexiCon is to take existing planning environments and impose temporal constraints on the states. These constrained problems are then translated into natural language and given to an LLM to solve. A key feature of LexiCon is its extensibility. That is, the set of supported environments can be extended with new (unconstrained) environment generators, for which temporal constraints are constructed automatically. This renders LexiCon future-proof: the hardness of the generated planning problems can be increased as the planning capabilities of LLMs improve. Our experiments reveal that the performance of state-of-the-art LLMs, including reasoning models like GPT-5, o3, and R1, deteriorates as the degree of constrainedness of the planning tasks increases. 

National Category
Artificial Intelligence
Identifiers
urn:nbn:se:oru:diva-125883 (URN)
Conference
39th Conference on Neural Information Processing Systems (NeurIPS 2025), San Diego Convention Center, San Diego, CA, United States, December 2-7, 2025
Available from: 2025-12-21 Created: 2025-12-21 Last updated: 2026-01-05Bibliographically approved
Hazra, R. (2025). Neurosymbolic Decision-Making with Large Language Models. (Doctoral dissertation). Örebro: Örebro University
Open this publication in new window or tab >>Neurosymbolic Decision-Making with Large Language Models
2025 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Reasoning and decision-making are foundational challenges in artificial intelligence (AI). These processes are closely linked – an intelligent agent must reason about its environment and goals in order to make decisions and select actions. Two principal frameworks for sequential decision-making are AI planning and reinforcement learning (RL). Planning assumes access to a known model of the environment and uses symbolic representations to compute a sequence of actions that leads from an initial state to a desired goal. In contrast, RL focuse son learning behavior through interaction, enabling agents to develop policies that maximize long-term reward under uncertainty. Despite methodological differences, both approaches aim to generate intelligent, goal-directed action sequences.

The rise of Large Language Models (LLMs) has sparked significant interest in their potential to perform reasoning, planning, and decision-making tasks. Despite their impressive performance in natural language understanding and generalization, there is growing skepticism about whether LLMs genuinely reason or merely leverage statistical correlations. This dissertation investigates this question through a principled evaluation grounded in computational theory, using 3-SAT – the canonical NP-complete problem – as a testbed. The findings demonstrate that LLMs fail to exhibit sound and complete reasoning, especially on complex instances where shallow heuristics fail, and that their apparent reasoning abilities often stem from overfitting to statistical patterns.

To address these limitations, this dissertation proposes a range of neurosymbolic architectures that combine the generative flexibility of LLMs with the rigor and reliability of symbolic methods. Empirical evaluations across planning, reward design, and plan verification tasks show that such integration yields systems that are more robust and accurate. This work advances our theoretical and practical understanding of LLM-based reasoning, provides concrete design principles for neurosymbolic systems, and charts a path toward AI agents that integrate world knowledge with logical precision.

Place, publisher, year, edition, pages
Örebro: Örebro University, 2025. p. 67
Series
Örebro Studies in Technology, ISSN 1650-8580 ; 106
National Category
Computer Sciences
Identifiers
urn:nbn:se:oru:diva-122456 (URN)9789175296869 (ISBN)
Public defence
2025-10-17, Örebro universitet, Långhuset, Hörsal L2, Fakultetsgatan 1, Örebro, 13:00 (English)
Opponent
Supervisors
Available from: 2025-07-22 Created: 2025-07-22 Last updated: 2025-09-04Bibliographically approved
Hazra, R., Sygkounas, A., Persson, A., Loutfi, A. & Zuidberg dos Martires, P. (2025). REvolve: Reward Evolution with Large Language Models using Human Feedback. In: 13th International Conference on Learning Representations (ICLR 2025): Proceedings. Paper presented at 13th International Conference on Learning Representations (ICLR 2025), Singapore, April 24-28, 2025 (pp. 25710-25751). International Conference on Learning Representations, ICLR
Open this publication in new window or tab >>REvolve: Reward Evolution with Large Language Models using Human Feedback
Show others...
2025 (English)In: 13th International Conference on Learning Representations (ICLR 2025): Proceedings, International Conference on Learning Representations, ICLR , 2025, p. 25710-25751Conference paper, Published paper (Refereed)
Abstract [en]

Designing effective reward functions is crucial to training reinforcement learning (RL) algorithms. However, this design is non-trivial, even for domain experts, due to the subjective nature of certain tasks that are hard to quantify explicitly. In recent works, large language models (LLMs) have been used for reward generation from natural language task descriptions, leveraging their extensive instruction tuning and commonsense understanding of human behavior. In this work, we hypothesize that LLMs, guided by human feedback, can be used to formulate reward functions that reflect human implicit knowledge. We study this in three challenging settings - autonomous driving, humanoid locomotion, and dexterous manipulation - wherein notions of “good” behavior are tacit and hard to quantify. To this end, we introduce REvolve, a truly evolutionary framework that uses LLMs for reward design in RL. REvolve generates and refines reward functions by utilizing human feedback to guide the evolution process, effectively translating implicit human knowledge into explicit reward functions for training (deep) RL agents. Experimentally, we demonstrate that agents trained on REvolve-designed rewards outperform other state-of-the-art baselines. 

Place, publisher, year, edition, pages
International Conference on Learning Representations, ICLR, 2025
National Category
Computer Sciences
Identifiers
urn:nbn:se:oru:diva-123277 (URN)10.48550/arXiv.2406.01309 (DOI)2-s2.0-105010222426 (Scopus ID)9798331320850 (ISBN)
Conference
13th International Conference on Learning Representations (ICLR 2025), Singapore, April 24-28, 2025
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)Knut and Alice Wallenberg Foundation
Available from: 2025-09-01 Created: 2025-09-01 Last updated: 2026-01-16Bibliographically approved
Hazra, R., Zuidberg dos Martires, P. & De Raedt, L. (2024). SayCanPay: Heuristic Planning with Large Language Models Using Learnable Domain Knowledge. In: Michael Wooldridge; Jennifer Dy; Sriraam Natarajan (Ed.), Proceedings of the 38th AAAI Conference on Artificial Intelligence: . Paper presented at 38th AAAI Conference on Artificial Intelligence (AAAI) / 36th Conference on Innovative Applications of Artificial Intelligence / 14th Symposium on Educational Advances in Artificial Intelligence, Vancouver, Canada, February 20-27, 2024 (pp. 20123-20133). AAAI Press, 38
Open this publication in new window or tab >>SayCanPay: Heuristic Planning with Large Language Models Using Learnable Domain Knowledge
2024 (English)In: Proceedings of the 38th AAAI Conference on Artificial Intelligence / [ed] Michael Wooldridge; Jennifer Dy; Sriraam Natarajan, AAAI Press, 2024, Vol. 38, p. 20123-20133Conference paper, Published paper (Refereed)
Abstract [en]

Large Language Models (LLMs) have demonstrated impressive planning abilities due to their vast "world knowledge". Yet, obtaining plans that are both feasible (grounded in affordances) and cost-effective (in plan length), remains a challenge, despite recent progress. This contrasts with heuristic planning methods that employ domain knowledge (formalized in action models such as PDDL) and heuristic search to generate feasible, optimal plans. Inspired by this, we propose to combine the power of LLMs and heuristic planning by leveraging the world knowledge of LLMs and the principles of heuristic search. Our approach, SayCanPay, employs LLMs to generate actions (Say) guided by learnable domain knowledge, that evaluates actions' feasibility (Can) and long-term reward/payoff (Pay), and heuristic search to select the best sequence of actions. Our contributions are (1) a novel framing of the LLM planning problem in the context of heuristic planning, (2) integrating grounding and cost-effective elements into the generated plans, and (3) using heuristic search over actions. Our extensive evaluations show that our model surpasses other LLM planning approaches.

Place, publisher, year, edition, pages
AAAI Press, 2024
Series
Proceedings of the AAAI Conference on Artificial Intelligence, ISSN 2159-5399, E-ISSN 2374-3468 ; 38:18
National Category
Computer Sciences
Identifiers
urn:nbn:se:oru:diva-115501 (URN)10.1609/aaai.v38i18.29991 (DOI)001241509500037 ()2-s2.0-85189544071 (Scopus ID)9781577358879 (ISBN)
Conference
38th AAAI Conference on Artificial Intelligence (AAAI) / 36th Conference on Innovative Applications of Artificial Intelligence / 14th Symposium on Educational Advances in Artificial Intelligence, Vancouver, Canada, February 20-27, 2024
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)Knut and Alice Wallenberg FoundationEU, Horizon 2020, 952215
Note

This work was supported by the Wallenberg AI, Autonomous Systems and Software Program (WASP) funded by the Knut and Alice Wallenberg Foundation, and is also part of the EU H2020 ICT48 project “TAILOR” under contract 952215, and the KU Leuven Research Fund (C14/18/062).

Available from: 2024-08-21 Created: 2024-08-21 Last updated: 2025-09-01Bibliographically approved
Hazra, R. & De Raedt, L. (2023). Deep Explainable Relational Reinforcement Learning: A Neuro-Symbolic Approach. In: Danai Koutra; Claudia Plant; Manuel Gomez Rodriguez; Elena Baralis; Francesco Bonchi (Ed.), Machine Learning and Knowledge Discovery in Databases: Research Track: European Conference, ECML PKDD 2023, Turin, Italy, September 18–22, 2023, Proceedings, Part IV. Paper presented at European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2023), Turin, Italy, September 18-22, 2023 (pp. 213-229). Springer, 14172
Open this publication in new window or tab >>Deep Explainable Relational Reinforcement Learning: A Neuro-Symbolic Approach
2023 (English)In: Machine Learning and Knowledge Discovery in Databases: Research Track: European Conference, ECML PKDD 2023, Turin, Italy, September 18–22, 2023, Proceedings, Part IV / [ed] Danai Koutra; Claudia Plant; Manuel Gomez Rodriguez; Elena Baralis; Francesco Bonchi, Springer, 2023, Vol. 14172, p. 213-229Conference paper, Published paper (Refereed)
Abstract [en]

Despite its successes, Deep Reinforcement Learning (DRL) yields non-interpretable policies. Moreover, since DRL does not exploit symbolic relational representations, it has difficulties in coping with structural changes in its environment (such as increasing the number of objects).  Meanwhile, Relational Reinforcement Learning inherits the relational representations from symbolic planning to learn reusable policies. However, it has so far been unable to scale up and exploit the power of deep neural networks. We propose Deep Explainable Relational Reinforcement Learning (DERRL), a framework that exploits the best of both -- neural and symbolic worlds. By resorting to a neuro-symbolic approach, DERRL combines relational representations and constraints from symbolic planning with deep learning to extract interpretable policies. These policies are in the form of logical rules that explain why each decision (or action) is arrived at. Through several experiments, in setups like the Countdown Game, Blocks World, Gridworld, Traffic, and Mingrid, we show that the policies learned by DERRL are adaptable to varying configurations and environmental changes.

Place, publisher, year, edition, pages
Springer, 2023
Series
Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349 ; 14172
Keywords
Neuro-Symbolic AI, Relational Reinforcement Learning, Deep Reinforcement Learning, Explainability
National Category
Computer Sciences
Research subject
Computer and Systems Science; Computer Science
Identifiers
urn:nbn:se:oru:diva-108100 (URN)10.48550/arXiv.2304.08349 (DOI)001156141200013 ()9783031434204 (ISBN)9783031434211 (ISBN)
Conference
European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2023), Turin, Italy, September 18-22, 2023
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Available from: 2023-09-05 Created: 2023-09-05 Last updated: 2025-09-01Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0003-3422-2085

Search in DiVA

Show all publications

Profile pages