To Örebro University

oru.seÖrebro University Publications
Change search
Link to record
Permanent link

Direct link
Bhatt, Mehul, ProfessorORCID iD iconorcid.org/0000-0002-6290-5492
Biography [eng]

 

 
Biography [swe]

 

 
Publications (10 of 142) Show all publications
Kondyli, V. & Bhatt, M. (2026). A Visuospatial Complexity Framework for Analysing Anticipation in Naturalistic Active Vision. In: EWIC 2026:  The 19th European Workshop on Imagery and Cognition: Book of abstracts. Paper presented at The 19th European Workshop on Imagery and Cognition (EWIC 2026), Leiden, The Netherlands, June 24-26, 2026 (pp. 9-9).
Open this publication in new window or tab >>A Visuospatial Complexity Framework for Analysing Anticipation in Naturalistic Active Vision
2026 (English)In: EWIC 2026:  The 19th European Workshop on Imagery and Cognition: Book of abstracts, 2026, p. 9-9Conference paper, Oral presentation with published abstract (Refereed)
Abstract [en]

Understanding active vision in naturalistic settings requires examining how the brain generates predictions and anticipates events under dynamic and uncertain conditions. The study of active vision necessitates investigation of coordinated interactions across distributed brain networks supporting diverse introspective and anticipatory functions while at the same time adopting a naturalistic stimulus design approach. Towards this, we propose a cognitive visuospatial complexity model that enables systematic parametrisation of dynamic visual stimuli for functional neuroimaging, behavioural experimentation, and psychophysics. The model supports controlled construction of graded complexity levels through parametric manipulation and is designed to investigate active vision and event‑based anticipation in ecologically valid scenarios. Our methodology defines an abstraction‑to‑realism axis capturing key prediction‑relevant dimensions of visuospatial complexity, including occlusions, contextual continuity, temporal regularity, event‑based anticipation, and constraints from commonsense and naïve physics. We additionally provide an accompanying stimulus dataset that demonstrates how the model can be implemented in practice for neuroimaging applications. The framework provides a standardised and reproducible stimulus‑design space, enhances cross‑study comparability, enables multimodal data integration, and supports the testing computational models of active vision and predictive processing. Our aim is to advance systematic methodological foundations for the neurocognitive and behavioural study of active vision under ecologically valid naturalistic conditions.

National Category
Psychology Computer Sciences
Research subject
Computer Science; Psychology
Identifiers
urn:nbn:se:oru:diva-129607 (URN)
Conference
The 19th European Workshop on Imagery and Cognition (EWIC 2026), Leiden, The Netherlands, June 24-26, 2026
Funder
Swedish Research Council, 2022-02960_VR
Available from: 2026-06-24 Created: 2026-06-24 Last updated: 2026-07-21Bibliographically approved
Nair, V., Bhatt, M., Suchan, J., Billing, E. & Hemeren, P. (2026). How do Naturalistic Visuo-Auditory Cues Guide Human Attention? Insights from Systematic Explorations in Visual Perception of Embodied Multimodal Interaction. ACM Transactions on Applied Perception
Open this publication in new window or tab >>How do Naturalistic Visuo-Auditory Cues Guide Human Attention? Insights from Systematic Explorations in Visual Perception of Embodied Multimodal Interaction
Show others...
2026 (English)In: ACM Transactions on Applied Perception, ISSN 1544-3558, E-ISSN 1544-3965Article in journal (Refereed) Epub ahead of print
Abstract [en]

Studies in visual cognition highlight the importance of visual, spatial, and auditory cues in influencing human attention. Such cues often tend to be indicative of actions or events, thereby serving as predictive indicators in both passive observation as well as in interactive engagement. Our research focuses on visual attention in passive observation, particularly examining the manner in which visual, spatial, and auditory cues —henceforth visuoauditory (shorthand) cues— influence attention on everyday multimodal interaction. We systematically develop a visuoauditory event model for investigating visual attention in naturalistic embodied settings. Rooted in this event model, we explore the influence of five select visuoauditory cues —namely, speaking, gaze, relative motion, hand action, and visibility— on visual attention. Our analysis utilizes eye-tracking data from (90) participants observing (27) carefully designed naturalistic event scenarios and correlating their attentional metrics with the select visuoauditory cues in the backdrop of the developed event model. Findings reveal strong associations between attention and both intra-modal (irrespective of other cues) and cross-modal (combined with other cues) cueing effects, thereby highlighting the nuanced interplay amongst the cues influencing attentional patterns. We develop a systematic and generalized method for analyzing interactions and behavioral parameters, thereby characterizing the impact of visuoauditory cues on attentional dynamics. Our methodology, combined with the obtained insights into the attentional cueing effects, provides an analytical framework explicating the manner in which everyday (interactive) events directly drive attention under naturalistic conditions. This facilitates not only the precise modeling of behavior and attention allocation but also offers a high-level ‘experimental lens’ for examining interactions in relation to behavioral parameters. Taken together, our methodological and behavioral findings are well-positioned to benefit multiple fields, particularly by advancing human-centered design across diverse application domains. Lastly, towards promoting open-science and for wider dissemination, the complete experimental basis of this research —e.g., event-scenarios, high-quality annotated data, data-set supplementary— have been independently documented together with instructions to use and access experimental data.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2026
Keywords
Visuoauditory cues, Attentional flow, Human-interaction, Eye-tracking, Semantic grounding
National Category
Computer Sciences Human Computer Interaction Psychology
Research subject
Computer Science; Media and Communication Studies
Identifiers
urn:nbn:se:oru:diva-128523 (URN)10.1145/3807941 (DOI)
Funder
Swedish Research Council
Available from: 2026-04-23 Created: 2026-04-23 Last updated: 2026-04-24Bibliographically approved
Suchan, J., Baloch, S. & Bhatt, M. (2026). Towards a VLM-Based Foundation for Generalised Neurosymbolic Visual Commonsense. In: Anni-Yasmin Turhan; Jonni Virtema (Ed.), Foundations of Information and Knowledge Systems: 14th International Symposium, FoIKS 2026, Hanover, Germany, March 23–26, 2026, Proceedings. Paper presented at 14th International Symposium (FoIKS 2026), Hanover, Germany, March 23–26, 2026 (pp. 359-365). Cham: Springer
Open this publication in new window or tab >>Towards a VLM-Based Foundation for Generalised Neurosymbolic Visual Commonsense
2026 (English)In: Foundations of Information and Knowledge Systems: 14th International Symposium, FoIKS 2026, Hanover, Germany, March 23–26, 2026, Proceedings / [ed] Anni-Yasmin Turhan; Jonni Virtema, Cham: Springer, 2026, p. 359-365Conference paper, Published paper (Refereed)
Abstract [en]

Commonsense visual sensemaking requires robust mechanisms to extract scene elements and other visual features from perceived imagery, as well as rich conceptual commonsense knowledge and reasoning to interpret the perceived interactive dynamic. In this context, we position ongoing work on using integrated Vision Language Models (VLM) and Answer Set Programming (ASP) based reasoning about (dynamic) spatial configurations using VLMs as a neural foundation for symbolic modeling of dynamic scene structures. We sketch preliminary results highlighting practical usability by application to the task of Visual Question Answering (VQA) with the STRIDE-QA driving dataset.

Place, publisher, year, edition, pages
Cham: Springer, 2026
Series
Lecture Notes in Computer Science (LNCS), ISSN 0302-9743, E-ISSN 1611-3349
Keywords
Visual Commonsense, Declarative AI, Neurosymbolic AI
National Category
Computer and Information Sciences Computer Vision and Learning Systems Human Computer Interaction Robotics and automation
Research subject
Computer Science
Identifiers
urn:nbn:se:oru:diva-128097 (URN)10.1007/978-3-032-21540-6_24 (DOI)001763366900024 ()9783032215406 (ISBN)9783032215390 (ISBN)
Conference
14th International Symposium (FoIKS 2026), Hanover, Germany, March 23–26, 2026
Projects
Counterfactual Commonsense
Funder
Swedish Research Council
Available from: 2026-03-22 Created: 2026-03-22 Last updated: 2026-07-23Bibliographically approved
Chaudhri, V. K., Baru, C., Bennett, B., Bhatt, M., Cassel, D., Cohn, A. G., . . . Witbrock, M. (2025). A community-driven vision for a new knowledge resource for AI. The AI Magazine, 46(4), Article ID e70035.
Open this publication in new window or tab >>A community-driven vision for a new knowledge resource for AI
Show others...
2025 (English)In: The AI Magazine, ISSN 0738-4602, E-ISSN 2371-9621, Vol. 46, no 4, article id e70035Article in journal (Refereed) Published
Abstract [en]

The long-standing goal of creating a comprehensive, multi-purpose knowledge resource, reminiscent of the 1984 Cyc project, still persists in AI. Despite the success of knowledge resources like WordNet, ConceptNet, Wolfram|Alpha and other commercial knowledge graphs, verifiable, general-purpose, widely available sources of knowledge remain a critical deficiency in AI infrastructure. Large language models struggle due to knowledge gaps; robotic planning lacks necessary world knowledge; and the detection of factually false information relies heavily on human expertise. What kind of knowledge resource is most needed in AI today? How can modern technology shape its development and evaluation? A recent AAAI workshop gathered over 50 researchers to explore these questions. This paper synthesizes our findings and outlines a community-driven vision for a new knowledge infrastructure. In addition to leveraging contemporary advances in knowledge representation and reasoning, one promising idea is to build an open engineering framework to exploit knowledge modules effectively within the context of practical applications. Such a framework should include sets of conventions and social structures that are adopted by contributors.

Place, publisher, year, edition, pages
John Wiley & Sons, 2025
National Category
Artificial Intelligence
Identifiers
urn:nbn:se:oru:diva-124792 (URN)10.1002/aaai.70035 (DOI)001599790300001 ()2-s2.0-105020256772 (Scopus ID)
Note

Funding agency: 

National Science Foundation (NSF) 2514820

Available from: 2025-11-05 Created: 2025-11-05 Last updated: 2026-01-23Bibliographically approved
Kondyli, V., Suchan, J. & Bhatt, M. (2025). A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision. In: ECAI 2025 - European Conference on Artificial Intelligence, part of AIC 2025, 10th Edition: Workshop on Artificial Intelligence and Cognition as part of ECAI 2025.: . Paper presented at 28th European Conference on Artificial Intelligence (ECAI 2025), Bologna, Italy, October 25-30, 2025.
Open this publication in new window or tab >>A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision
2025 (English)In: ECAI 2025 - European Conference on Artificial Intelligence, part of AIC 2025, 10th Edition: Workshop on Artificial Intelligence and Cognition as part of ECAI 2025., 2025Conference paper, Published paper (Refereed)
National Category
Computer Sciences Human Computer Interaction Applied Psychology Robotics and automation
Research subject
Computer Science
Identifiers
urn:nbn:se:oru:diva-125913 (URN)
Conference
28th European Conference on Artificial Intelligence (ECAI 2025), Bologna, Italy, October 25-30, 2025
Funder
Swedish Foundation for Strategic Research
Available from: 2025-12-27 Created: 2025-12-27 Last updated: 2026-01-05Bibliographically approved
Suchan, J., Bhatt, M. & Monsen, J. (2025). ASP-driven visual commonsense: A general framework for reasoning about embodied interaction in the wild. In: Magdalena Ortiz; Renata Wassermann; Torsten Schaub (Ed.), KR 2025: Proceedings of the 22nd International Conference on Principles of Knowledge Representation and Reasoning. Paper presented at 22nd International Conference on Principles of Knowledge Representation and Reasoning (KR 2025), Melbourne, Australia, November 11-17, 2025 (pp. 632-642). AAAI Press, International Joint Conferences on Artificial Intelligence (IJCAI)
Open this publication in new window or tab >>ASP-driven visual commonsense: A general framework for reasoning about embodied interaction in the wild
2025 (English)In: KR 2025: Proceedings of the 22nd International Conference on Principles of Knowledge Representation and Reasoning / [ed] Magdalena Ortiz; Renata Wassermann; Torsten Schaub, AAAI Press, International Joint Conferences on Artificial Intelligence (IJCAI) , 2025, p. 632-642Conference paper, Published paper (Refereed)
Abstract [en]

We present a general framework for declaratively grounded visual commonsense (reasoning) about embodied interaction in naturalistic, in-the-wild settings relevant to a range of AI application domains. The core computational capabilities of the framework pertaining visual commonsense are driven by a robust neurosymbolic architecture primarily consisting of: (1) answer set programming based modelling of foundational aspects pertaining spatio-temporal dynamics, encompassing space, time, events, action, motion; (2) modularly integrated visual computing techniques constituting the neural substrate linking quantitative perceptual features serving as low-level counterparts to high-level semantic characterisations of (inter)active visual commonsense.

Practically, we also present a first open-release of the developed framework with the aim to promote independent extensions and real-world applied KRR. The release comprises: (a) demonstrated case-studies in domains such as autonomous driving, psychology and media studies; (b) systematic evaluation mechanisms for community benchmarking; and (c) supporting material such as tutorials and datasets.

Place, publisher, year, edition, pages
AAAI Press, International Joint Conferences on Artificial Intelligence (IJCAI), 2025
Series
Proceedings of the Conference on Principles of Knowledge Representation and Reasoning (KR), ISSN 2334-1025, E-ISSN 2334-1033
National Category
Computer Sciences
Research subject
Computer Science
Identifiers
urn:nbn:se:oru:diva-125906 (URN)10.24963/kr.2025/61 (DOI)9781956792089 (ISBN)
Conference
22nd International Conference on Principles of Knowledge Representation and Reasoning (KR 2025), Melbourne, Australia, November 11-17, 2025
Projects
Counterfactual Commonsense
Funder
Swedish Research Council
Available from: 2025-12-23 Created: 2025-12-23 Last updated: 2026-01-13Bibliographically approved
Kondyli, V. & Bhatt, M. (2025). Change Blindness and Anticipation in Naturalistic Driving: The Role of Visuospatial and Temporal Complexity. In: TeaP 2025: Conference of Experimental Psychologists / Tagung experimentell arbeitender Psychologinnen, TeaP. Paper presented at 67th Conference of Experimental Psychologists (TeaP 2025), Frankfurt, Germany, March 9-12, 2025.
Open this publication in new window or tab >>Change Blindness and Anticipation in Naturalistic Driving: The Role of Visuospatial and Temporal Complexity
2025 (English)In: TeaP 2025: Conference of Experimental Psychologists / Tagung experimentell arbeitender Psychologinnen, TeaP, 2025Conference paper, Published paper (Refereed)
Abstract [en]

Everyday driving relies on cognitive functions such as visual search, change detection and predictive attention for maintaining a high-level mental model of the world. Such functions are especially constrained by limited cognitive resources under complex conditions. Understanding the manner in which environmental complexity influences cognitive driving performance is key to improving road safety and developing human-factors guided driver assistance systems. Towards this, we explore the relationship between visuospatial and temporal complexity, and drivers’ situational awareness and decision-making. Particularly, we address two research questions: (I) How environmental complexity (visuospatial and temporal) impacts (in)attentional blindness; and (II) How does visual cueing impact anticipatory attention. In a naturalistic VR study, 150 participants drove in an immersive urban environment with varying complexity levels while responding to conditions such as overtaking, pedestrian crossings, occlusions, safety-criticality. The experiment, rooted to real-world driving data, incorporates a systematic cognitive model of environmental complexity consisting of dynamic visual, spatial, multimodal interaction features. Furthermore, we also analyse multimodal behavioural data, including gaze patterns, head movements, steering and braking. Results show that: (1) High visuospatial complexity increases inattentional blindness, with drivers compensating by prioritizing critical events; (2) Temporal complexity impairs detection performance, especially when events occur in rapid succession (1 sec) or when they last 5-10 seconds; and (3) Visual cueing improves anticipation, detection performance, and reaction time. These findings highlight the importance of adaptive strategies in overcoming cognitive limitations, providing insights for designing safety systems and training protocols for improving or assessing driver performance in everyday driving environments.

Keywords
change blindness, visual search, anticipation, environmental complexity, naturalistic observation, attentional strategies, driver education
National Category
Applied Psychology Human Computer Interaction Computer Sciences
Identifiers
urn:nbn:se:oru:diva-125909 (URN)
Conference
67th Conference of Experimental Psychologists (TeaP 2025), Frankfurt, Germany, March 9-12, 2025
Funder
Swedish Research CouncilSwedish Foundation for Strategic Research
Available from: 2025-12-23 Created: 2025-12-23 Last updated: 2026-01-13Bibliographically approved
Monsen, J., Suchan, J. & Bhatt, M. (2025). Probabilistic Answer Set Programming Driven Ranking of Dynamic Space-Time Belief Models. In: Aidan Hogan; Ken Satoh; Hasan Dağ; Anni-Yasmin Turhan; Dumitru Roman; Ahmet Soylu (Ed.), Rules and Reasoning: 9th International Joint Conference, RuleML+RR 2025, Istanbul, Turkey, September 22-24, 2025, Proceedings. Paper presented at 9th International Joint Conference, RuleML+RR 2025, Istanbul, Turkey, September 22-24, 2025 (pp. 156-175). Springer, 16144
Open this publication in new window or tab >>Probabilistic Answer Set Programming Driven Ranking of Dynamic Space-Time Belief Models
2025 (English)In: Rules and Reasoning: 9th International Joint Conference, RuleML+RR 2025, Istanbul, Turkey, September 22-24, 2025, Proceedings / [ed] Aidan Hogan; Ken Satoh; Hasan Dağ; Anni-Yasmin Turhan; Dumitru Roman; Ahmet Soylu, Springer , 2025, Vol. 16144, p. 156-175Conference paper, Published paper (Refereed)
Abstract [en]

A key challenge in embodied, inter(active) vision is reasoning over alternative hypotheses about the dynamics of perceived objects and events, be it for real-time or even offline interpretation. Towards this, we address the problem of generating and ranking grounded visuospatial hypotheses based on a semantically encoded notion of hypothesis preference. Driven by probabilistic Answer Set Programming (ASP), we propose a general framework for modeling and reasoning about diverse preference types tailored to visuospatial interpretation tasks. The effectiveness of our probabilistic visuospatial hypotheses ranking method is demonstrated and evaluated with a community benchmark of Multi-Object Tracking (MOT17), where modeling uncertainty and preference is critical for robust scene interpretation. Furthermore, practical examples also showcase how semantically driven reasoning with preferences can be effectively used in real-world visual sensemaking tasks.

Place, publisher, year, edition, pages
Springer, 2025
Series
Lecture Notes in Computer Science (LNCS), ISSN 0302-9743, E-ISSN 1611-3349 ; 16144
Keywords
Probabilistic Answer Set Programming, Preferential Ranking, Visual Intelligence, Deep Semantics, Cognitive Vision
National Category
Computer Sciences Artificial Intelligence
Research subject
Computer Science
Identifiers
urn:nbn:se:oru:diva-125907 (URN)10.1007/978-3-032-08887-1_10 (DOI)001657534200012 ()9783032088864 (ISBN)9783032088871 (ISBN)
Conference
9th International Joint Conference, RuleML+RR 2025, Istanbul, Turkey, September 22-24, 2025
Funder
Swedish Research Council
Available from: 2025-12-23 Created: 2025-12-23 Last updated: 2026-02-05Bibliographically approved
Bhatt, M. (2025). Towards Responsible AI Foundations for Neurocognitive Analytics of Vision. In: ECVP 2025: 47th European Conference on Visual Perception, Mainz, Germany. Paper presented at 47th European Conference on Visual Perception (ECVP 2025), Mainz, Germany, August 24-28, 2025.
Open this publication in new window or tab >>Towards Responsible AI Foundations for Neurocognitive Analytics of Vision
2025 (English)In: ECVP 2025: 47th European Conference on Visual Perception, Mainz, Germany, 2025Conference paper, Published paper (Refereed)
National Category
Artificial Intelligence Neurosciences Psychology
Research subject
Computer Science
Identifiers
urn:nbn:se:oru:diva-125908 (URN)
Conference
47th European Conference on Visual Perception (ECVP 2025), Mainz, Germany, August 24-28, 2025
Funder
Swedish Foundation for Strategic Research
Available from: 2025-12-23 Created: 2025-12-23 Last updated: 2026-01-13Bibliographically approved
Kondyli, V. & Bhatt, M. (2024). Effects of Temporal Load on Attentional Engagement: Preliminary Outcomes with a Change Detection Task in a VR Setting. In: Rachel McDonnell; Lauren Buck; Julien Pettré; Manfred Lau; Göksu Yamaç (Ed.), ACM Symposium on Applied Perception 2024: Proceedings. Paper presented at ACM Symposium on Applied Perception (SAP'24), Dublin Ireland, August 30-31, 2024. Association for Computing Machinery, Article ID Article 18.
Open this publication in new window or tab >>Effects of Temporal Load on Attentional Engagement: Preliminary Outcomes with a Change Detection Task in a VR Setting
2024 (English)In: ACM Symposium on Applied Perception 2024: Proceedings / [ed] Rachel McDonnell; Lauren Buck; Julien Pettré; Manfred Lau; Göksu Yamaç, Association for Computing Machinery , 2024, article id Article 18Conference paper, Published paper (Refereed)
Abstract [en]

Situation awareness in driving involves detection of events and environmental changes. Failure in detection can be attributed to the density of these events in time, amongst other factors. In this research, we explore the effect of temporal proximity, and event duration in a change detection task during driving in VR. We replicate real-world interaction events in the streetscape and systematically manipulate temporal proximity among them. The results demonstrate that events occurring simultaneously deteriorate detection performance, while performance improves as the temporal gap increases. Moreover, attentional engagement to an event of 5-10 sec leads to compromised perception for the following event. We discuss the importance of naturalistic embodied perception studies for evaluating driving assistance and driver’s education.

Place, publisher, year, edition, pages
Association for Computing Machinery, 2024
Series
SAP ’24
National Category
Computer and Information Sciences Psychology
Research subject
Computer Science; Psychology; Human-Computer Interaction
Identifiers
urn:nbn:se:oru:diva-115594 (URN)10.1145/3675231.3687149 (DOI)001325273000018 ()2-s2.0-85205102252 (Scopus ID)9798400710612 (ISBN)
Conference
ACM Symposium on Applied Perception (SAP'24), Dublin Ireland, August 30-31, 2024
Funder
Swedish Research Council, 2022-02960
Available from: 2024-08-23 Created: 2024-08-23 Last updated: 2024-11-26Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-6290-5492

Search in DiVA

Show all publications