-
Action-oriented vision–language–action (VLA) driving models can produce structured decisions without exposing predicates usable by a state-based audit. We introduce Physical-Margin Chain-of-Contradiction (PhysMargin-CoC), a deterministic post-hoc visual-analyt…
crossref
Jiwoo Jung
2026-08-18T11:38:14Z
置信度 0.70
-
crossref
Jiaming Liu, Mengzhen Liu, Zhenyu Wang, Pengju An 等
2025-11-03T11:20:56Z
置信度 0.70
-
crossref
2010-02-16T14:03:11Z
置信度 0.70
-
crossref
2020-12-22T06:24:09Z
置信度 0.70
-
Amid growing efforts to leverage advances in large language models (LLMs) and vision-language models (VLMs) for robotics, Vision-Language-Action (VLA) models have recently gained significant attention. By unifying vision, language, and action data at scale, wh…
crossref
Kento Kawaharazuka, Jihoon Oh, Jun Yamada, Ingmar Posner 等
2025-08-12T19:39:17Z
置信度 0.70
-
crossref
Chak Fu Chan
2026-03-12T23:00:28Z
置信度 0.70
-
crossref
Feng Shi, Robert Laganière, Emil Petriu
2016-01-16T22:29:56Z
置信度 0.70
-
crossref
2010-02-16T14:03:11Z
置信度 0.70
-
Visual Dialogue Navigation (VDN) aims to enable agents to reach target locations through dialogue with humans. The integration of VDN into Unmanned Aerial Vehicle (UAV) systems enhances human-machine interaction by enabling intuitive, hands-free operation, the…
crossref
Jinyu Chen, Hongyu Li, Zongheng Tang, Xiaoduo Li 等
2026-03-18T01:02:36Z
置信度 0.70
-
crossref
Xiaojian Rao, Lin Fan, Yong Tian, Jingdong Tian
2026-08-20T18:47:21Z
置信度 0.70
-
Existing vision-and-language navigation models often deviate from the correct trajectory when executing instructions. However, these models lack effective error correction capability, hindering their recovery from errors. To address this challenge, we propose …
crossref
Zhuoyuan Yu, Yuxing Long, Zihan Yang, Chengyan Zeng 等
2026-03-18T01:02:49Z
置信度 0.70
-
Hizkuntzen ikas-irakaskuntza hezkuntza ekintzara bideratutako / ekintzan oinarritutako irakaskuntzara daraman egungo aldaketaren aurrekariak aurkezten ditu liburu honek, eta ekintzara bideratutako ikuspegiaren (EBI) teorizazioa eskaintzen du. EBIrako bidea ire…
crossref
Enrica Piccardo, Brian North
2025-05-29T08:46:12Z
置信度 0.70
-
crossref
2019-10-30T17:03:00Z
置信度 0.70
-
Vision-Language-Action (VLA) models integrate visual perception, natural language understanding, and embodied control into a unified framework, enabling end-to-end task execution from multimodal instructions. While such models have demonstrated impressive gene…
crossref
Jianfeng Pang
2025-09-24T09:26:12Z
置信度 0.70
-
crossref
Keke Tang
2026-06-15T19:49:11Z
置信度 0.70
-
Criticality has recently made its way into the field of English Language Teaching. It has mainly fostered the study of teachers’ individual commitments with their social context. A reflective account is offered here based on my praxis when I adopted a critical…
crossref
Enrique Alejandro Basabe
2019-08-25T12:37:40Z
置信度 0.70
-
crossref
Huaiyong Zhao, William H. Warren
2014-10-19T22:48:32Z
置信度 0.70
-
This paper proposes a novel neuro-symbolic framework that augments Vision-Language-Action (VLA) models to assist high-precision operations in industrial assembly. While Large Language Models (LLMs) demonstrate expertise in interpreting natural language instruc…
europepmc
Dario Antonelli, Xiaolang Yang, Leonardo Urbani, Liu Xumenei 等
2026
置信度 0.80
-
crossref
Adeel Yousaf, Mubarak Shah
2026-02-23T20:44:02Z
置信度 0.70
-
crossref
Hans-Hellmut Nagel
2004-12-24T01:26:07Z
置信度 0.70
-
crossref
Leanne Nortje
2025-03-31T15:26:58Z
置信度 0.70
-
Introduction GR-2 is a cutting-edge generative model designed for versatile and generalisable robot manipulation, developed by the Robotics Research Team at ByteDance Research. It represents a significant step forward in robotics, enabling robots to execute a …
crossref
Wendi Fan
2025-12-09T09:28:48Z
置信度 0.70
-
Introduction GR-2 is a cutting-edge generative model designed for versatile and generalisable robot manipulation, developed by the Robotics Research Team at ByteDance Research. It represents a significant step forward in robotics, enabling robots to execute a …
crossref
Wendi Fan
2024-10-26T17:52:19Z
置信度 0.70
-
crossref
Shiryu Ueno, Yoshikazu Hayashi, Shunsuke Nakatsuka, Yusei Yamada 等
2025-03-03T22:55:07Z
置信度 0.70
-
crossref
Glyn W. Humphreys, M. Jane Riddoch
2005-03-11T14:24:27Z
置信度 0.70
-
crossref
Kai Gui, Shun Gui, Yan Luximon
2026-08-27T19:13:53Z
置信度 0.70
-
crossref
Gobinda Chandra Sarker, AKM Azad, Sejuti Rahman, Md Mehedi Hasan
2025-04-26T16:38:00Z
置信度 0.70
-
Vision-language-action (VLA) models represent a promising direction for developing general-purpose robotic systems, demonstrating the ability to combine visual understanding, language comprehension, and action generation. However, systematic evaluation of thes…
preprints
Pranav Guruprasad, Harshvardhan Sikka, Jaewoo Song, Yangyue Wang 等
2024
置信度 0.74
-
crossref
Xiao Li
2026-06-05T19:40:57Z
置信度 0.70
-
crossref
Yan Li
2025-09-12T17:28:20Z
置信度 0.70
-
crossref
Catherine Glossop
2026-08-24T19:14:04Z
置信度 0.70
-
crossref
Yinghao Cai
2026-08-28T19:12:13Z
置信度 0.70
-
crossref
Bingqi Huang, Bingchuan Wei, Isaac Tze Yeung Hiew, Yingkai Cai 等
2026-06-30T04:45:26Z
置信度 0.70
-
Recent advances in vision-language-action (VLA) models have demonstrated impressive generalization for robotic manipulation. However, these models often operate by directly mapping visual and linguistic inputs to subsequent actions, lacking intermediate task p…
crossref
Xiang Li, Ya-Li Li, Yuan Wang, Huaqiang Wang 等
2026-03-17T23:28:16Z
置信度 0.70
-
Abstract Mapping surgery is fundamental to developing operative guidelines and enabling autonomous robotic surgery. Recent advances in artificial intelligence (AI) have shown promise in mapping the behaviour of surgeons from videos, yet current models remain n…
crossref
Dani Kiyasseh
2026-03-27T08:16:33Z
置信度 0.70
-
Computer vision is a very popular subfield of machine learning and data science, image classification is one of long-standing problems. Traditional single-modal algorithms using deep convolutional neural works have low generalization ability when dealing with …
crossref
Kunhong Yu, Piotr Wojcik
2024-04-15T10:20:54Z
置信度 0.70
-
In this chapter, we examine the push toward unified vision–language modeling, driven by the need for general-purpose systems that can perceive, reason, and generate across modalities. We review the evolution from alignment-centric pipelines to integrated archi…
crossref
Peng Wang
2026-05-11T09:26:42Z
置信度 0.70
-
crossref
ELENI KOUTSOMITOPOULOU
2007-06-08T05:49:12Z
置信度 0.70
-
Visually impaired individuals face great challenges with independently navigating dynamic environments because of their inability to fully comprehend the environment and actions of surrounding people. Conventional navigation approaches like Simultaneous Locali…
crossref
Jasmine Liu
2024-04-29T06:54:56Z
置信度 0.70
-
crossref
Mingfang Zhang, Ryo Yonetani, Yifei Huang, Liangyang Ouyang 等
2026-04-29T19:45:49Z
置信度 0.70
-
crossref
Akash S A, Hari Gokulan S, Harish S, Kailas Nath 等
2026-06-23T19:42:59Z
置信度 0.70
-
crossref
MARCO MIROLLI, DOMENICO PARISI
2007-06-08T01:49:12Z
置信度 0.70
-
crossref
Junjie Wen
2025-02-25T13:54:23Z
置信度 0.70
-
crossref
Gregor Miller, Steve Oldridge, Sidney Fels
2013-08-02T21:14:26Z
置信度 0.70
-
crossref
Mohammad Hovaidi Ardestani, Martin Giese
2017-08-31T16:32:51Z
置信度 0.70
-
crossref
Kay Harley
2012-09-11T03:29:24Z
置信度 0.70
-
crossref
Cuixin Yang, Rongkang Dong, Kin-Man Lam
2026-06-22T23:35:44Z
置信度 0.70
-
crossref
Yao Yeboah, Joseph Teye Ignatius Buertey, Kwabena Agyapong-Kodua, Zeyad Farisi 等
2026-05-27T19:40:54Z
置信度 0.70
-
crossref
Radha Seelaboyina, Naveen Rampa, Alle Shivanandha, Chamakura Reddy
2026-01-01T01:43:11Z
置信度 0.70
-
crossref
Roy Uziel, Oded Bialer
2025-04-08T17:08:13Z
置信度 0.70
-
crossref
Hidehisa Arai, Keita Miwa, Kento Sasaki, Kohei Watanabe 等
2025-04-08T13:08:13Z
置信度 0.70
-
High-throughput plant phenotyping, the quantitative measurement of observable plant traits, is critical for modern breeding but remains constrained by a "phenotyping bottleneck", where manual data collection is labor-intensive and prone to observer bias. Conve…
crossref
Abderrahmene Boudiaf, Sajid Javed
2026-02-21T13:42:17Z
置信度 0.70
-
Vision-language-action models have emerged as a crucial paradigm in robotic manipulation. However, existing VLA models exhibit notable limitations in handling ambiguous language instructions and unknown environmental states. Furthermore, their perception is la…
crossref
Helong Huang, Min Cen, Kai Tan, Xingyue Quan 等
2026-03-18T01:04:47Z
置信度 0.70
-
Let’s explore GR-2 together.
crossref
Wendi Fan
2025-07-08T05:21:23Z
置信度 0.70
-
crossref
Wenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang 等
2026-08-06T14:44:29Z
置信度 0.70
-
crossref
D.P. Benjamin
2002-11-27T13:07:58Z
置信度 0.70
-
crossref
Daniele Caligiore, Martin H. Fischer
2012-02-06T11:26:58Z
置信度 0.70
-
Reaching, grasping, and placing fresh produce is among the most labor-intensive operations in agricultural production, yet most robotic solutions remain task-specific, single-crop systems that generalize poorly across the morphological diversity of fresh produ…
crossref
Kapalik Khanal, Biswash Khatiwada, Stephen Afrifa, Sawyer Exum 等
2026-08-04T12:14:06Z
置信度 0.70
-
Abstract We present a novel, modular Graph based Vision-Language-Action (VLAG) framework designed for long-horizon robotic manipulation tasks. Our approach integrates a graph-based planner with dedicated vision, language, and action modules, enabling robust an…
crossref
Ardalan Aryashad, Yan Jin
2025-10-27T21:29:46Z
置信度 0.70
-
Recent advances in large-scale vision-language models (VLMs) have opened new pathways for robotic learning and control. Google DeepMind’s Robotics Transformer 2 (RT-2) represents a transformative step by integrating pre-trained internet-scale multimodal models…
crossref
Austin Zhou
2025-12-15T07:16:29Z
置信度 0.70
-
crossref
James W. Davis, Hui Gao
2003-09-12T05:06:58Z
置信度 0.70
-
crossref
Farida Asriani, azhari azhari, Wahyono Wahyono
2025-03-06T18:43:12Z
置信度 0.70
-
crossref
Farida Asriani, azhari azhari, Wahyono Wahyono
2025-04-16T08:37:45Z
置信度 0.70
-
crossref
Angel Martinez-Sanchez, Parthib Roy, Ross Greer
2026-07-30T19:05:05Z
置信度 0.70
-
crossref
T. Bultan
2002-11-07T23:28:47Z
置信度 0.70
-
crossref
Kaisen Hu
2025-12-18T18:31:46Z
置信度 0.70
-
crossref
Yingchao Zhang, Cheng Liu
2026-05-01T08:51:15Z
置信度 0.70
-
crossref
P. Matikainen, R. Sukthankar, M. Hebert
2012-08-01T11:05:51Z
置信度 0.70
-
crossref
Utkarsh Shandilya, Marsha Mariya Kappan, Sanyam Jain, Vijeta Sharma
2026-06-24T19:25:27Z
置信度 0.70
-
This study assesses the accuracy and consistency of a commercially available large language model (LLM) in extracting and interpreting sensitivity and reliability data from entire visual field (VF) test reports for the evaluation of glaucomatous defects. Singl…
crossref
Jeremy C. K. Tan
2025-04-11T06:45:31Z
置信度 0.70
-
crossref
Tung Le, Khoa Pho, Thong Bui, Huy Tien Nguyen 等
2022-02-16T19:49:00Z
置信度 0.70
-
crossref
2016-07-05T05:03:01Z
置信度 0.70
-
crossref
Timothy Lane
2014-04-21T11:51:34Z
置信度 0.70
-
Food preferences differ among individuals, and these variations reflect underlying personalities or mental tendencies. However, capturing and predicting these individual differences remains challenging. Here, we propose a novel method to predict individual foo…
crossref
Hiroki Kojima, Asako Toyama, Shinsuke Suzuki, Yuichi Yamashita
2025-03-28T20:55:49Z
置信度 0.70
-
crossref
Kuo Shi
2025-01-31T05:25:51Z
置信度 0.70
-
crossref
2025-10-24T14:57:49Z
置信度 0.70
-
crossref
Zhou Zhu, kai xu
2025-08-25T20:41:26Z
置信度 0.70
-
crossref
Siyu He, Shengsheng Wang, Sifan Long
2024-09-18T12:48:11Z
置信度 0.70
-
Abstract This study proposes a deep learning vision-language model for the automated diagnosis of pediatric dental diseases, with a focus on differentiating between caries and periapical infections. The model integrates visual features extracted from panoramic…
crossref
Tuan D. Pham
2025-05-22T13:45:25Z
置信度 0.70
-
crossref
Arthur Knauff
2025-12-02T15:31:36Z
置信度 0.70
-
We present HordeVision, an open-source Kazakh vision-language model designed for optical character recognition (OCR), image captioning, visual question answering (VQA), reasoning, and instruction following. Supporting local open-source initiatives, HordeVision…
crossref
Pavel Zubitskii, Vitaliy Morozov, Sanzhar Murzakhmetov, Beksultan Sagyndyk 等
2026-01-10T01:37:31Z
置信度 0.70
-
Abstract Malware family classification is crucial for threat detection, yet existing methods struggle with generalization, multi-task adaptability, and interpretability. We propose VIMAR, a unified vision–language model that supports classification, similarity…
crossref
Shiting Xu
2026-03-27T03:01:30Z
置信度 0.70
-
crossref
Weihao CAI, Yoshiki MORI, Nobutaka SHIMADA
2024-10-10T17:21:42Z
置信度 0.70
-
crossref
Tim Bary, Clément Fuchs, Benoît Macq
2026-02-17T21:05:43Z
置信度 0.70
-
crossref
Saher Elsayed
2026-08-06T19:10:04Z
置信度 0.70
-
crossref
Linfeng Wang, Deok Jin Lee
2025-07-24T05:36:34Z
置信度 0.70
-
crossref
Jenhao Hsiao, Yikang Li, Chiuman Ho
2021-11-24T20:40:09Z
置信度 0.70
-
crossref
Risa Shinoda, Nakamasa Inoue, Hirokatsu Kataoka, Masaki Onishi 等
2026-04-29T19:45:49Z
置信度 0.70
-
crossref
Tevfik Bultan
2004-02-03T15:47:51Z
置信度 0.70
-
crossref
C. Yu, D. H Ballard
2010-06-04T12:15:33Z
置信度 0.70
-
europepmc
2026
置信度 0.80
-
A safety margin derived from a vision–language verdict is only as valid as that verdict is recent, yet run-time safety filters model state noise, not the staleness of an intermittent label. We present AEGIS, a safe-action projection whose hold radius is an exp…
europepmc
Danial Zafaranchizadeh Moghaddam, Maryam Banitalebi Dehkordi, Hamed Rahimi Nohooji, Abolfazl Zaraki
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80