-
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment generalization, and execution smoothness. We present NebulaVLA, an asynchronous dual-frequency architecture that decouples…
datacite
Zhao, Cong, Tian, Shuai, Zhang, Xu, Ni, Baocheng 等
2026
置信度 0.66
Robotics (cs.RO)Artificial Intelligence (cs.AI)FOS: Computer and information sciencesI.2.9; I.2.10
-
Learning embodied urban navigation policies from real-world data is constrained by the cost of task-specific data collection and the limited coverage of rare yet safety-critical scenarios. To address these challenges, we present a scalable framework for learni…
datacite
Xia, Bingyi, Bao, Han, Chen, Zhewei, Ye, Hanjing 等
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
Artificial intelligence-assisted ultrasound scanning enhances diagnostic reliability and efficiency by providing real-time guidance for standardized image acquisition and reducing operator dependence. However, existing reinforcement learning and learning-assis…
datacite
Zhang, Cheng, Wu, Xingzheng, Yan, Guihao, Hu, Xifeng 等
2026
置信度 0.66
Robotics (cs.RO)Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit fr…
datacite
GigaBrain Team, Ye, Angen, Sun, Axiang, Jin, Can 等
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajectory states. Reinforcement learning can refine pretrained VLA policies, yet sparse success signals hinder exploration, wh…
datacite
Xu, Yijie, Jin, Haopeng, Zhou, Run, Liu, Shengbang 等
2026
置信度 0.66
Robotics (cs.RO)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cost and limited predicted action length hinder real-time deployment. Although Dadu-Corki, a dedicated accelerator for effic…
datacite
Qi, Chunyu, Song, Zhuoran, Weng, Jian, Jiang, Haozhe 等
2026
置信度 0.66
Robotics (cs.RO)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inference cost poses significant challenges for deployment in robotic systems. In practice, on-device inference is constrained by limited compute capacit…
datacite
Zhou, Ao, Dai, Bo, Yu, Le, Liu, Xingyu 等
2026
置信度 0.66
Artificial Intelligence (cs.AI)Robotics (cs.RO)FOS: Computer and information sciences
-
Quantized Vision-Language-Action (VLA) models expose a weight-fault surface: Rowhammer-style faults can corrupt deployed INT8 bits. We present the first bit-flip attack on a VLA: a few gradient-selected flips reduce closed-loop success to $0\%$, while hundreds…
datacite
Gao, Yudong, Chen, Linghan, Wu, Wenhan, Zhou, Mia 等
2026
置信度 0.66
Cryptography and Security (cs.CR)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Parameter-efficient fine-tuning (PEFT) is a natural way to adapt pretrained vision-language-action (VLA) policies, but most adapter designs apply temporally static updates throughout a control rollout, overlooking the phase-dependent nature of continuous-actio…
datacite
Guo, Yufei, Wu, Yinan, Duan, Haoran, Ding, Guiguang 等
2026
置信度 0.66
Robotics (cs.RO)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Embodied intelligent ultrasound scanning enables the automation and standardization of the ultrasound examination process by integrating perception, decision-making, and execution capabilities. However, existing methods suffer from loosely coupled modeling bet…
datacite
Wu, Xingzheng, Zhang, Cheng, Yan, Guihao, Hu, Xifeng 等
2026
置信度 0.66
Robotics (cs.RO)Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
Abstract The eradication of untouchability represents one of the most critical, contentious, and complex social reform agendas in modern Indian history. This article examines the multifaceted role of Mahatma Gandhi in conceptualizing, mobilizing, and leading t…
datacite
Dr. Venkatesha V
2019
置信度 0.66
Mahatma Gandhi, Untouchability, Caste System, Social Reform, B.R. Ambedkar, Sarvodaya, Hindu Tradition, Poona Pact
-
Abstract The eradication of untouchability represents one of the most critical, contentious, and complex social reform agendas in modern Indian history. This article examines the multifaceted role of Mahatma Gandhi in conceptualizing, mobilizing, and leading t…
datacite
Dr. Venkatesha V
2019
置信度 0.66
Mahatma Gandhi, Untouchability, Caste System, Social Reform, B.R. Ambedkar, Sarvodaya, Hindu Tradition, Poona Pact
-
Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure. In a fu…
datacite
HU, Haibo, Huang, Lianming, Li, Qiao, Guan, Nan 等
2026
置信度 0.66
Distributed, Parallel, and Cluster Computing (cs.DC)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
The Nexus Recursive Harmonic Framework (RHA) – A Triadic Harmonic Magnum Opus Driven by Dean A. Kulik November, 2025 1. Triangular Quantization and the Mark1 Harmonic Engine Triangular Archetypes 0–9: At the core of RHA is a triangular quantization model in wh…
datacite
Kulik, Dean
2025
置信度 0.66
-
The Nexus Recursive Harmonic Framework (RHA) – A Triadic Harmonic Magnum Opus Driven by Dean A. Kulik November, 2025 1. Triangular Quantization and the Mark1 Harmonic Engine Triangular Archetypes 0–9: At the core of RHA is a triangular quantization model in wh…
datacite
Kulik, Dean
2025
置信度 0.66
-
crossref
Hossein Rahmani, Mohammed Bennamoun
2017-12-25T16:51:45Z
置信度 0.70
-
Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention to target regions. I…
crossref
Wenxuan Song, Ziyang Zhou, Han Zhao, Jiayi Chen 等
2026-03-18T00:58:55Z
置信度 0.70
-
crossref
Berit Brogaard
2011-03-08T09:33:52Z
置信度 0.70
-
crossref
JinHai Li, Peng Chen, MoHan Li, LuYi Ren
2025-09-12T17:28:12Z
置信度 0.70
-
Continuous action recognition in video is more complicated compared with traditional isolated action recognition. Besides the high variability of postures and appearances of each action, the complex temporal dynamics of continuous action makes this problem cha…
crossref
Jun Lei, Guohui Li, Jun Zhang, Qiang Guo 等
2016-02-19T07:58:53Z
置信度 0.70
-
crossref
Sheng Guo, Ruiteng Zhao, Jinxian Zhou, Chenchen Liu 等
2026-01-21T21:06:47Z
置信度 0.70
-
Abstract This study evaluates two leading approaches for teaching construction robots new skills to understand their applicability for construction automation: a vision-language-action (VLA) model and reinforcement learning (RL) methods. The goal is to underst…
crossref
Zhaofeng Hu, Hongrui Yu, Vaidhyanathan Chandramouli, Ci-Jyun Liang
2026-08-15T13:16:25Z
置信度 0.70
-
Leaf disease identification often involves image classification. However, currently popular methods, such as pretraining the model on the ImageNet dataset, precisely a type of crop and limited data, or only using image or text classification without taking adv…
crossref
Khang Nguyen Quoc, Lan Le Thi Thu, Luyl-Da Quach
2025-02-26T23:15:16Z
置信度 0.70
-
crossref
Xuan Dong, Zhe Han, Tianhao Niu, Qingfu Zhu 等
2026-07-01T12:25:50Z
置信度 0.70
-
This study investigates the domain adaptation of the Vision-Language Model (VLM) for road damage assessment, focusing on a fine-tuning strategy optimized for resource-constrained engineering environments. Unlike conventional object detection models that operat…
crossref
Donghwi Kim, Heejung Youn
2026-03-13T20:54:38Z
置信度 0.70
-
Abstract Detrimental to individuals and society, online hateful messages have recently become a major social issue. Among them, one new type of hateful message, named “hateful meme”, has emerged and brought difficulties in traditional deep learning-based detec…
crossref
Yuyang Chen, Feng Pan
2022-04-11T15:01:46Z
置信度 0.70
-
crossref
Yeryeong Cho, Jaeyoung Choe, Joongheon Kim
2026-02-19T20:55:29Z
置信度 0.70
-
crossref
Yu Liu
2025-12-25T18:29:17Z
置信度 0.70
-
crossref
Mirsalar Kamari, Mahdi Ghafoori, Mohammadsoroush Tafazzoli
2026-08-27T10:05:07Z
置信度 0.70
-
crossref
Yuya Mukose, Tatsuya Sasaki, Tatsuhito Hasegawa
2026-05-18T08:53:16Z
置信度 0.70
-
crossref
Michal Nazarczuk, Krystian Mikolajczyk
2021-02-24T15:13:11Z
置信度 0.70
-
crossref
Ravi Ranjan, Agoritsa Polyzou
2026-07-01T12:25:50Z
置信度 0.70
-
World models are a transformative paradigm in embodied AI, enabling agents to learn efficiently and plan by simulating environmental dynamics. As vision-language action and navigation systems grow more sophisticated, integrating world models is crucial for bri…
crossref
Jingwen Sun, Hongjin Chen, Zezhi Liu, Xin Jin 等
2025-12-09T22:38:02Z
置信度 0.70
-
crossref
Gianna Hessel
2015-06-20T04:45:22Z
置信度 0.70
-
Generalist robots need to perform diverse tasks while operating in dynamic, uncertain, and unstructured environments, often around human beings. Vision-language-action (VLA) models have recently emerged as a promising and flexible framework for integrating per…
crossref
Umair Cheema, Youakim Badr, Thao Minh Le, Katie Fitzsimons
2026-08-17T13:42:17Z
置信度 0.70
-
crossref
2026-04-28T13:34:43Z
置信度 0.70
-
crossref
Gulchin Abdullayeva, Nigar Alishzade
2024-12-26T13:48:23Z
置信度 0.70
-
crossref
Chenfei Liu, Dong Lin, Pengyi Tian, Na Yang
2026-02-13T19:04:59Z
置信度 0.70
-
crossref
Robert Elliott, Janne Underriner
2022-10-12T11:34:09Z
置信度 0.70
-
crossref
2021-03-06T03:04:25Z
置信度 0.70
-
crossref
Bahram Mohammadi, Ehsan Abbasnejad, Yuankai Qi, Qi Wu 等
2025-09-29T16:12:07Z
置信度 0.70
-
The previous research that I conducted, the use of Team-Based Learning in Classroom Action Research (CAR) class is well accepted and gives a good influence on the students’ learning. To have a responsibility to teach Action Research and to keep being reflective …
crossref
Rusiana Rusiana
2017-08-30T05:27:20Z
置信度 0.70
-
crossref
Rita T. Denny
2008-12-19T09:05:44Z
置信度 0.70
-
crossref
Liang Chen, Ghazi Shazan Ahmad, Tianjun Yao, Lingqiao Liu 等
2026-04-29T19:45:49Z
置信度 0.70
-
crossref
Shiming Chen, Bowen Duan, Salman Khan, Fahad Shahbaz Khan
2026-04-29T19:45:49Z
置信度 0.70
-
crossref
Alaa Asfour, Christopher Indris, Leihan Chen, Tejas Vyas 等
2026-05-14T04:53:35Z
置信度 0.70
-
crossref
Xu Liu, Zhouhui Lian
2025-12-23T18:00:51Z
置信度 0.70
-
Abstract Milner and Goodale's Two Visual Systems Hypothesis (TVSH) is regarded as common ground in recent discussions of visual consciousness. A central part of TVSH is a functional model of vision and action (a functional perception‐action model, PAM for shor…
crossref
Thor Grünbaum
2017-07-14T10:52:36Z
置信度 0.70
-
Background: Lead optimization in small-molecule drug discovery is a constrained search problem: from a chemical starting point and the structural cues of its binding pose, propose analogs that preserve the pharmacophore, satisfy ADMET liabilities, and remain s…
crossref
Toluwanimi Odunewu
2026-06-15T11:03:08Z
置信度 0.70
-
crossref
Wenlong Chen, Zhen Tian, Zhou Zhou, Youhua Xia
2026-02-20T21:14:29Z
置信度 0.70
-
crossref
Jaehyuk Heo, Pilsung Kang
2025-01-10T16:41:59Z
置信度 0.70
-
Multimodal Large Language Models (MLLMs) have attracted significant attention in recent years due to their broad applicability across fields, including healthcare, autonomous vehicles (AVs), robotics, 3D reconstruction, and human-computer interaction. The usef…
crossref
Gulraiz Khan, Waqas Ahmed, K.Y. Wertheim, Kevin Pimbblet 等
2026-06-25T01:47:17Z
置信度 0.70
-
Large Vision-Language Models (LVLMs) have shown promising capabilities in understanding and generating information by integrating both visual and textual data. However, current models are still prone to hallucinations, which degrade the performance and greatly…
crossref
Robert Wijaya, Ngoc-Bao Nguyen, Ngai-Man Cheung
2026-07-01T18:30:17Z
置信度 0.70
-
crossref
Noor Oush, Reza Alirezaee, Saeed Mozaffari, Shahpour Alirezaee
2026-07-07T19:43:15Z
置信度 0.70
-
crossref
Pablo Valle, Sergio Segura, Shaukat Ali, Aitor Arrieta
2026-07-16T21:48:20Z
置信度 0.70
-
crossref
Ingook Jang, Samyeul Noh, Seonghyun Kim
2026-02-19T20:55:29Z
置信度 0.70
-
crossref
Jianxin Bi, Kevin Yuchen Ma, Ce Hao, Mike Shou Zheng 等
2026-05-11T19:48:18Z
置信度 0.70
-
crossref
Rhimesh Lwagun, John Yare, Justin Lukose, Olasubomi Rufai 等
2026-04-20T20:01:37Z
置信度 0.70
-
crossref
Alison Wilcox, Adam Bushnell
2026-06-19T09:22:57Z
置信度 0.70
-
crossref
Alison Wilcox, Adam Bushnell
2026-06-19T09:22:57Z
置信度 0.70
-
crossref
Fengze Yang, Bo Yu, Yang Zhou, Xuewen Luo 等
2026-05-26T12:19:49Z
置信度 0.70
-
crossref
Alison Wilcox, Adam Bushnell
2020-11-13T16:04:15Z
置信度 0.70
-
crossref
Alison Wilcox, Adam Bushnell
2020-11-13T21:04:15Z
置信度 0.70
-
crossref
Alison Wilcox, Adam Bushnell
2026-06-19T09:22:57Z
置信度 0.70
-
crossref
F. Fleischer, A. Casile, M. Giese
2010-06-04T12:30:50Z
置信度 0.70
-
crossref
Alison Wilcox, Adam Bushnell
2020-11-13T16:04:15Z
置信度 0.70
-
crossref
Alison Wilcox, Adam Bushnell
2020-11-13T21:04:15Z
置信度 0.70
-
crossref
Alison Wilcox, Adam Bushnell
2026-06-19T09:22:57Z
置信度 0.70
-
crossref
Redha Touati, Max Mignotte
2014-05-30T14:26:22Z
置信度 0.70
-
crossref
2013-10-07T13:04:52Z
置信度 0.70
-
crossref
Muaz Al Radi, Sajid Javed
2026-02-02T16:34:02Z
置信度 0.70
-
crossref
Sparsh Garg, Abhishek Aich
2026-02-23T20:44:02Z
置信度 0.70
-
crossref
Hongyu Bai, Haolong Wang, Genke Yang
2026-08-17T19:14:40Z
置信度 0.70
-
crossref
Peiran Xu, Xicheng Gong, Yadong Mu
2026-04-29T19:45:49Z
置信度 0.70
-
crossref
Feiyang Wang, Nan Luo, Wangyu Wu
2025-09-15T17:35:52Z
置信度 0.70
-
crossref
Philip Wootaek Shin, Ajay Narayanan Sridhar, Lakshmi Sivani Devarapalli, Rui Zhang 等
2026-07-01T12:25:50Z
置信度 0.70
-
The integration of visual perception, natural language understanding, and continuous motor control has led to the emergence of Vision-Language-Action models. These models demonstrate unprecedented generalization capabilities in embodied artificial intelligence…
crossref
Michael Reynolds, Noah Sato
2026-07-09T04:42:21Z
置信度 0.70
-
crossref
Eunseo Jeong, Gyunyeop Kim, Sangwoo Kang
2023-11-24T03:44:03Z
置信度 0.70
-
Ethically aligned design aims to ensure that intelligent systems remain human-centric, serving humanity’s values and ethical principles. Moderating content for achieving ethically aligned design is a critical step toward the realization of trustworthy AI. Howe…
crossref
Ning Zhang, Xuan Feng, Tianlong Gu, Liang Chang
2023-07-20T15:42:46Z
置信度 0.70
-
Abstract: This paper investigates the adaptation of pre-trained robotics foundation models, specifically RT-2 and PaLM-E, for construction automation applications. This research explores transfer learning methodologies across diverse construction tasks, develo…
crossref
Sai Kothapalli
2025-12-25T12:52:21Z
置信度 0.70
-
crossref
Ravi Varma Kumar Bevara Bevara
2025-02-15T09:36:16Z
置信度 0.70
-
The Vision-Language-Action (VLA) model faces basic challenges in actual physical deployment, including reasoning delay, opaque decision-making, and insufficient generalization ability. Existing research mainly focuses on data size and training strategies, and …
crossref
Zan Zheng, Jin Huang, Chao Wang, Xi Tan 等
2026-07-18T20:45:10Z
置信度 0.70
-
crossref
Shailaja Keyur Sampat, Pratyay Banerjee, Yezhou Yang, Chitta Baral
2023-08-04T16:21:02Z
置信度 0.70
-
<p>Ethically aligned design aims to ensure that intelligent systems remain human-centric, serving humanity’s values and ethical principles. Moderating content for achieving ethically aligned design is a critical step toward the realization of trustworthy…
crossref
Ning Zhang, Xuan Feng, Tianlong Gu, Liang Chang
2023-07-20T19:42:47Z
置信度 0.70
-
crossref
Lingyan Ran, Lidong Wang, Guangcong Wang, Peng Wang 等
2026-01-12T14:49:29Z
置信度 0.70
-
The relation between spatial vision and spatial language has always been a source of controversy. Three problems can be identified as in need of a solution. A first problem pertains to the nature of the minimal information units that make up spatial vision and…
crossref
Francesco-Alessio Ursini
2022-04-08T07:17:38Z
置信度 0.70
-
crossref
Yan Zeng
2024-01-08T14:42:04Z
置信度 0.70
-
crossref
Quek
2005-03-21T20:07:39Z
置信度 0.70
-
Abstract Medical Vision-Language Models (MVLMs) are emerging as powerful tools for tasks such as Visual Question Answering (VQA); however, they often struggle with hallucination and limited reasoning transparency, particularly in complex diagnostic scenarios. …
crossref
Kavya Dasaramoole Prakash, Kiseong Kim, Youngmahn Han
2025-08-01T23:45:19Z
置信度 0.70
-
crossref
Alex Postlmayr, Pamela Cosman, Sujit Dey
2025-07-02T14:37:38Z
置信度 0.70
-
We propose general visual inspection model using Vision-Language Model (VLM) with few-shot images of non-defective or defective products, along with explanatory texts that serve as inspection criteria. Although existing VLM exhibit high performance across vari…
crossref
Shiryu Ueno, Yoshikazu Hayashi, Shunsuke Nakatsuka, Yusei Yamada 等
2025-02-19T08:01:50Z
置信度 0.70
-
Vision-language foundation models (VLFMs), such as CLIP, have demonstrated remarkable generalizability across diverse downstream tasks, including both cross-modal and vision-centric tasks. Leveraging large-scale textual supervision, VLFMs capture a broad spect…
crossref
Zhenshi Li, Xueliang Zhang, Pengfeng Xiao, XiaoXiang Zhu
2026-03-14T01:11:56Z
置信度 0.70
-
crossref
Eleni Skourti
2025-02-18T10:03:19Z
置信度 0.70
-
crossref
G. Linda Rikard, Betty Beacham
2012-01-05T21:26:52Z
置信度 0.70
-
crossref
Shahid Alam
2013-07-03T18:22:14Z
置信度 0.70
-
crossref
Samuel A., Dorina Petriu, Pantanowitz Motshegw
2012-05-04T10:50:57Z
置信度 0.70
-
crossref
Tiejin Chen, Pingzhi Li, Kaixiong Zhou, Tianlong Chen 等
2025-08-04T09:54:03Z
置信度 0.70
-
crossref
Shijie Wang, Dahun Kim, Ali Taalimi, Chen Sun 等
2025-04-08T17:08:13Z
置信度 0.70
-
crossref
Guanxiong Wang
2026-07-26T16:14:39Z
置信度 0.70
-
crossref
Difei Gu, Yunhe Gao, Mu Zhou, Dimitris Metaxas
2026-05-05T19:59:32Z
置信度 0.70