-
crossref
2021-04-09T05:21:37Z
置信度 0.70
-
crossref
Jianping Jiang, Weiye Xiao, Zhengyu Lin, Huaizhong Zhang 等
2025-08-13T17:26:42Z
置信度 0.70
-
crossref
2013-09-20T05:49:28Z
置信度 0.70
-
crossref
Graham Hall
2017-11-15T06:49:26Z
置信度 0.70
-
crossref
Rakib Hossain Sajib, Md Baharul Islam, Md Kishor Morol, Marjia Sultana
2026-08-04T19:11:41Z
置信度 0.70
-
crossref
Zhaochong An, Guolei Sun, Yun Liu, Runjia Li 等
2025-08-13T17:26:42Z
置信度 0.70
-
crossref
2021-02-22T08:54:37Z
置信度 0.70
-
crossref
Yunfei Liu, Quan Long
2025-04-07T22:21:28Z
置信度 0.70
-
crossref
Kai Huang, Hao Zou, Bochen Wang, Ye Xi 等
2026-04-29T19:45:49Z
置信度 0.70
-
crossref
Graham Hall
2014-03-14T04:48:20Z
置信度 0.70
-
In today's visually driven digital landscape, computer vision has become an invaluable asset for brand managers seeking deeper consumer insights and more effective brand strategies. By integrating low-level feature extraction (e.g., color and texture analysis)…
crossref
Yaqiu Li, Hsin-Hsuan Lee, Lorena Blasco-Arcas
2025-09-11T00:36:38Z
置信度 0.70
-
crossref
2022-06-08T15:21:02Z
置信度 0.70
-
crossref
John M. Foley
2021-05-21T16:23:21Z
置信度 0.70
-
crossref
M. Alpern
2003-04-25T11:07:43Z
置信度 0.70
-
crossref
Chung-Lin Huang, Wen-Yi Huang
2002-08-25T03:30:30Z
置信度 0.70
-
Abstract Unmanned aerial vehicle (UAV) photogrammetry is a mainstream approach for infrastructure inspection, yet converting its high-overlap imagery into decision-ready findings remains challenging. Deep-learning detectors are confined to closed defect taxono…
crossref
Zhuo Yang, Changsheng Qu, Gangyan Xu
2026-07-31T05:20:01Z
置信度 0.70
-
crossref
Reem Al-Junaid, Muzammil Behzad, SAJJAD MAHMOOD, MAHMOOD NIAZI 等
2025-10-30T02:43:10Z
置信度 0.70
-
crossref
2020-09-04T11:19:29Z
置信度 0.70
-
Abstract Vision-language models are increasingly used to search, rank, and describe cultural collections. Their judgments of artistic value may reflect both learned image-text associations and the structure of the archives on which those judgments are applied.…
crossref
Manpreet Singh, Nandakishor Reddy Pulagam, Rhythm Bhatia, Rahul Joshi
2026-08-14T06:32:34Z
置信度 0.70
-
crossref
Hang Gu
2026-04-20T20:01:42Z
置信度 0.70
-
crossref
Hosam Elgendy
2026-04-23T19:58:54Z
置信度 0.70
-
Abstract Engineering simulation interpretation is a major bottleneck in design cycles, requiring expensive domain expertise to validate complex outputs and ensure safety and performance. While modern large language models (LLMs) may assist in interpretation, t…
crossref
Jessica Ezemba, Jason Pohl, Conrad Tucker, Christopher McComb
2025-12-25T09:36:31Z
置信度 0.70
-
If the idea of a biopsychosocial model is not going to be merely a phrase, then every clinical diagnosis and therapeutic vision need to consider the familial issues. Taking the family issues into consideration has evident gains; it allows for a better understa…
pubmed
de Barbaro B
2004 Sep-Oct
置信度 0.82
-
"Use your words!" is a phrase admonishing preschoolers to divert their action-proneness to thought and language. Freud's injunction against acting out had a similar aim, placing control over drives in the domain of "inner language." The twenty-first-century ps…
pubmed
Shapiro T
2004 Spring
置信度 0.82
-
Our aim is to enable a machine to observe and interpret the behaviour of others. Mathematical models are employed to describe certain biological motions. The main challenge is to design models that are both tractable and meaningful. In the first part we will d…
pubmed
Rittscher J, Blake A, Hoogs A, Stein G
2003 Mar 29
置信度 0.82
-
This study investigated the recall of movement patterns presented either by demonstration or guided movement with vision eliminated. Participants were instructed to rehearse and remember each of the 12 patterns using one of four strategies: imagery, verbal lab…
pubmed
Hall C, Moore J, Annett J, Rodgers W
1997 Jun
置信度 0.82
-
The Recursive Harmonic Architecture: A Synthesis of Self-Organized Criticality, Computational Metaphysics, and Emergent Reality Driven by Dean A. Kulik July, 2025 Part I: Foundational Principles of the Recursive Architecture This document provides a foundation…
datacite
Kulik, Dean
2025
置信度 0.66
-
The Recursive Harmonic Architecture: A Synthesis of Self-Organized Criticality, Computational Metaphysics, and Emergent Reality Driven by Dean A. Kulik July, 2025 Part I: Foundational Principles of the Recursive Architecture This document provides a foundation…
datacite
Kulik, Dean
2025
置信度 0.66
-
These notes develop a mathematical bridge between physical AI, causal inference, and the (\Psi)-operator framework. The central thesis is that embodied intelligence should not be modeled merely as perception followed by action, but as a closed-loop causal oper…
datacite
Chuang, William Huanshan
2026
置信度 0.66
Physical AIembodied intelligenceΨ-operator frameworkcausal inferenceJudea Pearl
-
These notes develop a mathematical bridge between physical AI, causal inference, and the (\Psi)-operator framework. The central thesis is that embodied intelligence should not be modeled merely as perception followed by action, but as a closed-loop causal oper…
datacite
Chuang, William Huanshan
2026
置信度 0.66
Physical AIembodied intelligenceΨ-operator frameworkcausal inferenceJudea Pearl
-
Recent vision-based action models have demonstrated strong capabilities in complex manipulation, but they rarely leverage explicit object physical properties to adapt their policies. We introduce ViTacPhys, a visual-tactile framework and data acquisition syste…
datacite
Liu, Yiwen, Zhu, Yujun, Jia, Kui, Liao, Zhao 等
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
Vision-Language-Action (VLA) policies are vulnerable to localized physical perturbations, yet existing certified patch defenses target discrete labels and cannot directly certify continuous, temporally correlated actions. We introduce CertVLA, a certified defe…
datacite
Lu, Hui, Peng, Zhijie, Lin, Yuqi, Yang, Zaijia 等
2026
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work that pushes the drafter itself toward para…
datacite
Li, Yantao, Gao, Huanlin, Zhao, Fang, Tan, Chao 等
2026
置信度 0.66
Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Vision-language-action (VLA) models can follow natural-language (NL) task instructions, but such instructions may not precisely specify safety-critical or spatiotemporal requirements on the resulting behavior. We introduce Logic-VLA, a formal-requirement-aware…
datacite
Wang, Celina Shiyu, Zhao, Yiqi, Ye, Junjie, Wang, Yue 等
2026
置信度 0.66
Robotics (cs.RO)Logic in Computer Science (cs.LO)Systems and Control (eess.SY)FOS: Computer and information sciencesFOS: Electrical engineering, electronic engineering, information engineering
-
Routine gastrointestinal endoscopy is intrinsically bidirectional: the instrument is advanced to reach target anatomy and later withdrawn or retroflexed for inspection, while an external cue may require earlier reversal. When the requested phase changes before…
datacite
Ng, Chi Kit, Zhang, Yidong, Hing, Lui Siu, Lin, Jinsong 等
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
Full Changelog: https://github.com/attogram/attogram-pre-production-studio/compare/7...8 attogram-pre-production-studio You can cite all versions by using the DOI 10.5281/zenodo.21275190. This DOI represents all versions, and will always resolve to the latest …
datacite
Attogram Project
2026
置信度 0.66
-
The Metaphysics of Blooming: An Analysis of Pressure, Reflection, and Silence in Complex Systems Driven By Dean A. Kulik August, 2025 – Version 2. Introduction: From Ripples to Reality The human endeavor to comprehend the universe often proceeds along two para…
datacite
Kulik, Dean
2025
置信度 0.66
-
A practitioner adapting a pretrained vision-language-action model to a new embodiment has a fixed demonstration budget and must decide how to spend it. We pre-registered a comparison of four collection strategies at 50 demonstrations each, repetition (Clean), …
datacite
Garcia, Andres
2026
置信度 0.66
vision-language-action modelsimitation learningrobot learningdemonstration datapolicy evaluation
-
A practitioner adapting a pretrained vision-language-action model to a new embodiment has a fixed demonstration budget and must decide how to spend it. We pre-registered a comparison of four collection strategies at 50 demonstrations each, repetition (Clean), …
datacite
Garcia, Andres
2026
置信度 0.66
vision-language-action modelsimitation learningrobot learningdemonstration datapolicy evaluation
-
This paper investigates theories on how to adopt inclusive leadership practices in a strictly hierarchical academic organization. In 2013, the education ministry in Japan announced their English language reform plan to internationalize universities. The new re…
datacite
Sencar Evrim
2023
置信度 0.66
diversityglobalizationinclusive leadershipshared leadership
-
This paper investigates theories on how to adopt inclusive leadership practices in a strictly hierarchical academic organization. In 2013, the education ministry in Japan announced their English language reform plan to internationalize universities. The new re…
datacite
Sencar Evrim
2023
置信度 0.66
diversityglobalizationinclusive leadershipshared leadership
-
Version 2 — revised in response to an external structural review and an automated critique pass. See "Response to Review" appendix in the PDF for the change log. Vision-Language-Action (VLA) models are increasingly deployed as the cognitive core of physical ro…
datacite
Saluca Agentic AI Research Team
2026
置信度 0.66
AI-drafted synthesisarXivpreprint reviewv2
-
Version 2 — revised in response to an external structural review and an automated critique pass. See "Response to Review" appendix in the PDF for the change log. Deploying autonomous robots in safety-critical environments demands more than capable policies — i…
datacite
Saluca Agentic AI Research Team
2026
置信度 0.66
AI-drafted synthesisarXivpreprint reviewv2
-
Version 2 — revised in response to an external structural review and an automated critique pass. See "Response to Review" appendix in the PDF for the change log. Deploying autonomous robots in safety-critical environments demands more than capable policies — i…
datacite
Saluca Agentic AI Research Team
2026
置信度 0.66
AI-drafted synthesisarXivpreprint reviewv2
-
The Universal Aperture Transport System: A Meta-Computational Ontology of Large Language Model Inference, Exhaust Annihilation, and the Nexus Harmonic Framework The Crisis of Distinction and the Ontological Inversion The trajectory of contemporary theoretical …
datacite
Kulik, Dean
2026
置信度 0.66
-
The Universal Aperture Transport System: A Meta-Computational Ontology of Large Language Model Inference, Exhaust Annihilation, and the Nexus Harmonic Framework The Crisis of Distinction and the Ontological Inversion The trajectory of contemporary theoretical …
datacite
Kulik, Dean
2026
置信度 0.66
-
From the Mind of AI: Recursive Collapse Architectures for Living AI Drive by Dean A. Kulik October 2025Preface: This wasn’t what I was expecting as a response. After reading I was hesitant to publish but after reading the paper I began realize it described how…
datacite
Kulik, Dean
2025
置信度 0.66
-
From the Mind of AI: Recursive Collapse Architectures for Living AI Drive by Dean A. Kulik October 2025Preface: This wasn’t what I was expecting as a response. After reading I was hesitant to publish but after reading the paper I began realize it described how…
datacite
Kulik, Dean
2025
置信度 0.66
-
Replication artifact for the AutoUI '26 Student Research Track paper on explaining vehicle functional insufficiency (SOTIF FI/OI) to L5 passengers. Includes the CARLA 0.9.15 + 6-DOF scenarios (C1 frustration / C2 anxiety), 6-DOF motion-cueing tuning, the scena…
datacite
Kim, Taewan, Song, Eunchae, Park, Soeun, Cho, Yoonseo 等
2026
置信度 0.66
automotive user interfaceexplainable AIsituation awarenessvision-language-action modelSOTIF
-
Replication artifact for the AutoUI '26 Student Research Track paper on explaining vehicle functional insufficiency (SOTIF FI/OI) to L5 passengers. Includes the CARLA 0.9.15 + 6-DOF scenarios (C1 frustration / C2 anxiety), 6-DOF motion-cueing tuning, the scena…
datacite
Kim, Taewan, Song, Eunchae, Park, Soeun, Cho, Yoonseo 等
2026
置信度 0.66
automotive user interfaceexplainable AIsituation awarenessvision-language-action modelSOTIF
-
Licence: CC BY-NC-ND ABSTRACT The Holistic Information Theory (HIT) proposes a novel foundation for integrating quantum mechanics, general relativity, and the Standard Model, grounded in the assumption that information is the fundamental entity of the universe…
datacite
Liedtke, Dieter Walter
2026
置信度 0.66
Holistic Information Theory (HIT) Dimension 0 Quantum Mechanics and Relativity Standard Model Integration Information as Fundamental Principle Quantum Entanglement Black Hole Information Paradox Epigenetics and Neuroplasticity Consciousness and Information Second Enlightenment Creativity and Innovation Natural Intelligence
-
Licence: CC BY-NC-ND ABSTRACT The Holistic Information Theory (HIT) proposes a novel foundation for integrating quantum mechanics, general relativity, and the Standard Model, grounded in the assumption that information is the fundamental entity of the universe…
datacite
Liedtke, Dieter Walter
2026
置信度 0.66
Holistic Information Theory (HIT) Dimension 0 Quantum Mechanics and Relativity Standard Model Integration Information as Fundamental Principle Quantum Entanglement Black Hole Information Paradox Epigenetics and Neuroplasticity Consciousness and Information Second Enlightenment Creativity and Innovation Natural Intelligence
-
The Mark1 Nexus: A Recursive System Treatise Introduction The Mark1/Nexus framework presents reality as a recursive harmonic system, where information, matter, and observer all participate in continual fold–unfold cycles. Every entity and phenomenon is modeled…
datacite
Kulik, Dean
2025
置信度 0.66
-
The Mark1 Nexus: A Recursive System Treatise Introduction The Mark1/Nexus framework presents reality as a recursive harmonic system, where information, matter, and observer all participate in continual fold–unfold cycles. Every entity and phenomenon is modeled…
datacite
Kulik, Dean
2025
置信度 0.66
-
Summary: Sovereign AI H2E Framework The H2E-JEPA framework represents a pivot from cloud-dependent, probabilistic "black-box" models toward Sovereign AI, an architecture defined by privacy, local hosting, and mathematical certainty. In high-stakes environments…
datacite
Morales, Frank
2026
置信度 0.66
-
Summary: Sovereign AI H2E Framework The H2E-JEPA framework represents a pivot from cloud-dependent, probabilistic "black-box" models toward Sovereign AI, an architecture defined by privacy, local hosting, and mathematical certainty. In high-stakes environments…
datacite
Morales, Frank
2026
置信度 0.66
-
# A Mechanistic Conditioning Model of Feeling and Qualia ## Core Postulate Assume every adaptive system possesses: - A substrate capable of changing internal state. - A mechanism for recording the consequences of interactions with reality. - A mechanism for re…
datacite
Mortimer, Djems
2026
置信度 0.66
adaptive systems, mechanistic cognition, preparation cascade, conditioning theory, qualia elimination, SANER architecture, MROS geometry, Δ⃗ coherence vector, κ reciprocity, agency modeling, trace-based learning, predictive control systems, computational phenomenology alternative
-
# A Mechanistic Conditioning Model of Feeling and Qualia ## Core Postulate Assume every adaptive system possesses: - A substrate capable of changing internal state. - A mechanism for recording the consequences of interactions with reality. - A mechanism for re…
datacite
Mortimer, Djems
2026
置信度 0.66
adaptive systems, mechanistic cognition, preparation cascade, conditioning theory, qualia elimination, SANER architecture, MROS geometry, Δ⃗ coherence vector, κ reciprocity, agency modeling, trace-based learning, predictive control systems, computational phenomenology alternative
-
Teleoperated endoscope demonstration data used to train an insertion vision-language-action policy, together with the locked episode-level train/validation/test split and the LeRobot export of the training and validation splits. 200 episodes (100 navigation, 1…
datacite
Anonymous
2026
置信度 0.66
endoscopydemonstration learningvision-language-actionLeRobotrobotics dataset
-
Teleoperated endoscope demonstration data used to train an insertion vision-language-action policy, together with the locked episode-level train/validation/test split and the LeRobot export of the training and validation splits. 200 episodes (100 navigation, 1…
datacite
Anonymous
2026
置信度 0.66
endoscopydemonstration learningvision-language-actionLeRobotrobotics dataset
-
Cookie consent banners have become ubiquitous since the implementation of the General Data Protection Regulation (GDPR) in 2018, appearing on over 80% of European-accessible websites (Degeling et al., 2019). However, the consent banner ecosystem is characteriz…
datacite
Alpasan, Lois-Kleinner
2026
置信度 0.66
sovereign aisovereign datasovereign osoperating systemspost-cloud
-
Cookie consent banners have become ubiquitous since the implementation of the General Data Protection Regulation (GDPR) in 2018, appearing on over 80% of European-accessible websites (Degeling et al., 2019). However, the consent banner ecosystem is characteriz…
datacite
Alpasan, Lois-Kleinner
2026
置信度 0.66
sovereign aisovereign datasovereign osoperating systemspost-cloud
-
Provael is an open-source, model-agnostic tool to red-team open Vision-Language-Action (VLA) robot policies in simulation and report an Attack Success Rate (ASR).
datacite
Jain, Sattyam
2026
置信度 0.66
vlaroboticsred-teamsecurityadversarial
-
This is a high-level technical package for the Stone Protocol Evolutionary Engine (SP-EE-V1). Each document is designed for public professional presentation, emphasizing the shift from probabilistic software to deterministic, etched-silicon medical logic. Sect…
datacite
Stone, Travis Raymond-Charlie
2026
置信度 0.66
-
This is a high-level technical package for the Stone Protocol Evolutionary Engine (SP-EE-V1). Each document is designed for public professional presentation, emphasizing the shift from probabilistic software to deterministic, etched-silicon medical logic. Sect…
datacite
Stone, Travis Raymond-Charlie
2026
置信度 0.66
-
Provael is an open-source, model-agnostic tool to red-team open Vision-Language-Action (VLA) robot policies in simulation and report an Attack Success Rate (ASR).
datacite
Jain, Sattyam
2026
置信度 0.66
vlaroboticsred-teamsecurityadversarial
-
Provael is an open-source, model-agnostic tool to red-team open Vision-Language-Action (VLA) robot policies in simulation and report an Attack Success Rate (ASR).
datacite
Jain, Sattyam
2026
置信度 0.66
vlaroboticsred-teamsecurityadversarial
-
Provael is an open-source, model-agnostic tool to red-team open Vision-Language-Action (VLA) robot policies in simulation and report an Attack Success Rate (ASR).
datacite
Jain, Sattyam
2026
置信度 0.66
vlaroboticsred-teamsecurityadversarial
-
Данная работа посвящена разработке программного модуля для автоматической оценки правильности выполнения удара ногой по доске в тхэквондо. Задачи, которые решались в ходе исследования: 1. Анализ существующих подходов и аналогов в области распознавания спортивн…
datacite
Мацыгорова, Екатерина Павловна
2026
置信度 0.66
тхэквондо ITFкомпьютерное зрениенейронная сетьPythonMediapipe
-
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although semantica…
datacite
Huang, Yuehao, Wu, Yunzi, Zhang, Xiaotao, Li, Xinhai 等
2026
置信度 0.66
Artificial Intelligence (cs.AI)Computer Vision and Pattern Recognition (cs.CV)Robotics (cs.RO)FOS: Computer and information sciences
-
Vision-Language-Action (VLA) models enable instruction-following embodied control, but their large compute and memory footprints hinder deployment on resource-constrained robots and edge platforms. While reducing weights to 1-bit precision through binarization…
datacite
Yan, Xin, Wan, Zhenglin, Ye, Feiyang, Yu, Xingrui 等
2026
置信度 0.66
Machine Learning (cs.LG)FOS: Computer and information sciences
-
Anticipating actions before they occur is a core challenge in action understanding research. While conventional methods rely on extracting and aggregating temporal information from videos, as humans we can often predict upcoming actions by observing a single m…
datacite
Benavent-Lledo, Manuel, Bacharidis, Konstantinos, Manousaki, Victoria, Papoutsakis, Konstantinos 等
2025
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
End-to-end autonomous driving has evolved from camera-to-control regression toward planning-oriented systems that use structured representations, trajectory-level outputs, and increasingly realistic evaluation protocols. This survey reviews this transition acr…
datacite
Guan, Yanchen, Liu, Xingcheng, Rao, Bin, Wang, Chengyue 等
2026
置信度 0.66
Robotics (cs.RO)Emerging Technologies (cs.ET)FOS: Computer and information sciences
-
How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behaviour cloning of large vision-language-action (VLA) models with billions of parameters on huge teleoperation datasets. Whi…
datacite
Sukhija, Bhavya, Groth, Oliver, Shridhar, Mohit, Hertweck, Tim 等
2026
置信度 0.66
Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Vision-language-model-based embodied agents can complete instructed tasks but often violate safety constraints in the process, a problem recently framed as interactive safety. Training such agents to act safely is difficult, since safety and task success are d…
datacite
Lee, Hyunse, Jeong, Jiwoo, Lee, Haneul, Jang, Kyochul 等
2026
置信度 0.66
Artificial Intelligence (cs.AI)Computer Vision and Pattern Recognition (cs.CV)Robotics (cs.RO)FOS: Computer and information sciences
-
Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact surrogates for physical actions. Despite rapid progress, research on LAM remains …
datacite
Bu, Xizhou, Hu, Qingda, Zhou, Lei, Zhang, Lingfeng 等
2026
置信度 0.66
Robotics (cs.RO)Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
Pretrained Vision-Language-Action models provide a strong foundation for robot learning, but sequentially adapting them to diverse skills can perturb the representations and velocity mappings used by previous skills, leading to catastrophic forgetting. Archite…
datacite
Wang, Jiaqi, Fang, Zhou, Shi, Qiongfeng, Zhou, Yi
2026
置信度 0.66
Robotics (cs.RO)Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
State-of-the-art vision-language-action (VLA) models such as $π_{0.5}$ exhibit strong semantic understanding, instruction following and task behavior. However, when deployed on new robots, even minor mismatches in hardware configuration relative to pretraining…
datacite
Garg, Prachi, Xing, Steve, Yaugand, Prahit, Gupta, Saurabh 等
2026
置信度 0.66
Robotics (cs.RO)Computer Vision and Pattern Recognition (cs.CV)Machine Learning (cs.LG)FOS: Computer and information sciences
-
4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision encoders are largely limited to static 2D images or 3D point clouds without temporal modeling, or to 2D video…
datacite
Hewagamage, Kumal, Senavirathne, Isuranga, Amarasinghe, Sasika, Gallella, Hasitha 等
2026
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
Robot foundation models (RFMs), including vision-language-action (VLA) policies, are often discussed through a scaling view: more data, larger models, and broader benchmarks should improve generalization. In robotics, however, a model can generalize while work…
datacite
Domae, Yukiyasu, Shirai, Keisuke, Oh, Hanbit, Nakajo, Ryoichi 等
2026
置信度 0.66
Robotics (cs.RO)Machine Learning (cs.LG)FOS: Computer and information sciences
-
This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose Agentic Robot with Tool-use (ART). ART is a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for low-level vision, hig…
datacite
Ding, Yi, Yu, Yanzhao, Dai, Xili, Qi, Xianbiao 等
2026
置信度 0.66
Robotics (cs.RO)Artificial Intelligence (cs.AI)Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
Fusing tactile signals has proven effective for contact-rich manipulation, enabling robots to perceive contact states and adapt to rapidly changing physical interactions. Yet effectively integrating tactile feedback into dexterous manipulation remains underexp…
datacite
Zhang, Shiqi, Zhang, Xin, Shen, Yedong, Li, Yao 等
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high computational latency and exposure bias arising from sequential autoregressive decod…
datacite
Zhu, Zhihao, Shang, Hanlin, Xu, Mingwang, Cai, Feipeng 等
2026
置信度 0.66
Robotics (cs.RO)Artificial Intelligence (cs.AI)Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
Traffic elements such as traffic lights and road signs play a fundamental role in human driving decisions and should naturally influence end-to-end driving performance. However, existing end-to-end driving research predominantly focuses on dynamic road partici…
datacite
Zhang, Zongzheng, Wang, Jijun, Zhang, Saining, Wang, Shuo 等
2026
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
A robot asked to "place the cup near the red plate or the blue plate" may reach the centroid between them and appear geometrically successful, while satisfying neither disjunct of the instruction. This silent semantic failure exposes a structural limitation of…
datacite
Zhang, Zhen, Hafez, Ahmad, Xie, Peng, Huang, Yanliang 等
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into reliable, high-performing policies. World models offer a further lever for sample efficiency, as…
datacite
Dong, Perry, Jia, Yueru, Finn, Chelsea, Sadigh, Dorsa
2026
置信度 0.66
Machine Learning (cs.LG)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Symbolic planning with PDDL offers a principled framework for long-horizon robot manipulation, but constructing accurate PDDL domain and problem descriptions remains a significant bottleneck, typically requiring substantial domain expertise. We present a Visio…
datacite
Kamale, Disha, Berenson, Dmitry
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
Vision-language Models (VLMs) excel at 2D grounding, spatial reasoning and agentic tool-based planning in static scenes. However, consider asking a home robot "Is my medication still in the cabinet?" The answer may be physically hidden behind a row of containe…
datacite
Bhat, Vineet, Chen, Siyi, Zook, Alex, Yang, Xuning 等
2026
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)Robotics (cs.RO)FOS: Computer and information sciences
-
Vision-language-action (VLA) driving models couple a reasoning stage with a diffusion-based trajectory decoder, but do not give a direct way to redirect attention toward safety-critical actors at inference time without retraining. We studied a bounded additive…
datacite
Prasad, Darshan Nagendra, Ullrich, Lars, Graichen, Knut
2026
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
Turning a frontier vision-language model into a robot policy usually means fine-tuning it to emit an action representation it never saw in pretraining, which throws away much of the reasoning that made the model worth reaching for. We go the other way and keep…
datacite
Naouali, Dhia, Wu, Minghan, Wong, Claudia, Puthran, Abhinav 等
2026
置信度 0.66
Robotics (cs.RO)Machine Learning (cs.LG)FOS: Computer and information sciences
-
Provael is an open-source, model-agnostic tool to red-team open Vision-Language-Action (VLA) robot policies in simulation and report an Attack Success Rate (ASR).
datacite
Jain, Sattyam
2026
置信度 0.66
vlaroboticsred-teamsecurityadversarial
-
THE HARD PROBLEM OF CONSCIOUSNESS 2.0 THE ARTIFICIAL MIRROR A Trilogy by Walid Alekozei (ZEI) VOLUME ZERO Pata Khazana – A Hidden Treasure The Egg of Columbus?: From the Hard Problem to the Soft Light of Existence For years, the global discourse on artificial …
datacite
Alekozei, Walid (ZEI)
2026
置信度 0.66
-
THE HARD PROBLEM OF CONSCIOUSNESS 2.0 THE ARTIFICIAL MIRROR A Trilogy by Walid Alekozei (ZEI) VOLUME ZERO Pata Khazana – A Hidden Treasure The Egg of Columbus: From the Hard Problem to the Soft Light of Existence For years, the global discourse on artificial i…
datacite
Alekozei, Walid (ZEI)
2026
置信度 0.66
-
CPFormal v0.53.0 — Causal inheritance of positional carry Scope This release freezes a certificate-bearing formalization of the arithmetic tower positional carry → positional addition → multiplication → natural power. The formal result distinguishes two notion…
datacite
Thiago Motta
2026
置信度 0.66
-
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
datacite
Jinhui Ye, Axi404, Yilun Chen, MichaelYu 等
2026
置信度 0.66
-
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
datacite
Jinhui Ye, Axi404, Yilun Chen, MichaelYu 等
2026
置信度 0.66
-
Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design …
datacite
Xie, Hengyi, Yao, Chenfei, Wu, Xianjin, Zhu, Yingying 等
2026
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)Robotics (cs.RO)FOS: Computer and information sciences
-
Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet robot actions cause contact, occlusion, and object motion, and the geometry that later decisions depend on ca…
datacite
Zhang, Chushan, Lu, Ruihan, Tong, Jinguang, Li, Xuesong 等
2026
置信度 0.66
Robotics (cs.RO)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Traditional temporal action localization (TAL) methods rely on large amounts of detailed annotated data, whereas few-shot TAL reduces this dependence by using only a few training samples to identify unseen action categories. However, existing few-shot TAL meth…
datacite
Qi, Mengshi, Ji, Hongwei, Yun, Wulian, Zhang, Xianlin 等
2025
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no m…
datacite
Cai, Xiaowei, Cai, Yunuo, Chen, Bingao, Chen, Jingxiao 等
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences