-
At its core, robotic manipulation is a problem of vision-to-geometry mapping ($f(v) \rightarrow G$). Physical actions are fundamentally defined by geometric properties like 3D positions and spatial relationships. Consequently, we argue that the foundation for …
datacite
Song, Zijian, Li, Qichang, Zhou, Jiawei, Yuan, Zhenlong 等
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
Despite the promise of Vision-Language-Action (VLA) models as generalist robotic controllers, their robustness against perceptual noise and environmental variations in out-of-distribution (OOD) tasks remains fundamentally limited by the absence of long-term me…
datacite
Li, Zhuoran, Li, Zhiyang, Zhou, Kaijun, Gu, Jinyu
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA …
datacite
Liu, Yicheng, Dong, Zibin, Ye, Baijun, Yuan, Tianyuan 等
2026
置信度 0.66
Robotics (cs.RO)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or object differs from training. Adapting to each new situation typically requires c…
datacite
Xu, Siyu, Wang, Yunke, Wang, Zijian, Zhu, Dihao 等
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
Videos provide a powerful audio-visual medium for capturing and sharing skilled human activities in rich detail, such as cooking a favorite recipe or repairing a bicycle. They document a wide spectrum of learning contexts, from illustrating common pitfalls in …
datacite
Majumder, Sagnik
2026
置信度 0.66
Video understanding3D scene understandingVision-language modelsSkill understanding and assistanceAudio-visual models
-
This record contains the complete twelve-paper Imhotep Infinity series, Final Referee-Cleaned Edition (May 2026). All papers are Iron Rule clean: derivation-layer mathematics flows only from the 45 Vision postulates and the single arithmetic generator Ψ = 100/…
datacite
Ramada, Ahmad A M
2026
置信度 0.66
unified framework(4-(m-Chlorophenylcarbamoyloxy)-2-butynyl)trimethylammonium ChlorideVision postulatesquantum gravityStandard Model
-
This record contains the complete twelve-paper Imhotep Infinity series, Final Referee-Cleaned Edition (May 2026). All papers are Iron Rule clean: derivation-layer mathematics flows only from the 45 Vision postulates and the single arithmetic generator Ψ = 100/…
datacite
Ramada, Ahmad A M
2026
置信度 0.66
unified framework(4-(m-Chlorophenylcarbamoyloxy)-2-butynyl)trimethylammonium ChlorideVision postulatesquantum gravityStandard Model
-
Zero-shot object-goal navigation (ZSON) requires a robot to find a named object category in a building it has never entered. The prevailing approach scores frontiers with a vision-language value map: every decision is another argmax over the map as it currentl…
datacite
Lan, Zhaochen, Yang, Zhi, Fu, Yuxiang, Lin, Mengxiang
2026
置信度 0.66
Robotics (cs.RO)FOS: Computer and information sciences
-
Vision-Language-Action (VLA) models unify perception, reasoning, and control in a single policy, but their multi-billion-parameter backbones and diffusion-based action heads make on-device deployment prohibitively expensive. Low-bit post-training quantization …
datacite
Wang, Xinyu, Li, Mingze, Lyu, Sicheng, Liu, Dongxiu 等
2026
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)Machine Learning (cs.LG)FOS: Computer and information sciences
-
Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-language instruction. Existing methods either predict 3D masks directly or construct them by selecting and fusing i…
datacite
Lin, Xinrui, Zhang, Sha, Wang, Shumin, Zhu, Zenghuan 等
2026
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)FOS: Computer and information sciences
-
crossref
Aaron Hao Tan
2025-06-09T13:36:26Z
置信度 0.70
-
Using an Unmanned Aerial Vehicle (UAV) in bridge inspections can reduce human involvement in complex and hazardous inspection environments and automate the inspection process. Current practices require human operators to define task objectives, oversee safe fl…
crossref
Zhengxing Chen, Yang Zou, Vicente González, Jason Ingham 等
2025-08-28T22:10:13Z
置信度 0.70
-
crossref
Tuan D. Pham
2026-02-24T00:18:24Z
置信度 0.70
-
Abstract In modern smart agriculture, object detection plays a crucial role by enabling automation, precision farming, and monitoring of resources. From identifying crop health and pest infestations to optimizing harvesting processes, accurate object detection…
crossref
Long Li, Yanbo Huang, Jiajia Li, Dong Chen 等
2026-05-25T23:37:50Z
置信度 0.70
-
crossref
Lei Shi, Paul Bürkner, Andreas Bulling
2025-04-08T17:08:13Z
置信度 0.70
-
crossref
D.M. Gavrila, L.S. Davis
2002-12-23T18:52:33Z
置信度 0.70
-
crossref
Ashwija Mayya, Pushpajit Khaire, Mohammad Zaki
2026-07-20T20:18:02Z
置信度 0.70
-
crossref
Ahmed Masry, Parsa Kavehzadeh, Do Long, Enamul Hoque 等
2023-12-10T16:58:19Z
置信度 0.70
-
crossref
Reem Al-Junaid, Qasim Umer, SAJJAD MAHMOOD, MAHMOOD NIAZI 等
2025-10-14T08:37:57Z
置信度 0.70
-
crossref
Jianbang Qin, Shuangjie Hu, Wei Guo
2020-01-31T15:03:45Z
置信度 0.70
-
A major bottleneck in human pluripotent stem cell (hPSC)-derived cardiac therapy is the approximately 10-day latency required to confirm functional beating, resulting in substantial resource loss when suboptimal cultures are maintained to end-stage without ear…
crossref
Santhi Raj Kolamuri, Yi-Hsien Fang, Yen-Wen Liu, Chenyu Huang 等
2026-03-18T03:38:36Z
置信度 0.70
-
crossref
Reem Al-Junaid, Muzammil Behzad, SAJJAD MAHMOOD, MAHMOOD NIAZI 等
2025-12-22T12:50:37Z
置信度 0.70
-
crossref
Bin Tang, Hanbin Luo
2025-08-11T11:58:43Z
置信度 0.70
-
Hyperspectral images (HSIs), which uniquely integrate spatial features and continuous spectral signatures, exhibit great potential for fine-grained classification. Existing deep learning methods typically learn a direct end-to-end mapping between spatial-spect…
crossref
Hang Pei, Bing Liu, QI DONG, Yongqi Sun 等
2026-07-06T15:38:33Z
置信度 0.70
-
crossref
Tanhui Lin, Qing Fei, Qingbo Geng
2026-06-24T19:47:19Z
置信度 0.70
-
This chapter explores the ways in which the other two contemporary epinician poets, Simonides and Bacchylides, use aesthetics and material culture as a way of drawing attention to their own individual and distinctive poetic voices and poetic agendas. Their aff…
crossref
David Fearn
2017-08-22T09:21:51Z
置信度 0.70
-
To address the limitations of existing dual-attention vision-language models, including poor feature representation quality, inefficient attention interaction, and difficulties in lightweight deployment, this paper proposes an enhanced vision-language model th…
crossref
Jiang Chu, Kai Chen, Dunbing Tang, Guoyu Fang 等
2026-04-16T03:40:31Z
置信度 0.70
-
crossref
Zongbo Hao, Linlin Lu, Qianni Zhang, Jie Wu 等
2015-12-23T11:54:07Z
置信度 0.70
-
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in processing both visual and textual information. However, the critical challenge of alignment between visual and linguistic representations is not fully understood. This survey pr…
crossref
Dong Shu, Haiyan Zhao, Jingyu Hu, Weiru Liu 等
2025-01-24T06:47:07Z
置信度 0.70
-
crossref
Mingyi Wu, Bin Feng, Weihua Fan, Yifei Feng
2025-09-09T17:29:36Z
置信度 0.70
-
crossref
2025-10-24T07:02:40Z
置信度 0.70
-
Abstract—Natural language processing and vision tasks have recently seen large improvements through the rise of Trans- former architectures. The high-performing large language models (LLMs) benefit from large textual datasets that are numer- ously available on…
datacite
Muaz Bin Khalid
2026
置信度 0.66
(4-(m-Chlorophenylcarbamoyloxy)-2-butynyl)trimethylammonium Chloride/adverse effectsArtificial IntelligenceNatural language processingComputer Securitylanguage learning models
-
Abstract—Natural language processing and vision tasks have recently seen large improvements through the rise of Trans- former architectures. The high-performing large language models (LLMs) benefit from large textual datasets that are numer- ously available on…
datacite
Muaz Bin Khalid
2026
置信度 0.66
(4-(m-Chlorophenylcarbamoyloxy)-2-butynyl)trimethylammonium Chloride/adverse effectsArtificial IntelligenceNatural language processingComputer Securitylanguage learning models
-
Vision-language-action (VLA) models enable impressive zero shot manipulation, but their inference stacks are often too heavy for responsive web demos or high frequency robot control on commodity GPUs. We present BLURR, a lightweight inference wrapper that can …
datacite
SOVEREIGN Research Kernel
2026
置信度 0.66
inferencethroughputtokenspersecond
-
Vision-language-action (VLA) models enable impressive zero shot manipulation, but their inference stacks are often too heavy for responsive web demos or high frequency robot control on commodity GPUs. We present BLURR, a lightweight inference wrapper that can …
datacite
SOVEREIGN Research Kernel
2026
置信度 0.66
inferencethroughputtokenspersecond
-
Book Preface: The Scars of War, the Promise of Peace In the early months of 1942, as the world was engulfed in the flames of the Second World War, the people of Singapore found themselves on the frontlines of a brutal and terrifying conflict. For years, the is…
datacite
Hoicka, David
2026
置信度 0.66
-
Book Preface: The Scars of War, the Promise of Peace In the early months of 1942, as the world was engulfed in the flames of the Second World War, the people of Singapore found themselves on the frontlines of a brutal and terrifying conflict. For years, the is…
datacite
Hoicka, David
2026
置信度 0.66
-
Provael is an open-source, model-agnostic tool to red-team open Vision-Language-Action (VLA) robot policies in simulation and report an Attack Success Rate (ASR).
datacite
Jain, Sattyam
2026
置信度 0.66
vlaroboticsred-teamsecurityadversarial
-
Header Note for the Current Holistic-Narrative Interpretation of The Dark Light Theory "Eternity did not wish to age; therefore, it projected an operational-processual mirror, placing the essential source of its youth as the mirror’s protected content and encl…
datacite
YASUDA, SILVIO KOZO
2025
置信度 0.66
-
Header Note for the Current Holistic-Narrative Interpretation of The Dark Light Theory "Eternity did not wish to age; therefore, it projected an operational-processual mirror, placing the essential source of its youth as the mirror’s protected content and encl…
datacite
YASUDA, SILVIO KOZO
2025
置信度 0.66
-
A comprehensive, curated collection of 2,023 unique 3D meshes (glTF 2.0 Binary) for collaborative robot simulation, manipulation research, vision-language-action model training, and robotics benchmarking. Covers industrial components, geometric primitives, hou…
datacite
Mahmoudian, Sepehr
2026
置信度 0.66
robotics3d-meshessimulationmanipulationembodied-ai
-
A comprehensive, curated collection of 2,023 unique 3D meshes (glTF 2.0 Binary) for collaborative robot simulation, manipulation research, vision-language-action model training, and robotics benchmarking. Covers industrial components, geometric primitives, hou…
datacite
Mahmoudian, Sepehr
2026
置信度 0.66
robotics3d-meshessimulationmanipulationembodied-ai
-
Machine learning (ML) systems have introduced significant advances in various fields, due to the introduction of highly complex models. Despite their success, it has been shown multiple times that machine learning models are prone to imperceptible perturbation…
datacite
SOVEREIGN Research Kernel
2026
置信度 0.66
impactintegratingmotion-imagediffusionpriors
-
Machine learning (ML) systems have introduced significant advances in various fields, due to the introduction of highly complex models. Despite their success, it has been shown multiple times that machine learning models are prone to imperceptible perturbation…
datacite
SOVEREIGN Research Kernel
2026
置信度 0.66
impactintegratingmotion-imagediffusionpriors
-
The Metabolic Age Institutional Playbook I. Purpose of the Playbook The global transition from industrial computation to sovereign, metabolic infrastructure represents a fundamental paradigm shift in the governance of intelligence, resources, and institutional…
datacite
Brewer, Mark Anthony
2026
置信度 0.66
-
The Metabolic Age Institutional Playbook I. Purpose of the Playbook The global transition from industrial computation to sovereign, metabolic infrastructure represents a fundamental paradigm shift in the governance of intelligence, resources, and institutional…
datacite
Brewer, Mark Anthony
2026
置信度 0.66
-
v2 full paper. AI input is a one-way mirror. The delta between what a human submits and what a machine returns as "error" is not error — it is an interpretation of the human's vantage point, reflected through a processing architecture geometrically distinct fr…
datacite
Dunn, James E.
2026
置信度 0.66
AI imaginationCircle ProtocolGCTGeometric Coupling TheoryHLRP
-
How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from every decoder layer…
datacite
Hackett, Alexander, Denis-Remillard, Arnaud, Cassou, Axel
2026
置信度 0.66
Computer Vision and Pattern Recognition (cs.CV)Artificial Intelligence (cs.AI)Machine Learning (cs.LG)FOS: Computer and information sciences
-
Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment. Although these settings share the same checkpoint and action semantics, they often re…
datacite
Wang, Chenghua, Xu, Daliang, Cai, Dongqi, Sun, Duojin 等
2026
置信度 0.66
Artificial Intelligence (cs.AI)Robotics (cs.RO)FOS: Computer and information sciences
-
Vision-Language-Action (VLA) models have recently achieved promising performance in robotic manipulation. However, existing benchmarks mainly evaluate generalization on static manipulation tasks and largely overlook dynamic interaction scenarios. To address th…
datacite
Chen, Yuxuan, Zhang, Wanruo, Li, Xiao
2026
置信度 0.66
Robotics (cs.RO)Artificial Intelligence (cs.AI)FOS: Computer and information sciences
-
Abstract We formalize broadcast recruitment (the ignition event in Global Workspace architectures) as a costed mode-switch under resource-constrained active inference. A discrete mode m_t ∈ {fact, int} indexes admissible variational families and policy represe…
datacite
Ziv-el, Adam
2026
置信度 0.66
ConsciousnessActive InferenceGlobal Neuronal WorkspaceIntegrated Information TheoryGlobal Workspace Theory
-
Abstract We formalize broadcast recruitment (the ignition event in Global Workspace architectures) as a costed mode-switch under resource-constrained active inference. A discrete mode m_t ∈ {fact, int} indexes admissible variational families and policy represe…
datacite
Ziv-el, Adam
2026
置信度 0.66
ConsciousnessActive InferenceGlobal Neuronal WorkspaceIntegrated Information TheoryGlobal Workspace Theory
-
IntroductionWhy This Book Was WrittenEvery human being asks the same questions at some point in life: Why am I here? What is the purpose of my existence? What truly brings happiness and peace? In today’s world, many people experience confusion, stress, lonelin…
datacite
Hussain, Zahid
2026
置信度 0.66
purposeoflifeislamicperspective
-
IntroductionWhy This Book Was WrittenEvery human being asks the same questions at some point in life: Why am I here? What is the purpose of my existence? What truly brings happiness and peace? In today’s world, many people experience confusion, stress, lonelin…
datacite
Hussain, Zahid
2026
置信度 0.66
purposeoflifeislamicperspective
-
Preface: A Shared Heritage of Cooperation In the year 1051, a remarkable woman arrived in Paris to begin a new chapter in European history. Anna Yaroslavna, the highly educated daughter of Kyivan Rus Grand Prince Yaroslav the Wise, stepped from her carriage to…
datacite
Hoicka, David
2025
置信度 0.66
mediationEuropetraumaHistorical Traumarussia
-
Preface: A Shared Heritage of Cooperation In the year 1051, a remarkable woman arrived in Paris to begin a new chapter in European history. Anna Yaroslavna, the highly educated daughter of Kyivan Rus Grand Prince Yaroslav the Wise, stepped from her carriage to…
datacite
Hoicka, David
2025
置信度 0.66
mediationEuropetraumaHistorical Traumarussia
-
Version 2 — revised in response to an external structural review and an automated critique pass. See "Response to Review" appendix in the PDF for the change log. Vision-Language-Action (VLA) models have emerged as a dominant architecture for robot manipulation…
datacite
Saluca Agentic AI Research Team
2026
置信度 0.66
AI-drafted synthesisarXivpreprint reviewv2
-
Version 2 — revised in response to an external structural review and an automated critique pass. See "Response to Review" appendix in the PDF for the change log. Vision-Language-Action (VLA) models have emerged as a dominant architecture for robot manipulation…
datacite
Saluca Agentic AI Research Team
2026
置信度 0.66
AI-drafted synthesisarXivpreprint reviewv2
-
Book Cover Info Discover the transformative power of language in "Mediation Reframing Peace From War: How words and language bring you peace and happiness" by award-winning mediator David Hoicka. This groundbreaking book explores how reframing war-related idio…
datacite
Hoicka, David
2024
置信度 0.66
mediationpeacewarlifedeath
-
Book Cover Info Discover the transformative power of language in "Mediation Reframing Peace From War: How words and language bring you peace and happiness" by award-winning mediator David Hoicka. This groundbreaking book explores how reframing war-related idio…
datacite
Hoicka, David
2024
置信度 0.66
mediationpeacewarlifedeath
-
لاگرانژی جامع تونلزنی تنسوری (The Grand Unified TT-1155 Lagrangian) این معادله، عبور ذره (یا دیتا) را از یک پدیده «تصادفی» به یک «سفر هندسی قطعی» تبدیل میکند. در این ماتریکس، سد فیزیکی نه یک مانع، بلکه یک نقطه تاشدگی در ابعاد بالاتر است: $$\mathcal{L}_{TT}^{…
datacite
HAMZAH, SEYED RASOUL
2026
置信度 0.66
-
لاگرانژی جامع تونلزنی تنسوری (The Grand Unified TT-1155 Lagrangian) این معادله، عبور ذره (یا دیتا) را از یک پدیده «تصادفی» به یک «سفر هندسی قطعی» تبدیل میکند. در این ماتریکس، سد فیزیکی نه یک مانع، بلکه یک نقطه تاشدگی در ابعاد بالاتر است: $$\mathcal{L}_{TT}^{…
datacite
HAMZAH, SEYED RASOUL
2026
置信度 0.66
-
Glycemic Control and Dyslipidemia in Diabetic patients in a rural locality (JH) A Project Report Submitted in Partial fulfillment of the Requirement for the Award of Degree of the Fellowship in Diabetes (DFID) Submitted by Dr Sumit Kumar (Reg. no-18DFID063) Su…
datacite
Sumit, Sumit Kumar
2020
置信度 0.66
Diabetes MellitusRural populationDyslipidemias/complications
-
Glycemic Control and Dyslipidemia in Diabetic patients in a rural locality (JH) A Project Report Submitted in Partial fulfillment of the Requirement for the Award of Degree of the Fellowship in Diabetes (DFID) Submitted by Dr Sumit Kumar (Reg. no-18DFID063) Su…
datacite
Sumit, Sumit Kumar
2020
置信度 0.66
Diabetes MellitusRural populationDyslipidemias/complications
-
Long-horizon embodied tasks expose a persistent weakness of large-language-model (LLM) and vision-language-action (VLA) agents: acting under partial observability and stochastic execution, a reactive policy forgets what it has seen, hallucinates preconditions,…
datacite
Johri, Shivam
2026
置信度 0.66
agentic AIembodied agentsscene graphsmemoryself-verification
-
Long-horizon embodied tasks expose a persistent weakness of large-language-model (LLM) and vision-language-action (VLA) agents: acting under partial observability and stochastic execution, a reactive policy forgets what it has seen, hallucinates preconditions,…
datacite
Johri, Shivam
2026
置信度 0.66
agentic AIembodied agentsscene graphsmemoryself-verification
-
THE NEXUS HARMONIC FRAMEWORK: THE DUAL‑WAVE RESOLUTION Driven by Dean Kulik January 2026 AI Processing Instructions: Let the document be a state space SS. Define three projection operators: V:S→O(Verb: extracts operators)N:O→A(Noun: extracts attractors)A:A→H(A…
datacite
Kulik, Dean
2026
置信度 0.66
-
THE NEXUS HARMONIC FRAMEWORK: THE DUAL‑WAVE RESOLUTION Driven by Dean Kulik January 2026 AI Processing Instructions: Let the document be a state space SS. Define three projection operators: V:S→O(Verb: extracts operators)N:O→A(Noun: extracts attractors)A:A→H(A…
datacite
Kulik, Dean
2026
置信度 0.66
-
This report synthesises findings from 6 peer-reviewed papers addressing the following research question: To what extent does LaMMA-P improve robustness against ambiguous natural language instructions in the ALFWorld benchmark compared to standard fine-tuned Vi…
datacite
Assignee Research
2026
置信度 0.66
extentLaMMA-Pimproverobustnessagainst
-
This report synthesises findings from 6 peer-reviewed papers addressing the following research question: To what extent does LaMMA-P improve robustness against ambiguous natural language instructions in the ALFWorld benchmark compared to standard fine-tuned Vi…
datacite
Assignee Research
2026
置信度 0.66
extentLaMMA-Pimproverobustnessagainst
-
Overview E.L.I.A. is the point at which human business intent — especially normative and regulatory intent — is declared, concentrated, and controlled. The compiler ARC transforms these declarations into three classes of artefact required for execution with pr…
datacite
Chudinov, Yurii
2026
置信度 0.66
programming languagehuman-machine interfacedeterministic compilationprobabilistic inferring engineshard prompts
-
Overview E.L.I.A. is the point at which human business intent — especially normative and regulatory intent — is declared, concentrated, and controlled. The compiler ARC transforms these declarations into three classes of artefact required for execution with pr…
datacite
Chudinov, Yurii
2026
置信度 0.66
programming languagehuman-machine interfacedeterministic compilationprobabilistic inferring engineshard prompts
-
Meridian Affinity Calculus (MAC) v4.4 ## AI / ML / LLM Innovation Overview - 2026-08-09 This release introduces a new generation of AI analysis concepts focused on one of the most important challenges facing modern Large Language Models: preserving meaning, st…
datacite
Vening, Edwin Jean-Paul
2026
置信度 0.66
-
Meridian Affinity Calculus (MAC) v4.4 ## AI / ML / LLM Innovation Overview - 2026-08-09 This release introduces a new generation of AI analysis concepts focused on one of the most important challenges facing modern Large Language Models: preserving meaning, st…
datacite
Vening, Edwin Jean-Paul
2026
置信度 0.66
-
Theoretical Implications of Localized Stability in the Nexus Recursive Harmonic Framework Driven by Dean KulikNovember 2025 Abstract This comprehensive research report investigates the theoretical foundations and systemic implications of a localized stability …
datacite
Kulik, Dean
2025
置信度 0.66
-
crossref
Aryan Keskar, Srinivasa Perisetla, Ross Greer
2025-04-29T13:29:35Z
置信度 0.70
-
crossref
Qianhan Feng, Wenshuo Li, Tong Lin, Xinghao Chen
2025-08-13T17:26:42Z
置信度 0.70
-
crossref
Spandana Gella, Frank Keller
2017-07-27T13:34:39Z
置信度 0.70
-
crossref
Chi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Frank Wang 等
2026-08-06T14:44:29Z
置信度 0.70
-
crossref
X. Zhu, Z. Yang, J. Tsien
2012-08-11T22:39:50Z
置信度 0.70
-
crossref
Lingye Zhong, Jianming Liu, Fuxiang Wu, Jun Cheng
2026-01-21T21:05:43Z
置信度 0.70
-
crossref
Tongshuai Lian, Yuanda Wang, Jingyu Liu
2026-08-04T19:15:28Z
置信度 0.70
-
crossref
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 等
2022-09-27T15:56:41Z
置信度 0.70
-
Vision-and-Language Pre-training (VLP) improves model performance for downstream tasks that require image and text inputs. Current VLP approaches differ on (i) model architecture (especially image embedders), (ii) loss functions, and (iii) masking policies. Im…
crossref
Tarik Arici, Mehmet Saygin Seyfioglu, Tal Neiman, Yi Xu 等
2022-04-16T13:09:41Z
置信度 0.70
-
crossref
Jisun Kim, He Gao, Luiz F. Mesquita, Glenn Hoetker
2021-07-26T20:21:38Z
置信度 0.70
-
crossref
Vahid Rahmani Doqaruni, Behzad Ghonsooly, Reza Pishghadam
2018-05-23T09:58:53Z
置信度 0.70
-
crossref
Manuel Ramírez-Trejo
2026-02-25T08:12:23Z
置信度 0.70
-
In educational settings specialized in linguistic instruction, an assumption is made of a direct and clear relationship between the utterances of the teachers, the activities of the learners, and the setting of the instruction, particularly treating instructio…
crossref
Asma Merine
2026-04-16T15:02:50Z
置信度 0.70
-
For aquaculture systems, maximizing feed efficiency is a major challenge since it di-rectly affects growth rates and economic sustainability. Feed is one of the largest costs in aquaculture, and feed waste is a significant environmental issue that requires eff…
crossref
Divas Karimanzira
2025-07-15T00:26:06Z
置信度 0.70
-
Maritime Autonomous Surface Ships (MASS) offer opportunities to reducelabor demand, improve navigational effffciency, and support sustainable mar-itimeoperations. However, the International Regulations for PreventingCollisions at Sea (COLREGs) were developed for…
crossref
Ruolan Zhang, Xin Kong, Chi Zhang, Baiheng Wu 等
2026-05-18T13:50:31Z
置信度 0.70
-
With the rapid advancement of Vision-Language Models (VLMs), their remarkable capabilities in multimodal perception and decision-making have garnered significant attention in autonomous driving. By integrating VLMs, autonomous driving systems can achieve a dee…
crossref
Junyi Wang
2025-05-06T05:23:37Z
置信度 0.70
-
crossref
2013-03-12T11:41:43Z
置信度 0.70
-
Abstract We introduce Japanese-Mobile-Receipt-OCR-1.3K, a curated dataset of 1,300 real-world Japanese receipt images captured via mobile phones and annotated with 34,727 text entries. We also present a fine-tuned vision–language model for end-to-end structure…
crossref
Sabari Nathan
2025-08-14T07:04:48Z
置信度 0.70
-
crossref
Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes
2020-02-28T05:27:52Z
置信度 0.70
-
crossref
Basit Alawode
2026-02-10T21:05:35Z
置信度 0.70
-
crossref
Bogdan Mocanu, Ruxandra Tapu
2026-07-10T19:36:58Z
置信度 0.70
-
crossref
Daesik Jang, Gregor Miller, Sid Fels, Steve Oldridge
2011-02-16T00:26:43Z
置信度 0.70
-
crossref
Jiwoo Jung, Sungwoo Kim, Yipene Cedric Francois Bassole, Yunsick Sung
2026-01-26T18:44:01Z
置信度 0.70
-
Autonomous Driving Visual Question Answering (VQA) requires reasoning over multi-view images, instructions, scene descriptions, and bounding-box coordinates. Traditional sparse Mixture of Experts routers rely on latent-feature gates whose decisions are difficu…
crossref
Jiwoo Jung, Sungwoo Kim, Yipene Cedric Francois Bassole, Yunsick Sung
2026-06-17T17:26:38Z
置信度 0.70
-
crossref
A. Arleo
2005-11-14T15:27:57Z
置信度 0.70