-
Abstract Autonomous agents have long been a research focus in academic and industry communities. Previous research often focuses on training agents with limited knowledge within isolated environments, which diverges significantly from human learning processes,…
openalex
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang 等
2024-03-22
置信度 0.72
Computer scienceArtificial intelligence
-
Vision Language Action (VLA) models represent a new frontier in robotics by unifying perception, reasoning, and control within a single multimodal learning framework. By integrating visual, linguistic, and action modalities, they enable multimodal fusion syste…
openalex
Muhayy Ud Din, Waseem Akram, Lyes Saad Saoud, Jan Rosell 等
2025-12-16
置信度 0.72
Computer scienceArtificial intelligenceMachine learningBenchmarkingRobustness (evolution)
-
The recent GPT-4 has demonstrated extraordinary multi-modal abilities, such as directly generating websites from handwritten text and identifying humorous elements within images. These features are rarely observed in previous vision-language models. However, t…
openalex
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等
2023-04-20
置信度 0.72
Computer scienceUsabilityRepetition (rhetorical device)ModalCode (set theory)
-
openalex
George Α. Papadopoulos, Farhad Arbab
1998-01-01
置信度 0.72
Computer scienceRotation formalisms in three dimensionsInterfacingComponent (thermodynamics)Set (abstract data type)
-
In modern healthcare, the demand for autonomous robotic assistants has grown significantly, particularly in the operating room, where surgical tasks require precision and reliability. Robotic scrub nurses have emerged as a promising solution to improve efficie…
openalex
Shunlei Li, Jin Wang, Rui Dai, Wanyu Ma 等
2025-10-19
置信度 0.72
Task (project management)Computer scienceHuman–computer interactionHandoverArtificial intelligence
-
The advancement of large Vision-Language-Action (VLA) models has significantly improved robotic manipulation in terms of language-guided task execution and generalization to unseen scenarios. While existing VLAs adapted from pretrained large Vision-Language-Mo…
openalex
Qixiu Li, Yaobo Liang, Zeyu Wang, Lin Luo 等
2024-11-29
置信度 0.72
Action (physics)Cognitive scienceCognitionCognitive roboticsPsychology
-
Pretrained large language models (LLMs) are widely used in many sub-fields of natural language processing (NLP) and generally known as excellent few-shot learners with task-specific exemplars. Notably, chain of thought (CoT) prompting, a recent technique for e…
openalex
Takeshi Kojima, Shixiang Gu, Machel Reid, Yutaka Matsuo 等
2022-05-24
置信度 0.72
Shot (pellet)Task (project management)Benchmark (surveying)Computer scienceZero (linguistics)
-
This paper reviews a selection of research from the field of foreign and second language teaching into what is referred to here as teacher cognition – what teachers think, know, and believe and the relationships of these mental constructs to what teachers do i…
openalex
Simon Borg
2003-04-01
置信度 0.72
CognitionPsychologyMainstreamLanguage educationPerspective (graphical)
-
“Force dynamics” refers to a previously neglected semantic category—how entities interact with respect to force. This category includes such concepts as: the exertion of force, resistance to such exertion and the overcoming of such resistance, blockage of a fo…
openalex
Léonard Talmy
1988-01-01
置信度 0.72
Notional amountDynamics (music)CognitionCognitive scienceCognitive psychology
-
PASSIVE VISION AND ACTIVE VISION 1.1 Introduction 1.2 Passive vision 1.3 Visual attention 1.4 Active vision 1.5 Active vision and vision for action 1.6 Outline of the book BACKGROUND TO ACTIVE VISION 2.1 Introduction 2.2 The inhomogeneity of the visual project…
openalex
2004-04-01
置信度 0.72
AestheticsPsychologyHistoryEpistemologyPhilosophy
-
BACKGROUND: Health literacy concerns the knowledge and competences of persons to meet the complex demands of health in modern society. Although its importance is increasingly recognised, there is no consensus about the definition of health literacy or about it…
openalex
Kristine Sørensen, Stephan Van den Broucke, James Fullam, Gerardine Doyle 等
2012-01-25
置信度 0.72
Health literacyPublic healthHealth promotionConceptual modelHealth care
-
Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human vision. Computer vision foundation models, which are trained on diverse, large-scale …
openalex
Lu Yuan, Dongdong Chen, Yi‐Ling Chen, Noel Codella 等
2021-11-22
置信度 0.72
Computer scienceArtificial intelligenceObject (grammar)Computer visionRepresentation (politics)
-
We present OpenDriveVLA, a Vision-Language Action (VLA) model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially-grounded driving actions by leveraging multimodal inputs, including both 2D…
openalex
Xingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma 等
2026-03-14
置信度 0.72
Computer scienceArtificial intelligenceAction (physics)TrajectoryBridge (graph theory)
-
Developing versatile quadruped robots that can smoothly perform various actions and tasks in real-world environments remains a significant challenge. This paper introduces a novel vision-language-action (VLA) model, mixture of robotic experts (MoRE), for quadr…
openalex
Han Zhao, Wenxuan Song, Donglin Wang, Xinyang Tong 等
2025-05-19
置信度 0.72
Reinforcement learningScalabilityComputer scienceAction (physics)Artificial intelligence
-
In recent years, many accurate decision support systems have been constructed as black boxes, that is as systems that hide their internal logic to the user. This lack of explanation constitutes both a practical and an ethical issue. The literature reports many…
openalex
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini 等
2019-01-01
置信度 0.72
InterpretabilityBlack boxComputer sciencePerspective (graphical)Data science
-
Abstract Transformer-based large language models are making significant strides in various fields, such as natural language processing 1–5 , biology 6,7 , chemistry 8–10 and computer programming 11,12 . Here, we show the development and capabilities of Coscien…
openalex
Daniil A. Boiko, Robert MacKnight, Ben Kline, Gabriel dos Passos Gomes
2023-12-20
置信度 0.72
Computer scienceDocumentationAutomationTransformerArtificial intelligence
-
This second edition of Norton’s classic text on language learning and identity will bring her ground-breaking ideas to a new generation of students, teachers and researchers. Featuring a comprehensive Introduction and an Afterword by Claire Kramsch, this new e…
openalex
Bonny Norton
2013-12-31
置信度 0.72
Identity (music)LinguisticsLanguage acquisitionMathematics educationComputer science
-
Large-scale pre-training and instruction tuning have been successful at creating general-purpose language models with broad competence. However, building general-purpose vision-language models is challenging due to the rich input distributions and task diversi…
openalex
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等
2023-05-11
置信度 0.72
Computer scienceLanguage modelArtificial intelligenceTransformerNatural language processing
-
Physical inactivity is the fourth leading cause of death worldwide. We summarise present global efforts to counteract this problem and point the way forward to address the pandemic of physical inactivity. Although evidence for the benefits of physical activity…
openalex
Harold W. Kohl, Cora L. Craig, Estelle V. Lambert, Shigeru Inoue 等
2012-07-01
置信度 0.72
PandemicPublic healthHealth promotionAction (physics)Workforce
-
Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes. However, current VLA models face significant challenges: they are slow during inference and require e…
openalex
Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu 等
2024-09-19
置信度 0.72
Computer scienceAction (physics)Artificial intelligenceData manipulation languageComputer vision
-
Coot is a molecular-graphics application for model building and validation of biological macromolecules. The program displays electron-density maps and atomic models and allows model manipulations such as idealization, real-space refinement, manual rotation/tr…
openalex
Paul Emsley, Bernhard Lohkamp, W. G. Scott, Kevin Cowtan
2010-03-23
置信度 0.72
Computer scienceScripting languageSoftwareInterface (matter)Human–computer interaction
-
The rapid advancements in artificial intelligence (AI) have led to the development of sophisticated large language models (LLMs) such as GPT-4 and Bard. The potential implementation of LLMs in healthcare settings has already garnered considerable attention bec…
openalex
Bertalan Meskó, Eric J. Topol
2023-07-06
置信度 0.72
Transformative learningContext (archaeology)Health careHarmEngineering ethics
-
Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training c…
openalex
Yihao Wang, Pengxiang Ding, Lingxiao Li, Can Cui 等
2026-03-14
置信度 0.72
Bridging (networking)Computer scienceBridge (graph theory)Artificial intelligenceInference
-
openalex
Xinghang Li, Peiyan Li, Long Qian, Minghuan Liu 等
2026-02-11
置信度 0.72
Computer scienceRobotKey (lock)Foundation (evidence)Focus (optics)
-
openalex
Melvyn A. Goodale
2010-08-05
置信度 0.72
Vision for perception and vision for actionAction (physics)PerceptionPsychologyControl (management)
-
Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling. Such a trend underscored the potential of training LLMs with unlimited language data, …
openalex
Gengze Zhou, Yicong Hong, Qi Wu
2024-03-24
置信度 0.72
Computer scienceLinguisticsCognitive scienceArtificial intelligencePsychology
-
The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language understanding, and control within a single policy. Researchers in autonomous driving…
openalex
Shan Jiang, Zilin Huang, Kangan Qian, Ziang Luo 等
2025-06-30
置信度 0.72
Computer scienceTRACE (psycholinguistics)Control (management)Measure (data warehouse)Human–computer interaction
-
This paper formulates the action of psychedelics by integrating the free-energy principle and entropic brain hypothesis. We call this formulation relaxed beliefs under psychedelics (REBUS) and the anarchic brain, founded on the principle that—via their entropi…
openalex
Robin Carhart‐Harris, Karl Friston
2019-06-20
置信度 0.72
Prior probabilityAction (physics)PsychologyConsciousnessPsilocybin
-
Learning to navigate in a visual environment following natural-language instructions is a challenging task, because the multimodal inputs to the agent are highly variable, and the training data on a new task is often limited. In this paper, we present the firs…
openalex
Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin 等
2020-06-01
置信度 0.72
Computer scienceTask (project management)Artificial intelligenceNatural languageBenchmark (surveying)
-
Responsible innovation on large-scale Language Models (LMs) requires foresight into and in-depth understanding of the risks these models may pose. This paper develops a comprehensive taxonomy of ethical and social risks associated with LMs. We identify twenty-…
openalex
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin 等
2022-06-20
置信度 0.72
MisinformationTaxonomy (biology)HarmRisk analysis (engineering)Futures studies
-
In robotic, task goals can be conveyed through various modalities, such as language, goal images, and goal videos. However, natural language can be ambiguous, while images or videos may offer overly detailed specifications. To tackle these challenges, we intro…
openalex
Xiaoqi Li, Jingyun Xu, Mingxu Zhang, Jiaming Liu 等
2025-06-10
置信度 0.72
Computer scienceAction (physics)Artificial intelligenceObject (grammar)Computer vision
-
Recently, some studies have integrated Multimodal Large Language Models into robotic manipulation, constructing vision-language-action models (VLAs) to interpret multimodal information and predict SE(3) poses. While VLAs have shown promising progress, they may…
openalex
Chenxuan Li, Jiaming Liu, Guanqun Wang, Li, Xiaoqi 等
2024-05-27
置信度 0.72
End-to-end principleComputer scienceRobotArtificial intelligence
-
In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated impressive results for end-to-end robot control, it remains an open question how f…
openalex
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian 等
2025-04-22
置信度 0.72
Computer scienceGeneralizationRobotArtificial intelligenceObject (grammar)
-
Task-oriented grasping (TOG) aims to predict the appropriate pose for grasping based on a specific task. While recent approaches have incorporated semantic knowledge into TOG models to enable robots to understand linguistic commands, they lack the ability to l…
openalex
Jianwei Zhu, Xueying Sun, Qiang Zhang, Mingmin Liu
2025-05-06
置信度 0.72
GRASPComputational intelligenceModality (human–computer interaction)Action (physics)Task (project management)
-
openalex
Gérard Berry, Georges Gonthier
1992-11-01
置信度 0.72
Computer scienceProgramming languageSemantics (computer science)Asynchronous communicationOperational semantics
-
A decade of unprecedented progress in artificial intelligence (AI) has demonstrated the potential for many fields-including medicine-to benefit from the insights that AI techniques can extract from data. Here we survey recent progress in the development of mod…
openalex
Andre Esteva, Katherine Chou, Serena Yeung, Nikhil Naik 等
2021-01-08
置信度 0.72
Software deploymentWorkflowDeep learningConvolutional neural networkComputer science
-
Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking…
openalex
Wenxuan Song, Jiayi Chen, Pengxiang Ding, Han Zhao 等
2025-10-19
置信度 0.72
Decoding methodsChunking (psychology)Computer scienceAction (physics)Inference
-
The anatomy of language has been investigated with PET or fMRI for more than 20 years. Here I attempt to provide an overview of the brain areas associated with heard speech, speech production and reading. The conclusions of many hundreds of studies were consid…
openalex
Cathy J. Price
2012-05-12
置信度 0.72
Reading (process)Articulation (sociology)Speech productionComputer scienceSpoken language
-
Brains, it has recently been argued, are essentially prediction machines. They are bundles of cells that support perception and action by constantly attempting to match incoming sensory inputs with top-down expectations or predictions. This is achieved using a…
openalex
Andy Clark
2013-05-10
置信度 0.72
SituatedCognitive scienceCognitionPsychologyComputer science
-
Vision-Language Navigation (VLN) is a task where an agent learns to navigate following a natural language instruction. The key to this task is to perceive both the visual scene and natural language sequentially. Conventional approaches fully exploit vision and…
openalex
Fengda Zhu, Yi Zhu, Xiaojun Chang, Xiaodan Liang
2020-06-01
置信度 0.72
Computer scienceTask (project management)Benchmark (surveying)Natural languageArtificial intelligence
-
Large Language Models (LLMs) showcase impressive capabilities but encounter challenges like hallucination, outdated knowledge, and non-transparent, untraceable reasoning processes. Retrieval-Augmented Generation (RAG) has emerged as a promising solution by inc…
openalex
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia 等
2023-12-18
置信度 0.72
CredibilityComputer scienceModular designBenchmark (surveying)Data science
-
All this work could be conceivable
openalex
Zoltán Dörnyei
1998-07-01
置信度 0.72
Content (measure theory)PsychologyForeign languageMathematics educationCognitive psychology
-
openalex
Natalie Sebanz, H. Bekkering, Günther Knoblich
2006-01-11
置信度 0.72
Action (physics)CognitionPsychologyContext (archaeology)Perception
-
This paper introduces Shake-VLA, a Vision-Language-Action (VLA) model-based system designed to enable bimanual robotic manipulation for automated cocktail preparation. The system integrates a vision module for detecting ingredient bottles and reading labels, a…
openalex
Mohammad Salman Khan, Selamawit Asfaw, Dmitrii Iarchuk, Miguel Altamirano Cabrera 等
2025-03-04
置信度 0.72
ShakeMixing (physics)Action (physics)Computer scienceArtificial intelligence
-
openalex
Arthur C. Graesser, Danielle S. McNamara, Max M. Louwerse, Zhiqiang Cai
2004-05-01
置信度 0.72
Cohesion (chemistry)Computer scienceReadabilityNatural language processingArtificial intelligence
-
Abstract Foundation Vision Language Models (VLMs) exhibit strong capabilities in multi-modal representation learning, comprehension, and reasoning. By injecting action components into the VLMs, Vision-Language-Action models (VLAs) can be naturally formed and a…
openalex
Huaping Liu, Xinghang Li, Peiyan Li, Minghuan Liu 等
2025-02-25
置信度 0.72
Generalist and specialist speciesAction (physics)RobotComputer scienceArtificial intelligence
-
In endoscopic procedures, autonomous tracking of abnormal regions and following circumferential cutting markers can significantly reduce the cognitive burden on endoscopists. However, conventional model-based pipelines are fragile for each component (e.g., det…
openalex
Chi Kit Ng, Long Bai, Guankun Wang, Yupeng Wang 等
2025-05-21
置信度 0.72
Artificial intelligenceComputer visionComputer scienceGeneralizationComponent (thermodynamics)
-
Vision-language-action (VLA) models hold promise as generalist robotics solutions by translating visual and linguistic inputs into robot actions, yet they lack reliability due to their black-box nature and sensitivity to environmental changes. In contrast, cog…
openalex
Hong Lu, Hengxu Li, Prithviraj Singh Shahani, Stephanie Herbers 等
2025-06-24
置信度 0.72
Computer scienceArchitectureAction (physics)Cognitive architectureCognition
-
Grasping tasks are crucial for the interaction between the robot and environment, requiring the robot to make autonomous decisions based on environmental conditions and instructions. Recently, multimodal Vision Language Models (VLMs) have demonstrated signific…
openalex
Lingling Fan, Kang Chen, Zhezhuang Xu, Meng Yuan 等
2024-11-01
置信度 0.72
Computer scienceAction (physics)Artificial intelligenceLanguage modelHuman–computer interaction
-
Building on the advancements of Large Language Models (LLMs) and Vision Language Models (VLMs), recent research has introduced Vision-Language-Action (VLA) models as an integrated solution for robotic manipulation tasks. These models take camera images and nat…
openalex
Zhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang 等
2024-10-07
置信度 0.72
Action (physics)Computer scienceHuman–computer interactionArtificial intelligenceQuantum mechanics
-
Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object interactions. We adapt vision-language models (VLMs) to this challenging domain with …
arxiv
Hao Zheng, Jinyi Huang, Tiantian Zheng, Xun Xu 等
2026-07-12T15:09:33Z
置信度 0.78
cs.CV
-
The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks. However, existing approaches often treat vision encoders and large language models (LLMs) as independent modules, lim…
arxiv
Eugene Lee, Ting-Yu Chang, Jui-Huang Tsai, Jiajie Diao 等
2026-03-31T18:00:39Z
置信度 0.78
cs.CVcs.AIcs.CLcs.LG
-
In the past few years, the emergence of pre-training models has brought uni-modal fields such as computer vision (CV) and natural language processing (NLP) to a new era. Substantial works have shown they are beneficial for downstream uni-modal tasks and avoid …
arxiv
Feilong Chen, Duzhen Zhang, Minglun Han, Xiuyi Chen 等
2022-02-18T07:54:02Z
置信度 0.78
cs.CVcs.CL
-
Vision-Language Model (VLM) have gained widespread adoption in Open-Vocabulary (OV) object detection and segmentation tasks. Despite they have shown promise on OV-related tasks, their effectiveness in conventional vision tasks has thus far been unevaluated. In…
arxiv
Yongchao Feng, Yajie Liu, Shuai Yang, Wenrui Cai 等
2025-04-13T08:28:13Z
置信度 0.78
cs.CVcs.AI
-
The fusion of language and vision in large vision-language models (LVLMs) has revolutionized deep learning-based object detection by enhancing adaptability, contextual reasoning, and generalization beyond traditional architectures. This in-depth review present…
arxiv
Ranjan Sapkota, Manoj Karkee
2025-08-25T17:21:00Z
置信度 0.78
cs.CVcs.AIcs.CL
-
Building state-of-the-art Vision-Language Models (VLMs) with strong captioning capabilities typically necessitates training on billions of high-quality image-text pairs, requiring millions of GPU hours. This paper introduces the Vision-Language-Vision (VLV) au…
arxiv
Tiezheng Zhang, Yitong Li, Yu-cheng Chou, Jieneng Chen 等
2025-07-09T17:59:04Z
置信度 0.78
cs.CV
-
Vision-Language-Action (VLA) models have demonstrated strong potential for predicting semantic actions in navigation tasks, demonstrating the ability to reason over complex linguistic instructions and visual contexts. However, they are fundamentally hindered b…
arxiv
Jaehwan Jeong, Evelyn Zhu, Jinying Lin, Emmanuel Jaimes 等
2026-03-14T06:26:11Z
置信度 0.78
cs.ROcs.CV
-
This paper surveys vision-language pre-training (VLP) methods for multimodal intelligence that have been developed in the last few years. We group these approaches into three categories: ($i$) VLP for image-text tasks, such as image captioning, image-text retr…
arxiv
Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang 等
2022-10-17T17:11:36Z
置信度 0.78
cs.CVcs.CL
-
Vision-language-action (VLA) models enable impressive zero shot manipulation, but their inference stacks are often too heavy for responsive web demos or high frequency robot control on commodity GPUs. We present BLURR, a lightweight inference wrapper that can …
arxiv
Xiaoyu Ma, Zhengqing Yuan, Zheyuan Zhang, Kaiwen Shi 等
2025-12-12T18:30:45Z
置信度 0.78
cs.RO
-
Mechanistic interpretability seeks to understand the neural mechanisms that enable specific behaviors in Large Language Models (LLMs) by leveraging causality-based methods. While these approaches have identified neural circuits that copy spans of text, capture…
arxiv
Vedant Palit, Rohan Pandey, Aryaman Arora, Paul Pu Liang
2023-08-27T18:46:47Z
置信度 0.78
cs.CLcs.AIcs.CV
-
Artificial Intelligence makes great advances today and starts to bridge the gap between vision and language. However, we are still far from understanding, explaining and controlling explicitly the visual content from a linguistic perspective, because we still …
arxiv
Mihai Masala, Nicolae Cudlenco, Traian Rebedea, Marius Leordeanu
2023-08-29T07:25:06Z
置信度 0.78
cs.AIcs.CLcs.CV
-
Existing vision-based action recognition is susceptible to occlusion and appearance variations, while wearable sensors can alleviate these challenges by capturing human motion with one-dimensional time-series signal. For the same action, the knowledge learned …
arxiv
Yang Liu, Keze Wang, Guanbin Li, Liang Lin
2020-09-01T03:38:31Z
置信度 0.78
cs.CV
-
In this paper, we introduce an open-source Korean-English vision-language model (VLM), VARCO-VISION. We incorporate a step-by-step training strategy that allows a model learn both linguistic and visual information while preserving the backbone model's knowledg…
arxiv
Jeongho Ju, Daeyoung Kim, SunYoung Park, Youngjune Kim
2024-11-28T12:38:42Z
置信度 0.78
cs.CVcs.CL
-
Large language models accumulate vast knowledge during pre-training, yet the dynamics governing this acquisition remain poorly understood. This work investigates the learning dynamics of language models on a synthetic factual recall task, uncovering three key …
arxiv
Nicolas Zucchet, Jörg Bornschein, Stephanie Chan, Andrew Lampinen 等
2025-03-27T16:43:45Z
置信度 0.78
cs.CLcs.LG
-
Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly reflect reliance on visual evidence. Motivated by a surprising observation that re…
arxiv
Zixuan Lan, Luzhe Sun, Matthew R. Walter, Jiawei Zhou
2026-05-21T17:35:04Z
置信度 0.78
cs.CVcs.AIcs.CL
-
Why do pretrained diffusion or flow-matching policies fail when the same task is performed near an obstacle, on a shifted support surface, or amid mild clutter? Such failures rarely reflect missing motor skills; instead, they expose a limitation of imitation l…
arxiv
Shuo Liu, Ishneet Sukhvinder Singh, Yiqing Xu, Jiafei Duan 等
2026-02-03T19:50:16Z
置信度 0.78
cs.ROcs.CV
-
Similar to LLMs, the development of vision language models is mainly driven by English datasets and models trained in English and Chinese language, whereas support for other languages, even those considered high-resource languages such as German, remains signi…
arxiv
René Peinl, Vincent Tischler
2025-04-15T11:55:24Z
置信度 0.78
cs.CL
-
Medical Vision-Language Models (Med-VLMs) achieve strong expert-level performance, yet their ability to generate patient-accessible descriptions remains underexplored. With the 21st Century Cures Act now mandating immediate patient access to diagnostic imaging…
arxiv
Han Jang, Junhyeok Lee, Songsoo Kim, Chae Young Lim 等
2026-06-19T08:06:09Z
置信度 0.78
cs.CVcs.AIcs.CL
-
The emergence of new wearable technologies such as action cameras and smart-glasses has increased the interest of computer vision scientists in the First Person perspective. Nowadays, this field is attracting attention and investments of companies aiming to de…
arxiv
Alejandro Betancourt, Pietro Morerio, Carlo S. Regazzoni, Matthias Rauterberg
2014-09-04T16:38:43Z
置信度 0.78
cs.CV
-
Vision-Language Models (VLMs) and Multi-Modal Language models (MMLMs) have become prominent in autonomous driving research, as these models can provide interpretable textual reasoning and responses for end-to-end autonomous driving safety tasks using traffic s…
arxiv
Akshay Gopalkrishnan, Ross Greer, Mohan Trivedi
2024-03-28T21:18:33Z
置信度 0.78
cs.CVcs.AI
-
Vision-language models, which integrate computer vision and natural language processing capabilities, have demonstrated significant advancements in tasks such as image captioning and visual question and answering. However, similar to traditional models, they a…
arxiv
Ashwin Ramesh Babu, Sajad Mousavi, Vineet Gundecha, Sahand Ghorbanpour 等
2025-06-05T08:09:05Z
置信度 0.78
cs.CVcs.AIcs.CLcs.LG
-
Large vision-language models have achieved outstanding performance, but their size and computational requirements make their deployment on resource-constrained devices and time-sensitive tasks impractical. Model distillation, the process of creating smaller, f…
arxiv
Xuanlin Li, Yunhao Fang, Minghua Liu, Zhan Ling 等
2023-07-06T17:05:26Z
置信度 0.78
cs.CVcs.AIcs.CLcs.LG
-
Despite their promise to perform complex reasoning, large language models (LLMs) have been shown to have limited effectiveness in end-to-end planning. This has inspired an intriguing question: if these models cannot plan well, can they still contribute to the …
arxiv
Mohamed Aghzal, Xiang Yue, Erion Plaku, Ziyu Yao
2024-11-27T19:32:03Z
置信度 0.78
cs.CVcs.CL
-
Generalist robot policies must follow user instructions while reasoning about how objects, cameras, and robot actions interact in the 3D physical world. Recent vision-language-action models (VLAs) and video world-action models (WAMs) inherit strong semantic or…
arxiv
Jisang Han, Seonghu Jeon, Jaewoo Jung, René Zurbrügg 等
2026-06-15T17:58:03Z
置信度 0.78
cs.ROcs.CVcs.LG
-
Vision-Language-Action (VLA) models have shown promising capabilities for embodied intelligence, but most existing approaches rely on text-based chain-of-thought reasoning where visual inputs are treated as static context. This limits the ability of the model …
arxiv
Chaoyang Wang, Wenrui Bao, Sicheng Gao, Bingxin Xu 等
2026-03-15T17:59:51Z
置信度 0.78
cs.CVcs.AIcs.RO
-
Vision-Language-Action models have emerged as essential generalist robot policies for diverse manipulation tasks, conventionally relying on directly translating multimodal inputs into actions via Vision-Language Model embeddings. Recent advancements have intro…
arxiv
Linqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong 等
2026-01-16T16:17:06Z
置信度 0.78
cs.RO
-
Event-driven reactive functionalities are an urgent need in nowadays distributed service-oriented applications and (Semantic) Web-based environments. An important problem to be addressed is how to correctly and efficiently capture and process the event-based b…
arxiv
Adrian Paschke
2006-09-26T14:36:47Z
置信度 0.78
cs.AIcs.LOcs.SE
-
We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models. While advances in VLAs have introduced robot policies that are both generalizable and semantically grounded, these model…
arxiv
Charlotte Morissette, Amin Abyaneh, Wei-Di Chang, Anas Houssaini 等
2026-03-15T20:57:51Z
置信度 0.78
cs.ROcs.CVcs.LG
-
In this paper we present the Women in Computer Vision Workshop - WiCV 2019, organized in conjunction with CVPR 2019. This event is meant for increasing the visibility and inclusion of women researchers in the computer vision field. Computer vision and machine …
arxiv
Irene Amerini, Elena Balashova, Sayna Ebrahimi, Kathryn Leonard 等
2019-09-23T08:52:33Z
置信度 0.78
cs.CV
-
Contrastive language-image pretraining (CLIP) links vision and language modalities into a unified embedding space, yielding the tremendous potential for vision-language (VL) tasks. While early concurrent works have begun to study this potential on a subset of …
arxiv
Zhecan Wang, Noel Codella, Yen-Chun Chen, Luowei Zhou 等
2022-01-15T01:54:01Z
置信度 0.78
cs.CVcs.AIcs.CLcs.LGcs.MM
-
Medical vision-and-language pre-training (Med-VLP) has received considerable attention owing to its applicability to extracting generic vision-and-language representations from medical images and texts. Most existing methods mainly contain three elements: uni-…
arxiv
Zhihong Chen, Guanbin Li, Xiang Wan
2022-09-15T08:00:01Z
置信度 0.78
cs.CLcs.CV
-
Accurately identifying, understanding and describing traffic safety-critical events (SCEs), including crashes, tire strikes, and near-crashes, is crucial for advanced driver assistance systems, automated driving systems, and traffic safety. As SCEs are rare ev…
arxiv
Liang Shi, Boyu Jiang, Tong Zeng, Feng Guo
2024-10-01T18:10:23Z
置信度 0.78
cs.CV
-
Vision-and-Language Navigation (VLN) requires an agent to find a path to a remote location on the basis of natural-language instructions and a set of photo-realistic panoramas. Most existing methods take the words in the instructions and the discrete views of …
arxiv
Yuankai Qi, Zizheng Pan, Yicong Hong, Ming-Hsuan Yang 等
2021-04-09T02:44:39Z
置信度 0.78
cs.CLcs.CV
-
A minimalist vision system uses the smallest number of pixels needed to solve a vision task. While traditional cameras use a large grid of square pixels, a minimalist camera uses freeform pixels that can take on arbitrary shapes to increase their information c…
arxiv
Jeremy Klotz, Shree K. Nayar
2024-12-30T21:27:07Z
置信度 0.78
cs.CVeess.IV
-
The advent of Large Language Models (LLMs) has significantly reshaped the trajectory of the AI revolution. Nevertheless, these LLMs exhibit a notable limitation, as they are primarily adept at processing textual information. To address this constraint, researc…
arxiv
Akash Ghosh, Arkadeep Acharya, Sriparna Saha, Vinija Jain 等
2024-02-20T18:57:34Z
置信度 0.78
cs.CVcs.AIcs.CL
-
We introduce Granite Vision, a lightweight large language model with vision capabilities, specifically designed to excel in enterprise use cases, particularly in visual document understanding. Our model is trained on a comprehensive instruction-following datas…
arxiv
Granite Vision Team, Leonid Karlinsky, Assaf Arbelle, Abraham Daniels 等
2025-02-14T05:36:32Z
置信度 0.78
cs.CVcs.AI
-
We present an analysis of the performance of machine learning classifiers on discriminating between similar languages and language varieties. We carried out a number of experiments using the results of the two editions of the Discriminating between Similar Lan…
arxiv
Cyril Goutte, Serge Léger, Shervin Malmasi, Marcos Zampieri
2016-09-30T20:57:52Z
置信度 0.78
cs.CL
-
Large multimodal language models (LMMs) have achieved significant success in general domains. However, due to the significant differences between medical images and text and general web content, the performance of LMMs in medical scenarios is limited. In ophth…
arxiv
Weihao Gao, Zhuo Deng, Zhiyuan Niu, Fuju Rong 等
2023-06-21T11:09:48Z
置信度 0.78
cs.CV
-
To reduce the inference cost of large language models, model compression is increasingly used to create smaller scalable models. However, little is known about their robustness to minority subgroups defined by the labels and attributes of a dataset. In this pa…
arxiv
Leonidas Gee, Andrea Zugarini, Novi Quadrianto
2024-03-26T15:50:37Z
置信度 0.78
cs.LGcs.CL
-
Vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they often fail on tasks that require fine-grained visual perception, even when the required information is still present in their internal rep…
arxiv
Haz Sameen Shahgir, Xiaofu Chen, Yu Fu, Erfan Shayegani 等
2026-04-02T19:40:56Z
置信度 0.78
cs.CVcs.CL
-
This paper presents a number of experiments to model changes in a historical Portuguese corpus composed of literary texts for the purpose of temporal text classification. Algorithms were trained to classify texts with respect to their publication date taking i…
arxiv
Marcos Zampieri, Shervin Malmasi, Mark Dras
2016-09-30T20:57:01Z
置信度 0.78
cs.CL
-
When language models (LMs) are trained to forget (or "unlearn'') a skill, how precisely does their behavior change? We study the behavior of transformer LMs in which tasks have been forgotten via fine-tuning on randomized labels. Such LMs learn to generate nea…
arxiv
Eric Zhang, Leshem Choshen, Jacob Andreas
2024-09-03T18:55:54Z
置信度 0.78
cs.LGcs.CL
-
Building on recent advances in language-based reasoning models, we explore multimodal reasoning that integrates vision and text. Existing multimodal benchmarks primarily test visual extraction combined with text-based reasoning, lacking true visual reasoning w…
arxiv
Mert Unsal, Aylin Akkus
2025-06-13T09:03:33Z
置信度 0.78
cs.CVcs.LG
-
Medical vision-and-language pre-training provides a feasible solution to extract effective vision-and-language representations from medical images and texts. However, few studies have been dedicated to this field to facilitate medical vision-and-language under…
arxiv
Zhihong Chen, Yuhao Du, Jinpeng Hu, Yang Liu 等
2022-09-15T07:26:43Z
置信度 0.78
cs.CVcs.CL
-
We present VoxelPrompt, an end-to-end image analysis agent that tackles free-form radiological tasks. Given any number of volumetric medical images and a natural language prompt, VoxelPrompt integrates a language model that generates executable code to invoke …
arxiv
Andrew Hoopes, Neel Dey, Victor Ion Butoi, John V. Guttag 等
2024-10-10T22:11:43Z
置信度 0.78
eess.IVcs.AIcs.CV
-
Vision-Language-Action (VLA) models have shown promise in robot manipulation but often struggle to generalize to new instructions or complex multi-task scenarios. We identify a critical pathology in current training paradigms where goal-driven data collection …
arxiv
Shijie Lian, Bin Yu, Xiaopeng Lin, Laurence T. Yang 等
2026-01-21T17:15:22Z
置信度 0.78
cs.AIcs.CLcs.CVcs.RO
-
Large Vision-Language Models (LVLMs) offer remarkable benefits for a variety of vision-language tasks. However, a challenge hindering their application in real-world scenarios, particularly regarding safety, robustness, and reliability, is their constrained se…
arxiv
Jiaying Lu, Jinmeng Rao, Kezhen Chen, Xiaoyuan Guo 等
2023-09-07T22:59:56Z
置信度 0.78
cs.CVcs.CL
-
Large Vision-Language Models (LVLMs) evolve rapidly as Large Language Models (LLMs) was equipped with vision modules to create more human-like models. However, we should carefully evaluate their applications in different domains, as they may possess undesired …
arxiv
Yuhang Xiao, Yudi Lin, Ming-Chang Chiu
2024-09-23T17:54:47Z
置信度 0.78
cs.CLcs.AI
-
The Stellar Imager (SI) is a UV-Optical, Space-Based Interferometer designed to enable 0.1 milli-arcsecond (mas) spectral imaging of stellar surfaces and of the Universe in general and asteroseismic imaging of stellar interiors. SI is identified as a "Flagship…
arxiv
Kenneth G. Carpenter, Carolus J. Schrijver, Margarita Karovska, SI Vision Mission Team
2006-06-16T19:05:51Z
置信度 0.78
astro-ph
-
Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplored. We introduce VLG, a vision-language-garment model that synthesizes garments f…
arxiv
Jan Ackermann, Kiyohiro Nakayama, Guandao Yang, Tong Wu 等
2025-06-05T16:22:17Z
置信度 0.78
cs.CV