-
Autonomous driving in complex real-world environments requires robust perception, reasoning, and physically feasible planning, which remain challenging for current end-to-end approaches. This paper introduces VLA-MP, a unified vision-language-action framework …
europepmc
Maoning Ge, Kento Ohtani, Yingjie Niu, Yuxiao Zhang 等
2025
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
Considering how to make the model accurately understand and follow natural language instructions and perform actions consistent with world knowledge is a key challenge in robot manipulation. This mainly includes human fuzzy instruction reasoning and the follow…
pubmed
Ren P, Zhang K, Zheng H, Li Z 等
2025
置信度 0.82
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
With the rapid growth of electronic health records, medical imaging, and high-throughput omics data, precision oncology faces increasing demands for cross-modal information integration and complex clinical decision support. In recent years, large language mode…
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
The expanding role of intelligent systems in biomedical science marks a shift from passive analysis towards active participation in discovery and care. Recent scholarship has begun to frame these systems not merely as models, but as agents capable of planning,…
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2025
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
With the rapid advancement of artificial intelligence and robotics, the integration of Large Language Models (LLMs) with 3D vision is emerging as a transformative approach to enhancing robotic sensing technologies. This convergence enables machines to perceive…
pubmed
Mehta V, Sharma C, Thiyagarajan K
2025
置信度 0.82
-
europepmc
2026
置信度 0.80
-
europepmc
2025
置信度 0.80
-
europepmc
2026
置信度 0.80
-
This systematic literature review investigates the integration of deep learning (DL), vision-language models (VLMs), and multiagent systems in the analysis of pathology images and automated report generation. The rapid advancement of whole-slide imaging (WSI) …
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2025
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
General intelligence enables flexible problem solving across diverse contexts by minimizing uncertainty. Symbolic systems such as language extend this capacity, allowing humans to build social groups and construct world models beyond typical biological constra…
europepmc
2026
置信度 0.80
-
europepmc
2026
置信度 0.80
-
europepmc
2025
置信度 0.80
-
europepmc
2026
置信度 0.80
-
The digital transformation of rehabilitation training has become a public health imperative driven by a global demand that outstrips professional medical resources and is compounded by a deficit in public rehabilitation literacy. As traditional hospital-centri…
pubmed
Ye P, Li Y, Qu M, Liu J 等
2026
置信度 0.82
-
Analyzing hand-object interaction in egocentric vision facilitates VR/AR applications and human-robot policy transfer. Existing research has mostly focused on modeling the behavior paradigm of interactive actions (i.e., "how to interact"). However, the more ch…
pubmed
Ma J, Zhang E, Zheng YD, Xie Y 等
2026 Jul 29
置信度 0.82
-
Cross-domain few-shot facial expression recognition (CF-FER) aims to adapt models trained on basic expressions to recognize novel compound expressions using only a few annotated examples. Although vision-language models (VLMs) have shown promise in few-shot le…
pubmed
Wang K, Ding R, Wang H, Yan Y
2026
置信度 0.82
-
Offline action prediction does not by itself guarantee reliable closed-loop flight for unmanned aerial vehicle (UAV) vision-language-action (VLA) models. We study a controlled AirSim Blocks waypoint task using 100 expert episodes and 3385 RGB-D, instruction, s…
pubmed
Xu Y, Lin H, Yang Y, Xia J 等
2026 Jul 22
置信度 0.82
-
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlik…
pubmed
Phan DT, Nguyen HD
2026 Jul 17
置信度 0.82
-
Artificial intelligence (AI) is rapidly reshaping orthopaedic surgery, supported by advances in data science, computational power, and perioperative digitalization. Within this evolving landscape, five "AI companions" structure the surgeon's workflow. The "AI …
pubmed
Gauci MO, Duval G, Rony L
2026 Jul 16
置信度 0.82
-
Facial movements can be key indicators of neurological health and emotional state, offering insights into motor and neuropsychiatric functions that are disrupted in neurologic disorders. Neurological disease can present with characteristic differences in facia…
pubmed
Nylander A, Poole S, Hsu NS, Henderson K 等
2025 Dec
置信度 0.82
-
Developing generalizable robotic policies that balance inference efficiency, manipulation accuracy, and robustness remains a formidable challenge. Existing Vision-Language-Action models demand prohibitive data scales, while keyframe-based approaches struggle t…
pubmed
Wang S, Wang L, Huo H, Zhou S 等
2026 Jul 13
置信度 0.82
-
Pathology vision-language models (VLMs) are promising for building the clinical decision support systems. However, a key barrier to real-world clinical deployment lies in the lack of rigorous and clinically meaningful model evaluation. Existing pathology visua…
pubmed
Chen K, Wei L, Rui S, Yuan Y 等
2026 Sep
置信度 0.82
-
Tuberculosis (TB) is the leading global cause of death from a single infectious agent. Recent reductions in global health funding have threatened TB control, making comprehensive assessment of TB, HIV-related TB, and drug-resistant TB burdens before these disr…
pubmed
GBD 2023 TB HIV Collaborators
2026 Jul 1
置信度 0.82
-
Embodied AI is widely recognized as a cornerstone of artificial general intelligence (AGI) because it involves controlling embodied agents to perform tasks in the physical world. Building on the success of large language models (LLMs) and vision-language model…
pubmed
Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao 等
2026
置信度 0.82
Embodied cognitionAction (physics)Computer scienceCognitive scienceArtificial intelligence
-
We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both lear…
openalex
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar 等
2023-07-28
置信度 0.72
Computer scienceNatural languageArtificial intelligenceRobotGeneralization
-
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is needed to specify any other visual concep…
openalex
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等
2021-02-26
置信度 0.72
Computer scienceArtificial intelligenceGeneralityTransfer of learningTask (project management)
-
Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes. However, current VLA models face significant challenges: they are slow during inference and require e…
openalex
Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu 等
2025-02-24
置信度 0.72
Action (physics)Computer scienceArtificial intelligenceHuman–computer interactionComputer vision
-
Large policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can fine-tune such vision-langua…
openalex
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao 等
2024-06-13
置信度 0.72
Open sourceAction (physics)Computer scienceArtificial intelligenceProgramming language
-
Despite progress in perceptual tasks such as image classification, computers still perform poorly on cognitive tasks such as image description and question answering. Cognition is core to tasks that involve not just recognizing, but reasoning about our visual …
openalex
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson 等
2017-02-06
置信度 0.72
Artificial intelligenceComputer scienceNatural language processingGenomeImage (mathematics)
-
Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behaviors.However, such models require us to choose a tokenization of our continuous …
openalex
Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess 等
2025-06-21
置信度 0.72
Computer scienceAction (physics)Artificial intelligenceNatural language processingFeature (linguistics)
-
Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robo…
openalex
Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu 等
2025-06-10
置信度 0.72
Computer scienceAction (physics)Cognitive scienceChain (unit)Artificial intelligence
-
Amid growing efforts to leverage advances in large language models (LLMs) and visionlanguage models (VLMs) for robotics, Vision-Language-Action (VLA) models have recently gained significant attention. By unifying vision, language, and action data at scale, whi…
openalex
Kento Kawaharazuka, Jihoon Oh, Jun Yamada, Ingmar Posner 等
2025-01-01
置信度 0.72
Computer scienceLeverage (statistics)Software deploymentScalabilityRobot
-
openalex
Jayavardhana Gubbi, Rajkumar Buyya, Slaven Marusic, Marimuthu Palaniswami
2013-02-24
置信度 0.72
Computer scienceInternet of ThingsThe InternetMultimediaWorld Wide Web
-
openalex
Arthur M. Glenberg, Michael P. Kaschak
2002-09-01
置信度 0.72
SentenceEmbodied cognitionComprehensionPsychologyMeaning (existential)
-
openalex
Kevin Black, Noah Brown, Danny Driess, A. Esmail 等
2025-06-21
置信度 0.72
Computer scienceControl theory (sociology)RobotControl engineeringControl (management)
-
Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned…
openalex
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida 等
2022-03-04
置信度 0.72
Computer scienceLanguage modelSet (abstract data type)Simple (philosophy)Reinforcement learning
-
In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) c…
openalex
Karen Simonyan, Andrew Zisserman
2014-09-04
置信度 0.72
Computer scienceConvolution (computer science)Convolutional neural networkArtificial intelligenceDeep learning
-
openalex
Pengxiang Ding, Han Zhao, Wenjie Zhang, Wenxuan Song 等
2024-10-29
置信度 0.72
Computer scienceRobotAction (physics)Artificial intelligenceComputer vision
-
NVIDIA https://navila-bot.github.ioMove forward out of the room.Turn right at the end.Proceed to the grass and stop in front of the soccers.Walk forward along the way.Turn a little left and keep going straight.Stop in front of the red door.Turn left immediatel…
openalex
An‐Chieh Cheng, Yandong Ji, Zhaojing Yang, Zaitian Gongye 等
2025-06-21
置信度 0.72
Legged robotComputer scienceRobotArtificial intelligenceComputer vision
-
Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to ad…
openalex
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi 等
2023-07-12
置信度 0.72
Computer scienceBenchmark (surveying)Language modelComprehensionArtificial intelligence
-
openalex
Shunlei Li, Longsen Gao, Jin Wang, Chang Che 等
2026-01-29
置信度 0.72
Computer scienceArtificial intelligenceVisual reasoningRobotTransformer
-
This research introduces the Bi-VLA (Vision-Language-Action) model, a novel system designed for bimanual robotic dexterous manipulation that seamlessly integrates vision for scene understanding, language comprehension for translating human instructions into ex…
openalex
Koffivi Fidèle Gbagbe, Miguel Altamirano Cabrera, Ali Alabbas, Oussama Alyunes 等
2024-10-06
置信度 0.72
Computer scienceAction (physics)Computer visionArtificial intelligenceHuman–computer interaction
-
The job demands-resources (JD-R) model proposes that working conditions can be categorized into 2 broad categories, job demands and job resources. that are differentially related to specific outcomes. A series of LISREL analyses using self-reports as well as o…
openalex
Evangelia Demerouti, Arnold B. Bakker, Friedhelm Nachreiner, Wilmar B. Schaufeli
2001-06-01
置信度 0.72
LISRELDisengagement theoryBurnoutPsychologyOccupational burnout
-
We propose VisualBERT, a simple and flexible framework for modeling a broad range of vision-and-language tasks. VisualBERT consists of a stack of Transformer layers that implicitly align elements of an input text and regions in an associated input image with s…
openalex
Liunian Harold Li, Mark Yatskar, Da Yin, Cho‐Jui Hsieh 等
2019-08-09
置信度 0.72
Computer scienceTransformerImage (mathematics)Language understandingBaseline (sea)
-
Building Graphical User Interface (GUI) assistants holds significant promise for enhancing human workflow productivity. While most agents are language-based, relying on closed-source API with text-rich meta-information (e.g., HTML or accessibility tree), they …
openalex
Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang 等
2025-06-10
置信度 0.72
Computer scienceAction (physics)Artificial intelligenceHuman–computer interactionProgramming language
-
Recent studies have successfully integrated large vision-language models (VLMs) into low-level robotic control by supervised fine-tuning (SFT) with expert robotic datasets, resulting in what we term vision-language-action (VLA) models. Although the VLA models …
openalex
Yanjiang Guo, Jianke Zhang, Xiaoyu Chen, Xiang Ji 等
2025-05-19
置信度 0.72
Reinforcement learningComputer scienceAction (physics)Artificial intelligenceHuman–computer interaction
-
The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language understanding, and control within a single policy. Researchers in autonomous driving…
openalex
Sicong Jiang, Menglin Kong, Yihong Tang, Lijun Sun 等
2025-10-19
置信度 0.72
Computer scienceTRACE (psycholinguistics)Control (management)Measure (data warehouse)Human–computer interaction
-
Mobile manipulation is the fundamental challenge for robotics to assist humans with diverse tasks and environments in everyday life. However, conventional mobile manipulation approaches often struggle to generalize across different tasks and environments becau…
openalex
Zhenyu Wu, Yuheng Zhou, Xiuwei Xu, Ziwei Wang 等
2025-06-10
置信度 0.72
Computer scienceAction (physics)Human–computer interactionArtificial intelligenceComputer vision
-
Vision-Language-Action (VLA) models excel at robotic tasks by leveraging large-scale 2D vision-language pretraining, but their reliance on RGB images limits spatial reasoning critical for real-world interaction. Retraining these models with 3D data is computat…
openalex
Chengmeng Li, Junjie Wen, Yaxin Peng, Yan Peng 等
2026-01-12
置信度 0.72
Computer sciencePoint cloudArtificial intelligenceKey (lock)Table (database)
-
openalex
Daniele Miorandi, Sabrina Sicari, Francesco De Pellegrini, Imrich Chlamtac
2012-04-21
置信度 0.72
Software deploymentThe InternetRealmComputer scienceInternet of Things
-
Goal-conditioned policies for robotic navigation can be trained on large, unannotated datasets, providing for good generalization to real-world settings. However, particularly in vision-based settings where specifying goals requires an image, this makes for an…
openalex
Dhruv Shah, Błażej Osiński, Brian Ichter, Sergey Levine
2022-07-10
置信度 0.72
Computer scienceGeneralizationInterface (matter)Modality (human–computer interaction)Code (set theory)
-
A practical navigation agent must be capable of handling a wide range of interaction demands, such as following instructions, searching objects, answering questions, tracking people, and more. Existing models for embodied navigation fall short of serving as pr…
openalex
Jiazhao Zhang, Kunyu Wang, Shaoan Wang, Minghan Li 等
2025-06-21
置信度 0.72
Computer scienceEmbodied cognitionArtificial intelligenceHuman–computer interactionField (mathematics)
-
Abstract The rapid evolution of large language models (LLMs) has driven a transformative shift in artificial intelligence (AI), reshaping both research paradigms and practical applications. Distinguished from their predecessors by unprecedented scale and advan…
openalex
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang 等
2026-05-09
置信度 0.72
Language modelComputer scienceMainstreamScale (ratio)Artificial intelligence
-
The rapid advancement of generative AI and multi-modal foundation models has shown significant potential in advancing robotic manipulation. Vision-language-action (VLA) models, in particular, have emerged as a promising approach for visuomotor control by lever…
openalex
Zhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang 等
2025-06-19
置信度 0.72
Computer scienceRobustness (evolution)Artificial intelligenceSoftware deploymentMachine learning
-
We focus on the task of language-conditioned grasping in clutter, in which a robot is supposed to grasp the target object based on a language instruction. Previous works separately conduct visual grounding to localize the target object, and generate a grasp fo…
openalex
Kechun Xu, Shuqi Zhao, Zhongxiang Zhou, Zizhang Li 等
2023-05-29
置信度 0.72
Computer scienceGRASPClutterArtificial intelligenceObject (grammar)
-
A fundamental objective in robot manipulation is to enable models to comprehend visual scenes and execute actions. Although existing Vision-Language-Action (VLA) models for robots can handle a range of basic tasks, they still face challenges in two areas: (1) …
openalex
Jiaming Liu, Mengzhen Liu, Zhenyu Wang, Pengju An 等
2024
置信度 0.72
State spaceSpace (punctuation)RobotComputer scienceState (computer science)
-
Zhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen, Ning Liu, Zhiyuan Xu, Weibin Meng, Yaxin Peng, Chaomin Shen, Feifei Feng, Yi Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
openalex
Zhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen 等
2025-01-01
置信度 0.72
Computer scienceArtificial intelligenceControl (management)RobotNatural language
-
I created a basic proof of concept of LMVM (Language Model Virtual Machine), a command line toolchain that English instructions into native binaries without intermediary high-level language compilation. The command line uses a large language model (LLM) as IR …
openalex
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan 等
2021-07-07
置信度 0.72
CorrectnessPython (programming language)Computer scienceCode (set theory)Programming language
-
We present Unified-IO 2,the. first autoregressive multi-modal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we tokenize inputs and outputs - images, text, audio, action, bounding boxes etc.…
openalex
Jiasen Lu, Christopher M. Clark, Sang-Ho Lee, Zichen Zhang 等
2024-06-16
置信度 0.72
Computer scienceAutoregressive modelAction (physics)Speech recognitionAudio visual
-
Humans can naturally and effectively find salient regions in complex scenes. Motivated by this observation, attention mechanisms were introduced into computer vision with the aim of imitating this aspect of the human visual system. Such an attention mechanism …
openalex
Meng-Hao Guo, Tian-Xing Xu, Jiangjiang Liu, Zheng-Ning Liu 等
2022-03-15
置信度 0.72
Computer scienceCategorizationArtificial intelligenceComputer graphicsProcess (computing)
-
Multilayer neural networks trained with the back-propagation algorithm constitute the best example of a successful gradient based learning technique. Given an appropriate network architecture, gradient-based learning algorithms can be used to synthesize a comp…
openalex
Yann LeCun, Léon Bottou, Yoshua Bengio, Patrick Haffner
1998-01-01
置信度 0.72
Computer scienceArtificial intelligenceConvolutional neural networkIntelligent character recognitionHandwriting recognition
-
This paper presents a unified Vision-Language Pre-training (VLP) model. The model is unified in that (1) it can be fine-tuned for either vision-language generation (e.g., image captioning) or understanding (e.g., visual question answering) tasks, and (2) it us…
openalex
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等
2020-04-03
置信度 0.72
Closed captioningComputer scienceTransformerQuestion answeringLanguage model
-
Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Models (VLM) with this ability. We propose k…
openalex
Alayrac, Jean-Baptiste, Jeff Donahue, Pauline Luc, Antoine Miech 等
2022-04-29
置信度 0.72
Computer scienceClosed captioningContext (archaeology)Flexibility (engineering)Variety (cybernetics)
-
BACKGROUND: Global and regional prevalence estimates for blindness and vision impairment are important for the development of public health policies. We aimed to provide global estimates, trends, and projections of global blindness and vision impairment. METHO…
openalex
Rupert Bourne, Seth Flaxman, Tasanee Braithwaite, Maria Vittoria Cicinelli 等
2017-08-03
置信度 0.72
Visual impairmentVisual acuityPresbyopiaMedicineMeta-analysis
-
Navigation guided by natural language instructions presents a challenging reasoning problem for instruction followers. Natural language instructions typically identify only a few high-level decisions and landmarks rather than complete low-level motor behaviors…
openalex
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach 等
2018-06-07
置信度 0.72
Computer scienceHuman–computer interactionLinguisticsSpeech recognitionArtificial intelligence
-
This article seeks to develop Translanguaging as a theory of language and discuss the theoretical motivations behind and the added values of the concept. I contextualize Translanguaging in the linguistic realities of the 21st century, especially the fluid and …
openalex
Li Wei
2017-10-06
置信度 0.72
TranslanguagingLinguisticsSociologySociocultural evolutionBiosemiotics