-
Accurately identifying, understanding and describing traffic safety-critical events (SCEs), including crashes, tire strikes, and near-crashes, is crucial for advanced driver assistance systems, automated driving systems, and traffic safety. As SCEs are rare ev…
arxiv
Liang Shi, Boyu Jiang, Tong Zeng, Feng Guo
2024-10-01T18:10:23Z
置信度 0.78
cs.CV
-
Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language understanding and reasoning from generative pretraining. While it has long been conje…
arxiv
Valentin Gabeur, Shangbang Long, Songyou Peng, Paul Voigtlaender 等
2026-04-22T08:23:48Z
置信度 0.78
cs.CVcs.AI
-
This paper presents a novel approach that integrates vision foundation models with reinforcement learning to enhance object interaction capabilities in simulated environments. By combining the Segment Anything Model (SAM) and YOLOv5 with a Proximal Policy Opti…
arxiv
Ahmad Farooq, Kamran Iqbal
2025-08-07T20:29:01Z
置信度 0.78
cs.ROcs.AIcs.CVcs.LGeess.SY
-
When it comes to clinical images, automatic segmentation has a wide variety of applications and a considerable diversity of input domains, such as different types of Magnetic Resonance Images (MRIs) and Computerized Tomography (CT) scans. This heterogeneity is…
arxiv
Matteo Bastico, David Ryckelynck, Laurent Corté, Yannick Tillier 等
2023-10-09T09:51:44Z
置信度 0.78
eess.IVcs.CVcs.LG
-
Reasoning about fine-grained spatial relationships in warehouse-scale environments poses a significant challenge for existing vision-language models (VLMs), which often struggle to comprehend 3D layouts, object arrangements, and multimodal cues in real-world i…
arxiv
Vinh-Thuan Ly, Hoang M. Truong, Xuan-Huong Nguyen
2025-08-25T01:36:22Z
置信度 0.78
cs.CV
-
The NIPS 2018 Adversarial Vision Challenge is a competition to facilitate measurable progress towards robust machine vision models and more generally applicable adversarial attacks. This document is an updated version of our competition proposal that was accep…
arxiv
Wieland Brendel, Jonas Rauber, Alexey Kurakin, Nicolas Papernot 等
2018-08-06T16:13:43Z
置信度 0.78
cs.LGcs.CVstat.ML
-
Vision transformers have shown great success on numerous computer vision tasks. However, its central component, softmax attention, prohibits vision transformers from scaling up to high-resolution images, due to both the computational complexity and memory foot…
arxiv
Weixuan Sun, Zhen Qin, Hui Deng, Jianyuan Wang 等
2022-06-21T17:33:53Z
置信度 0.78
cs.CV
-
Over the past few years, there has been growing interest in developing a broad, universal, and general-purpose computer vision system. Such systems have the potential to address a wide range of vision tasks simultaneously, without being limited to specific pro…
arxiv
Feng Lin, Wenze Hu, Yaowei Wang, Yonghong Tian 等
2022-12-19T12:40:13Z
置信度 0.78
cs.CV
-
Computer vision is a growing field with a lot of new applications in automation and robotics, since it allows the analysis of images and shapes for the generation of numerical or analytical information. One of the most used method of information extraction is …
arxiv
Dominique Beaini, Sofiane Achiche, Yann-Seing Law-Kam Cio, Maxime Raison
2018-06-20T21:31:00Z
置信度 0.78
cs.CV
-
Building state-of-the-art Vision-Language Models (VLMs) with strong captioning capabilities typically necessitates training on billions of high-quality image-text pairs, requiring millions of GPU hours. This paper introduces the Vision-Language-Vision (VLV) au…
arxiv
Tiezheng Zhang, Yitong Li, Yu-cheng Chou, Jieneng Chen 等
2025-07-09T17:59:04Z
置信度 0.78
cs.CV
-
A core objective of the TERRA-REF project was to generate an open-access reference dataset for the evaluation of sensing technologies to study plants under field conditions. The TERRA-REF program deployed a suite of high-resolution, cutting edge technology sen…
arxiv
David LeBauer, Max Burnette, Noah Fahlgren, Rob Kooper 等
2021-07-29T15:01:29Z
置信度 0.78
cs.CVcs.LG
-
Computer vision is hard because of a large variability in lighting, shape, and texture; in addition the image signal is non-additive due to occlusion. Generative models promised to account for this variability by accurately modelling the image formation proces…
arxiv
Varun Jampani, Sebastian Nowozin, Matthew Loper, Peter V. Gehler
2014-02-04T20:52:26Z
置信度 0.78
cs.CVcs.LGstat.ML
-
Automatic dietary assessment based on food images remains a challenge, requiring precise food detection, segmentation, and classification. Vision-Language Models (VLMs) offer new possibilities by integrating visual and textual reasoning. In this study, we eval…
arxiv
Sergio Romero-Tapiador, Ruben Tolosana, Blanca Lacruz-Pleguezuelos, Laura Judith Marcos Zambrano 等
2025-04-09T14:33:59Z
置信度 0.78
cs.CVcs.AI
-
Facial video-based remote physiological measurement is a promising research area for detecting human vital signs (e.g., heart rate, respiration frequency) in a non-contact way. Conventional approaches are mostly supervised learning, requiring extensive collect…
arxiv
Zijie Yue, Miaojing Shi, Hanli Wang, Shuai Ding 等
2024-07-11T13:45:50Z
置信度 0.78
cs.CV
-
Transformers are widely used as generic backbones in computer vision, despite initially introduced for natural language processing. Recently, the Long Short-Term Memory (LSTM) has been extended to a scalable and performant architecture - the xLSTM - which over…
arxiv
Benedikt Alkin, Maximilian Beck, Korbinian Pöppel, Sepp Hochreiter 等
2024-06-06T17:49:21Z
置信度 0.78
cs.CVcs.AIcs.LG
-
Large vision--language models (VLMs) often use a frozen vision backbone, whose image features are mapped into a large language model through a lightweight connector. While transformer-based encoders are the standard visual backbone, we ask whether state space …
arxiv
Shang-Jui Ray Kuo, Paola Cascante-Bonilla
2026-03-19T17:56:32Z
置信度 0.78
cs.CVcs.LG
-
Proceedings of the Second Croatian Computer Vision Workshop (CCVW 2013, http://www.fer.unizg.hr/crv/ccvw2013) held September 19, 2013, in Zagreb, Croatia. Workshop was organized by the Center of Excellence for Computer Vision of the University of Zagreb.
arxiv
Sven Lončarić, Siniša Šegvić
2013-10-01T14:26:29Z
置信度 0.78
cs.CV
-
While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have helped address in language, but which remains underexplored in vision. We address this gap by training SAEs on CLIP'…
arxiv
Sonia Joseph, Praneet Suresh, Ethan Goldfarb, Lorenz Hufe 等
2025-04-11T17:56:09Z
置信度 0.78
cs.CVcs.AIcs.LG
-
Recently, we have witnessed the great success of the generalist model in natural language processing. The generalist model is a general framework trained with massive data and is able to process various downstream tasks simultaneously. Encouraged by their impr…
arxiv
Ziyi Wang, Yongming Rao, Shuofeng Sun, Xinrun Liu 等
2025-06-11T17:23:41Z
置信度 0.78
cs.CVcs.AI
-
Deep neural networks can be converted to multi-exit architectures by inserting early exit branches after some of their intermediate layers. This allows their inference process to become dynamic, which is useful for time critical IoT applications with stringent…
arxiv
Arian Bakhtiarnia, Qi Zhang, Alexandros Iosifidis
2021-06-29T09:01:13Z
置信度 0.78
cs.CV
-
The systematic evaluation and understanding of computer vision models under varying conditions require large amounts of data with comprehensive and customized labels, which real-world vision datasets rarely satisfy. While current synthetic data generators offe…
arxiv
Yunhao Ge, Yihe Tang, Jiashu Xu, Cem Gokmen 等
2024-05-15T17:57:56Z
置信度 0.78
cs.CV
-
Computer vision systems require large amounts of manually annotated data to properly learn challenging visual concepts. Crowdsourcing platforms offer an inexpensive method to capture human knowledge and understanding, for a vast number of visual perception tas…
arxiv
Adriana Kovashka, Olga Russakovsky, Li Fei-Fei, Kristen Grauman
2016-11-07T16:11:19Z
置信度 0.78
cs.CVcs.HC
-
We attempt to reduce the computational costs in vision transformers (ViTs), which increase quadratically in the token number. We present a novel training paradigm that trains only one ViT model at a time, but is capable of providing improved image recognition …
arxiv
Mingbao Lin, Mengzhao Chen, Yuxin Zhang, Chunhua Shen 等
2022-05-23T15:42:12Z
置信度 0.78
cs.CV
-
Federated Learning (FL) is a distributed learning paradigm that can learn a global or personalized model from decentralized datasets on edge devices. However, in the computer vision domain, model performance in FL is far behind centralized training due to the …
arxiv
Chaoyang He, Alay Dilipbhai Shah, Zhenheng Tang, Di Fan1Adarshan Naiynar Sivashunmugam 等
2021-11-22T09:26:08Z
置信度 0.78
cs.CVcs.AIcs.LG
-
Egocentric vision aims to capture and analyse the world from the first-person perspective. We explore the possibilities for egocentric wearable devices to improve and enhance industrial use cases w.r.t. data collection, annotation, labelling and downstream app…
arxiv
Vivek Chavan, Oliver Heimann, Jörg Krüger
2024-06-11T21:48:20Z
置信度 0.78
cs.CV
-
Vision transformers require a huge amount of labeled data to outperform convolutional neural networks. However, labeling a huge dataset is a very expensive process. Self-supervised learning techniques alleviate this problem by learning features similar to supe…
arxiv
Sachin Chhabra, Prabal Bijoy Dutta, Hemanth Venkateswara, Baoxin Li
2022-10-27T18:55:12Z
置信度 0.78
cs.CV
-
Large Vision-Language Models (LVLMs) typically follow a two-stage training paradigm-pretraining and supervised fine-tuning. Recently, preference optimization, derived from the language domain, has emerged as an effective post-training reinforcement strategy to…
arxiv
Yufei Zhan, Yousong Zhu, Shurong Zheng, Hongyin Zhao 等
2025-03-23T10:21:14Z
置信度 0.78
cs.CVcs.AI
-
Building on recent advances in language-based reasoning models, we explore multimodal reasoning that integrates vision and text. Existing multimodal benchmarks primarily test visual extraction combined with text-based reasoning, lacking true visual reasoning w…
arxiv
Mert Unsal, Aylin Akkus
2025-06-13T09:03:33Z
置信度 0.78
cs.CVcs.LG
-
Window-based attention has become a popular choice in vision transformers due to its superior performance, lower computational complexity, and less memory footprint. However, the design of hand-crafted windows, which is data-agnostic, constrains the flexibilit…
arxiv
Qiming Zhang, Jing Zhang, Yufei Xu, Dacheng Tao
2023-03-27T11:13:50Z
置信度 0.78
cs.CV
-
In this report, we present our optical flow approach, MS-RAFT+, that won the Robust Vision Challenge 2022. It is based on the MS-RAFT method, which successfully integrates several multi-scale concepts into single-scale RAFT. Our approach extends this method by…
arxiv
Azin Jahedi, Maximilian Luz, Lukas Mehl, Marc Rivinius 等
2022-10-30T17:48:11Z
置信度 0.78
cs.CV
-
Domain adaptation has been extensively investigated in computer vision but still requires access to target data at the training time, which might be difficult to obtain in real-world autonomous driving scenarios, especially under rare or adverse conditions. In…
arxiv
Mohammad Fahes, Tuan-Hung Vu, Andrei Bursuc, Patrick Pérez 等
2024-10-28T17:59:53Z
置信度 0.78
cs.CVcs.LG
-
Finding the minimum spanning tree (MST) of a graph is an important task in computer vision, as it enables a sparse and low-cost representation of connectivity among elements (such as superpixels, points, or regions), which is useful for tasks such as segmentat…
arxiv
Guilherme E. L. Pexe, Lucas A. M. Rattighieri, Leandro A. Passos, Douglas Rodrigues 等
2026-03-21T07:34:07Z
置信度 0.78
quant-ph
-
The emergence of COVID-19 has had a global and profound impact, not only on society as a whole, but also on the lives of individuals. Various prevention measures were introduced around the world to limit the transmission of the disease, including face masks, m…
arxiv
Fevziye Irem Eyiokur, Alperen Kantarcı, Mustafa Ekrem Erakın, Naser Damer 等
2022-11-07T17:20:39Z
置信度 0.78
cs.CVeess.IV
-
Vision graph neural networks (ViG) offer a new avenue for exploration in computer vision. A major bottleneck in ViGs is the inefficient k-nearest neighbor (KNN) operation used for graph construction. To solve this issue, we propose a new method for designing V…
arxiv
Mustafa Munir, William Avery, Md Mostafijur Rahman, Radu Marculescu
2024-05-10T23:21:16Z
置信度 0.78
cs.CVcs.AIcs.LG
-
The graph matching optimization problem is an essential component for many tasks in computer vision, such as bringing two deformable objects in correspondence. Naturally, a wide range of applicable algorithms have been proposed in the last decades. Since a com…
arxiv
Stefan Haller, Lorenz Feineis, Lisa Hutschenreiter, Florian Bernard 等
2022-07-01T09:37:34Z
置信度 0.78
cs.CVmath.OC
-
Eye contact is a crucial non-verbal interaction modality and plays an important role in our everyday social life. While humans are very sensitive to eye contact, the capabilities of machines to capture a person's gaze are still mediocre. We tackle this challen…
arxiv
Thorsten Hempel, Magnus Jung, Ahmed A. Abdelrahman, Ayoub Al-Hamadi
2023-11-08T07:42:31Z
置信度 0.78
cs.CV
-
Egocentric videos can bring a lot of information about how humans perceive the world and interact with the environment, which can be beneficial for the analysis of human behaviour. The research in egocentric video analysis is developing rapidly thanks to the i…
arxiv
Ivan Rodin, Antonino Furnari, Dimitrios Mavroedis, Giovanni Maria Farinella
2021-07-28T14:58:13Z
置信度 0.78
cs.CV
-
Vision Transformers (ViTs) have attracted a lot of popularity in recent years, due to their exceptional capabilities in modeling long-range spatial dependencies and scalability for large scale training. Although the training parallelism of self-attention mecha…
arxiv
Ali Hatamizadeh, Michael Ranzinger, Shiyi Lan, Jose M. Alvarez 等
2023-10-30T16:55:50Z
置信度 0.78
cs.CVcs.AIcs.LG
-
Prompt learning has emerged as an efficient and effective method for fine-tuning vision-language models such as CLIP. While many studies have explored generalisation abilities of these models in few-shot classification tasks and a few studies have addressed fa…
arxiv
Myong Chol Jung, Joanna Dipnall, Belinda Gabbe, He Zhao
2024-05-25T06:46:16Z
置信度 0.78
cs.CV
-
In recent years, there has been an explosion of 2D vision models for numerous tasks such as semantic segmentation, style transfer or scene editing, enabled by large-scale 2D image datasets. At the same time, there has been renewed interest in 3D scene represen…
arxiv
Mukund Varma T, Peihao Wang, Zhiwen Fan, Zhangyang Wang 等
2024-03-27T18:13:16Z
置信度 0.78
cs.CV
-
Vision transformers have demonstrated the potential to outperform CNNs in a variety of vision tasks. But the computational and memory requirements of these models prohibit their use in many applications, especially those that depend on high-resolution images, …
arxiv
Yue Liu, Christos Matsoukas, Fredrik Strand, Hossein Azizpour 等
2022-08-10T14:08:55Z
置信度 0.78
cs.CVcs.LG
-
Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision mechanistic interpretability has been hindered by the lack of accessible framewo…
arxiv
Sonia Joseph, Praneet Suresh, Lorenz Hufe, Edward Stevinson 等
2025-04-28T04:31:24Z
置信度 0.78
cs.CVcs.AIcs.LG
-
The number of road vehicles significantly increased in recent decades. This trend accompanied a build-up of road infrastructure and development of various control systems to increase road traffic safety, road capacity and travel comfort. In traffic safety sign…
arxiv
Kristian Kovačić, Edouard Ivanjko, Hrvoje Gold
2013-10-01T14:19:11Z
置信度 0.78
cs.CV
-
Omnidirectional vision, using 360-degree vision to understand the environment, has become increasingly critical across domains like robotics, industrial inspection, and environmental monitoring. Compared to traditional pinhole vision, omnidirectional vision pr…
arxiv
Xu Zheng, Chenfei Liao, Ziqiao Weng, Kaiyu Lei 等
2025-09-16T11:54:37Z
置信度 0.78
cs.CV
-
Foundation models have transformed vision and language processing by providing rich, reusable representations that transfer across diverse tasks. Sheet music, as a visual encoding of musical language, lacks such a strong domain-specific backbone. We introduce …
arxiv
Carlos Penarrubia, Antonio Rios-Vila, Eliseo Fuentes-Martinez, Juan C. Martinez-Sevilla 等
2026-06-30T15:27:12Z
置信度 0.78
cs.CV
-
Egocentric vision holds great promises for increasing access to visual information and improving the quality of life for people with visual impairments, with object recognition being one of the daily challenges for this population. While we strive to improve r…
arxiv
Kyungjun Lee, Abhinav Shrivastava, Hernisa Kacorri
2020-02-28T05:32:36Z
置信度 0.78
cs.CVcs.HC
-
The remarkable success of Vision Transformers in Artificial Neural Networks (ANNs) has led to a growing interest in incorporating the self-attention mechanism and transformer-based architecture into Spiking Neural Networks (SNNs). While existing methods propos…
arxiv
Xinyu Shi, Zecheng Hao, Zhaofei Yu
2024-03-21T11:16:42Z
置信度 0.78
cs.NEcs.CVcs.LG
-
Scene text recognition (STR) enables computers to recognize and read the text in various real-world scenes. Recent STR models benefit from taking linguistic information in addition to visual cues into consideration. We propose a novel Masked Vision-Language Tr…
arxiv
Jie Wu, Ying Peng, Shengming Zhang, Weigang Qi 等
2022-11-09T10:28:23Z
置信度 0.78
cs.CV
-
Deep Learning has pushed the limits of what was possible in the domain of Digital Image Processing. However, that is not to say that the traditional computer vision techniques which had been undergoing progressive development in years prior to the rise of DL h…
arxiv
Niall O' Mahony, Sean Campbell, Anderson Carvalho, Suman Harapanahalli 等
2019-10-30T12:25:10Z
置信度 0.78
cs.CVcs.LG
-
Vision Transformer (ViT) has achieved excellent performance and demonstrated its promising potential in various computer vision tasks. The wide deployment of ViT in real-world tasks requires a thorough understanding of the societal impact of the model. However…
arxiv
Bowei Tian, Ruijie Du, Yanning Shen
2024-07-20T08:10:37Z
置信度 0.78
cs.CVcs.CY
-
Vision Graph Neural Networks (ViGs) offer a new direction for advancements in vision architectures. While powerful, ViGs often face substantial computational challenges stemming from their graph construction phase, which can hinder their efficiency. To address…
arxiv
Mustafa Munir, Md Mostafijur Rahman, Radu Marculescu
2025-11-13T04:16:20Z
置信度 0.78
cs.CVcs.AIcs.LG
-
Vision backbone networks play a central role in modern computer vision. Enhancing their efficiency directly benefits a wide range of downstream applications. To measure efficiency, many publications rely on MACs (Multiply Accumulate operations) as a predictor …
arxiv
Moritz Nottebaum, Matteo Dunnhofer, Christian Micheloni
2026-03-27T16:05:20Z
置信度 0.78
cs.CVcs.AI
-
In the past few years, the emergence of pre-training models has brought uni-modal fields such as computer vision (CV) and natural language processing (NLP) to a new era. Substantial works have shown they are beneficial for downstream uni-modal tasks and avoid …
arxiv
Feilong Chen, Duzhen Zhang, Minglun Han, Xiuyi Chen 等
2022-02-18T07:54:02Z
置信度 0.78
cs.CVcs.CL
-
The ImageNet hierarchy provides a structured taxonomy of object categories, offering a valuable lens through which to analyze the representations learned by deep vision models. In this work, we conduct a comprehensive analysis of how vision models encode the I…
arxiv
Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff, Changbai Li 等
2025-05-21T19:38:48Z
置信度 0.78
cs.CVcs.LG
-
Visual object tracking and segmentation are becoming fundamental tasks for understanding human activities in egocentric vision. Recent research has benchmarked state-of-the-art methods and concluded that first person egocentric vision presents challenges compa…
arxiv
Matteo Dunnhofer, Zaira Manigrasso, Christian Micheloni
2025-07-21T19:25:50Z
置信度 0.78
cs.CV
-
Vision-Language Model (VLM) have gained widespread adoption in Open-Vocabulary (OV) object detection and segmentation tasks. Despite they have shown promise on OV-related tasks, their effectiveness in conventional vision tasks has thus far been unevaluated. In…
arxiv
Yongchao Feng, Yajie Liu, Shuai Yang, Wenrui Cai 等
2025-04-13T08:28:13Z
置信度 0.78
cs.CVcs.AI
-
Transformers have revolutionized computer vision and natural language processing, but their high computational complexity limits their application in high-resolution image processing and long-context analysis. This paper introduces Vision-RWKV (VRWKV), a model…
arxiv
Yuchen Duan, Weiyun Wang, Zhe Chen, Xizhou Zhu 等
2024-03-04T18:46:20Z
置信度 0.78
cs.CV
-
Despite their promise to perform complex reasoning, large language models (LLMs) have been shown to have limited effectiveness in end-to-end planning. This has inspired an intriguing question: if these models cannot plan well, can they still contribute to the …
arxiv
Mohamed Aghzal, Xiang Yue, Erion Plaku, Ziyu Yao
2024-11-27T19:32:03Z
置信度 0.78
cs.CVcs.CL
-
Vision-language models normally execute the same complete vision encoder for every question, even when OCR, counting, object, attribute, and spatial queries may not require identical computation. We study whether fixed-budget combinations of vision blocks can …
arxiv
Tarun Tomar
2026-07-19T03:43:26Z
置信度 0.78
cs.CV
-
Reliable perception is fundamental for safety critical decision making in autonomous driving. Yet, vision based object detector neural networks remain vulnerable to uncertainty arising from issues such as data bias and distributional shifts. In this paper, we …
arxiv
Nishad Sahu, Shounak Sural, Aditya Satish Patil, Ragunathan 等
2025-10-17T18:04:31Z
置信度 0.78
cs.CV
-
Sparse coding has been incorporated in models of the visual cortex for its computational advantages and connection to biology. But how the level of sparsity contributes to performance on visual tasks is not well understood. In this work, sparse coding has been…
arxiv
Joshua Bowren, Luis Sanchez-Giraldo, Odelia Schwartz
2021-08-03T14:55:33Z
置信度 0.78
cs.CVcs.LG
-
The extension of convolutional neural networks (CNNs) to non-Euclidean geometries has led to multiple frameworks for studying manifolds. Many of those methods have shown design limitations resulting in poor modelling of long-range associations, as the generali…
arxiv
Simon Dahan, Logan Z. J. Williams, Abdulah Fawaz, Daniel Rueckert 等
2022-05-31T14:41:01Z
置信度 0.78
cs.CVcs.LGq-bio.NC
-
Multimodal Large Language Models (MLLMs) still struggle with fine-grained visual understanding, where answers often depend on small but decisive evidence in the full image. We observe a regional-to-global perception gap: the same MLLM answers fine-grained ques…
arxiv
Qianhao Yuan, Jie Lou, Xing Yu, Hongyu Lin 等
2026-05-18T17:57:04Z
置信度 0.78
cs.CVcs.AIcs.CLcs.LG
-
Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two challenges: a semantic-geometric misalignm…
arxiv
Zuntao Liu, Yi Du, Taimeng Fu, Shaoshu Su 等
2025-11-25T18:59:02Z
置信度 0.78
cs.CV
-
We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models. While advances in VLAs have introduced robot policies that are both generalizable and semantically grounded, these model…
arxiv
Charlotte Morissette, Amin Abyaneh, Wei-Di Chang, Anas Houssaini 等
2026-03-15T20:57:51Z
置信度 0.78
cs.ROcs.CVcs.LG
-
This paper surveys vision-language pre-training (VLP) methods for multimodal intelligence that have been developed in the last few years. We group these approaches into three categories: ($i$) VLP for image-text tasks, such as image captioning, image-text retr…
arxiv
Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang 等
2022-10-17T17:11:36Z
置信度 0.78
cs.CVcs.CL
-
Convolutional Neural Networks (CNNs), architectures consisting of convolutional layers, have been the standard choice in vision tasks. Recent studies have shown that Vision Transformers (VTs), architectures based on self-attention modules, achieve comparable p…
arxiv
Kishaan Jeeveswaran, Senthilkumar Kathiresan, Arnav Varma, Omar Magdy 等
2022-01-21T13:18:16Z
置信度 0.78
cs.CV
-
In recent years, vision-centric Bird's Eye View (BEV) perception has garnered significant interest from both industry and academia due to its inherent advantages, such as providing an intuitive representation of the world and being conducive to data fusion. Th…
arxiv
Yuexin Ma, Tai Wang, Xuyang Bai, Huitong Yang 等
2022-08-04T17:53:17Z
置信度 0.78
cs.CV
-
The widespread use of cameras in our society has created an overwhelming amount of video data, far exceeding the capacity for human monitoring. This presents a critical challenge for public safety and security, as the timely detection of anomalous or criminal …
arxiv
Pascal Benschop, Cristian Meo, Justin Dauwels, Jelte P. Mense
2025-10-27T10:27:02Z
置信度 0.78
cs.CV
-
In recent years, a large amount of multi-disciplinary research has been conducted on sparse models and their applications. In statistics and machine learning, the sparsity principle is used to perform model selection---that is, automatically selecting a simple…
arxiv
Julien Mairal, Francis Bach, Jean Ponce
2014-11-12T16:33:37Z
置信度 0.78
cs.CV
-
Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends leverage Vision-Language Models (VLMs) for semantic supervision, these multimodal approaches often demand prohi…
arxiv
Mika Feng, Pierre Gallin-Martel, Koichi Ito, Takafumi Aoki
2026-04-21T08:05:21Z
置信度 0.78
cs.CV
-
Computer vision methods that explicitly detect object parts and reason on them are a step towards inherently interpretable models. Existing approaches that perform part discovery driven by a fine-grained classification task make very restrictive assumptions on…
arxiv
Ananthu Aniraj, Cassio F. Dantas, Dino Ienco, Diego Marcos
2024-07-05T14:24:37Z
置信度 0.78
cs.CVcs.AIcs.LG
-
The Convolutional Neural Networks (CNNs) have been the dominant and effective approach for general computer vision tasks. Recently, Kolmogorov-Arnold neural networks (KANs), based on the Kolmogorov-Arnold representation theorem, have shown potential to replace…
arxiv
Zhaoxiang Liu, Zhicheng Ma, Kaikai Zhao, Kai Wang 等
2026-04-25T14:14:41Z
置信度 0.78
cs.CV
-
Domain generalisation involves pooling knowledge from source domain(s) into a single model that can generalise to unseen target domain(s). Recent research in domain generalisation has faced challenges when using deep learning models as they interact with data …
arxiv
Hamza Riaz, Alan F. Smeaton
2023-07-16T17:50:37Z
置信度 0.78
cs.CVcs.LG
-
Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond question answering remains unclear. In this paper, we benchmark popular MFMs (GPT-4o, o4-mini, Gemini 1.5 Pro and G…
arxiv
Rahul Ramachandran, Ali Garjani, Roman Bachmann, Andrei Atanov 等
2025-07-02T17:59:07Z
置信度 0.78
cs.CVcs.AIcs.LG
-
crossref
Matthew Petoe, Lauren Ayton, Mohit Shivdasani
2025-08-26T06:19:03Z
置信度 0.70
-
crossref
Valtteri Heiskanen, Kalle Marjanen, Pasi Kallio
2009-01-03T09:13:01Z
置信度 0.70
-
Introduction: There are nearly 40 million cases of blindness in the worldwide, and 124 million people have been affected by low vision. For this reason, researchers are intent on developing new ways to restore vision. One of these efforts was the offer of an i…
crossref
Khadijeh moulaei
2017-06-19T09:31:02Z
置信度 0.70
-
Bionic vision technology is a quickly expanding topic that includes creating tools to help people with visual impairments see again.The technology replaces damaged cells and stimulates the visual system using a variety of methods, including retinal implants, c…
crossref
2023-05-15T18:41:06Z
置信度 0.70
-
crossref
Zhao-jun Yang, Fei Chen, Ji Zhao, Xiao-jie Wu
2009-04-13T04:10:42Z
置信度 0.70
-
crossref
Julien R. Serres, Franck Ruffier
2015-01-15T23:30:13Z
置信度 0.70
-
crossref
Yu-zhang Gu, Makoto Sato, Xiao-lin Zhang
2008-01-02T22:36:46Z
置信度 0.70
-
One’s vision impacts sport performance through an ability to perceive, track, hit or shoot more proficiently. Achieving improvements in sight can tend to lead to greater sporting success. Methods or strategies for improving one’s eyesight are legal in sport. T…
crossref
Cheryl Mallen
2019-02-18T17:16:33Z
置信度 0.70
-
The burgeoning fields of the Internet of things (IoT) and artificial intelligence (AI) have escalated the demands for image sensing technologies, necessitating advancements in sensor efficiency and functionality. Traditional image sensors, structured on von Ne…
crossref
2025-03-27T23:10:13Z
置信度 0.70
-
How can we return a functional form of sight to people who are living with incurable blindness? Despite recent advances in the development of visual neuroprostheses, the quality of current prosthetic vision is still rudimentary and does not differ much across …
crossref
Michael Beyeler, Melani Sanchez Garcia
2022-06-17T15:12:06Z
置信度 0.70
-
crossref
Xiao-guang Li, Jing-long Wu, Sadao Kawamura
2018-02-08T09:32:47Z
置信度 0.70
-
Among all of the skills required in sports, batting a ball is one of the most difficult. In this paper, we propose a vision-based batting training system for batting practice. Our goal is to predict the flying trajectory of a batted ball by using a high-speed …
crossref
Hsiu-Min Chuang, Yang Liu, Akio Namiki
2018-01-22T17:39:22Z
置信度 0.70
-
Currently, the prevailing methods to design and manufacture fish-like robots focused on articulated way with multi-degree-of-freedom structure. For increasing the reliability and reducing the cost of manufacturing and controlling, a 1-DOF beam-driven bionic ro…
crossref
Jialin Xu, Xingsong Wang
2009-01-16T11:43:02Z
置信度 0.70
-
crossref
Yuan Li
2025-06-06T13:43:49Z
置信度 0.70
-
Aiming at the low efficiency of information processing in a robot vision system, our work studies how to extract fixed points during tasks so that significant and interesting target areas in the scene can be located. First, the architecture of a gaze extractio…
crossref
Yingni Duan, Guozhu Li, Yanzi Deng
2023-01-04T18:55:16Z
置信度 0.70
-
The International Association for Hydro-Environment Engineering and Research (IAHR), founded in 1935, is a worldwide independent organisation of engineers and water specialists working in fields related to the hydro-environmental sciences and their practical a…
crossref
Hairong Gao, Zihan Liu, Yu Han
2024-06-18T12:00:55Z
置信度 0.70
-
ABSTRACT Retinitis pigmentosa refers to a family of inherited photoreceptor degenerations resulting in blindness. During and after photoreceptor loss, neurons of the inner retina are known to undergo plastic changes. Here, we have investigated in detail whethe…
crossref
E.E. O'Brien, U. Greferath, E.L. Fletcher
2013-10-31T11:03:57Z
置信度 0.70
-
crossref
Yong Wang, Hongqi Liu, Xiaoguang Wang
2021-12-10T06:48:00Z
置信度 0.70
-
Bionic retinal implants are gaining acceptance in the treatment of blindness from degenerative diseases including retinitis pigmentosa and macular degeneration. A current obstacle to the improved performance of such implants is the difficulty of comparing the …
crossref
Daniel Rathbun, Nima Ghorbani, Hamed Shabani, Eberhart Zrenner 等
2018-07-02T07:42:37Z
置信度 0.70
-
Flexible polymers have gained much attention in the development of low cost, magnetic resonance compatible, and nonfragile implantable medical devices. However, efficacy of the conventional polymer encapsulations containing hybrid interfaces is limited due to …
openalex
Seung Woo Lee, Kyou Sik Min, Joonsoo Jeong, Jung‐Hoon Kim 等
2011-04-06
置信度 0.72
Materials scienceEncapsulation (networking)PolymerPolyimideParylene
-
One of the most sought-after applications of neuroengineering is the communication between the arm and an artificial prosthetic device for the replacement of an amputated hand or the treatment of peripheral nerve injuries. For that, an electrode is placed arou…
openalex
Ignacio Delgado, Jordi Badía, Arán Pascual‐Font, Alfonso Rodríguez‐Baeza 等
2016-07-01
置信度 0.72
NeuroprostheticsFascicleMedicineNeural ProsthesisForearm
-
Neuroprosthetics that combine a brain computer interface (BCI) with functional electrical stimulation (FES) can restore voluntary control of a patients' own paralyzed limbs. To date, human studies have demonstrated an "all-or-none" type of control for a fixed …
openalex
David A. Friedenberg, Michael A. Schwemmer, Andrew J. Landgraf, Nicholas V. Annetta 等
2017-08-10
置信度 0.72
NeuroprostheticsFunctional electrical stimulationBrain–computer interfacePhysical medicine and rehabilitationWrist
-
openalex
Vivek R. Athalye, Karunesh Ganguly, Rui M. Costa, Jose M. Carmena
2017-02-01
置信度 0.72
NeuroscienceMotor learningMotor cortexMotor controlPopulation
-
Layer-by-layer (LBL) assembly has been used to prepare single-walled carbon nanotube (SWNT) freestanding structures that can be used for implantable devices with unique mechanical and electrical properties (see Figure). The thin LBL membranes prepared are bioc…
openalex
Muhammed K. Gheith, Vladimir A. Sinani, James P. Wicksted, Robert L. Matts 等
2005-10-11
置信度 0.72
Materials scienceCarbon nanotubeBiocompatible materialNanotechnologyPolyelectrolyte
-
Reducing the mechanical mismatch between the stiffness of a neural implant and the softness of the neural tissue is still an open challenge in neuroprosthetics. The emergence of conductive hydrogels in the last few years has considerably widened the spectrum o…
openalex
Laura Ferlauto, Antonio Nunzio D’Angelo, Paola Vagni, Marta Jole Ildelfonsa Airaghi Leccardi 等
2018-09-19
置信度 0.72
NeuroprostheticsMicroelectrodePEDOT:PSSCharacterization (materials science)Materials science