| family | paper | links |
| general | Towards Source-Free Machine Unlearning | pdf |
| general | Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video | pdf |
| general | Hyperbolic Category Discovery | pdf |
| beings | The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion | pdf |
| general | CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models | pdf |
| general | Words or Vision: Do Vision-Language Models Have Blind Faith in Text? | pdf |
| general | Learning to Detect Objects from Multi-Agent LiDAR Scans without Manual Labels | pdf |
| general | DeepLA-Net: Very Deep Local Aggregation Networks for Point Cloud Analysis | pdf |
| general | Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices | pdf |
| general | APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers | pdf |
| general | AdaptCMVC: Robust Adaption to Incremental Views in Continual Multi-view Clustering | pdf |
| general | UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References | pdf |
| general | Binarized Mamba-Transformer for Lightweight Quad Bayer HybridEVS Demosaicing | pdf |
| general | Interpretable Image Classification via Non-parametric Part Prototype Learning | pdf |
| beings | DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh | pdf |
| worlds | Estimating Body and Hand Motion in an Ego-sensed World | pdf |
| general | Evaluating Vision-Language Models as Evaluators in Path Planning | pdf |
| general | Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM | pdf |
| general | SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI Detection | pdf |
| general | Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding | pdf |
| general | SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback Optimization | pdf |
| motion | Exploring Timeline Control for Facial Motion Generation | pdf |
| gsplat | GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion | pdf |
| beings | AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness Benchmark | pdf |
| general | Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation | pdf |
| general | De^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze Estimation | pdf |
| general | ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning | pdf |
| general | Self-Expansion of Pre-trained Models with Mixture of Adapters for Continual Learning | pdf |
| general | Brain-Inspired Spiking Neural Networks for Energy-Efficient Object Detection | pdf |
| general | Medusa: A Multi-Scale High-order Contrastive Dual-Diffusion Approach for Multi-View Clustering | pdf |
| general | MambaOut: Do We Really Need Mamba for Vision? | pdf |
| general | Seurat: From Moving Points to Depth | pdf |
| general | Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models | pdf |
| general | The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation | pdf |
| general | DnLUT: Ultra-Efficient Color Image Denoising via Channel-Aware Lookup Tables | pdf |
| motion | BiM-VFI: Bidirectional Motion Field-Guided Frame Interpolation for Video with Non-uniform Motions | pdf |
| general | SATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision Transformers | pdf |
| general | Nested Diffusion Models Using Hierarchical Latent Priors | pdf |
| general | A Theory of Learning Unified Model via Knowledge Integration from Label Space Varying Domains | pdf |
| general | HiLoTs: High-Low Temporal Sensitive Representation Learning for Semi-Supervised LiDAR Segmentation in Autonomous Driving | pdf |
| general | DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any Architecture | pdf |
| general | SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization | pdf |
| general | Debiasing Multimodal Large Language Models via Noise-Aware Preference Optimization | pdf |
| gsplat | Feat2GS: Probing Visual Foundation Models with Gaussian Splatting | pdf |
| general | LSNet: See Large, Focus Small | pdf |
| general | DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes | pdf |
| general | DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding | pdf |
| general | EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation | pdf |
| general | Handling Spatial-Temporal Data Heterogeneity for Federated Continual Learning via Tail Anchor | pdf |
| gsplat | DeSiRe-GS: 4D Street Gaussians for Static-Dynamic Decomposition and Surface Reconstruction for Urban Driving Scenes | pdf |
| beings | REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning | pdf |
| general | DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles | pdf |
| gsplat | Gaussian Splashing: Unified Particles for Versatile Motion Synthesis and Rendering | pdf |
| general | Improve Representation for Imbalanced Regression through Geometric Constraints | pdf |
| general | PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model | pdf |
| general | DiffFNO: Diffusion Fourier Neural Operator | pdf |
| general | Zero-Shot Styled Text Image Generation, but Make It Autoregressive | pdf |
| general | Leveraging Perturbation Robustness to Enhance Out-of-Distribution Detection | pdf |
| motion | SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing | pdf |
| general | LookingGlass: Generative Anamorphoses via Laplacian Pyramid Warping | pdf |
| general | ShowMak3r: Compositional TV Show Reconstruction | pdf |
| general | CADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature Leveraging | pdf |
| general | VideoDirector: Precise Video Editing via Text-to-Video Models | pdf |
| general | VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation | pdf |
| general | GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context Encoding | pdf |
| gsplat | RigGS: Rigging of 3D Gaussians for Modeling Articulated Objects in Videos | pdf |
| general | Noise Modeling in One Hour: Minimizing Preparation Efforts for Self-supervised Low-Light RAW Image Denoising | pdf |
| general | High Dynamic Range Video Compression: A Large-Scale Benchmark Dataset and A Learned Bit-depth Scalable Compression Algorithm | pdf |
| general | DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models | pdf |
| general | 3D-AVS: LiDAR-based 3D Auto-Vocabulary Segmentation | pdf |
| beings | MEGA: Masked Generative Autoencoder for Human Mesh Recovery | pdf |
| general | Disentangling Safe and Unsafe Image Corruptions via Anisotropy and Locality | pdf |
| worlds | Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation | pdf |
| general | SphereUFormer: A U-Shaped Transformer for Spherical 360 Perception | pdf |
| general | Beyond Clean Training Data: A Versatile and Model-Agnostic Framework for Out-of-Distribution Detection with Contaminated Training Data | pdf |
| general | FreeUV: Ground-Truth-Free Realistic Facial UV Texture Recovery via Cross-Assembly Inference Strategy | pdf |
| general | HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization | pdf |
| general | StyleMaster: Stylize Your Video with Artistic Generation and Translation | pdf |
| general | Unsupervised Continual Domain Shift Learning with Multi-Prototype Modeling | pdf |
| general | OmniGuard: Hybrid Manipulation Localization via Augmented Versatile Deep Image Watermarking | pdf |
| general | Open-Canopy: Towards Very High Resolution Forest Monitoring | pdf |
| general | Vision-Language Model IP Protection via Prompt-based Learning | pdf |
| general | Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation | pdf |
| general | Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content | pdf |
| general | VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification | pdf |
| general | SPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Models | pdf |
| general | Erase Diffusion: Empowering Object Removal Through Calibrating Diffusion Pathways | pdf |
| general | Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis | pdf |
| general | Instruction-based Image Manipulation by Watching How Things Move | pdf |
| general | Ferret: An Efficient Online Continual Learning Framework under Varying Memory Constraints | pdf |
| general | VidComposition: Can MLLMs Analyze Compositions in Compiled Videos? | pdf |
| general | Self-Supervised Learning for Color Spike Camera Reconstruction | pdf |
| general | From Elements to Design: A Layered Approach for Automatic Graphic Design Composition | pdf |
| general | SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis | pdf |
| general | DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers | pdf |
| general | Towards Lossless Implicit Neural Representation via Bit Plane Decomposition | pdf |
| gsplat | iSegMan: Interactive Segment-and-Manipulate 3D Gaussians | pdf |
| general | BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices | pdf |
| general | Unraveling Normal Anatomy via Fluid-Driven Anomaly Randomization | pdf |
| general | Taming Teacher Forcing for Masked Autoregressive Video Generation | pdf |
| general | Revisiting Backdoor Attacks against Large Vision-Language Models from Domain Shift | pdf |
| general | TCFG: Tangential Damping Classifier-free Guidance | pdf |
| general | MatAnyone: Stable Video Matting with Consistent Memory Propagation | pdf |
| general | Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models | pdf |
| general | MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception | pdf |
| general | T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation | pdf |
| general | Multimodal Autoregressive Pre-training of Large Vision Encoders | pdf |
| general | AKiRa: Augmentation Kit on Rays for Optical Video Generation | pdf |
| general | TFCustom: Customized Image Generation with Time-Aware Frequency Feature Guidance | pdf |
| general | SketchFusion: Learning Universal Sketch Features through Fusing Foundation Models | pdf |
| general | Bridging the Vision-Brain Gap with an Uncertainty-Aware Blur Prior | pdf |
| general | AffordDP: Generalizable Diffusion Policy with Transferable Affordance | pdf |
| general | HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation | pdf |
| general | DKC: Differentiated Knowledge Consolidation for Cloth-Hybrid Lifelong Person Re-identification | pdf |
| general | Enhancing Facial Privacy Protection via Weakening Diffusion Purification | pdf |
| general | ORIDa: Object-centric Real-world Image Composition Dataset | pdf |
| general | Image Generation Diversity Issues and How to Tame Them | pdf |
| general | Annotation Ambiguity Aware Semi-Supervised Medical Image Segmentation | pdf |
| beings | CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models | pdf |
| beings | CORE4D: A 4D Human-Object-Human Interaction Dataset for Collaborative Object REarrangement | pdf |
| gsplat | POp-GS: Next Best View in 3D-Gaussian Splatting with P-Optimality | pdf |
| general | Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning | pdf |
| general | MaRI: Material Retrieval Integration across Domains | pdf |
| general | Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs | pdf |
| general | Glossy Object Reconstruction with Cost-effective Polarized Acquisition | pdf |
| general | L-SWAG: Layer-Sample Wise Activation with Gradients Information for Zero-Shot NAS on Vision Transformers | pdf |
| general | Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning | pdf |
| general | Lifelong Knowledge Editing for Vision Language Models with Low-Rank Mixture-of-Experts | pdf |
| general | PartGen: Part-level 3D Generation and Reconstruction with Multi-view Diffusion Models | pdf |
| general | SINR: Sparsity Driven Compressed Implicit Neural Representations | pdf |
| general | ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning | pdf |
| general | Shining Yourself: High-Fidelity Ornaments Virtual Try-on with Diffusion Model | pdf |
| general | Universal Domain Adaptation for Semantic Segmentation | pdf |
| gsplat | HyperGS: Hyperspectral 3D Gaussian Splatting | pdf |
| general | LMO: Linear Mamba Operator for MRI Reconstruction | pdf |
| general | AnomalyNCD: Towards Novel Anomaly Class Discovery in Industrial Scenarios | pdf |
| general | Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding | pdf |
| general | Ges3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding | pdf |
| general | Progressive Focused Transformer for Single Image Super-Resolution | pdf |
| general | VladVA: Discriminative Fine-tuning of LVLMs | pdf |
| beings | HumanMM: Global Human Motion Recovery from Multi-shot Videos | pdf |
| general | Removing Reflections from RAW Photos | pdf |
| general | AMR-Transformer: Enabling Efficient Long-range Interaction for Complex Neural Fluid Simulation | pdf |
| general | Blurry-Edges: Photon-Limited Depth Estimation from Defocused Boundaries | pdf |
| general | MICAS: Multi-grained In-Context Adaptive Sampling for 3D Point Cloud Processing | pdf |
| general | GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving | pdf |
| motion | ColabSfM: Collaborative Structure-from-Motion by Point Cloud Registration | pdf |
| general | MangaNinja: Line Art Colorization with Precise Reference Following | pdf |
| gsplat | Nonisotropic Gaussian Diffusion for Realistic 3D Human Motion Prediction | pdf |
| general | PICO: Reconstructing 3D People In Contact with Objects | pdf |
| general | Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition | pdf |
| general | Scaling up Image Segmentation across Data and Tasks | pdf |
| general | Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning | pdf |
| general | Blood Flow Speed Estimation with Optical Coherence Tomography Angiography Images | pdf |
| general | DreamTrack: Dreaming the Future for Multimodal Visual Object Tracking | pdf |
| general | OmniStyle: Filtering High Quality Style Transfer Data at Scale | pdf |
| general | Cross-View Completion Models are Zero-shot Correspondence Estimators | pdf |
| general | Multi-party Collaborative Attention Control for Image Customization | pdf |
| general | HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos | pdf |
| general | DELT: A Simple Diversity-driven EarlyLate Training for Dataset Distillation | pdf |
| general | RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete | pdf |
| general | Beyond Image Classification: A Video Benchmark and Dual-Branch Hybrid Discrimination Framework for Compositional Zero-Shot Learning | pdf |
| general | ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image Objects | pdf |
| general | Do Your Best and Get Enough Rest for Continual Learning | pdf |
| general | Enhancing Few-Shot Class-Incremental Learning via Training-Free Bi-Level Modality Calibration | pdf |
| general | MUSt3R: Multi-view Network for Stereo 3D Reconstruction | pdf |
| general | Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Models | pdf |
| general | A New Statistical Model of Star Speckles for Learning to Detect and Characterize Exoplanets in Direct Imaging Observations | pdf |
| general | CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models | pdf |
| general | WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models | pdf |
| general | CXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Dataset | pdf |
| general | Event-Equalized Dense Video Captioning | pdf |
| general | EDCFlow: Exploring Temporally Dense Difference Maps for Event-based Optical Flow Estimation | pdf |
| general | LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions | pdf |
| general | Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues | pdf |
| general | Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition | pdf |
| gsplat | FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video | pdf |
| general | Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstruction | pdf |
| general | VITED: Video Temporal Evidence Distillation | pdf |
| general | Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts | pdf |
| motion | Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise | pdf |
| general | Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks | pdf |
| general | 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer | pdf |
| general | Boost the Inference with Co-training: A Depth-guided Mutual Learning Framework for Semi-supervised Medical Polyp Segmentation | pdf |
| worlds | From Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-Identification | pdf |
| general | 4Deform: Neural Surface Deformation for Robust Shape Interpolation | pdf |
| general | Dense Match Summarization for Faster Two-view Estimation | pdf |
| general | Align-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video Editing | pdf |
| general | LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant | pdf |
| motion | Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level | pdf |
| general | Toward Robust Neural Reconstruction from Sparse Point Sets | pdf |
| gsplat | GPAvatar: High-fidelity Head Avatars by Learning Efficient Gaussian Projections | pdf |
| general | PIAD: Pose and Illumination agnostic Anomaly Detection | pdf |
| general | Two is Better than One: Efficient Ensemble Defense for Robust and Compact Models | pdf |
| general | Tiled Diffusion | pdf |
| general | Descriptor-In-Pixel : Point-Feature Tracking For Pixel Processor Arrays | pdf |
| gsplat | UVGS: Reimagining Unstructured 3D Gaussian Splatting using UV Mapping | pdf |
| beings | InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation | pdf |
| general | TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition | pdf |
| general | BioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational Pathology | pdf |
| gsplat | GauCho: Gaussian Distributions with Cholesky Decomposition for Oriented Object Detection | pdf |
| general | No Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse Weather | pdf |
| general | Mind the Gap: Detecting Black-box Adversarial Attacks in the Making through Query Update Analysis | pdf |
| worlds | GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction | pdf |
| general | ICP: Immediate Compensation Pruning for Mid-to-high Sparsity | pdf |
| general | VinaBench: Benchmark for Faithful and Consistent Visual Narratives | pdf |
| general | Dual Diffusion for Unified Image Generation and Understanding | pdf |
| general | WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation | pdf |
| gsplat | 4DGC: Rate-Aware 4D Gaussian Compression for Efficient Streamable Free-Viewpoint Video | pdf |
| gsplat | GASP: Gaussian Avatars with Synthetic Priors | pdf |
| general | COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation | pdf |
| general | High-quality Point Cloud Oriented Normal Estimation via Hybrid Angular and Euclidean Distance Encoding | pdf |
| general | Prior-free 3D Object Tracking | pdf |
| general | Progressive Correspondence Regenerator for Robust 3D Registration | pdf |
| general | Cross-Modal 3D Representation with Multi-View Images and Point Clouds | pdf |
| general | Decompositional Neural Scene Reconstruction with Generative Diffusion Prior | pdf |
| general | Learning Visual Generative Priors without Text | pdf |
| general | Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference | pdf |
| general | Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach | pdf |
| general | Mr. DETR: Instructive Multi-Route Training for Detection Transformers | pdf |
| general | Hearing Hands: Generating Sounds from Physical Interactions in 3D Scenes | pdf |
| general | AirRoom: Objects Matter in Room Reidentification | pdf |
| general | DefMamba: Deformable Visual State Space Model | pdf |
| general | HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation | pdf |
| gsplat | VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction | pdf |
| beings | ControlFace: Harnessing Facial Parametric Control for Face Rigging | pdf |
| general | Imputation-free and Alignment-free: Incomplete Multi-view Clustering Driven by Consensus Semantic Learning | pdf |
| general | Sensitivity-Aware Efficient Fine-Tuning via Compact Dynamic-Rank Adaptation | pdf |
| general | A Selective Re-learning Mechanism for Hyperspectral Fusion Imaging | pdf |
| general | Autoregressive Sequential Pretraining for Visual Tracking | pdf |
| beings | PromptHMR: Promptable Human Mesh Recovery | pdf |
| general | VISTREAM: Improving Computation Efficiency of Visual Streaming Perception via Law-of-Charge-Conservation Inspired Spiking Neural Network | pdf |
| general | STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding | pdf |
| general | Rashomon Sets for Prototypical-Part Networks: Editing Interpretable Models in Real-Time | pdf |
| beings | EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty Modeling | pdf |
| general | Tuning the Frequencies: Robust Training for Sinusoidal Neural Networks | pdf |
| beings | Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures | pdf |
| general | Large Self-Supervised Models Bridge the Gap in Domain Adaptive Object Detection | pdf |
| general | Evaluating Model Perception of Color Illusions in Photorealistic Scenes | pdf |
| general | Do Visual Imaginations Improve Vision-and-Language Navigation Agents? | pdf |
| general | HotSpot: Signed Distance Function Optimization with an Asymptotically Sufficient Condition | pdf |
| general | Libra-Merging: Importance-redundancy and Pruning-merging Trade-off for Acceleration Plug-in in Large Vision-Language Model | pdf |
| general | LEDiff: Latent Exposure Diffusion for HDR Generation | pdf |
| general | VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation | pdf |
| general | Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion | pdf |
| general | Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach | pdf |
| general | SemiDAViL: Semi-supervised Domain Adaptation with Vision-Language Guidance for Semantic Segmentation | pdf |
| general | dFLMoE: Decentralized Federated Learning via Mixture of Experts for Medical Data Analysis | pdf |
| beings | Reconstructing Humans with a Biomechanically Accurate Skeleton | pdf |
| general | AdaCM^2: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction | pdf |
| general | VGGT: Visual Geometry Grounded Transformer | pdf |
| general | Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models | pdf |
| general | Visual Consensus Prompting for Co-Salient Object Detection | pdf |
| general | Quantization without Tears | pdf |
| worlds | PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric Videos | pdf |
| general | Towards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific Parameters | pdf |
| worlds | SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generation | pdf |
| general | HiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusion | pdf |
| general | FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis | pdf |
| general | RAD: Region-Aware Diffusion Models for Image Inpainting | pdf |
| general | Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space | pdf |
| gsplat | TexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian Splatting | pdf |
| general | Concept Replacer: Replacing Sensitive Concepts in Diffusion Models via Precision Localization | pdf |
| general | A Regularization-Guided Equivariant Approach for Image Restoration | pdf |
| general | Deep Fair Multi-View Clustering with Attention KAN | pdf |
| general | LineArt: A Knowledge-guided Training-free High-quality Appearance Transfer for Design Drawing with Diffusion Model | pdf |
| general | VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding | pdf |
| general | Zero-Shot Image Restoration Using Few-Step Guidance of Consistency Models (and Beyond) | pdf |
| general | Similarity-Guided Layer-Adaptive Vision Transformer for UAV Tracking | pdf |
| general | LidarGait++: Learning Local Features and Size Awareness from LiDAR Point Clouds for 3D Gait Recognition | pdf |
| general | Coarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Model | pdf |
| general | FoundationStereo: Zero-Shot Stereo Matching | pdf |
| general | UniNet: A Contrastive Learning-guided Unified Framework with Feature Selection for Anomaly Detection | pdf |
| general | MoEdit: On Learning Quantity Perception for Multi-object Image Editing | pdf |
| beings | Seeing More with Less: Human-like Representations in Vision Models | pdf |
| beings | Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identification | pdf |
| general | AeroGen: Enhancing Remote Sensing Object Detection with Diffusion-Driven Data Generation | pdf |
| general | Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning | pdf |
| general | Style Quantization for Data-Efficient GAN Training | pdf |
| general | Localizing Events in Videos with Multimodal Queries | pdf |
| general | PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability | pdf |
| general | CleanDIFT: Diffusion Features without Noise | pdf |
| general | MAD: Memory-Augmented Detection of 3D Objects | pdf |
| general | Doppelgangers and Adversarial Vulnerability | pdf |
| general | Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalance | pdf |
| general | Steady Progress Beats Stagnation: Mutual Aid of Foundation and Conventional Models in Mixed Domain Semi-Supervised Medical Image Segmentation | pdf |
| worlds | ARKit LabelMaker: A New Scale for Indoor 3D Scene Understanding | pdf |
| general | StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text | pdf |
| general | AFL: A Single-Round Analytic Approach for Federated Learning with Pre-trained Models | pdf |
| general | BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding | pdf |
| general | Reference-Based 3D-Aware Image Editing with Triplanes | pdf |
| general | One is Plenty: A Polymorphic Feature Interpreter for Immutable Heterogeneous Collaborative Perception | pdf |
| beings | SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories | pdf |
| general | SceneCrafter: Controllable Multi-View Driving Scene Editing | pdf |
| beings | HeatFormer: A Neural Optimizer for Multiview Human Mesh Recovery | pdf |
| general | GPS as a Control Signal for Image Generation | pdf |
| general | CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology | pdf |
| gsplat | MAGiC-SLAM: Multi-Agent Gaussian Globally Consistent SLAM | pdf |
| general | NTClick: Achieving Precise Interactive Segmentation With Noise-tolerant Clicks | pdf |
| general | MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model | pdf |
| gsplat | HiMoR: Monocular Deformable Gaussian Reconstruction with Hierarchical Motion Representation | pdf |
| general | Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute Recognition | pdf |
| beings | HSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene Interaction | pdf |
| beings | Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior | pdf |
| worlds | RoomPainter: View-Integrated Diffusion for Consistent Indoor Scene Texturing | pdf |
| worlds | IRIS: Inverse Rendering of Indoor Scenes from Low Dynamic Range Images | pdf |
| gsplat | RoGSplat: Learning Robust Generalizable Human Gaussian Splatting from Sparse Multi-View Images | pdf |
| general | EnliveningGS: Active Locomotion of 3DGS | pdf |
| general | Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition | pdf |
| general | LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic Segmentation | pdf |
| worlds | PhysGen3D: Crafting a Miniature Interactive World from a Single Image | pdf |
| general | Docopilot: Improving Multimodal Models for Document-Level Understanding | pdf |
| general | Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution | pdf |
| general | LATTE-MV: Learning to Anticipate Table Tennis Hits from Monocular Videos | pdf |
| general | Distilling Spatially-Heterogeneous Distortion Perception for Blind Image Quality Assessment | pdf |
| general | 2DMamba: Efficient State Space Model for Image Representation with Applications on Giga-Pixel Whole Slide Image Classification | pdf |
| general | Unboxed: Geometrically and Temporally Consistent Video Outpainting | pdf |
| beings | K-Sort Arena: Efficient and Reliable Benchmarking for Generative Models via K-wise Human Preferences | pdf |
| motion | Dense-SfM: Structure from Motion with Dense Consistent Matching | pdf |
| general | Sketchy Bounding-box Supervision for 3D Instance Segmentation | pdf |
| general | StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models | pdf |
| beings | Learning Person-Specific Animatable Face Models from In-the-Wild Images via a Shared Base Model | pdf |
| beings | TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion Generation | pdf |
| general | Hybrid Reciprocal Transformer with Triplet Feature Alignment for Scene Graph Generation | pdf |
| general | Gradient Inversion Attacks on Parameter-Efficient Fine-Tuning | pdf |
| general | UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation | pdf |
| general | MBQ: Modality-Balanced Quantization for Large Vision-Language Models | pdf |
| general | VideoDPO: Omni-Preference Alignment for Video Diffusion Generation | pdf |
| general | Associative Transformer | pdf |
| beings | ChatGarment: Garment Estimation, Generation and Editing via Large Language Models | pdf |
| general | RDD: Robust Feature Detector and Descriptor using Deformable Transformer | pdf |
| general | Building Vision Models upon Heat Conduction | pdf |
| general | LT3SD: Latent Trees for 3D Scene Diffusion | pdf |
| meshing | CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry Refiner | pdf |
| beings | GIF: Generative Inspiration for Face Recognition at Scale | pdf |
| general | SKDream: Controllable Multi-view and 3D Generation with Arbitrary Skeletons | pdf |
| general | Optimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policy | pdf |
| worlds | Classic Video Denoising in a Machine Learning World: Robust, Fast, and Controllable | pdf |
| general | Population Normalization for Federated Learning | pdf |
| general | RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety | pdf |
| general | ESCAPE: Equivariant Shape Completion via Anchor Point Encoding | pdf |
| general | Satellite to GroundScape - Large-scale Consistent Ground View Generation from Satellite Views | pdf |
| general | Variance-Based Membership Inference Attacks Against Large-Scale Image Captioning Models | pdf |
| general | Learning Dynamic Collaborative Network for Semi-supervised 3D Vessel Segmentation | pdf |
| general | Temporal Alignment-Free Video Matching for Few-shot Action Recognition | pdf |
| general | OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP | pdf |
| general | VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary | pdf |
| general | Forensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language Models | pdf |
| general | Forensic Self-Descriptions Are All You Need for Zero-Shot Detection, Open-Set Source Attribution, and Clustering of AI-generated Images | pdf |
| gsplat | FlexDrive: Toward Trajectory Flexibility in Driving Scene Gaussian Splatting Reconstruction and Rendering | pdf |
| gsplat | Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs | pdf |
| general | Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering | pdf |
| general | DifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolution | pdf |
| gsplat | 3D-GSW: 3D Gaussian Splatting for Robust Watermarking | pdf |
| general | OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation | pdf |
| general | Dual Exposure Stereo for Extended Dynamic Range 3D Imaging | pdf |
| general | PAVE: Patching and Adapting Video Large Language Models | pdf |
| general | Generative Image Layer Decomposition with Visual Effects | pdf |
| general | AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion | pdf |
| general | Revisiting Audio-Visual Segmentation with Vision-Centric Transformer | pdf |
| general | HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language Models | pdf |
| general | Taxonomy-Aware Evaluation of Vision-Language Models | pdf |
| general | Active Event-based Stereo Vision | pdf |
| motion | SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning | pdf |
| general | Video Language Model Pretraining with Spatio-temporal Masking | pdf |
| general | Synthetic Data is an Elegant GIFT for Continual Vision-Language Models | pdf |
| general | Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model | pdf |
| general | Hypergraph Vision Transformers: Images are More than Nodes, More than Edges | pdf |
| general | Binarized Neural Network for Multi-spectral Image Fusion | pdf |
| gsplat | GaussianIP: Identity-Preserving Realistic 3D Human Generation via Human-Centric Diffusion Prior | pdf |
| general | FineVQ: Fine-Grained User Generated Content Video Quality Assessment | pdf |
| general | Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly | pdf |
| general | MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation | pdf |
| general | METASCENES: Towards Automated Replica Creation for Real-world 3D Scans | pdf |
| general | Robust Multimodal Survival Prediction with Conditional Latent Differentiation Variational AutoEncoder | pdf |
| general | Zero-Shot Blind-spot Image Denoising via Implicit Neural Sampling | pdf |
| general | Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation | pdf |
| beings | InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing | pdf |
| general | Wonderland: Navigating 3D Scenes from a Single Image | pdf |
| general | Towards Satellite Image Road Graph Extraction: A Global-Scale Dataset and A Novel Method | pdf |
| general | SuperLightNet: Lightweight Parameter Aggregation Network for Multimodal Brain Tumor Segmentation | pdf |
| general | Self-Supervised Spatial Correspondence Across Modalities | pdf |
| general | MOS-Attack: A Scalable Multi-objective Adversarial Attack Framework | pdf |
| motion | Motion Modes: What Could Happen Next? | pdf |
| general | Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation | pdf |
| general | The Change You Want To Detect: Semantic Change Detection In Earth Observation With Hybrid Data Generationf | pdf |
| general | Weakly Supervised Semantic Segmentation via Progressive Confidence Region Expansion | pdf |
| general | RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations | pdf |
| general | HOTFormerLoc: Hierarchical Octree Transformer for Versatile Lidar Place Recognition Across Ground and Aerial Views | pdf |
| general | UniK3D: Universal Camera Monocular 3D Estimation | pdf |
| motion | ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer | pdf |
| general | AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification | pdf |
| general | EBS-EKF: Accurate and High Frequency Event-based Star Tracking | pdf |
| worlds | Benchmarking Object Detectors under Real-World Distribution Shifts in Satellite Imagery | pdf |
| general | SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image Pretraining | pdf |
| general | Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models | pdf |
| general | CoLLM: A Large Language Model for Composed Image Retrieval | pdf |
| general | GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning | pdf |
| general | MMVU: Measuring Expert-Level Multi-Discipline Video Understanding | pdf |
| motion | EgoLM: Multi-Modal Language Model of Egocentric Motions | pdf |
| general | Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation | pdf |
| general | Learning Heterogeneous Tissues with Mixture of Experts for Gigapixel Whole Slide Images | pdf |
| general | Disentangled Pose and Appearance Guidance for Multi-Pose Generation | pdf |
| general | Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-Mismatch | pdf |
| general | Electromyography-Informed Facial Expression Reconstruction for Physiological-Based Synthesis and Analysis | pdf |
| beings | Improving the Transferability of Adversarial Attacks on Face Recognition with Diverse Parameters Augmentation | pdf |
| general | Adapting to Observation Length of Trajectory Prediction via Contrastive Learning | pdf |
| general | Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation | pdf |
| general | NitroFusion: High-Fidelity Single-Step Diffusion through Dynamic Adversarial Training | pdf |
| general | CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based Framework | pdf |
| general | RC-AutoCalib: An End-to-End Radar-Camera Automatic Calibration Network | pdf |
| general | Argus: A Compact and Versatile Foundation Model for Vision | pdf |
| general | Sampling Innovation-Based Adaptive Compressive Sensing | pdf |
| motion | MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models | pdf |
| general | ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation | pdf |
| general | ArcPro: Architectural Programs for Structured 3D Abstraction of Sparse Points | pdf |
| gsplat | Hardware-Rasterized Ray-Based Gaussian Splatting | pdf |
| general | FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasing | pdf |
| general | Multi-subject Open-set Personalization in Video Generation | pdf |
| motion | Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation | pdf |
| general | FG^2: Fine-Grained Cross-View Localization by Fine-Grained Feature Matching | pdf |
| general | Distilled Prompt Learning for Incomplete Multimodal Survival Prediction | pdf |
| general | Learning Conditional Space-Time Prompt Distributions for Video Class-Incremental Learning | pdf |
| general | Hyperbolic Safety-Aware Vision-Language Models | pdf |
| gsplat | SinGS: Animatable Single-Image Human Gaussian Splats with Kinematic Priors | pdf |
| general | Multi-Modal Synergistic Implicit Image Enhancement for Efficient Optical Flow Estimation | pdf |
| general | Generating Multimodal Driving Scenes via Next-Scene Prediction | pdf |
| general | Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generation | pdf |
| general | PosterMaker: Towards High-Quality Product Poster Generation with Accurate Text Rendering | pdf |
| general | Rethinking Personalized Aesthetics Assessment: Employing Physique Aesthetics Assessment as An Exemplification | pdf |
| general | You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale | pdf |
| general | PEACE: Empowering Geologic Map Holistic Understanding with MLLMs | pdf |
| general | ConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigation | pdf |
| general | MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation | pdf |
| general | ERUPT: Efficient Rendering with Unposed Patch Transformer | pdf |
| gsplat | Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splatting | pdf |
| general | Quad-Pixel Image Defocus Deblurring: A New Benchmark and Model | pdf |
| general | GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency | pdf |
| general | Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics | pdf |
| general | ESC: Erasing Space Concept for Knowledge Deletion | pdf |
| general | Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks | pdf |
| general | Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference Systems | pdf |
| general | Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution | pdf |
| general | Interpretable Generative Models through Post-hoc Concept Bottlenecks | pdf |
| general | Watermarking One for All: A Robust Watermarking Scheme Against Partial Image Theft | pdf |
| general | Minding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion Segmentation | pdf |
| motion | IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Manner | pdf |
| general | Link-based Contrastive Learning for One-Shot Unsupervised Domain Adaptation | pdf |
| general | UniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object Detection | pdf |
| general | Rethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense Game | pdf |
| general | VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide | pdf |
| general | OmniStereo: Real-time Omnidireactional Depth Estimation with Multiview Fisheye Cameras | pdf |
| gsplat | DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone Imagery | pdf |
| general | SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction | pdf |
| beings | Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency | pdf |
| general | IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation | pdf |
| general | LoTUS: Large-Scale Machine Unlearning with a Taste of Uncertainty | pdf |
| general | SleeperMark: Towards Robust Watermark against Fine-Tuning Text-to-image Diffusion Models | pdf |
| general | Illumination Spectrum Estimation for Multispectral Images via Surface Reflectance Modeling and Spatial-Spectral Feature Generation | pdf |
| general | Towards Explainable and Unprecedented Accuracy in Matching Challenging Finger Crease Patterns | pdf |
| general | Neural Hierarchical Decomposition for Single Image Plant Modeling | pdf |
| general | Dual-Agent Optimization framework for Cross-Domain Few-Shot Segmentation | pdf |
| general | SACB-Net: Spatial-awareness Convolutions for Medical Image Registration | pdf |
| general | Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps | pdf |
| general | DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion | pdf |
| beings | AIpparel: A Multimodal Foundation Model for Digital Garments | pdf |
| general | PO3AD: Predicting Point Offsets toward Better 3D Point Cloud Anomaly Detection | pdf |
| general | ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models | pdf |
| general | PreciseCam: Precise Camera Control for Text-to-Image Generation | pdf |
| general | SET: Spectral Enhancement for Tiny Object Detection | pdf |
| general | Differentiable Inverse Rendering with Interpretable Basis BRDFs | pdf |
| general | EquiPose: Exploiting Permutation Equivariance for Relative Camera Pose Estimation | pdf |
| beings | Face Forgery Video Detection via Temporal Forgery Cue Unraveling | pdf |
| general | Temporally Consistent Object-Centric Learning by Contrasting Slots | pdf |
| general | MC^2: Multi-concept Guidance for Customized Multi-concept Generation | pdf |
| general | Multi-modal Vision Pre-training for Medical Image Analysis | pdf |
| general | STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training | pdf |
| general | LIM: Large Interpolator Model for Dynamic Reconstruction | pdf |
| general | AutoPresent: Designing Structured Visuals from Scratch | pdf |
| worlds | VisionArena: 230k Real World User-VLM Conversations with Preference Labels | pdf |
| general | FAM Diffusion: Frequency and Attention Modulation for High-Resolution Image Generation with Stable Diffusion | pdf |
| beings | MultiGO: Towards Multi-level Geometry Learning for Monocular 3D Textured Human Reconstruction | pdf |
| general | Generative Photomontage | pdf |
| general | Multi-view Reconstruction via SfM-guided Monocular Depth Estimation | pdf |
| beings | HuMoCon: Concept Discovery for Human Motion Understanding | pdf |
| worlds | FreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free Prompts | pdf |
| general | Rethinking Correspondence-based Category-Level Object Pose Estimation | pdf |
| general | Curriculum Direct Preference Optimization for Diffusion and Consistency Models | pdf |
| general | Personalized Preference Fine-tuning of Diffusion Models | pdf |
| general | NN-Former: Rethinking Graph Structure in Neural Architecture Representation | pdf |
| general | A Unified Image-Dense Annotation Generation Model for Underwater Scenes | pdf |
| gsplat | NTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on Thermodynamics | pdf |
| general | FSHNet: Fully Sparse Hybrid Network for 3D Object Detection | pdf |
| general | JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV Systems | pdf |
| worlds | HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos | pdf |
| general | High Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous Flight | pdf |
| worlds | Generative Gaussian Splatting for Unbounded 3D City Generation | pdf |
| general | GeoMM: On Geodesic Perspective for Multi-modal Learning | pdf |
| general | VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning | pdf |
| general | Adaptive Dropout: Unleashing Dropout across Layers for Generalizable Image Super-Resolution | pdf |
| general | Breaking the Memory Barrier of Contrastive Loss via Tile-Based Strategy | pdf |
| general | Learning Phase Distortion with Selective State Space Models for Video Turbulence Mitigation | pdf |
| general | RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training | pdf |
| general | Distraction is All You Need for Multimodal Large Language Model Jailbreaking | pdf |
| general | Learning to Normalize on the SPD Manifold under Bures-Wasserstein Geometry | pdf |
| general | SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation | pdf |
| general | BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance | pdf |
| beings | FinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal Guidance | pdf |
| general | SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration | pdf |
| general | Beyond Generation: A Diffusion-based Low-level Feature Extractor for Detecting AI-generated Images | pdf |
| general | Optical-Flow Guided Prompt Optimization for Coherent Video Generation | pdf |
| general | Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection | pdf |
| general | Rethinking Reconstruction and Denoising in the Dark: New Perspective, General Architecture and Beyond | pdf |
| general | Federated Learning with Domain Shift Eraser | pdf |
| general | DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation | pdf |
| beings | Link to the Past: Temporal Propagation for Fast 3D Human Reconstruction from Monocular Video | pdf |
| general | Deterministic Certification of Graph Neural Networks against Graph Poisoning Attacks with Arbitrary Perturbations | pdf |
| general | A3: Few-shot Prompt Learning of Unlearnable Examples with Cross-Modal Adversarial Feature Alignment | pdf |
| general | MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3D | pdf |
| general | ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding | pdf |
| motion | MotiF: Making Text Count in Image Animation with Motion Focal Loss | pdf |
| general | Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse Weather | pdf |
| gsplat | MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks | pdf |
| general | SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding | pdf |
| general | CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation | pdf |
| general | Navigating Image Restoration with VAR's Distribution Alignment Prior | pdf |
| general | Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability | pdf |
| general | Graph Neural Network Combining Event Stream and Periodic Aggregation for Low-Latency Event-based Vision | pdf |
| general | ArtFormer: Controllable Generation of Diverse 3D Articulated Objects | pdf |
| general | Bridging Gait Recognition and Large Language Models Sequence Modeling | pdf |
| beings | DiSRT-In-Bed: Diffusion-Based Sim-to-Real Transfer Framework for In-Bed Human Mesh Recovery | pdf |
| general | COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts | pdf |
| general | HOT: Hadamard-based Optimized Training | pdf |
| general | TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation | pdf |
| general | SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device | pdf |
| general | Adapting Dense Matching for Homography Estimation with Grid-based Acceleration | pdf |
| general | CoSDH: Communication-Efficient Collaborative Perception via Supply-Demand Awareness and Intermediate-Late Hybridization | pdf |
| general | Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail | pdf |
| general | Order-Robust Class Incremental Learning: Graph-Driven Dynamic Similarity Grouping | pdf |
| general | Where the Devil Hides: Deepfake Detectors Can No Longer Be Trusted | pdf |
| general | CaMuViD: Calibration-Free Multi-View Detection | pdf |
| general | Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing | pdf |
| general | HVI: A New Color Space for Low-light Image Enhancement | pdf |
| general | DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction | pdf |
| general | One Diffusion to Generate Them All | pdf |
| general | CoSER: Towards Consistent Dense Multiview Text-to-Image Generator for 3D Creation | pdf |
| general | UNEM: UNrolled Generalized EM for Transductive Few-Shot Learning | pdf |
| general | G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation | pdf |
| general | Explaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classification | pdf |
| gsplat | Textured Gaussians for Enhanced 3D Scene Appearance Modeling | pdf |
| general | NeighborRetr: Balancing Hub Centrality in Cross-Modal Retrieval | pdf |
| worlds | Global-Local Tree Search in VLMs for 3D Indoor Scene Generation | pdf |
| general | GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow Networks | pdf |
| general | MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining | pdf |
| general | Segment Any-Quality Images with Generative Latent Space Enhancement | pdf |
| worlds | CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos | pdf |
| general | Learning Visual Composition through Improved Semantic Guidance | pdf |
| general | JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation | pdf |
| general | Visual Prompting for One-shot Controllable Video Editing without Inversion | pdf |
| general | AVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual Learning | pdf |
| general | Flash-Split: 2D Reflection Removal with Flash Cues and Latent Diffusion Separation | pdf |
| general | Attention IoU: Examining Biases in CelebA using Attention Maps | pdf |
| general | HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator | pdf |
| motion | Segment Any Motion in Videos | pdf |
| general | Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding | pdf |
| general | PatchDEMUX: A Certifiably Robust Framework for Multi-label Classifiers Against Adversarial Patches | pdf |
| general | EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering | pdf |
| general | Token Cropr: Faster ViTs for Quite a Few Tasks | pdf |
| general | STCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Prediction | pdf |
| general | Resilient Sensor Fusion Under Adverse Sensor Failures via Multi-Modal Expert Fusion | pdf |
| general | MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing | pdf |
| worlds | IndoorGS: Geometric Cues Guided Gaussian Splatting for Indoor Scene Reconstruction | pdf |
| general | Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis | pdf |
| general | Stop Walking in Circles! Bailing Out Early in Projected Gradient Descent | pdf |
| general | MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation | pdf |
| general | Stable-SCore: A Stable Registration-based Framework for 3D Shape Correspondence | pdf |
| general | Beyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and Harmonization | pdf |
| general | Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Large Model Enhancement | pdf |
| general | Pose Priors from Language Models | pdf |
| general | LogoSP: Local-global Grouping of Superpoints for Unsupervised Semantic Segmentation of 3D Point Clouds | pdf |
| general | Exploring Intrinsic Normal Prototypes within a Single Image for Universal Anomaly Detection | pdf |
| general | Augmenting Perceptual Super-Resolution via Image Quality Predictors | pdf |
| general | TurboFill: Adapting Few-step Text-to-image Model for Fast Image Inpainting | pdf |
| beings | Stochastic Human Motion Prediction with Memory of Action Transition and Action Characteristic | pdf |
| general | Perception Tokens Enhance Visual Reasoning in Multimodal Language Models | pdf |
| beings | X-Dyna: Expressive Dynamic Human Image Animation | pdf |
| general | Towards Effective and Sparse Adversarial Attack on Spiking Neural Networks via Breaking Invisible Surrogate Gradients | pdf |
| general | Towards Understanding and Quantifying Uncertainty for Text-to-Image Generation | pdf |
| general | LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding | pdf |
| general | Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models | pdf |
| general | MaIR: A Locality- and Continuity-Preserving Mamba for Image Restoration | pdf |
| general | RSAR: Restricted State Angle Resolver and Rotated SAR Benchmark | pdf |
| motion | Continuous Space-Time Video Resampling with Invertible Motion Steganography | pdf |
| general | ProtoDepth: Unsupervised Continual Depth Completion with Prototypes | pdf |
| beings | ParaHome: Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactions | pdf |
| general | Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds | pdf |
| beings | OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation | pdf |
| general | Track Any Anomalous Object:A Granular Video Anomaly Detection Pipeline | pdf |
| general | Object-aware Sound Source Localization via Audio-Visual Scene Understanding | pdf |
| general | SerialGen: Personalized Image Generation by First Standardization Then Personalization | pdf |
| general | Augmented Deep Contexts for Spatially Embedded Video Coding | pdf |
| general | Proximal Algorithm Unrolling: Flexible and Efficient Reconstruction Networks for Single-Pixel Imaging | pdf |
| beings | Image Quality Assessment: From Human to Machine Preference | pdf |
| general | Context-Aware Multimodal Pretraining | pdf |
| general | Task-driven Image Fusion with Learnable Fusion Loss | pdf |
| general | LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant | pdf |
| general | CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation | pdf |
| general | MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision | pdf |
| general | Beyond Background Shift: Rethinking Instance Replay in Continual Semantic Segmentation | pdf |
| general | ScaleLSD: Scalable Deep Line Segment Detection Streamlined | pdf |
| general | Revisiting MAE Pre-training for 3D Medical Image Segmentation | pdf |
| beings | ChatHuman: Chatting about 3D Humans with Tools | pdf |
| general | Scalable Autoregressive Monocular Depth Estimation | pdf |
| general | Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval | pdf |
| general | Camouflage Anything: Learning to Hide using Controlled Out-painting and Representation Engineering | pdf |
| general | Test-Time Fine-Tuning of Image Compression Models for Multi-Task Adaptability | pdf |
| beings | DynPose: Largely Improving the Efficiency of Human Pose Estimation by a Simple Dynamic Framework | pdf |
| general | VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos | pdf |
| general | Distinguish Then Exploit: Source-free Open Set Domain Adaptation via Weight Barcode Estimation and Sparse Label Assignment | pdf |
| general | Multi-Sensor Object Anomaly Detection: Unifying Appearance, Geometry, and Internal Properties | pdf |
| meshing | Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation | pdf |
| gsplat | PlanarSplatting: Accurate Planar Surface Reconstruction in 3 Minutes | pdf |
| general | Omni-ID: Holistic Identity Representation Designed for Generative Tasks | pdf |
| general | MIRE: Matched Implicit Neural Representations | pdf |
| general | AeSPa : Attention-guided Self-supervised Parallel Imaging for MRI Reconstruction | pdf |
| general | RobSense: A Robust Multi-modal Foundation Model for Remote Sensing with Static, Temporal, and Incomplete Data Adaptability | pdf |
| gsplat | MAC-Ego3D: Multi-Agent Gaussian Consensus for Real-Time Collaborative Ego-Motion and Photorealistic 3D Reconstruction | pdf |
| general | Online Video Understanding: OVBench and VideoChat-Online | pdf |
| general | LightLoc: Learning Outdoor LiDAR Localization at Light Speed | pdf |
| general | Accurate Differential Operators for Hybrid Neural Fields | pdf |
| general | FeedEdit: Text-Based Image Editing with Dynamic Feedback Regulation | pdf |
| general | Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification | pdf |
| beings | UMotion: Uncertainty-driven Human Motion Estimation from Inertial and Ultra-wideband Units | pdf |
| general | Scene Map-based Prompt Tuning for Navigation Instruction Generation | pdf |
| gsplat | DropoutGS: Dropping Out Gaussians for Better Sparse-view Rendering | pdf |
| general | Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language Models | pdf |
| general | Enhancing Dataset Distillation via Non-Critical Region Refinement | pdf |
| gsplat | PUP 3D-GS: Principled Uncertainty Pruning for 3D Gaussian Splatting | pdf |
| worlds | ScribbleLight: Single Image Indoor Relighting with Scribbles | pdf |
| general | InsightEdit: Towards Better Instruction Following for Image Editing | pdf |
| general | One-for-More: Continual Diffusion Model for Anomaly Detection | pdf |
| general | Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks | pdf |
| general | EDM: Equirectangular Projection-Oriented Dense Kernelized Feature Matching | pdf |
| general | EZSR: Event-based Zero-Shot Recognition | pdf |
| beings | SVFR: A Unified Framework for Generalized Video Face Restoration | pdf |
| general | Decoupling Fine Detail and Global Geometry for Compressed Depth Map Super-Resolution | pdf |
| general | Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset | pdf |
| meshing | MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation | pdf |
| beings | DeClotH: Decomposable 3D Cloth and Human Body Reconstruction from a Single Image | pdf |
| general | Neuron: Learning Context-Aware Evolving Representations for Zero-Shot Skeleton Action Recognition | pdf |
| motion | High-Fidelity Relightable Monocular Portrait Animation with Lighting-Controllable Video Diffusion Model | pdf |
| general | Plug-and-Play PPO: An Adaptive Point Prompt Optimizer Making SAM Greater | pdf |
| general | EchoONE: Segmenting Multiple Echocardiography Planes in One Model | pdf |
| general | EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild | pdf |
| general | Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception | pdf |
| worlds | PosterO: Structuring Layout Trees to Enable Language Models in Generalized Content-Aware Layout Generation | pdf |
| general | One2Any: One-Reference 6D Pose Estimation for Any Object | pdf |
| general | Contextual AD Narration with Interleaved Multimodal Sequence | pdf |
| general | MNE-SLAM: Multi-Agent Neural SLAM for Mobile Robots | pdf |
| general | TensoFlow: Tensorial Flow-based Sampler for Inverse Rendering | pdf |
| general | FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering | pdf |
| general | LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences | pdf |
| general | Exploring Temporally-Aware Features for Point Tracking | pdf |
| general | V^2Dial: Unification of Video and Visual Dialog via Multimodal Experts | pdf |
| general | Detail-Preserving Latent Diffusion for Stable Shadow Removal | pdf |
| general | CrossOver: 3D Scene Cross-Modal Alignment | pdf |
| general | Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction | pdf |
| general | Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents | pdf |
| general | Tokenize Image Patches: Global Context Fusion for Effective Haze Removal in Large Images | pdf |
| general | ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction | pdf |
| general | MET3R: Measuring Multi-View Consistency in Generated Images | pdf |
| general | Segmenting Maxillofacial Structures in CBCT Volumes | pdf |
| general | 3D Dental Model Segmentation with Geometrical Boundary Preserving | pdf |
| general | VideoGigaGAN: Towards Detail-rich Video Super-Resolution | pdf |
| general | GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation | pdf |
| general | Towards RAW Object Detection in Diverse Conditions | pdf |
| general | FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training | pdf |
| general | Adapter Merging with Centroid Prototype Mapping for Scalable Class-Incremental Learning | pdf |
| worlds | OpenSDI: Spotting Diffusion-Generated Images in the Open World | pdf |
| general | Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset | pdf |
| general | DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention | pdf |
| gsplat | Monocular and Generalizable Gaussian Talking Head Animation | pdf |
| general | Locally Orderless Images for Optimization in Differentiable Rendering | pdf |
| general | Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Control | pdf |
| general | Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation Models | pdf |
| general | DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models | pdf |
| beings | ALIEN: Implicit Neural Representations for Human Motion Prediction under Arbitrary Latency | pdf |
| general | Sufficient Invariant Learning for Distribution Shift | pdf |
| general | Domain Generalization in CLIP via Learning with Diverse Text Prompts | pdf |
| general | IterIS: Iterative Inference-Solving Alignment for LoRA Merging | pdf |
| general | Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacement | pdf |
| general | PrEditor3D: Fast and Precise 3D Shape Editing | pdf |
| general | ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices | pdf |
| general | LOCORE: Image Re-ranking with Long-Context Sequence Modeling | pdf |
| general | LiVOS: Light Video Object Segmentation with Gated Linear Matching | pdf |
| general | Polarized Color Screen Matting | pdf |
| general | GOAL: Global-local Object Alignment Learning | pdf |
| general | Post-pre-training for Modality Alignment in Vision-Language Foundation Models | pdf |
| beings | SynthLight: Portrait Relighting with Diffusion Model by Learning to Re-render Synthetic Faces | pdf |
| general | Pseudo Visible Feature Fine-Grained Fusion for Thermal Object Detection | pdf |
| general | NVILA: Efficient Frontier Visual Language Models | pdf |
| general | SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting | pdf |
| general | NoPain: No-box Point Cloud Attack via Optimal Transport Singular Boundary | pdf |
| beings | FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation | pdf |
| gsplat | Geometry Field Splatting with Gaussian Surfels | pdf |
| general | PS-EIP: Robust Photometric Stereo Based on Event Interval Profile | pdf |
| general | GenPC: Zero-shot Point Cloud Completion via 3D Generative Priors | pdf |
| general | Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis | pdf |
| general | Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking Transformers | pdf |
| general | HeMoRa: Unsupervised Heuristic Consensus Sampling for Robust Point Cloud Registration | pdf |
| general | Reducing Class-wise Confusion for Incremental Learning with Disentangled Manifolds | pdf |
| motion | AniMo: Species-Aware Model for Text-Driven Animal Motion Generation | pdf |
| general | EditAR: Unified Conditional Generation with Autoregressive Models | pdf |
| general | Instance-wise Supervision-level Optimization in Active Learning | pdf |
| general | BHViT: Binarized Hybrid Vision Transformer | pdf |
| general | Pathways on the Image Manifold: Image Editing via Video Generation | pdf |
| gsplat | DeSplat: Decomposed Gaussian Splatting for Distractor-Free Rendering | pdf |
| general | Stable Flow: Vital Layers for Training-Free Image Editing | pdf |
| beings | TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation | pdf |
| general | CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Tools | pdf |
| general | Conditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generation | pdf |
| beings | KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation | pdf |
| general | Context-Enhanced Memory-Refined Transformer for Online Action Detection | pdf |
| general | Data-Free Group-Wise Fully Quantized Winograd Convolution via Learnable Scales | pdf |
| worlds | GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control | pdf |
| general | A Dataset for Semantic Segmentation in the Presence of Unknowns | pdf |
| general | HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding | pdf |
| general | CASP: Compression of Large Multimodal Models Based on Attention Sparsity | pdf |
| general | UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation | pdf |
| general | Towards Cost-Effective Learning: A Synergy of Semi-Supervised and Active Learning | pdf |
| general | Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis | pdf |
| general | Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 Dataset | pdf |
| gsplat | EnvGS: Modeling View-Dependent Appearance with Environment Gaussian | pdf |
| general | Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning | pdf |
| general | MonoDGP: Monocular 3D Object Detection with Decoupled-Query and Geometry-Error Priors | pdf |
| general | Flexible Group Count Enables Hassle-Free Structured Pruning | pdf |
| beings | EasyCraft: A Robust and Efficient Framework for Automatic Avatar Crafting | pdf |
| meshing | MeshArt: Generating Articulated Meshes with Structure-Guided Transformers | pdf |
| general | Adaptive Non-Uniform Timestep Sampling for Accelerating Diffusion Model Training | pdf |
| general | Explainable Saliency: Articulating Reasoning with Contextual Prioritization | pdf |
| general | Compass Control: Multi Object Orientation Control for Text-to-Image Generation | pdf |
| general | Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation | pdf |
| general | VideoGEM: Training-free Action Grounding in Videos | pdf |
| motion | Structure-from-Motion with a Non-Parametric Camera Model | pdf |
| beings | LAL: Enhancing 3D Human Motion Prediction with Latency-aware Auxiliary Learning | pdf |
| general | RefPose: Leveraging Reference Geometric Correspondences for Accurate 6D Pose Estimation of Unseen Objects | pdf |
| general | Relation3D : Enhancing Relation Modeling for Point Cloud Instance Segmentation | pdf |
| beings | Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions | pdf |
| worlds | Beyond Human Perception: Understanding Multi-Object World from Monocular View | pdf |
| gsplat | LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fields | pdf |
| general | Latent Space Super-Resolution for Higher-Resolution Image Generation with Diffusion Models | pdf |
| general | Rethinking Noisy Video-Text Retrieval via Relation-aware Alignment | pdf |
| general | Scaling Vision Pre-Training to 4K Resolution | pdf |
| beings | GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation | pdf |
| general | Improving Editability in Image Generation with Layer-wise Memory | pdf |
| general | Simplification Is All You Need against Out-of-Distribution Overconfidence | pdf |
| general | SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input | pdf |
| gsplat | LOD-GS: Achieving Levels of Detail using Scalable Gaussian Soup | pdf |
| general | The Devil is in Low-Level Features for Cross-Domain Few-Shot Segmentation | pdf |
| general | Discrete to Continuous: Generating Smooth Transition Poses from Sign Language Observations | pdf |
| general | Unified Medical Lesion Segmentation via Self-referring Indicator | pdf |
| general | Unveiling Differences in Generative Models: A Scalable Differential Clustering Approach | pdf |
| general | Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentation | pdf |
| general | PhyS-EdiT: Physics-aware Semantic Image Editing with Text Description | pdf |
| worlds | SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model | pdf |
| gsplat | Gaussian Splatting Feature Fields for (Privacy-Preserving) Visual Localization | pdf |
| beings | Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities | pdf |
| general | Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal Transport | pdf |
| general | DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding | pdf |
| general | From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration | pdf |
| beings | HERA: Hybrid Explicit Representation for Ultra-Realistic Head Avatars | pdf |
| general | CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification | pdf |
| general | Hierarchical Adaptive Filtering Network for Text Image Specular Highlight Removal | pdf |
| general | Learning Extremely High Density Crowds as Active Matters | pdf |
| beings | EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation | pdf |
| gsplat | Gaussian Splatting for Efficient Satellite Image Photogrammetry | pdf |
| general | Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long Text | pdf |
| general | Parallel Sequence Modeling via Generalized Spatial Propagation Network | pdf |
| general | NADER: Neural Architecture Design via Multi-Agent Collaboration | pdf |
| general | Fortifying Federated Learning Towards Trustworthiness via Auditable Data Valuation and Verifiable Client Contribution | pdf |
| gsplat | UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting | pdf |
| general | Charm: The Missing Piece in ViT Fine-Tuning for Image Aesthetic Assessment | pdf |
| general | SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding | pdf |
| general | Linear Attention Modeling for Learned Image Compression | pdf |
| general | Asynchronous Collaborative Graph Representation for Frames and Events | pdf |
| worlds | ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration | pdf |
| general | GenFusion: Closing the Loop between Reconstruction and Generation via Videos | pdf |
| general | The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour Recognition | pdf |
| motion | Sonic: Shifting Focus to Global Audio Perception in Portrait Animation | pdf |
| worlds | Multitwine: Multi-Object Compositing with Text and Layout Control | pdf |
| general | Video Depth without Video Models | pdf |
| general | PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning | pdf |
| beings | HumanRig: Learning Automatic Rigging for Humanoid Character in a Large Scale Dataset | pdf |
| gsplat | GaussHDR: High Dynamic Range Gaussian Splatting via Learning Unified 3D and 2D Local Tone Mapping | pdf |
| worlds | Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor Scenes | pdf |
| general | Targeted Forgetting of Image Subgroups in CLIP Models | pdf |
| general | SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model | pdf |
| general | DSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot Segmentation | pdf |
| general | SemAlign3D: Semantic Correspondence between RGB-Images through Aligning 3D Object-Class Representations | pdf |
| general | DPU: Dynamic Prototype Updating for Multimodal Out-of-Distribution Detection | pdf |
| general | Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation | pdf |
| general | Boosting Domain Incremental Learning: Selecting the Optimal Parameters is All You Need | pdf |
| general | Perceptual Inductive Bias Is What You Need Before Contrastive Learning | pdf |
| beings | FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs | pdf |
| general | EffiDec3D: An Optimized Decoder for High-Performance and Efficient 3D Medical Image Segmentation | pdf |
| general | Exploring Historical Information for RGBE Visual Tracking with Mamba | pdf |
| worlds | ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary | pdf |
| general | Improving Sound Source Localization with Joint Slot Attention on Image and Audio | pdf |
| meshing | Feature-Preserving Mesh Decimation for Normal Integration | pdf |
| general | DexHandDiff: Interaction-aware Diffusion Planning for Adaptive Dexterous Manipulation | pdf |
| general | VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness | pdf |
| general | Memories of Forgotten Concepts | pdf |
| general | PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection | pdf |
| motion | Phoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action Correction | pdf |
| general | Multirate Neural Image Compression with Adaptive Lattice Vector Quantization | pdf |
| general | EventFly: Event Camera Perception from Ground to the Sky | pdf |
| general | CH3Depth: Efficient and Flexible Depth Foundation Model with Flow Matching | pdf |
| general | Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors | pdf |
| general | MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking | pdf |
| general | Adaptive Parameter Selection for Tuning Vision-Language Models | pdf |
| general | Learning-enabled Polynomial Lyapunov Function Synthesis via High-Accuracy Counterexample-Guided Framework | pdf |
| general | Patient-Level Anatomy Meets Scanning-Level Physics: Personalized Federated Low-Dose CT Denoising Empowered by Large Language Model | pdf |
| gsplat | Exploiting Deblurring Networks for Radiance Fields | pdf |
| general | Rethinking Lanes and Points in Complex Scenarios for Monocular 3D Lane Detection | pdf |
| general | PSA-SSL: Pose and Size-aware Self-Supervised Learning on LiDAR Point Clouds | pdf |
| general | SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training | pdf |
| general | Community Forensics: Using Thousands of Generators to Train Fake Image Detectors | pdf |
| motion | ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling | pdf |
| general | Quaffure: Real-Time Quasi-Static Neural Hair Simulation | pdf |
| worlds | LumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relighting | pdf |
| general | DiC: Rethinking Conv3x3 Designs in Diffusion Models | pdf |
| gsplat | MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds | pdf |
| general | Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models | pdf |
| beings | TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization | pdf |
| general | ShapeShifter: 3D Variations Using Multiscale and Sparse Point-Voxel Diffusion | pdf |
| beings | FRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few Images | pdf |
| gsplat | ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning | pdf |
| worlds | WonderWorld: Interactive 3D Scene Generation from a Single Image | pdf |
| general | A Lightweight UDF Learning Framework for 3D Reconstruction Based on Local Shape Functions | pdf |
| general | DiffCAM: Data-Driven Saliency Maps by Capturing Feature Differences | pdf |
| motion | From Sparse Signal to Smooth Motion: Real-Time Motion Generation with Rolling Prediction Models | pdf |
| general | Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy | pdf |
| general | EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions | pdf |
| general | Reversing Flow for Image Restoration | pdf |
| general | Shadow Generation Using Diffusion Model with Geometry Prior | pdf |
| general | Rethinking Epistemic and Aleatoric Uncertainty for Active Open-Set Annotation: An Energy-Based Approach | pdf |
| general | Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask Tracking | pdf |
| general | FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image Editing | pdf |
| general | MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling | pdf |
| general | ROS-SAM: High-Quality Interactive Segmentation for Remote Sensing Moving Object | pdf |
| general | MI-DETR: An Object Detection Model with Multi-time Inquiries Mechanism | pdf |
| general | Synthetic Visual Genome | pdf |
| general | Stop Learning it all to Mitigate Visual Hallucination, Focus on the Hallucination Target. | pdf |
| general | Seeing the Abstract: Translating the Abstract Language for Vision Language Models | pdf |
| general | One-Step Event-Driven High-Speed Autofocus | pdf |
| general | PanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial Augmentation | pdf |
| beings | Towards High-fidelity 3D Talking Avatar with Personalized Dynamic Texture | pdf |
| worlds | Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model | pdf |
| worlds | JiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World Data | pdf |
| general | OSMamba: Omnidirectional Spectral Mamba with Dual-Domain Prior Generator for Exposure Correction | pdf |
| general | RandAR: Decoder-only Autoregressive Visual Generation in Random Orders | pdf |
| general | Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation | pdf |
| general | Type-R: Automatically Retouching Typos for Text-to-Image Generation | pdf |
| general | Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding | pdf |
| general | Single Domain Generalization for Few-Shot Counting via Universal Representation Matching | pdf |
| general | Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning | pdf |
| general | Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild | pdf |
| general | A General Adaptive Dual-level Weighting Mechanism for Remote Sensing Pansharpening | pdf |
| general | RASP: Revisiting 3D Anamorphic Art for Shadow-Guided Packing of Irregular Objects | pdf |
| general | Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image Captioning | pdf |
| general | MODfinity: Unsupervised Domain Adaptation with Multimodal Information Flow Intertwining | pdf |
| general | Towards Universal Soccer Video Understanding | pdf |
| general | Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models | pdf |
| general | NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images | pdf |
| general | Efficient Personalization of Quantized Diffusion Model without Backpropagation | pdf |
| general | Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space | pdf |
| general | KVQ: Boosting Video Quality Assessment via Saliency-guided Local Perception | pdf |
| general | Learning Flow Fields in Attention for Controllable Person Image Generation | pdf |
| general | Early-Bird Diffusion: Investigating and Leveraging Timestep-Aware Early-Bird Tickets in Diffusion Models for Efficient Training | pdf |
| general | DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos | pdf |
| general | EfficientLLaVA: Generalizable Auto-Pruning for Large Vision-language Models | pdf |
| general | OnlineAnySeg: Online Zero-Shot 3D Segmentation by Visual Foundation Model Guided 2D Mask Merging | pdf |
| general | Tora: Trajectory-oriented Diffusion Transformer for Video Generation | pdf |
| gsplat | Morpheus: Text-Driven 3D Gaussian Splat Shape and Color Stylization | pdf |
| general | Generalized Diffusion Detector: Mining Robust Features from Diffusion Models for Domain-Generalized Detection | pdf |
| motion | Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model | pdf |
| general | Neuro-Symbolic Evaluation of Text-to-Video Models using Formal Verification | pdf |
| general | Spherical Manifold Guided Diffusion Model for Panoramic Image Generation | pdf |
| general | Rethinking Query-based Transformer for Continual Image Segmentation | pdf |
| general | SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding | pdf |
| general | Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding | pdf |
| general | Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature Space | pdf |
| general | ECVC: Exploiting Non-Local Correlations in Multiple Frames for Contextual Video Compression | pdf |
| general | LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity | pdf |
| general | Towards Open-Vocabulary Audio-Visual Event Localization | pdf |
| gsplat | S2Gaussian: Sparse-View Super-Resolution 3D Gaussian Splatting | pdf |
| general | HIIF: Hierarchical Encoding based Implicit Image Function for Continuous Super-resolution | pdf |
| motion | Motion Prompting: Controlling Video Generation with Motion Trajectories | pdf |
| general | VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Models | pdf |
| general | CCIN: Compositional Conflict Identification and Neutralization for Composed Image Retrieval | pdf |
| general | OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels | pdf |
| general | Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Locality | pdf |
| beings | Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and Harmonization | pdf |
| gsplat | MAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse Views | pdf |
| general | Extreme Rotation Estimation in the Wild | pdf |
| general | Traversing Distortion-Perception Tradeoff using a Single Score-Based Generative Model | pdf |
| general | Task-Agnostic Guided Feature Expansion for Class-Incremental Learning | pdf |
| motion | Let's Chorus: Partner-aware Hybrid Song-Driven 3D Head Animation | pdf |
| general | Twinner: Shining Light on Digital Twins in a Few Snaps | pdf |
| general | MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular Detection | pdf |
| gsplat | SGCR: Spherical Gaussians for Efficient 3D Curve Reconstruction | pdf |
| general | Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers | pdf |
| general | On the Out-Of-Distribution Generalization of Large Multimodal Models | pdf |
| general | Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report Generation | pdf |
| general | Scaling Inference Time Compute for Diffusion Models | pdf |
| general | Chat-based Person Retrieval via Dialogue-Refined Cross-Modal Alignment | pdf |
| general | AVF-MAE++: Scaling Affective Video Facial Masked Autoencoders via Efficient Audio-Visual Self-Supervised Learning | pdf |
| motion | DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion | pdf |
| beings | GroundingFace: Fine-grained Face Understanding via Pixel Grounding Multimodal Large Language Model | pdf |
| general | DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding | pdf |
| general | V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents | pdf |
| general | ID-Patch: Robust ID Association for Group Photo Personalization | pdf |
| gsplat | iG-6DoF: Model-free 6DoF Pose Estimation for Unseen Object via Iterative 3D Gaussian Splatting | pdf |
| general | ForestLPR: LiDAR Place Recognition in Forests Attentioning Multiple BEV Density Images | pdf |
| general | AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization | pdf |
| general | SLVR: Super-Light Visual Reconstruction via Blueprint Controllable Convolutions and Exploring Feature Diversity Representation | pdf |
| general | Layered Image Vectorization via Semantic Simplification | pdf |
| general | Hearing Anywhere in Any Environment | pdf |
| general | Automated Proof of Polynomial Inequalities via Reinforcement Learning | pdf |
| general | SAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Cost | pdf |
| motion | How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions | pdf |
| general | Joint Vision-Language Social Bias Removal for CLIP | pdf |
| general | MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds | pdf |
| general | Explicit Depth-Aware Blurry Video Frame Interpolation Guided by Differential Curves | pdf |
| general | MonSter: Marry Monodepth to Stereo Unleashes Power | pdf |
| general | A Tale of Two Classes: Adapting Supervised Contrastive Learning to Binary Imbalanced Datasets | pdf |
| general | Learning Class Prototypes for Unified Sparse-Supervised 3D Object Detection | pdf |
| worlds | Open-World Amodal Appearance Completion | pdf |
| general | RivuletMLP: An MLP-based Architecture for Efficient Compressed Video Quality Enhancement | pdf |
| general | Reanimating Images using Neural Representations of Dynamic Stimuli | pdf |
| general | DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos | pdf |
| beings | Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector | pdf |
| general | Active Hyperspectral Imaging Using an Event Camera | pdf |
| gsplat | Bridging the Gap between Gaussian Diffusion Models and Universal Quantization for Image Compression | pdf |
| general | SAMBLE: Shape-Specific Point Cloud Sampling for an Optimal Trade-Off Between Local Detail and Global Uniformity | pdf |
| general | Driving by the Rules: A Benchmark for Integrating Traffic Sign Regulations into Vectorized HD Map | pdf |
| general | S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representation | pdf |
| general | Science-T2I: Addressing Scientific Illusions in Image Synthesis | pdf |
| general | MoST: Efficient Monarch Sparse Tuning for 3D Representation Learning | pdf |
| general | Re-thinking Temporal Search for Long-Form Video Understanding | pdf |
| general | When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven Approach | pdf |
| general | BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence | pdf |
| general | Query Efficient Black-Box Visual Prompting with Subspace Learning | pdf |
| general | Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction | pdf |
| beings | Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans? | pdf |
| worlds | Solving Instance Detection from an Open-World Perspective | pdf |
| worlds | Percept, Memory, and Imagine: World Feature Simulating for Open-Domain Unknown Object Detection | pdf |
| general | Efficient Depth Estimation for Unstable Stereo Camera Systems on AR Glasses | pdf |
| gsplat | LiDAR-RT: Gaussian-based Ray Tracing for Dynamic LiDAR Re-simulation | pdf |
| general | Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator | pdf |
| gsplat | Flow-NeRF: Joint Learning of Geometry, Poses, and Dense Flow within Unified Neural Representations | pdf |
| motion | Consistent and Controllable Image Animation with Motion Diffusion Models | pdf |
| general | AA-CLIP: Enhancing Zero-Shot Anomaly Detection via Anomaly-Aware CLIP | pdf |
| gsplat | HybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splatting | pdf |
| general | Channel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image Deraining | pdf |
| general | MobileMamba: Lightweight Multi-Receptive Visual Mamba Network | pdf |
| general | SimLTD: Simple Supervised and Semi-Supervised Long-Tailed Object Detection | pdf |
| general | HyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual Perceiver | pdf |
| general | Diffusion-based Event Generation for High-Quality Image Deblurring | pdf |
| general | Balanced Rate-Distortion Optimization in Learned Image Compression | pdf |
| general | Bridge the Gap: From Weak to Full Supervision for Temporal Action Localization with PseudoFormer | pdf |
| family | paper | links |
| general | DynScene: Scalable Generation of Dynamic Robotic Manipulation Scenes for Embodied AI | pdf |
| general | DiffLocks: Generating 3D Hair from a Single Image using Diffusion Models | pdf |
| general | Harnessing Frequency Spectrum Insights for Image Copyright Protection Against Diffusion Models | pdf |
| general | IDEA-Bench: How Far are Generative Models from Professional Designing? | pdf |
| general | PhD: A ChatGPT-Prompted Visual Hallucination Evaluation Dataset | pdf |
| worlds | ClimbingCap: Multi-Modal Dataset and Method for Rock Climbing in World Coordinate | pdf |
| general | A Bias-Free Training Paradigm for More General AI-generated Image Detection | pdf |
| general | FALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene Understanding | pdf |
| beings | Certified Human Trajectory Prediction | pdf |
| general | Transformers without Normalization | pdf |
| beings | HiPART: Hierarchical Pose AutoRegressive Transformer for Occluded 3D Human Pose Estimation | pdf |
| beings | From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech | pdf |
| general | DFM: Differentiable Feature Matching for Anomaly Detection | pdf |
| general | PointSR: Self-Regularized Point Supervision for Drone-View Object Detection | pdf |
| worlds | v-CLR: View-Consistent Learning for Open-World Instance Segmentation | pdf |
| general | Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization | pdf |
| general | Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation | pdf |
| general | MagicArticulate: Make Your 3D Models Articulation-Ready | pdf |
| general | Dual Prompting Image Restoration with Diffusion Transformers | pdf |
| general | DepthCues: Evaluating Monocular Depth Perception in Large Vision Models | pdf |
| gsplat | SpecTRe-GS: Modeling Highly Specular Surfaces with Reflected Nearby Objects by Tracing Rays in 3D Gaussian Splatting | pdf |
| general | AuraFusion360: Augmented Unseen Region Alignment for Reference-based 360deg Unbounded Scene Inpainting | pdf |
| general | Language-Guided Image Tokenization for Generation | pdf |
| general | Hyperbolic Uncertainty-Aware Few-Shot Incremental Point Cloud Segmentation | pdf |
| beings | D^3-Human: Dynamic Disentangled Digital Human from Monocular Video | pdf |
| general | Curriculum Coarse-to-Fine Selection for High-IPC Dataset Distillation | pdf |
| general | BADGR: Bundle Adjustment Diffusion Conditioned by Gradients for Wide-Baseline Floor Plan Reconstruction | pdf |
| general | Three Cars Approaching within 100m! Enhancing Distant Geometry by Tri-Axis Voxel Scanning for Camera-based Semantic Scene Completion | pdf |
| general | DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehension | pdf |
| general | Spiking Transformer with Spatial-Temporal Attention | pdf |
| general | Perceptual Video Compression with Neural Wrapping | pdf |
| general | ViKIENet: Towards Efficient 3D Object Detection with Virtual Key Instance Enhanced Network | pdf |
| general | Data-free Universal Adversarial Perturbation with Pseudo-semantic Prior | pdf |
| beings | FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video | pdf |
| general | Generalized Zero-Shot Classification via Semantics-Free Inter-Class Feature Generation | pdf |
| general | Multi-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEs | pdf |
| general | MaDCoW: Marginal Distortion Correction for Wide-Angle Photography with Arbitrary Objects | pdf |
| general | Any6D: Model-free 6D Pose Estimation of Novel Objects | pdf |
| general | DrVideo: Document Retrieval Based Long Video Understanding | pdf |
| general | Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors | pdf |
| beings | PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing | pdf |
| general | Hiding Images in Diffusion Models by Editing Learned Score Functions | pdf |
| general | WeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba Diffusion | pdf |
| general | MUST: The First Dataset and Unified Framework for Multispectral UAV Single Object Tracking | pdf |
| general | Tightening Robustness Verification of MaxPool-based Neural Networks via Minimizing the Over-Approximation Zone | pdf |
| general | PhysicsGen: Can Generative Models Learn from Images to Predict Complex Physical Relations? | pdf |
| general | Spectral Informed Mamba for Robust Point Cloud Processing | pdf |
| general | BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations | pdf |
| general | D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective Recognition. | pdf |
| general | LaVin-DiT: Large Vision Diffusion Transformer | pdf |
| general | CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image | pdf |
| general | Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization | pdf |
| gsplat | BARD-GS: Blur-Aware Reconstruction of Dynamic Scenes via Gaussian Splatting | pdf |
| general | DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy Labels | pdf |
| beings | S^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion Priors | pdf |
| general | FSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understanding | pdf |
| general | Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic Segmentation | pdf |
| general | LLM-driven Multimodal and Multi-Identity Listening Head Generation | pdf |
| meshing | OffsetOPT: Explicit Surface Reconstruction without Normals | pdf |
| general | Any-Resolution AI-Generated Image Detection by Spectral Learning | pdf |
| general | STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding | pdf |
| motion | TimeTracker: Event-based Continuous Point Tracking for Video Frame Interpolation with Non-linear Motion | pdf |
| worlds | Shading Meets Motion: Self-supervised Indoor 3D Reconstruction Via Simultaneous Shape-from-Shading and Structure-from-Motion | pdf |
| general | Believing is Seeing: Unobserved Object Detection using Generative Models | pdf |
| general | NLPrompt: Noise-Label Prompt Learning for Vision-Language Models | pdf |
| gsplat | PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields | pdf |
| general | No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image Recognition | pdf |
| general | ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large Language Models | pdf |
| general | TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusion | pdf |
| general | Physical Plausibility-aware Trajectory Prediction via Locomotion Embodiment | pdf |
| beings | AvatarArtist: Open-Domain 4D Avatarization | pdf |
| general | Using Powerful Prior Knowledge of Diffusion Model in Deep Unfolding Networks for Image Compressive Sensing | pdf |
| general | UniGoal: Towards Universal Zero-shot Goal-oriented Navigation | pdf |
| general | Noise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentation | pdf |
| general | DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection | pdf |
| general | Less is More: Efficient Image Vectorization with Adaptive Parameterization | pdf |
| general | FedMIA: An Effective Membership Inference Attack Exploiting "All for One" Principle in Federated Learning | pdf |
| general | DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid Framework | pdf |
| general | DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning | pdf |
| general | Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling | pdf |
| general | ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language Models | pdf |
| general | Masking meets Supervision: A Strong Learning Alliance | pdf |
| worlds | DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creation | pdf |
| general | Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question Answering | pdf |
| general | UniRestore: Unified Perceptual and Task-Oriented Image Restoration Model Using Diffusion Prior | pdf |
| general | Condensing Action Segmentation Datasets via Generative Network Inversion | pdf |
| general | Can Generative Video Models Help Pose Estimation? | pdf |
| general | DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving | pdf |
| meshing | High-Fidelity Lightweight Mesh Reconstruction from Point Clouds | pdf |
| general | MDP: Multidimensional Vision Model Pruning with Latency Constraint | pdf |
| beings | OSDFace: One-Step Diffusion Model for Face Restoration | pdf |
| general | Task Singular Vectors: Reducing Task Interference in Model Merging | pdf |
| general | Self-Evolving Visual Concept Library using Vision-Language Critics | pdf |
| general | Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal Transport | pdf |
| general | Effective Cloud Removal for Remote Sensing Images by an Improved Mean-Reverting Denoising Model with Elucidated Design Space | pdf |
| general | OpticalNet: An Optical Imaging Dataset and Benchmark Beyond the Diffraction Limit | pdf |
| general | Empowering Vector Graphics with Consistently Arbitrary Viewing and View-dependent Visibility | pdf |
| general | DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID | pdf |
| general | HyperPose: Hypernetwork-Infused Camera Pose Localization and an Extended Cambridge Landmarks Dataset | pdf |
| general | Mono3DVLT: Monocular-Video-Based 3D Visual Language Tracking | pdf |
| general | Towards Universal Dataset Distillation via Task-Driven Diffusion | pdf |
| meshing | Parametric Point Cloud Completion for Polygonal Surface Reconstruction | pdf |
| general | SyncSDE: A Probabilistic Framework for Diffusion Synchronization | pdf |
| general | MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation | pdf |
| general | Dual Semantic Guidance for Open Vocabulary Semantic Segmentation | pdf |
| general | Generalizable Object Keypoint Localization from Generative Priors | pdf |
| general | FedCALM: Conflict-aware Layer-wise Mitigation for Selective Aggregation in Deeper Personalized Federated Learning | pdf |
| general | CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Booth | pdf |
| gsplat | FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splatting | pdf |
| general | Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning | pdf |
| general | T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation | pdf |
| general | Make It Count: Text-to-Image Generation with an Accurate Number of Objects | pdf |
| general | TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent Perception | pdf |
| general | DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching | pdf |
| general | FlexUOD: The Answer to Real-world Unsupervised Image Outlier Detection | pdf |
| general | Focusing on Tracks for Online Multi-Object Tracking | pdf |
| general | Identity-preserving Distillation Sampling by Fixed-Point Iterator | pdf |
| general | WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild | pdf |
| general | BiomedCoOp: Learning to Prompt for Biomedical Vision-Language Models | pdf |
| beings | MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation | pdf |
| general | AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities | pdf |
| worlds | OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? | pdf |
| gsplat | GuardSplat: Efficient and Robust Watermarking for 3D Gaussian Splatting | pdf |
| general | RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives | pdf |
| worlds | Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation | pdf |
| gsplat | LookCloser: Frequency-aware Radiance Field for Tiny-Detail Scene | pdf |
| worlds | Convex Relaxation for Robust Vanishing Point Estimation in Manhattan World | pdf |
| gsplat | FruitNinja: 3D Object Interior Texture Generation with Gaussian Splatting | pdf |
| general | Take the Bull by the Horns: Learning to Segment Hard Samples | pdf |
| general | EIDT-V: Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video Generation | pdf |
| general | Reproducible Vision-Language Models Meet Concepts Out of Pre-Training | pdf |
| motion | Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation | pdf |
| general | MAGE : Single Image to Material-Aware 3D via the Multi-View G-Buffer Estimation Model | pdf |
| general | MESC-3D:Mining Effective Semantic Cues for 3D Reconstruction from a Single Image | pdf |
| general | Advancing Multiple Instance Learning with Continual Learning for Whole Slide Imaging | pdf |
| general | Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric Tasks | pdf |
| general | TAET: Two-Stage Adversarial Equalization Training on Long-Tailed Distributions | pdf |
| general | Few-shot Personalized Scanpath Prediction | pdf |
| general | Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models | pdf |
| general | Mamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic Segmentation | pdf |
| beings | Vision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D Scenes | pdf |
| beings | ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation | pdf |
| general | CLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair Loss | pdf |
| general | ObjectMover: Generative Object Movement with Video Prior | pdf |
| beings | MLLM-as-a-Judge for Image Safety without Human Labeling | pdf |
| general | Learning to Filter Outlier Edges in Global SfM | pdf |
| beings | Forensics Adapter: Adapting CLIP for Generalizable Face Forgery Detection | pdf |
| general | KAC: Kolmogorov-Arnold Classifier for Continual Learning | pdf |
| general | BOOTPLACE: Bootstrapped Object Placement with Detection Transformers | pdf |
| general | FASTer: Focal token Acquiring-and-Scaling Transformer for Long-term 3D Objection Detection | pdf |
| general | Geometry-guided Online 3D Video Synthesis with Multi-View Temporal Consistency | pdf |
| worlds | Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances | pdf |
| gsplat | CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Images | pdf |
| general | Semantic and Sequential Alignment for Referring Video Object Segmentation | pdf |
| general | Continual SFT Matches Multimodal RLHF with Negative Supervision | pdf |
| general | Semantic-guided Cross-Modal Prompt Learning for Skeleton-based Zero-shot Action Recognition | pdf |
| general | ChatGen: Automatic Text-to-Image Generation From FreeStyle Chatting | pdf |
| general | VEU-Bench: Towards Comprehensive Understanding of Video Editing | pdf |
| general | Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image Compression | pdf |
| general | Yo'Chameleon: Personalized Vision and Language Generation | pdf |
| general | PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution | pdf |
| general | FluxSpace: Disentangled Semantic Editing in Rectified Flow Models | pdf |
| general | Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization | pdf |
| general | ShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware Prompts | pdf |
| general | Auto-Encoded Supervision for Perceptual Image Super-Resolution | pdf |
| general | Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation | pdf |
| worlds | Iterative Predictor-Critic Code Decoding for Real-World Image Dehazing | pdf |
| general | Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning | pdf |
| general | Gradient-Guided Annealing for Domain Generalization | pdf |
| general | MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research | pdf |
| general | Event-based Video Super-Resolution via State Space Models | pdf |
| general | Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene Understanding | pdf |
| general | VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding | pdf |
| general | CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction | pdf |
| general | Paint by Inpaint: Learning to Add Image Objects by Removing Them First | pdf |
| general | PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba Adapter | pdf |
| general | LC-Mamba: Local and Continuous Mamba with Shifted Windows for Frame Interpolation | pdf |
| worlds | Zero-Shot Head Swapping in Real-World Scenarios | pdf |
| general | CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment | pdf |
| general | COBRA: COmBinatorial Retrieval Augmentation for Few-Shot Adaptation | pdf |
| meshing | ProbeSDF: Light Field Probes For Neural Surface Reconstruction | pdf |
| general | Hybrid Concept Bottleneck Models | pdf |
| general | Dual Consolidation for Pre-Trained Model-Based Domain-Incremental Learning | pdf |
| beings | RORem: Training a Robust Object Remover with Human-in-the-Loop | pdf |
| general | All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages | pdf |
| beings | Video-Bench: Human-Aligned Video Generation Benchmark | pdf |
| general | MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization | pdf |
| general | Anyattack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models | pdf |
| gsplat | Joint Optimization of Neural Radiance Fields and Continuous Camera Motion from a Monocular Video | pdf |
| gsplat | IRGS: Inter-Reflective Gaussian Splatting with 2D Gaussian Ray Tracing | pdf |
| beings | InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions | pdf |
| general | Efficient Test-time Adaptive Object Detection via Sensitivity-Guided Pruning | pdf |
| general | A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning | pdf |
| general | Visual Agentic AI for Spatial Reasoning with a Dynamic API | pdf |
| general | Feature Spectrum Learning for Remote Sensing Change Detection | pdf |
| worlds | DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation | pdf |
| general | LoKi: Low-dimensional KAN for Efficient Fine-tuning Image Models | pdf |
| gsplat | Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding Registration | pdf |
| general | Consistent Normal Orientation for 3D Point Clouds via Least Squares on Delaunay Graph | pdf |
| general | ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpainting | pdf |
| general | Optimizing for the Shortest Path in Denoising Diffusion Model | pdf |
| general | Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception | pdf |
| general | Dynamic Pseudo Labeling via Gradient Cutting for High-Low Entropy Exploration | pdf |
| general | VODiff: Controlling Object Visibility Order in Text-to-Image Generation | pdf |
| general | CAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model Generation | pdf |
| general | Leveraging SD Map to Augment HD Map-based Trajectory Prediction | pdf |
| general | ONDA-Pose: Occlusion-Aware Neural Domain Adaptation for Self-Supervised 6D Object Pose Estimation | pdf |
| general | Q-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction Regression | pdf |
| general | Composing Parts for Expressive Object Generation | pdf |
| general | CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-Scale Reinforcement Learning in Autonomous Driving | pdf |
| general | Apply Hierarchical-Chain-of-Generation to Complex Attributes Text-to-3D Generation | pdf |
| general | SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving | pdf |
| general | Shift the Lens: Environment-Aware Unsupervised Camouflaged Object Detection | pdf |
| general | DriveScape: High-Resolution Driving Video Generation by Multi-View Feature Fusion | pdf |
| general | Training-free Neural Architecture Search through Variance of Knowledge of Deep Network Weights | pdf |
| general | Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond | pdf |
| general | RAEncoder: A Label-Free Reversible Adversarial Examples Encoder for Dataset Intellectual Property Protection | pdf |
| motion | LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand Diffusion. | pdf |
| general | Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs | pdf |
| general | Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly Detection | pdf |
| general | Full-DoF Egomotion Estimation for Event Cameras Using Geometric Solvers | pdf |
| general | Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distribution | pdf |
| general | Your Scale Factors are My Weapon: Targeted Bit-Flip Attacks on Vision Transformers via Scale Factor Manipulation | pdf |
| general | Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding | pdf |
| general | Mamba-Reg: Vision Mamba Also Needs Registers | pdf |
| beings | Visual Persona: Foundation Model for Full-Body Human Customization | pdf |
| gsplat | SOGS: Second-Order Anchor for Advanced 3D Gaussian Splatting | pdf |
| general | MExD: An Expert-Infused Diffusion Model for Whole-Slide Image Classification | pdf |
| general | Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of Samples | pdf |
| gsplat | MoDec-GS: Global-to-Local Motion Decomposition and Temporal Interval Adjustment for Compact Dynamic 3D Gaussian Splatting | pdf |
| general | DaCapo: Score Distillation as Stacked Bridge for Fast and High-quality 3D Editing | pdf |
| general | Number it: Temporal Grounding Videos like Flipping Manga | pdf |
| general | SyncVP: Joint Diffusion for Synchronous Multi-Modal Video Prediction | pdf |
| general | HUSH: Holistic Panoramic 3D Scene Understanding using Spherical Harmonics | pdf |
| general | SkillMimic: Learning Basketball Interaction Skills from Demonstrations | pdf |
| gsplat | RGBAvatar: Reduced Gaussian Blendshapes for Online Modeling of Head Avatars | pdf |
| general | EEE-Bench: A Comprehensive Multimodal Electrical And Electronics Engineering Benchmark | pdf |
| general | A Unified Framework for Heterogeneous Semi-supervised Learning | pdf |
| gsplat | Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed Views | pdf |
| general | Open Ad-hoc Categorization with Contextualized Feature Learning | pdf |
| general | Dynamic Updates for Language Adaptation in Visual-Language Tracking | pdf |
| general | Multi-focal Conditioned Latent Diffusion for Person Image Synthesis | pdf |
| worlds | Uncertainty Meets Diversity: A Comprehensive Active Learning Framework for Indoor 3D Object Detection | pdf |
| general | Identity-Clothing Similarity Modeling for Unsupervised Clothing Change Person Re-Identification | pdf |
| general | OccMamba: Semantic Occupancy Prediction with State Space Models | pdf |
| general | Cheb-GR: Rethinking K-nearest Neighbor Search in Re-ranking for Person Re-identification | pdf |
| general | Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving | pdf |
| general | Is `Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning | pdf |
| general | GCC: Generative Color Constancy via Diffusing a Color Checker | pdf |
| general | On Denoising Walking Videos for Gait Recognition | pdf |
| general | Conformal Prediction for Zero-Shot Models | pdf |
| motion | PhysAnimator: Physics-Guided Generative Cartoon Animation | pdf |
| general | FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation | pdf |
| general | BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs | pdf |
| general | VasTSD: Learning 3D Vascular Tree-state Space Diffusion Model for Angiography Synthesis | pdf |
| gsplat | PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting | pdf |
| general | WISNet: Pseudo Label Generation on Unbalanced and Patch Annotated Waste Images | pdf |
| beings | MixerMDM: Learnable Composition of Human Motion Diffusion Models | pdf |
| general | Hand-held Object Reconstruction from RGB Video with Dynamic Interaction | pdf |
| beings | AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers | pdf |
| general | Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces | pdf |
| general | A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs | pdf |
| general | SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Models | pdf |
| beings | Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidance | pdf |
| general | Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes | pdf |
| general | Structure from Collision | pdf |
| general | Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation | pdf |
| general | Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection | pdf |
| general | OralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware Mining | pdf |
| gsplat | SplatAD: Real-Time Lidar and Camera Rendering with 3D Gaussian Splatting for Autonomous Driving | pdf |
| general | Audio-Visual Instance Segmentation | pdf |
| general | UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation | pdf |
| general | RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression | pdf |
| general | Recognition-Synergistic Scene Text Editing | pdf |
| beings | WildAvatar: Learning In-the-wild 3D Avatars from the Web | pdf |
| general | Rectified Diffusion Guidance for Conditional Generation | pdf |
| general | IAAO: Interactive Affordance Learning for Articulated Objects in 3D Environments | pdf |
| general | RaSS: Improving Denoising Diffusion Samplers with Reinforced Active Sampling Scheduler | pdf |
| general | OSV: One Step is Enough for High-Quality Image to Video Generation | pdf |
| general | Fuzzy Multimodal Learning for Trusted Cross-modal Retrieval | pdf |
| general | GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration | pdf |
| general | Few-Shot Recognition via Stage-Wise Retrieval-Augmented Finetuning | pdf |
| gsplat | RestorGS: Depth-aware Gaussian Splatting for Efficient 3D Scene Restoration | pdf |
| general | 4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion | pdf |
| general | Z-Magic: Zero-shot Multiple Attributes Guided Image Creator | pdf |
| general | On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free Approach | pdf |
| beings | Towards General Visual-Linguistic Face Forgery Detection | pdf |
| general | Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts | pdf |
| general | LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos | pdf |
| general | Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key | pdf |
| general | Simpler Diffusion: 1.5 FID on ImageNet512 with Pixel-space Diffusion | pdf |
| worlds | STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection | pdf |
| general | Not All Parameters Matter: Masking Diffusion Models for Enhancing Generation Ability | pdf |
| general | Complexity Experts are Task-Discriminative Learners for Any Image Restoration | pdf |
| general | Generative Omnimatte: Learning to Decompose Video into Layers | pdf |
| general | 5%>100%: Breaking Performance Shackles of Full Fine-Tuning on Visual Recognition Tasks | pdf |
| worlds | Real-IAD D3: A Real-World 2D/Pseudo-3D/3D Dataset for Industrial Anomaly Detection | pdf |
| general | ATP: Adaptive Threshold Pruning for Efficient Data Encoding in Quantum Neural Networks | pdf |
| motion | Decoupled Motion Expression Video Segmentation | pdf |
| general | K-LoRA: Unlocking Training-Free Fusion of Any Subject and Style LoRAs | pdf |
| general | WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model | pdf |
| general | XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery? | pdf |
| general | Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks | pdf |
| general | StyleSSP: Sampling StartPoint Enhancement for Training-free Diffusion-based Method for Style Transfer | pdf |
| beings | Motions as Queries: One-Stage Multi-Person Holistic Human Motion Capture | pdf |
| general | AMO Sampler: Enhancing Text Rendering with Overshooting | pdf |
| general | ImViD: Immersive Volumetric Videos for Enhanced VR Engagement | pdf |
| general | I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video Models | pdf |
| general | Saliuitl: Ensemble Salience Guided Recovery of Adversarial Patches against CNNs | pdf |
| general | OPTICAL: Leveraging Optimal Transport for Contribution Allocation in Dataset Distillation | pdf |
| general | Show and Segment: Universal Medical Image Segmentation via In-Context Learning | pdf |
| general | CADCrafter: Generating Computer-Aided Design Models from Unconstrained Images | pdf |
| general | Generative Multiview Relighting for 3D Reconstruction under Extreme Illumination Variation | pdf |
| general | DyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective Scheduling | pdf |
| general | GENIUS: A Generative Framework for Universal Multimodal Search | pdf |
| meshing | SF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement | pdf |
| general | Towards Precise Embodied Dialogue Localization via Causality Guided Diffusion | pdf |
| general | EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality | pdf |
| general | A4A: Adapter for Adapter Transfer via All-for-All Mapping for Cross-Architecture Models | pdf |
| general | ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation | pdf |
| general | A Universal Scale-Adaptive Deformable Transformer for Image Restoration across Diverse Artifacts | pdf |
| general | Towards Precise Scaling Laws for Video Diffusion Transformers | pdf |
| general | SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking | pdf |
| general | AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual Videos | pdf |
| general | Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision | pdf |
| general | Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and Understanding | pdf |
| general | Towards Efficient Foundation Model for Zero-shot Amodal Segmentation | pdf |
| general | Scaling Properties of Diffusion Models For Perceptual Tasks | pdf |
| general | Exact: Exploring Space-Time Perceptive Clues for Weakly Supervised Satellite Image Time Series Semantic Segmentation | pdf |
| general | PolarFree: Polarization-based Reflection-Free Imaging | pdf |
| general | Seeking Consistent Flat Minima for Better Domain Generalization via Refining Loss Landscapes | pdf |
| general | MultimodalStudio: A Heterogeneous Sensor Dataset and Framework for Neural Rendering across Multiple Imaging Modalities | pdf |
| general | MuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image Translation | pdf |
| general | Image Quality Assessment: Investigating Causal Perceptual Effects with Abductive Counterfactual Inference | pdf |
| general | Pos3R: 6D Pose Estimation for Unseen Objects Made Easy | pdf |
| general | RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusion | pdf |
| general | Understanding Multi-Task Activities from Single-Task Videos | pdf |
| motion | Co-Speech Gesture Video Generation with Implicit Motion-Audio Entanglement | pdf |
| general | TransPixeler: Advancing Text-to-Video Generation with Transparency | pdf |
| general | What's in the Image? A Deep-Dive into the Vision of Vision Language Models | pdf |
| general | FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes | pdf |
| general | Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding | pdf |
| general | GPVK-VL: Geometry-Preserving Virtual Keyframes for Visual Localization under Large Viewpoint Changes | pdf |
| general | Enhancing Online Continual Learning with Plug-and-Play State Space Model and Class-Conditional Mixture of Discretization | pdf |
| general | GRAPHGPT-O: Synergistic Multimodal Comprehension and Generation on Graphs | pdf |
| general | Model Poisoning Attacks to Federated Learning via Multi-Round Consistency | pdf |
| gsplat | TaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian Splatting | pdf |
| beings | Stacking Brick by Brick: Aligned Feature Isolation for Incremental Face Forgery Detection | pdf |
| general | CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AI | pdf |
| general | GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation | pdf |
| general | Localized Concept Erasure for Text-to-Image Diffusion Models Using Training-Free Gated Low-Rank Adaptation | pdf |
| general | Camera Resection from Known Line Pencils and a Radially Distorted Scanline | pdf |
| worlds | SPC-GS: Gaussian Splatting with Semantic-Prompt Consistency for Indoor Open-World Free-view Synthesis from Sparse Inputs | pdf |
| general | M3amba: Memory Mamba is All You Need for Whole Slide Image Classification | pdf |
| general | Redefining <Creative> in Dictionary: Towards an Enhanced Semantic Understanding of Creative Generation | pdf |
| general | Dinomaly: The Less Is More Philosophy in Multi-Class Unsupervised Anomaly Detection | pdf |
| general | Detect-and-Guide: Self-regulation of Diffusion Models for Safe Text-to-Image Generation via Guideline Token Optimization | pdf |
| worlds | MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World | pdf |
| gsplat | EAP-GS: Efficient Augmentation of Pointcloud for 3D Gaussian Splatting in Few-shot Scene Reconstruction | pdf |
| general | Empowering Large Language Models with 3D Situation Awareness | pdf |
| general | EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights | pdf |
| general | Interactive Medical Image Segmentation: A Benchmark Dataset and Baseline | pdf |
| general | GigaHands: A Massive Annotated Dataset of Bimanual Hand Activities | pdf |
| general | AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing | pdf |
| gsplat | FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned Priors | pdf |
| general | Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning | pdf |
| gsplat | CompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D Gaussians | pdf |
| general | FIRE: Robust Detection of Diffusion-Generated Images via Frequency-Guided Reconstruction Error | pdf |
| general | Assessing and Learning Alignment of Unimodal Vision and Language Models | pdf |
| general | Action Detail Matters: Refining Video Recognition with Local Action Queries | pdf |
| general | Generative Map Priors for Collaborative BEV Semantic Segmentation | pdf |
| general | Coherent 3D Portrait Video Reconstruction via Triplane Fusion | pdf |
| general | ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping | pdf |
| general | FedCS: Coreset Selection for Federated Learning | pdf |
| general | Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General Pansharpening | pdf |
| general | OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts | pdf |
| general | SOAP: Vision-Centric 3D Semantic Scene Completion with Scene-Adaptive Decoder and Occluded Region-Aware View Projection | pdf |
| worlds | SimVS: Simulating World Inconsistencies for Robust View Synthesis | pdf |
| general | From Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral Perspective | pdf |
| general | COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training | pdf |
| worlds | Lifting Motion to the 3D World via 2D Diffusion | pdf |
| general | TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language Models | pdf |
| general | Active Data Curation Effectively Distills Large-Scale Multimodal Models | pdf |
| general | SCSA: A Plug-and-Play Semantic Continuous-Sparse Attention for Arbitrary Semantic Style Transfer | pdf |
| general | Can't Slow Me Down: Learning Robust and Hardware-Adaptive Object Detectors against Latency Attacks for Edge Devices | pdf |
| general | SAM2Object: Consolidating View Consistency via SAM2 for Zero-Shot 3D Instance Segmentation | pdf |
| general | CDI: Copyrighted Data Identification in Diffusion Models | pdf |
| general | CRISP: Object Pose and Shape Estimation with Test-Time Adaptation | pdf |
| gsplat | Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian Splatting | pdf |
| general | Sim-to-Real Causal Transfer: A Metric Learning Approach to Causally-Aware Interaction Representations | pdf |
| general | Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental Learning | pdf |
| general | PerLA: Perceptive 3D Language Assistant | pdf |
| general | PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation | pdf |
| general | Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation | pdf |
| general | JamMa: Ultra-lightweight Local Feature Matching with Joint Mamba | pdf |
| general | DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models | pdf |
| general | MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alps | pdf |
| motion | Diffusion-based Realistic Listening Head Generation via Hybrid Motion Modeling | pdf |
| meshing | SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokens | pdf |
| worlds | UniScene: Unified Occupancy-centric Driving Scene Generation | pdf |
| general | Learning from Streaming Video with Orthogonal Gradients | pdf |
| general | Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers | pdf |
| beings | An Image-like Diffusion Method for Human-Object Interaction Detection | pdf |
| gsplat | COB-GS: Clear Object Boundaries in 3DGS Segmentation Based on Boundary-Adaptive Gaussian Splitting | pdf |
| general | PEER Pressure: Model-to-Model Regularization for Single Source Domain Generalization | pdf |
| general | Revisiting Fairness in Multitask Learning: A Performance-Driven Approach for Variance Reduction | pdf |
| motion | VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models | pdf |
| general | Compositional Caching for Training-free Open-vocabulary Attribute Detection | pdf |
| general | VI^3NR: Variance Informed Initialization for Implicit Neural Representations | pdf |
| general | M-LLM Based Video Frame Selection for Efficient Video Understanding | pdf |
| general | Search and Detect: Training-Free Long Tail Object Detection via Web-Image Retrieval | pdf |
| general | Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation | pdf |
| general | Diffusion Model is Effectively Its Own Teacher | pdf |
| general | UnCommon Objects in 3D | pdf |
| worlds | Learning Textual Prompts for Open-World Semi-Supervised Learning | pdf |
| general | LongDiff: Training-Free Long Video Generation in One Go | pdf |
| general | Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation | pdf |
| general | MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving | pdf |
| general | ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way | pdf |
| general | Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding | pdf |
| general | On the Generalization of Handwritten Text Recognition Models | pdf |
| general | InsTaG: Learning Personalized 3D Talking Head from Few-Second Video | pdf |
| general | Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning | pdf |
| general | Rotation-Equivariant Self-Supervised Method in Image Denoising | pdf |
| general | FlashSloth : Lightning Multimodal Large Language Models via Embedded Visual Compression | pdf |
| general | T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving | pdf |
| general | RealEdit: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations | pdf |
| general | VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step | pdf |
| gsplat | 3D-HGS: 3D Half-Gaussian Splatting | pdf |
| general | Scale Efficient Training for Large Datasets | pdf |
| general | Decoder Gradient Shield: Provable and High-Fidelity Prevention of Gradient-Based Box-Free Watermark Removal | pdf |
| general | Convex Combination Star Shape Prior for Data-driven Image Semantic Segmentation | pdf |
| general | Parameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentation | pdf |
| general | Relative Pose Estimation through Affine Corrections of Monocular Depth Priors | pdf |
| beings | Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion | pdf |
| worlds | Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object Recognition | pdf |
| general | Conical Visual Concentration for Efficient Large Vision-Language Models | pdf |
| general | Foundations of the Theory of Performance-Based Ranking | pdf |
| gsplat | BIGS: Bimanual Category-agnostic Interaction Reconstruction from Monocular Videos via 3D Gaussian Splatting | pdf |
| general | Frequency-Biased Synergistic Design for Image Compression and Compensation | pdf |
| gsplat | Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field Rendering | pdf |
| general | MambaIC: State Space Models for High-Performance Learned Image Compression | pdf |
| gsplat | Instant Gaussian Stream: Fast and Generalizable Streaming of Dynamic Scene Reconstruction via Gaussian Splatting | pdf |
| beings | Locality-Aware Zero-Shot Human-Object Interaction Detection | pdf |
| general | Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation | pdf |
| general | SGFormer: Satellite-Ground Fusion for 3D Semantic Scene Completion | pdf |
| general | Random Conditioning for Diffusion Model Compression with Distillation | pdf |
| gsplat | Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generation | pdf |
| general | Heterogeneous Skeleton-Based Action Representation Learning | pdf |
| motion | AnyMap: Learning a General Camera Model for Structure-from-Motion with Unknown Distortion in Dynamic Scenes | pdf |
| general | Language Guided Concept Bottleneck Models for Interpretable Continual Learning | pdf |
| worlds | Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model | pdf |
| general | Odd-One-Out: Anomaly Detection by Comparing with Neighbors | pdf |
| general | D^3CTTA: Domain-Dependent Decorrelation for Continual Test-Time Adaption of 3D LiDAR Segmentation | pdf |
| general | A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training | pdf |
| general | Empowering LLMs to Understand and Generate Complex Vector Graphics | pdf |
| gsplat | PanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene Understanding | pdf |
| general | MLVU: Benchmarking Multi-task Long Video Understanding | pdf |
| general | Recovering Dynamic 3D Sketches from Videos | pdf |
| gsplat | EigenGS Representation: From Eigenspace to Gaussian Image Space | pdf |
| general | MaSS13K: A Matting-level Semantic Segmentation Benchmark | pdf |
| general | Enhancing Testing-Time Robustness for Trusted Multi-View Classification in the Wild | pdf |
| general | ROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language Models | pdf |
| general | nnWNet: Rethinking the Use of Transformers in Biomedical Image Segmentation and Calling for a Unified Evaluation Benchmark | pdf |
| general | VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment | pdf |
| general | Seeing is Not Believing: Adversarial Natural Object Optimization for Hard-Label 3D Scene Attacks | pdf |
| general | Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual Recognition | pdf |
| beings | Pippo: High-Resolution Multi-View Humans from a Single Image | pdf |
| general | H2ST: Hierarchical Two-Sample Tests for Continual Out-of-Distribution Detection | pdf |
| general | MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders | pdf |
| general | CamFreeDiff: Camera-free Image to Panorama Generation with Diffusion Model | pdf |
| general | Improving Visual and Downstream Performance of Low-Light Enhancer with Vision Foundation Models Collaboration | pdf |
| general | FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs | pdf |
| general | Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation | pdf |
| general | Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models | pdf |
| general | Video-Guided Foley Sound Generation with Multimodal Controls | pdf |
| general | F^3OCUS - Federated Finetuning of Vision-Language Foundation Models with Optimal Client Layer Updating Strategy via Multi-objective Meta-Heuristics | pdf |
| general | 3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation | pdf |
| general | Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves? | pdf |
| general | g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks | pdf |
| general | UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics | pdf |
| general | Exploring Contextual Attribute Density in Referring Expression Counting | pdf |
| general | SegMAN: Omni-scale Context Modeling with State Space Models and Local Attention for Semantic Segmentation | pdf |
| general | OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows | pdf |
| general | SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos | pdf |
| beings | SemGeoMo: Dynamic Contextual Human Motion Generation with Semantic and Geometric Guidance | pdf |
| worlds | Detecting Open World Objects via Partial Attribute Assignment | pdf |
| general | Neural Inverse Rendering from Propagating Light | pdf |
| gsplat | DecoupledGaussian: Object-Scene Decoupling for Physics-Based Interaction | pdf |
| gsplat | DashGaussian: Optimizing 3D Gaussian Splatting in 200 Seconds | pdf |
| general | PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models | pdf |
| general | R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localization | pdf |
| general | Style Evolving along Chain-of-Thought for Unknown-Domain Object Detection | pdf |
| gsplat | OmniSplat: Taming Feed-Forward 3D Gaussian Splatting for Omnidirectional Images with Editable Capabilities | pdf |
| general | UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detection | pdf |
| worlds | Remote Photoplethysmography in Real-World and Extreme Lighting Scenarios | pdf |
| general | Multi-Modal Contrastive Masked Autoencoders: A Two-Stage Progressive Pre-training Approach for RGBD Datasets | pdf |
| general | Font-Agent: Enhancing Font Understanding with Large Language Models | pdf |
| general | Secret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution Analysis | pdf |
| general | Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding | pdf |
| general | SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation | pdf |
| general | Mixture of Submodules for Domain Adaptive Person Search | pdf |
| general | SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation | pdf |
| general | EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with Events | pdf |
| worlds | Seeing A 3D World in A Grain of Sand | pdf |
| beings | MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillation | pdf |
| general | Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval | pdf |
| general | UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping | pdf |
| general | Apollo: An Exploration of Video Understanding in Large Multimodal Models | pdf |
| general | Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves | pdf |
| general | PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation | pdf |
| motion | MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos | pdf |
| general | Robust-MVTON: Learning Cross-Pose Feature Alignment and Fusion for Robust Multi-View Virtual Try-On | pdf |
| general | Identity-Preserving Text-to-Video Generation by Frequency Decomposition | pdf |
| gsplat | FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity | pdf |
| general | MOS: Modeling Object-Scene Associations in Generalized Category Discovery | pdf |
| general | Test-time Augmentation Improves Efficiency in Conformal Prediction | pdf |
| general | StoryGPT-V: Large Language Models as Consistent Story Visualizers | pdf |
| general | Edge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional Conditioning | pdf |
| general | INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations | pdf |
| gsplat | EVPGS: Enhanced View Prior Guidance for Splatting-based Extrapolated View Synthesis | pdf |
| general | GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding | pdf |
| general | Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal Perception | pdf |
| general | MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations | pdf |
| general | UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing | pdf |
| general | Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph Learning | pdf |
| general | FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement | pdf |
| general | End-to-End Implicit Neural Representations for Classification | pdf |
| general | UNICL-SAM: Uncertainty-Driven In-Context Segmentation with Part Prototype Discovery | pdf |
| motion | Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos | pdf |
| general | Diffusion Self-Distillation for Zero-Shot Customized Image Generation | pdf |
| general | Uncertainty-guided Perturbation for Image Super-Resolution Diffusion Model | pdf |
| beings | Towards Human-Understandable Multi-Dimensional Concept Discovery | pdf |
| general | ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval | pdf |
| general | Perturb-and-Revise: Flexible 3D Editing with Generative Trajectories | pdf |
| general | Learning Compatible Multi-Prize Subnetworks for Asymmetric Retrieval | pdf |
| general | LoRACLR: Contrastive Adaptation for Customization of Diffusion Models | pdf |
| general | Opportunistic Single-Photon Time of Flight | pdf |
| general | Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought | pdf |
| general | Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations | pdf |
| general | Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis | pdf |
| general | Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis | pdf |
| general | Improving Personalized Search with Regularized Low-Rank Parameter Updates | pdf |
| general | HyperLoRA: Parameter-Efficient Adaptive Generation for Portrait Synthesis | pdf |
| general | EchoMatch: Partial-to-Partial Shape Matching via Correspondence Reflection | pdf |
| general | Cross-Modal Interactive Perception Network with Mamba for Lung Tumor Segmentation in PET-CT Images | pdf |
| general | SuperPC: A Single Diffusion Model for Point Cloud Completion, Upsampling, Denoising, and Colorization | pdf |
| general | Maintaining Consistent Inter-Class Topology in Continual Test-Time Adaptation | pdf |
| gsplat | Generalized Gaussian Entropy Model for Point Cloud Attribute Compression with Dynamic Likelihood Intervals | pdf |
| general | Self-Learning Hyperspectral and Multispectral Image Fusion via Adaptive Residual Guided Subspace Diffusion Model | pdf |
| beings | StickMotion: Generating 3D Human Motions by Drawing a Stickman | pdf |
| general | Enduring, Efficient and Robust Trajectory Prediction Attack in Autonomous Driving via Optimization-Driven Multi-Frame Perturbation Framework | pdf |
| general | Toward Generalized Image Quality Assessment: Relaxing the Perfect Reference Quality Assumption | pdf |
| general | Hazy Low-Quality Satellite Video Restoration Via Learning Optimal Joint Degradation Patterns and Continuous-Scale Super-Resolution Reconstruction | pdf |
| general | Overcoming Shortcut Problem in VLM for Robust Out-of-Distribution Detection | pdf |
| general | RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics | pdf |
| gsplat | BG-Triangle: Bezier Gaussian Triangle for 3D Vectorization and Rendering | pdf |
| general | TKG-DM: Training-free Chroma Key Content Generation Diffusion Model | pdf |
| general | Lift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulation | pdf |
| general | Multi-View Pose-Agnostic Change Localization with Zero Labels | pdf |
| general | Accelerating Diffusion Transformer via Increment-Calibrated Caching with Channel-Aware Singular Value Decomposition | pdf |
| worlds | A Simple yet Effective Layout Token in Large Language Models for Document Understanding | pdf |
| general | Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models | pdf |
| beings | MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention | pdf |
| general | Free Lunch Enhancements for Multi-modal Crowd Counting | pdf |
| gsplat | EVolSplat: Efficient Volume-based Gaussian Splatting for Urban View Synthesis | pdf |
| general | PIDSR: Complementary Polarized Image Demosaicing and Super-Resolution | pdf |
| general | MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data | pdf |
| general | HandOS: 3D Hand Reconstruction in One Stage | pdf |
| general | All-Day Multi-Camera Multi-Target Tracking | pdf |
| beings | EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space | pdf |
| general | StarVector: Generating Scalable Vector Graphics Code from Images and Text | pdf |
| general | Explaining in Diffusion: Explaining a Classifier with Diffusion Semantics | pdf |
| general | Attention Distillation: A Unified Approach to Visual Characteristics Transfer | pdf |
| general | From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing | pdf |
| general | DreamRelation: Bridging Customization and Relation Generation | pdf |
| gsplat | Depth-Guided Bundle Sampling for Efficient Generalizable Neural Radiance Field Reconstruction | pdf |
| general | TinyFusion: Diffusion Transformers Learned Shallow | pdf |
| gsplat | SVG-IR: Spatially-Varying Gaussian Splatting for Inverse Rendering | pdf |
| meshing | Scaling Mesh Generation via Compressive Tokenization | pdf |
| general | Towards Optimizing Large-Scale Multi-Graph Matching in Bioimaging | pdf |
| general | PS-Diffusion: Photorealistic Subject-Driven Image Editing with Disentangled Control and Attention | pdf |
| general | Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models | pdf |
| general | Few-shot Implicit Function Generation via Equivariance | pdf |
| general | vesselFM: A Foundation Model for Universal 3D Blood Vessel Segmentation | pdf |
| general | Classifier-Free Guidance Inside the Attraction Basin May Cause Memorization | pdf |
| general | Magma: A Foundation Model for Multimodal AI Agents | pdf |
| general | Volume Tells: Dual Cycle-Consistent Diffusion for 3D Fluorescence Microscopy De-noising and Super-Resolution | pdf |
| general | Matrix3D: Large Photogrammetry Model All-in-One | pdf |
| general | 3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement | pdf |
| general | Investigating the Role of Weight Decay in Enhancing Nonconvex SGD | pdf |
| general | MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures | pdf |
| general | Detecting Backdoor Attacks in Federated Learning via Direction Alignment Inspection | pdf |
| motion | BlockDance: Reuse Structurally Similar Spatio-Temporal Features to Accelerate Diffusion Transformers | pdf |
| general | Mamba-Adaptor: State Space Model Adaptor for Visual Recognition | pdf |
| general | Robust Message Embedding via Attention Flow-Based Steganography | pdf |
| general | Compositional Targeted Multi-Label Universal Perturbations | pdf |
| general | PatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo Anomalies | pdf |
| general | Neural Video Compression with Context Modulation | pdf |
| general | On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only Events | pdf |
| general | Learning with Noisy Triplet Correspondence for Composed Image Retrieval | pdf |
| general | Parallelized Autoregressive Visual Generation | pdf |
| general | CGMatch: A Different Perspective of Semi-supervised Learning | pdf |
| general | FIction: 4D Future Interaction Prediction from Video | pdf |
| general | D^2iT: Dynamic Diffusion Transformer for Accurate Image Generation | pdf |
| motion | AniDoc: Animation Creation Made Easier | pdf |
| general | LiSu: A Dataset and Method for LiDAR Surface Normal Estimation | pdf |
| motion | Spk2SRImgNet: Super-Resolve Dynamic Scene from Spike Stream via Motion Aligned Collaborative Filtering | pdf |
| general | VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in Videos | pdf |
| general | ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping | pdf |
| general | PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic Manipulation | pdf |
| general | Dense Dispersed Structured Light for Hyperspectral 3D Imaging of Dynamic Scenes | pdf |
| general | MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments | pdf |
| general | Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders | pdf |
| general | Once-Tuning-Multiple-Variants: Tuning Once and Expanded as Multiple Vision-Language Model Variants | pdf |
| general | Reconstructing Animals and the Wild | pdf |
| general | DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving | pdf |
| general | DVHGNN: Multi-Scale Dilated Vision HGNN for Efficient Vision Recognition | pdf |
| beings | Reconstructing In-the-Wild Open-Vocabulary Human-Object Interactions | pdf |
| general | GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill | pdf |
| gsplat | GauSTAR: Gaussian Surface Tracking and Reconstruction | pdf |
| general | Training-free Dense-Aligned Diffusion Guidance for Modular Conditional Image Synthesis | pdf |
| general | TADFormer: Task-Adaptive Dynamic TransFormer for Efficient Multi-Task Learning | pdf |
| general | CamPoint: Boosting Point Cloud Segmentation with Virtual Camera | pdf |
| general | MERGE: Multi-faceted Hierarchical Graph-based GNN for Gene Expression Prediction from Whole Slide Histopathology Images | pdf |
| general | SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Models | pdf |
| general | PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion Model | pdf |
| general | Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation | pdf |
| general | Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration | pdf |
| general | FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation | pdf |
| general | Visual Lexicon: Rich Image Features in Language Space | pdf |
| general | Test-Time Visual In-Context Tuning | pdf |
| general | Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge Models | pdf |
| general | SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Images | pdf |
| general | Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models? | pdf |
| general | Harnessing Global-Local Collaborative Adversarial Perturbation for Anti-Customization | pdf |
| general | Acc3D: Accelerating Single Image to 3D Diffusion Models via Edge Consistency Guided Score Distillation | pdf |
| general | Soft Self-labeling and Potts Relaxations for Weakly-supervised Segmentation | pdf |
| general | MVSAnywhere: Zero-Shot Multi-View Stereo | pdf |
| general | Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision | pdf |
| general | BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature | pdf |
| general | Structure-Aware Correspondence Learning for Relative Pose Estimation | pdf |
| general | PyTorchGeoNodes: Enabling Differentiable Shape Programs for 3D Shape Reconstruction | pdf |
| general | FIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze Estimation | pdf |
| general | Shape Abstraction via Marching Differentiable Support Functions | pdf |
| general | Scaling Down Text Encoders of Text-to-Image Diffusion Models | pdf |
| general | POT: Prototypical Optimal Transport for Weakly Supervised Semantic Segmentation | pdf |
| worlds | SKE-Layout: Spatial Knowledge Enhanced Layout Generation with LLMs | pdf |
| gsplat | Gaussian Eigen Models for Human Heads | pdf |
| general | 4D-Fly: Fast 4D Reconstruction from a Single Monocular Video | pdf |
| general | Complementary Advantages: Exploiting Cross-Field Frequency Correlation for NIR-Assisted Image Denoising | pdf |
| general | Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation | pdf |
| general | DiffLO: Semantic-Aware LiDAR Odometry with Diffusion-Based Refinement | pdf |
| general | Style-Editor: Text-driven Object-centric Style Editing | pdf |
| general | Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene | pdf |
| general | FastVLM: Efficient Vision Encoding for Vision Language Models | pdf |
| general | VISTA3D: A Unified Segmentation Foundation Model For 3D Medical Imaging | pdf |
| general | S2D-LFE: Sparse-to-Dense Light Field Event Generation | pdf |
| general | The Art of Deception: Color Visual Illusions and Diffusion Models | pdf |
| meshing | Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Data | pdf |
| beings | Do Computer Vision Foundation Models Learn the Low-level Characteristics of the Human Visual System? | pdf |
| general | Online Task-Free Continual Learning via Dynamic Expansionable Memory Distribution | pdf |
| general | Rethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level Tasks | pdf |
| general | SVDC: Consistent Direct Time-of-Flight Video Depth Completion with Frequency Selective Fusion | pdf |
| general | Rethinking Training for De-biasing Text-to-Image Generation: Unlocking the Potential of Stable Diffusion | pdf |
| general | Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects | pdf |
| general | STDD: Spatio-Temporal Dual Diffusion for Video Generation | pdf |
| general | Implicit Correspondence Learning for Image-to-Point Cloud Registration | pdf |
| general | ILIAS: Instance-Level Image retrieval At Scale | pdf |
| general | GeoDepth: From Point-to-Depth to Plane-to-Depth Modeling for Self-Supervised Monocular Depth Estimation | pdf |
| general | SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split Optimization | pdf |
| gsplat | USP-Gaussian: Unifying Spike-based Image Reconstruction, Pose Correction and Gaussian Splatting | pdf |
| general | Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity | pdf |
| general | QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge | pdf |
| general | ReWind: Understanding Long Videos with Instructed Learnable Memory | pdf |
| gsplat | DirectTriGS: Triplane-based Gaussian Splatting Field Representation for 3D Generation | pdf |
| motion | Make-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Characters | pdf |
| general | Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning | pdf |
| worlds | NeRFPrior: Learning Neural Radiance Field as a Prior for Indoor Scene Reconstruction | pdf |
| general | Towards Training-free Anomaly Detection with Vision and Language Foundation Models | pdf |
| beings | Dynamic Content Prediction with Motion-aware Priors for Blind Face Video Restoration | pdf |
| general | Exploring Sparse MoE in GANs for Text-conditioned Image Synthesis | pdf |
| general | Efficient Event-Based Object Detection: A Hybrid Neural Network with Spatial and Temporal Attention | pdf |
| general | HUNet: Homotopy Unfolding Network for Image Compressive Sensing | pdf |
| general | See Further When Clear: Curriculum Consistency Model | pdf |
| general | PassionSR: Post-Training Quantization with Adaptive Scale in One-Step Diffusion based Image Super-Resolution | pdf |
| gsplat | RainyGS: Efficient Rain Synthesis with Physically-Based Gaussian Splatting | pdf |
| general | Three-view Focal Length Recovery From Homographies | pdf |
| general | RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models | pdf |
| general | Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation | pdf |
| general | Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos | pdf |
| general | FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation | pdf |
| general | InterDyn: Controllable Interactive Dynamics with Video Diffusion Models | pdf |
| general | LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models | pdf |
| general | MagicQuill: An Intelligent Interactive Image Editing System | pdf |
| worlds | Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces | pdf |
| general | Boosting Adversarial Transferability through Augmentation in Hypothesis Space | pdf |
| meshing | ViiNeuS: Volumetric Initialization for Implicit Neural Surface Reconstruction of Urban Scenes with Limited Image Overlap | pdf |
| general | Model Diagnosis and Correction via Linguistic and Implicit Attribute Editing | pdf |
| general | UniPhy: Learning a Unified Constitutive Model for Inverse Physics Simulation | pdf |
| general | STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural Networks | pdf |
| general | Knowledge Memorization and Rumination for Pre-trained Model-based Class-Incremental Learning | pdf |
| general | Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval | pdf |
| general | Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning | pdf |
| general | VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling | pdf |
| beings | Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification | pdf |
| general | Exposure-slot: Exposure-centric Representations Learning with Slot-in-Slot Attention for Region-aware Exposure Correction | pdf |
| general | EdgeDiff: Edge-aware Diffusion Network for Building Reconstruction from Point Clouds | pdf |
| general | DeNVeR: Deformable Neural Vessel Representations for Unsupervised Video Vessel Segmentation | pdf |
| general | Task-Aware Clustering for Prompting Vision-Language Models | pdf |
| general | FSboard: Over 3 Million Characters of ASL Fingerspelling Collected via Smartphones | pdf |
| general | Light Transport-aware Diffusion Posterior Sampling for Single-View Reconstruction of 3D Volumes | pdf |
| general | STiL: Semi-supervised Tabular-Image Learning for Comprehensive Task-Relevant Information Exploration in Multimodal Classification | pdf |
| general | Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language | pdf |
| general | ReRAW: RGB-to-RAW Image Reconstruction via Stratified Sampling for Efficient Object Detection on the Edge | pdf |
| motion | HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation | pdf |
| general | Zero-shot 3D Question Answering via Voxel-based Dynamic Token Compression | pdf |
| general | Enhanced then Progressive Fusion with View Graph for Multi-View Clustering | pdf |
| general | Correcting Deviations from Normality: A Reformulated Diffusion Model for Multi-Class Unsupervised Anomaly Detection | pdf |
| general | Continuous 3D Perception Model with Persistent State | pdf |
| worlds | LP-Diff: Towards Improved Restoration of Real-World Degraded License Plate | pdf |
| general | FilmComposer: LLM-Driven Music Production for Silent Film Clips | pdf |
| general | EventPSR: Surface Normal and Reflectance Estimation from Photometric Stereo Using an Event Camera | pdf |
| general | CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional Video | pdf |
| beings | Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual Foraging | pdf |
| general | FOCUS: Knowledge-enhanced Adaptive Visual Compression for Few-shot Whole Slide Image Classification | pdf |
| general | GRAE-3DMOT: Geometry Relation-Aware Encoder for Online 3D Multi-Object Tracking | pdf |
| general | Automatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compression | pdf |
| general | MAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generation | pdf |
| beings | Synthetic Prior for Few-Shot Drivable Head Avatar Inversion | pdf |
| general | Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems Approach | pdf |
| general | DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation | pdf |
| beings | GroomLight: Hybrid Inverse Rendering for Relightable Human Hair Appearance Modeling | pdf |
| general | Sea-ing in Low-light | pdf |
| general | Generative Modeling of Class Probability for Multi-Modal Representation Learning | pdf |
| general | VisionZip: Longer is Better but Not Necessary in Vision Language Models | pdf |
| general | BlenderGym: Benchmarking Foundational Model Systems for Graphics Editing | pdf |
| general | VoteFlow: Enforcing Local Rigidity in Self-Supervised Scene Flow | pdf |
| general | Uncertainty Weighted Gradients for Model Calibration | pdf |
| gsplat | FFaceNeRF: Few-shot Face Editing in Neural Radiance Fields | pdf |
| general | Minimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance Segmentation | pdf |
| general | Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformers | pdf |
| general | Zero-shot RGB-D Point Cloud Registration with Pre-trained Large Vision Model | pdf |
| general | DistinctAD: Distinctive Audio Description Generation in Contexts | pdf |
| general | CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering | pdf |
| general | Point Cloud Upsampling Using Conditional Diffusion Module with Adaptive Noise Suppression | pdf |
| general | Trajectory Mamba: Efficient Attention-Mamba Forecasting Model Based on Selective SSM | pdf |
| beings | VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis | pdf |
| general | DEIM: DETR with Improved Matching for Fast Convergence | pdf |
| beings | Human Motion Instruction Tuning | pdf |
| general | A Flag Decomposition for Hierarchical Datasets | pdf |
| general | RCP-Bench: Benchmarking Robustness for Collaborative Perception Under Diverse Corruptions | pdf |
| general | Olympus: A Universal Task Router for Computer Vision Tasks | pdf |
| general | Circumventing Shortcuts in Audio-visual Deepfake Detection Datasets with Unsupervised Learning | pdf |
| general | Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching | pdf |
| general | Improving Semi-Supervised Semantic Segmentation with Sliced-Wasserstein Feature Alignment and Uniformity | pdf |
| general | 3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning | pdf |
| worlds | Navigation World Models | pdf |
| general | Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation | pdf |
| beings | Reconstructing Close Human Interaction with Appearance and Proxemics Reasoning | pdf |
| general | Scenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environments | pdf |
| general | Poly-Autoregressive Prediction for Modeling Interactions | pdf |
| beings | PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation | pdf |
| general | Decision SpikeFormer: Spike-Driven Transformer for Decision Making | pdf |
| general | Theory-Inspired Deep Multi-View Multi-Label Learning with Incomplete Views and Noisy Labels | pdf |
| general | EMOE: Modality-Specific Enhanced Dynamic Emotion Experts | pdf |
| general | Generative Video Propagation | pdf |
| general | From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons | pdf |
| general | Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation | pdf |
| general | T-CIL: Temperature Scaling using Adversarial Perturbation for Calibration in Class-Incremental Learning | pdf |
| general | LoRA Subtraction for Drift-Resistant Space in Exemplar-Free Continual Learning | pdf |
| general | AniMer: Animal Pose and Shape Estimation Using Family Aware Transformer | pdf |
| general | Co-op: Correspondence-based Novel Object Pose Estimation | pdf |
| general | CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-Resolution | pdf |
| general | RayFlow: Instance-Aware Diffusion Acceleration via Adaptive Flow Trajectories | pdf |
| general | Self-Supervised Large Scale Point Cloud Completion for Archaeological Site Restoration | pdf |
| general | Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks | pdf |
| general | Rate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty Estimation | pdf |
| general | Thin-Shell-SfT: Fine-Grained Monocular Non-rigid 3D Surface Tracking with Neural Deformation Fields | pdf |
| general | DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception | pdf |
| general | SocialGesture: Delving into Multi-person Gesture Understanding | pdf |
| general | Multi-modal Topology-embedded Graph Learning for Spatially Resolved Genes Prediction from Pathology Images with Prior Gene Similarity Information | pdf |
| gsplat | Question-Aware Gaussian Experts for Audio-Visual Question Answering | pdf |
| general | Adaptive Rectangular Convolution for Remote Sensing Pansharpening | pdf |
| general | UIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion Models | pdf |
| general | FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation | pdf |
| general | AdMiT: Adaptive Multi-Source Tuning in Dynamic Environments | pdf |
| general | Unbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasing | pdf |
| general | Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation | pdf |
| general | Revisiting Generative Replay for Class Incremental Object Detection | pdf |
| general | Bridging Viewpoint Gaps: Geometric Reasoning Boosts Semantic Correspondence | pdf |
| general | Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis | pdf |
| beings | MobileH2R: Learning Generalizable Human to Mobile Robot Handover Exclusively from Scalable and Diverse Synthetic Data | pdf |
| beings | PERSE: Personalized 3D Generative Avatars from A Single Portrait | pdf |
| general | Dynamic Stereotype Theory Induced Micro-expression Recognition with Oriented Deformation | pdf |
| general | ACL: Activating Capability of Linear Attention for Image Restoration | pdf |
| general | MARBLE: Material Recomposition and Blending in CLIP-Space | pdf |
| general | Efficient Visual State Space Model for Image Deblurring | pdf |
| general | Enhancing 3D Gaze Estimation in the Wild using Weak Supervision with Gaze Following Labels | pdf |
| general | Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward | pdf |
| general | Detecting Out-of-Distribution Through the Lens of Neural Collapse | pdf |
| gsplat | Sparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse Views | pdf |
| general | VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation | pdf |
| gsplat | Efficient Decoupled Feature 3D Gaussian Splatting via Hierarchical Compression | pdf |
| general | CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model | pdf |
| general | SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images | pdf |
| general | Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation | pdf |
| general | Label Shift Meets Online Learning: Ensuring Consistent Adaptation with Universal Dynamic Regret | pdf |
| general | A Physics-Informed Blur Learning Framework for Imaging Systems | pdf |
| general | Towards Practical Real-Time Neural Video Compression | pdf |
| gsplat | DepthSplat: Connecting Gaussian Splatting and Depth | pdf |
| general | Dynamic Camera Poses and Where to Find Them | pdf |
| general | OmniGen: Unified Image Generation | pdf |
| general | QuCOOP: A Versatile Framework for Solving Composite and Binary-Parametrised Problems on Quantum Annealers | pdf |
| meshing | Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured Meshes | pdf |
| general | SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation | pdf |
| general | Calibrated Multi-Preference Optimization for Aligning Diffusion Models | pdf |
| gsplat | Advancing Adversarial Robustness in GNeRFs: The IL2-NeRF Attack | pdf |
| general | PolarNeXt: Rethink Instance Segmentation with Polar Representation | pdf |
| general | SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything Model | pdf |
| general | DarkIR: Robust Low-Light Image Restoration | pdf |
| general | R2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action Planner | pdf |
| general | From Prototypes to General Distributions: An Efficient Curriculum for Masked Image Modeling | pdf |
| general | Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation | pdf |
| general | MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting | pdf |
| general | Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions | pdf |
| general | Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation | pdf |
| general | Evolving High-Quality Rendering and Reconstruction in a Unified Framework with Contribution-Adaptive Regularization | pdf |
| general | NoiseCtrl: A Sampling-Algorithm-Agnostic Conditional Generation Method for Diffusion Models | pdf |
| general | KMD: Koopman Multi-modality Decomposition for Generalized Brain Tumor Segmentation under Incomplete Modalities | pdf |
| general | DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-Resolution | pdf |
| general | Fractal Calibration for Long-tailed Object Detection | pdf |
| general | M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world Settings | pdf |
| general | FedSPA: Generalizable Federated Graph Learning under Homophily Heterogeneity | pdf |
| general | GazeGene: Large-scale Synthetic Gaze Dataset with 3D Eyeball Annotations | pdf |
| general | VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors | pdf |
| gsplat | GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding | pdf |
| general | Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions | pdf |
| general | SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment | pdf |
| general | Improved Video VAE for Latent Video Diffusion Model | pdf |
| general | Efficient Video Super-Resolution for Real-time Rendering with Decoupled G-buffer Guidance | pdf |
| general | Learned Image Compression with Dictionary-based Entropy Model | pdf |
| general | FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model | pdf |
| general | DL2G: Degradation-guided Local-to-Global Restoration for Eyeglass Reflection Removal | pdf |
| general | MFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and Forecasting | pdf |
| general | The Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion Models | pdf |
| general | Leveraging Global Stereo Consistency for Category-Level Shape and 6D Pose Estimation from Stereo Images | pdf |
| general | AlphaPre: Amplitude-Phase Disentanglement Model for Precipitation Nowcasting | pdf |
| general | Detection-Friendly Nonuniformity Correction: A Union Framework for Infrared UAV Target Detection | pdf |
| general | Articulated Kinematics Distillation from Video Diffusion Models | pdf |
| general | ExpertAF: Expert Actionable Feedback from Video | pdf |
| gsplat | Volumetrically Consistent 3D Gaussian Rasterization | pdf |
| general | The Impact Label Noise and Choice of Threshold has on Cross-Entropy and Soft-Dice in Image Segmentation | pdf |
| general | LLaVA-Critic: Learning to Evaluate Multimodal Models | pdf |
| general | VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge | pdf |
| general | Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation | pdf |
| general | Large-scale Multi-view Tensor Clustering with Implicit Linear Kernels | pdf |
| general | Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content | pdf |
| general | Dual Focus-Attention Transformer for Robust Point Cloud Registration | pdf |
| general | Forming Auxiliary High-confident Instance-level Loss to Promote Learning from Label Proportions | pdf |
| general | Progress-Aware Video Frame Captioning | pdf |
| general | SMTPD: A New Benchmark for Temporal Prediction of Social Media Popularity | pdf |
| general | Learning on Model Weights using Tree Experts | pdf |
| general | Image Reconstruction from Readout-Multiplexed Single-Photon Detector Arrays | pdf |
| general | Towards Transformer-Based Aligned Generation with Self-Coherence Guidance | pdf |
| general | Accurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-Distillation | pdf |
| general | DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report Generation | pdf |
| general | On the Consistency of Video Large Language Models in Temporal Comprehension | pdf |
| general | Less is More: Efficient Model Merging with Binary Task Switch | pdf |
| general | One-Minute Video Generation with Test-Time Training | pdf |
| general | InteractionMap: Improving Online Vectorized HDMap Construction with Interaction | pdf |
| worlds | ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting | pdf |
| general | RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness | pdf |
| gsplat | EditSplat: Multi-View Fusion and Attention-Guided Optimization for View-Consistent 3D Scene Editing with 3D Gaussian Splatting | pdf |
| general | One-shot 3D Object Canonicalization based on Geometric and Semantic Consistency | pdf |
| general | Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning | pdf |
| general | ProHOC: Probabilistic Hierarchical Out-of-Distribution Classification via Multi-Depth Networks | pdf |
| general | CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP | pdf |
| general | Graph-Embedded Structure-Aware Perceptual Hashing for Neural Network Protection and Piracy Detection | pdf |
| general | Interleaved-Modal Chain-of-Thought | pdf |
| general | Enhancing Adversarial Transferability with Checkpoints of a Single Model's Training | pdf |
| general | O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models | pdf |
| general | Analyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose Estimation | pdf |
| gsplat | Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields | pdf |
| general | Hyperspectral Pansharpening via Diffusion Models with Iteratively Zero-Shot Guidance | pdf |
| general | EASEMVC:Efficient Dual Selection Mechanism for Deep Multi-View Clustering | pdf |
| general | DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering | pdf |
| general | IceDiff: High Resolution and High-Quality Arctic Sea Ice Forecasting with Generative Diffusion Prior | pdf |
| general | DTOS: Dynamic Time Object Sensing with Large Multimodal Model | pdf |
| general | How to Merge Your Multimodal Models Over Time? | pdf |
| general | Identifying and Mitigating Position Bias of Multi-image Vision-Language Models | pdf |
| general | Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation | pdf |
| general | ShowUI: One Vision-Language-Action Model for GUI Visual Agent | pdf |
| general | Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis | pdf |
| beings | HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation | pdf |
| general | ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos | pdf |
| general | ArtiFade: Learning to Generate High-quality Subject from Blemished Images | pdf |
| general | Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation | pdf |
| general | GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery | pdf |
| general | Test-Time Domain Generalization via Universe Learning: A Multi-Graph Matching Approach for Medical Image Segmentation | pdf |
| general | DeDe: Detecting Backdoor Samples for SSL Encoders via Decoders | pdf |
| beings | Towards Scalable Human-aligned Benchmark for Text-guided Image Editing | pdf |
| general | Coeff-Tuning: A Graph Filter Subspace View for Tuning Attention-Based Large Models | pdf |
| general | Self-Supervised Cross-View Correspondence with Predictive Cycle Consistency | pdf |
| beings | CryptoFace: End-to-End Encrypted Face Recognition | pdf |
| general | Relation-Rich Visual Document Generator for Visual Information Extraction | pdf |
| general | PromptHash:Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrieval | pdf |
| general | Universal Scene Graph Generation | pdf |
| general | Split Adaptation for Pre-trained Vision Transformers | pdf |
| general | SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models | pdf |
| general | Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking | pdf |
| general | Plug-and-Play Versatile Compressed Video Enhancement | pdf |
| general | UltraFusion: Ultra High Dynamic Imaging using Exposure Fusion | pdf |
| general | Noise-Resistant Video Anomaly Detection via RGB Error-Guided Multiscale Predictive Coding and Dynamic Memory | pdf |
| general | GroupMamba: Efficient Group-Based Visual State Space Model | pdf |
| general | Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spaces | pdf |
| general | ActiveGAMER: Active GAussian Mapping through Efficient Rendering | pdf |
| general | Positive2Negative: Breaking the Information-Lossy Barrier in Self-Supervised Single Image Denoising | pdf |
| general | SeaLion: Semantic Part-Aware Latent Point Diffusion Models for 3D Generation | pdf |
| gsplat | Toward Real-world BEV Perception: Depth Uncertainty Estimation via Gaussian Splatting | pdf |
| general | Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models | pdf |
| gsplat | WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments | pdf |
| beings | RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance | pdf |
| worlds | CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning | pdf |
| general | Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method | pdf |
| general | AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM | pdf |
| general | Autoregressive Distillation of Diffusion Transformers | pdf |
| general | OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints | pdf |
| general | DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval | pdf |
| general | Visual-Instructed Degradation Diffusion for All-in-One Image Restoration | pdf |
| general | Insightful Instance Features for 3D Instance Segmentation | pdf |
| general | EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing | pdf |
| general | A Hubness Perspective on Representation Learning for Graph-Based Multi-View Clustering | pdf |
| general | Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation | pdf |
| general | ZeroVO: Visual Odometry with Minimal Assumptions | pdf |
| general | VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM | pdf |
| general | HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding | pdf |
| general | Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects | pdf |
| worlds | SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling | pdf |
| general | AdaDARE-gamma: Balancing Stability and Plasticity in Multi-modal LLMs through Efficient Adaptation | pdf |
| general | LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis | pdf |
| beings | SapiensID: Foundation for Human Recognition | pdf |
| motion | Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think | pdf |
| beings | FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling | pdf |
| gsplat | InstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perception | pdf |
| general | CSC-PA: Cross-image Semantic Correlation via Prototype Attentions for Single-network Semi-supervised Breast Tumor Segmentation | pdf |
| general | VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving | pdf |
| general | Detecting Adversarial Data Using Perturbation Forgery | pdf |
| general | CoA: Towards Real Image Dehazing via Compression-and-Adaptation | pdf |
| general | TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model | pdf |
| general | Learned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus Cues | pdf |
| beings | MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices | pdf |
| motion | Light3R-SfM: Towards Feed-forward Structure-from-Motion | pdf |
| general | Robotic Visual Instruction | pdf |
| general | MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors | pdf |
| general | Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learning | pdf |
| general | Cross-modal Information Flow in Multimodal Large Language Models | pdf |
| general | Keyframe-Guided Creative Video Inpainting | pdf |
| general | EdgeTAM: On-Device Track Anything Model | pdf |
| general | EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues | pdf |
| general | Video Summarization with Large Language Models | pdf |
| general | Sketchtopia: A Dataset and Foundational Agents for Benchmarking Asynchronous Multimodal Communication with Iconic Feedback | pdf |
| general | Consistency-aware Self-Training for Iterative-based Stereo Matching | pdf |
| general | MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts | pdf |
| general | Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model | pdf |
| family | paper | links |
| general | Deterministic Image-to-Image Translation via Denoising Brownian Bridge Models with Dual Approximators | pdf |
| general | Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment | pdf |
| general | Cross-modal Causal Relation Alignment for Video Question Grounding | pdf |
| general | Diffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion Models | pdf |
| gsplat | Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction | pdf |
| general | 3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion | pdf |
| worlds | Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval | pdf |
| general | DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation | pdf |
| general | Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions | pdf |
| general | CARL: A Framework for Equivariant Image Registration | pdf |
| gsplat | FlashGS: Efficient 3D Gaussian Splatting for Large-scale and High-resolution Rendering | pdf |
| general | Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models | pdf |
| general | Inference-Scale Complexity in ANN-SNN Conversion for High-Performance and Low-Power Applications | pdf |
| beings | MVDoppler-Pose: Multi-Modal Multi-View mmWave Sensing for Long-Distance Self-Occluded Human Walking Pose Estimation | pdf |
| general | TopNet: Transformer-Efficient Occupancy Prediction Network for Octree-Structured Point Cloud Geometry Compression | pdf |
| general | Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring Classes | pdf |
| general | M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation | pdf |
| general | Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment | pdf |
| general | Multi-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain Adaptation | pdf |
| motion | A Polarization-Aided Transformer for Image Deblurring via Motion Vector Decomposition | pdf |
| general | CocoER: Aligning Multi-Level Feature by Competition and Coordination for Emotion Recognition | pdf |
| general | Enhancing Creative Generation on Stable Diffusion-based Models | pdf |
| general | Denoising Functional Maps: Diffusion Models for Shape Correspondence | pdf |
| general | ProReflow: Progressive Reflow with Decomposed Velocity | pdf |
| general | Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention | pdf |
| general | MetaShadow: Object-Centered Shadow Detection, Removal, and Synthesis | pdf |
| general | TANGO: Training-free Embodied AI Agents for Open-world Tasks | pdf |
| general | Stealthy Backdoor Attack in Self-Supervised Learning Vision Encoders for Large Vision Language Models | pdf |
| general | SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes | pdf |
| general | GIVEPose: Gradual Intra-class Variation Elimination for RGB-based Category-Level Object Pose Estimation | pdf |
| beings | Sketch Down the FLOPs: Towards Efficient Networks for Human Sketch | pdf |
| general | Rethinking Decoder Design: Improving Biomarker Segmentation Using Depth-to-Space Restoration and Residual Linear Attention | pdf |
| general | SynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesis | pdf |
| beings | Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative Editing | pdf |
| general | Improving Accuracy and Calibration via Differentiated Deep Mutual Learning | pdf |
| general | Infighting in the Dark: Multi-Label Backdoor Attack in Federated Learning | pdf |
| general | Tartan IMU: A Light Foundation Model for Inertial Positioning in Robotics | pdf |
| general | Event Ellipsometer: Event-based Mueller-Matrix Video Imaging | pdf |
| general | End-to-End HOI Reconstruction Transformer with Graph-based Encoding | pdf |
| beings | Disco4D: Disentangled 4D Human Generation and Animation from a Single Image | pdf |
| beings | IDOL: Instant Photorealistic 3D Human Creation from a Single Image | pdf |
| general | SketchVideo: Sketch-based Video Generation and Editing | pdf |
| general | Taste More, Taste Better: Diverse Data and Strong Model Boost Semi-Supervised Crowd Counting | pdf |
| beings | AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models | pdf |
| general | Latent Space Imaging | pdf |
| general | Balanced Direction from Multifarious Choices: Arithmetic Meta-Learning for Domain Generalization | pdf |
| general | Anatomical Consistency and Adaptive Prior-informed Transformation for Multi-contrast MR Image Synthesis via Diffusion Model | pdf |
| general | SeCap: Self-Calibrating and Adaptive Prompts for Cross-view Person Re-Identification in Aerial-Ground Networks | pdf |
| general | Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving | pdf |
| worlds | Neural Motion Simulator Pushing the Limit of World Models in Reinforcement Learning | pdf |
| worlds | Adversarial Diffusion Compression for Real-World Image Super-Resolution | pdf |
| general | DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery | pdf |
| general | SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters | pdf |
| general | EntropyMark: Towards More Harmless Backdoor Watermark via Entropy-based Constraint for Open-source Dataset Copyright Protection | pdf |
| general | Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding | pdf |
| general | Towards Universal AI-Generated Image Detection by Variational Information Bottleneck Network | pdf |
| general | HSI: A Holistic Style Injector for Arbitrary Style Transfer | pdf |
| general | V2V3D: View-to-View Denoised 3D Reconstruction for Light Field Microscopy | pdf |
| gsplat | Splatter-360: Generalizable 360 Gaussian Splatting for Wide-baseline Panoramic Images | pdf |
| general | Towards Understanding How Knowledge Evolves in Large Vision-Language Models | pdf |
| general | A Unified, Resilient, and Explainable Adversarial Patch Detector | pdf |
| general | Structured 3D Latents for Scalable and Versatile 3D Generation | pdf |
| general | Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects | pdf |
| general | Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial Attacks | pdf |
| general | Fish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Images | pdf |
| general | PCM : Picard Consistency Model for Fast Parallel Sampling of Diffusion Models | pdf |
| gsplat | CoMapGS: Covisibility Map-based Gaussian Splatting for Sparse Novel View Synthesis | pdf |
| general | Training Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training? | pdf |
| general | Improving the Training of Data-Efficient GANs via Quality Aware Dynamic Discriminator Rejection Sampling | pdf |
| motion | MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation | pdf |
| general | Advancing Generalizable Tumor Segmentation with Anomaly-Aware Open-Vocabulary Attention Maps and Frozen Foundation Diffusion Models | pdf |
| general | Towards Generalizable Scene Change Detection | pdf |
| general | Incomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space Model | pdf |
| general | FedAWA: Adaptive Optimization of Aggregation Weights in Federated Learning Using Client Vectors | pdf |
| beings | Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression | pdf |
| general | Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget | pdf |
| beings | Guiding Human-Object Interactions with Rich Geometry and Relations | pdf |
| general | CADDreamer: CAD Object Generation from Single-view Images | pdf |
| general | Where's the Liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated Content | pdf |
| general | DiTASK: Multi-Task Fine-Tuning with Diffeomorphic Transformations | pdf |
| worlds | OW-OVD: Unified Open World and Open Vocabulary Object Detection | pdf |
| general | Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing | pdf |
| general | DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models | pdf |
| general | SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE | pdf |
| general | Dual-Interrelated Diffusion Model for Few-Shot Anomaly Image Generation | pdf |
| general | Interactive Medical Image Analysis with Concept-based Similarity Reasoning | pdf |
| general | h-Edit: Effective and Flexible Diffusion-Based Editing via Doob's h-Transform | pdf |
| beings | Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized? | pdf |
| general | Spectral State Space Model for Rotation-Invariant Visual Representation Learning | pdf |
| general | Sharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and Manipulation | pdf |
| general | URWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image Restoration | pdf |
| general | Functionality Understanding and Segmentation in 3D Scenes | pdf |
| general | Dragin3D: Image Editing by Dragging in 3D Space | pdf |
| general | Towards Stable and Storage-efficient Dataset Distillation: Matching Convexified Trajectory | pdf |
| general | TSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual Segmentation | pdf |
| general | Invisible Backdoor Attack against Self-supervised Learning | pdf |
| meshing | Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation Metrics | pdf |
| general | BWFormer: Building Wireframe Reconstruction from Airborne LiDAR Point Cloud with Transformer | pdf |
| general | Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models | pdf |
| general | OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning | pdf |
| gsplat | MeGA: Hybrid Mesh-Gaussian Head Avatar for High-Fidelity Rendering and Head Editing | pdf |
| general | Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers | pdf |
| general | Dataset Distillation with Neural Characteristic Function: A Minmax Perspective | pdf |
| beings | Free-viewpoint Human Animation with Pose-correlated Reference Selection | pdf |
| general | PillarHist: A Quantization-aware Pillar Feature Encoder based on Height-aware Histogram | pdf |
| general | Semantic and Expressive Variations in Image Captions Across Languages | pdf |
| general | ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models | pdf |
| general | ADD: Attribution-Driven Data Augmentation Framework for Boosting Image Super-Resolution | pdf |
| general | CroCoDL: Cross-device Collaborative Dataset for Localization | pdf |
| general | CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCR | pdf |
| general | What Makes a Good Dataset for Knowledge Distillation? | pdf |
| general | Rectification-specific Supervision and Constrained Estimator for Online Stereo Rectification | pdf |
| general | Shape and Texture: What Influences Reliable Optical Flow Estimation? | pdf |
| general | Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters | pdf |
| beings | HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation | pdf |
| general | Order-One Rolling Shutter Cameras | pdf |
| general | Animate and Sound an Image | pdf |
| general | Foveated Instance Segmentation | pdf |
| general | Emphasizing Discriminative Features for Dataset Distillation in Complex Scenarios | pdf |
| general | Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation | pdf |
| general | Task-Specific Gradient Adaptation for Few-Shot One-Class Classification | pdf |
| gsplat | 3D Gaussian Inpainting with Depth-Guided Cross-View Consistency | pdf |
| general | Floxels: Fast Unsupervised Voxel Based Scene Flow Estimation | pdf |
| general | LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale | pdf |
| general | FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute | pdf |
| general | HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation | pdf |
| beings | FSFM: A Generalizable Face Security Foundation Model via Self-Supervised Facial Representation Learning | pdf |
| general | AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment | pdf |
| general | VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models | pdf |
| general | One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion | pdf |
| general | Can Text-to-Video Generation help Video-Language Alignment? | pdf |
| general | Weakly Supervised Contrastive Adversarial Training for Learning Robust Features from Semi-supervised Data | pdf |
| general | From Poses to Identity: Training-Free Person Re-Identification via Feature Centralization | pdf |
| general | MIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal Output | pdf |
| general | Bias for Action: Video Implicit Neural Representations with Bias Modulation | pdf |
| general | Segment Anything, Even Occluded | pdf |
| general | LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot Learning | pdf |
| general | Universal Actions for Enhanced Embodied Foundation Models | pdf |
| general | FaithDiff: Unleashing Diffusion Priors for Faithful Image Super-resolution | pdf |
| general | Scene-agnostic Pose Regression for Visual Localization | pdf |
| general | Divide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial Purification | pdf |
| general | SEC-Prompt:SEmantic Complementary Prompting for Few-Shot Class-Incremental Learning | pdf |
| general | LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes | pdf |
| beings | PI-HMR: Towards Robust In-bed Temporal Human Shape Reconstruction with Contact Pressure Sensing | pdf |
| general | CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation | pdf |
| general | SEEN-DA: SEmantic ENtropy guided Domain-aware Attention for Domain Adaptive Object Detection | pdf |
| general | Blind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion Model | pdf |
| general | Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking | pdf |
| worlds | GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control | pdf |
| general | Scene-Centric Unsupervised Panoptic Segmentation | pdf |
| general | Learning Physics From Video: Unsupervised Physical Parameter Estimation for Continuous Dynamical Systems | pdf |
| general | ProAPO: Progressively Automatic Prompt Optimization for Visual Classification | pdf |
| general | Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events | pdf |
| gsplat | RNG: Relightable Neural Gaussians | pdf |
| gsplat | Towards Realistic Example-based Modeling via 3D Gaussian Stitching | pdf |
| gsplat | Generative Sparse-View Gaussian Splatting | pdf |
| general | Generative Inbetweening through Frame-wise Conditions-Driven Video Generation | pdf |
| general | DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awareness | pdf |
| general | CustAny: Customizing Anything from A Single Example | pdf |
| general | PoseTraj: Pose-Aware Trajectory Control in Video Diffusion | pdf |
| general | VL2Lite: Task-Specific Knowledge Distillation from Large Vision-Language Models to Lightweight Networks | pdf |
| general | StageDesigner: Artistic Stage Generation for Scenography via Theater Scripts | pdf |
| general | Interpreting Object-level Foundation Models via Visual Precision Search | pdf |
| general | Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows | pdf |
| general | All-directional Disparity Estimation for Real-world QPD Images | pdf |
| general | Using Diffusion Priors for Video Amodal Segmentation | pdf |
| motion | Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera | pdf |
| general | The Scene Language: Representing Scenes with Programs, Words, and Embeddings | pdf |
| beings | Learning Physics-Based Full-Body Human Reaching and Grasping from Brief Walking References | pdf |
| general | EmoEdit: Evoking Emotions through Image Manipulation | pdf |
| general | SparseAlign: a Fully Sparse Framework for Cooperative Object Detection | pdf |
| general | Data Distributional Properties As Inductive Bias for Systematic Generalization | pdf |
| general | TopoCellGen: Generating Histopathology Cell Topology with a Diffusion Model | pdf |
| general | Meta-Learning Hyperparameters for Parameter Efficient Fine-Tuning | pdf |
| meshing | TriTex: Learning Texture from a Single Mesh via Triplane Semantic Features | pdf |
| general | Wavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect Detection | pdf |
| general | Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering | pdf |
| general | Language-Guided Audio-Visual Learning for Long-Term Sports Assessment | pdf |
| motion | Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation | pdf |
| beings | PersonaHOI: Effortlessly Improving Face Personalization in Human-Object Interaction Generation | pdf |
| general | Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models | pdf |
| gsplat | MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models | pdf |
| general | Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation | pdf |
| general | Probability Density Geodesics in Image Diffusion Latent Space | pdf |
| general | EgoLife: Towards Egocentric Life Assistant | pdf |
| general | BrepGiff: Lightweight Generation of Complex B-rep with 3D GAT Diffusion | pdf |
| general | Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency Partition | pdf |
| general | Joint Scheduling of Causal Prompts and Tasks for Multi-Task Learning | pdf |
| gsplat | 3DGUT: Enabling Distorted Cameras and Secondary Rays in Gaussian Splatting | pdf |
| general | It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data | pdf |
| general | Open Set Label Shift with Test Time Out-of-Distribution Reference | pdf |
| gsplat | GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction | pdf |
| general | Flexible Frame Selection for Efficient Video Reasoning | pdf |
| general | EventGPT: Event Stream Understanding with Multimodal Large Language Models | pdf |
| general | MITracker: Multi-View Integration for Visual Object Tracking | pdf |
| general | Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models | pdf |
| general | Exploring Scene Affinity for Semi-Supervised LiDAR Semantic Segmentation | pdf |
| general | Minority-Focused Text-to-Image Generation via Prompt Optimization | pdf |
| general | MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny Objects | pdf |
| general | SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in Structures | pdf |
| general | Collaborative Tree Search for Enhancing Embodied Multi-Agent Collaboration | pdf |
| general | Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual Abduction | pdf |
| general | Adapting Text-to-Image Generation with Feature Difference Instruction for Generic Image Restoration | pdf |
| general | ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence Learning | pdf |
| general | Preconditioners for the Stochastic Training of Neural Fields | pdf |
| worlds | ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark | pdf |
| gsplat | SfM-Free 3D Gaussian Splatting via Hierarchical Training | pdf |
| general | CASAGPT: Cuboid Arrangement and Scene Assembly for Interior Design | pdf |
| general | MINIMA: Modality Invariant Image Matching | pdf |
| gsplat | 3D Convex Splatting: Radiance Field Rendering with 3D Smooth Convexes | pdf |
| general | 3D Prior Is All You Need: Cross-Task Few-shot 2D Gaze Estimation | pdf |
| general | SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding | pdf |
| general | Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models | pdf |
| general | GliaNet: Adaptive Neural Network Structure Learning with Glia-Driven | pdf |
| general | EntitySAM: Segment Everything in Video | pdf |
| general | GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction | pdf |
| general | Video Depth Anything: Consistent Depth Estimation for Super-Long Videos | pdf |
| general | InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption | pdf |
| gsplat | Luminance-GS: Adapting 3D Gaussian Splatting to Challenging Lighting Conditions with View-Adaptive Curve Adjustment | pdf |
| gsplat | EventSplat: 3D Gaussian Splatting from Moving Event Cameras for Real-time Rendering | pdf |
| gsplat | 3D Student Splatting and Scooping | pdf |
| worlds | World-consistent Video Diffusion with Explicit 3D Modeling | pdf |
| general | Learning Partonomic 3D Reconstruction from Image Collections | pdf |
| general | ODA-GAN: Orthogonal Decoupling Alignment GAN Assisted by Weakly-supervised Learning for Virtual Immunohistochemistry Staining | pdf |
| general | EVOS: Efficient Implicit Neural Training via EVOlutionary Selector | pdf |
| general | MEET: Towards Memory-Efficient Temporal Sparse Deep Neural Networks | pdf |
| general | Probabilistic Prompt Distribution Learning for Animal Pose Estimation | pdf |
| general | Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention | pdf |
| general | UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines | pdf |
| gsplat | Mani-GS: Gaussian Splatting Manipulation with Triangular Mesh | pdf |
| general | BooW-VTON: Boosting In-the-Wild Virtual Try-On via Mask-Free Pseudo Data Training | pdf |
| general | Supervising Sound Localization by In-the-wild Egomotion | pdf |
| general | AutoLUT: LUT-Based Image Super-Resolution with Automatic Sampling and Adaptive Residual Learning | pdf |
| gsplat | AniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian Reconstruction | pdf |
| general | IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular VideosC | pdf |
| gsplat | DynaMoDe-NeRF: Motion-aware Deblurring Neural Radiance Field for Dynamic Scenes | pdf |
| general | UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene Simulation | pdf |
| general | Diff-Palm: Realistic Palmprint Generation with Polynomial Creases and Intra-Class Variation Controllable Diffusion Models | pdf |
| gsplat | Chain of Semantics Programming in 3D Gaussian Splatting Representation for 3D Vision Grounding | pdf |
| motion | MVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait Animation | pdf |
| general | Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction | pdf |
| general | Matrix-Free Shared Intrinsics Bundle Adjustment | pdf |
| general | Uncertainty-Instructed Structure Injection for Generalizable HD Map Construction | pdf |
| general | Color Alignment in Diffusion | pdf |
| general | LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living | pdf |
| general | Language-Guided Salient Object Ranking | pdf |
| general | Towards More General Video-based Deepfake Detection through Facial Component Guided Adaptation for Foundation Model | pdf |
| general | SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic Prompts | pdf |
| general | VoCo-LLaMA: Towards Vision Compression with Large Language Models | pdf |
| general | Focal Split: Untethered Snapshot Depth from Differential Defocus | pdf |
| general | PURA: Parameter Update-Recovery Test-Time Adaption for RGB-T Tracking | pdf |
| general | Towards All-in-One Medical Image Re-Identification | pdf |
| general | Integral Fast Fourier Color Constancy | pdf |
| general | ResCLIP: Residual Attention for Training-free Dense Vision-language Inference | pdf |
| general | Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction | pdf |
| general | Bayesian Test-Time Adaptation for Vision-Language Models | pdf |
| general | Causal Composition Diffusion Model for Closed-loop Traffic Generation | pdf |
| general | Change3D: Revisiting Change Detection and Captioning from A Video Modeling Perspective | pdf |
| general | Attribute-formed Class-specific Concept Space: Endowing Language Bottleneck Model with Better Interpretability and Scalability | pdf |
| general | Customized Condition Controllable Generation for Video Soundtrack | pdf |
| beings | ProjAttacker: A Configurable Physical Adversarial Attack for Face Recognition via Projector | pdf |
| general | WISE: A Framework for Gigapixel Whole-Slide-Image Lossless Compression | pdf |
| general | Gromov-Wasserstein Problem with Cyclic Symmetry | pdf |
| beings | SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing | pdf |
| general | Test-Time Backdoor Detection for Object Detection Models | pdf |
| general | SDBF: Steep-Decision-Boundary Fingerprinting for Hard-Label Tampering Detection of DNN Models | pdf |
| general | Distilling Multi-modal Large Language Models for Autonomous Driving | pdf |
| general | HD-EPIC: A Highly-Detailed Egocentric Video Dataset | pdf |
| general | Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training | pdf |
| beings | H-MoRe: Learning Human-centric Motion Representation for Action Analysis | pdf |
| general | Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric Learning | pdf |
| general | Effortless Active Labeling for Long-Term Test-Time Adaptation | pdf |
| general | Leveraging Temporal Cues for Semi-Supervised Multi-View 3D Object Detection | pdf |
| general | Logits DeConfusion with CLIP for Few-Shot Learning | pdf |
| general | Pay Attention to the Foreground in Object-Centric Learning | pdf |
| general | FluidNexus: 3D Fluid Reconstruction and Prediction from a Single Video | pdf |
| general | DeformCL: Learning Deformable Centerline Representation for Vessel Extraction in 3D Medical Image | pdf |
| worlds | OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation Triad | pdf |
| general | SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction | pdf |
| beings | VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation | pdf |
| general | Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos | pdf |
| general | Adaptive Keyframe Sampling for Long Video Understanding | pdf |
| general | Person De-reidentification: A Variation-guided Identity Shift Modeling | pdf |
| general | DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer | pdf |
| general | Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion | pdf |
| general | Realistic Test-Time Adaptation of Vision-Language Models | pdf |
| gsplat | SelfSplat: Pose-Free and 3D Prior-Free Generalizable 3D Gaussian Splatting | pdf |
| general | Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise Scheduling | pdf |
| general | Exploring Simple Open-Vocabulary Semantic Segmentation | pdf |
| general | MP-GUI: Modality Perception with MLLMs for GUI Understanding | pdf |
| general | Improving Adversarial Transferability on Vision Transformers via Forward Propagation Refinement | pdf |
| general | Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection | pdf |
| general | Erasing Undesirable Influence in Diffusion Models | pdf |
| general | Closest Neighbors are Harmful for Lightweight Masked Auto-encoders | pdf |
| general | Decouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learning | pdf |
| worlds | HELVIPAD: A Real-World Dataset for Omnidirectional Stereo Depth Estimation | pdf |
| general | Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency | pdf |
| general | Practical Solutions to the Relative Pose of Three Calibrated Cameras | pdf |
| general | PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models | pdf |
| general | RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins | pdf |
| motion | AnimateAnything: Consistent and Controllable Animation for Video Generation | pdf |
| general | PRaDA: Projective Radial Distortion Averaging | pdf |
| general | GenAssets: Generating in-the-wild 3D Assets in Latent Space | pdf |
| general | Low-Rank Adaptation in Multilinear Operator Networks for Security-Preserving Incremental Learning | pdf |
| general | FiRe: Fixed-points of Restoration Priors for Solving Inverse Problems | pdf |
| general | CoMBO: Conflict Mitigation via Branched Optimization for Class Incremental Segmentation | pdf |
| general | Recurrent Feature Mining and Keypoint Mixup Padding for Category-Agnostic Pose Estimation | pdf |
| general | Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks | pdf |
| general | Embodied Scene Understanding for Vision Language Models via MetaVQA | pdf |
| general | Learning Temporally Consistent Video Depth from Video Diffusion Priors | pdf |
| general | Samba: A Unified Mamba-based Framework for General Salient Object Detection | pdf |
| beings | LesionLocator: Zero-Shot Universal Tumor Segmentation and Tracking in 3D Whole-Body Imaging | pdf |
| gsplat | DOF-GS: Adjustable Depth-of-Field 3D Gaussian Splatting for Post-Capture Refocusing, Defocus Rendering and Blur Removal | pdf |
| general | The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers | pdf |
| motion | Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation | pdf |
| general | GraphI2P: Image-to-Point Cloud Registration with Exploring Pattern of Correspondence via Graph Learning | pdf |
| general | SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer | pdf |
| general | DPC: Dual-Prompt Collaboration for Tuning Vision-Language Models | pdf |
| general | AIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic Data | pdf |
| general | Robust Multi-Object 4D Generation for In-the-wild Videos | pdf |
| general | Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training | pdf |
| general | FLAVC: Learned Video Compression with Feature Level Attention | pdf |
| general | An End-to-End Robust Point Cloud Semantic Segmentation Network with Single-Step Conditional Diffusion Models | pdf |
| general | PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors | pdf |
| general | Your ViT is Secretly an Image Segmentation Model | pdf |
| general | Cross-Rejective Open-Set SAR Image Registration | pdf |
| gsplat | SplineGS: Robust Motion-Adaptive Spline for Real-Time Dynamic 3D Gaussians from Monocular Video | pdf |
| beings | Multi-modal Knowledge Distillation-based Human Trajectory Forecasting | pdf |
| general | ShiftwiseConv: Small Convolutional Kernel with Large Kernel Effect | pdf |
| general | Object-Shot Enhanced Grounding Network for Egocentric Video | pdf |
| general | Ev-3DOD: Pushing the Temporal Boundaries of 3D Object Detection with Event Cameras | pdf |
| general | Nearly Zero-Cost Protection Against Mimicry by Personalized Diffusion Models | pdf |
| general | The Devil is in Temporal Token: High Quality Video Reasoning Segmentation | pdf |
| gsplat | LITA-GS: Illumination-Agnostic Novel View Synthesis via Reference-Free 3D Gaussian Splatting and Physical Priors | pdf |
| general | T-FAKE: Synthesizing Thermal Images for Facial Landmarking | pdf |
| general | Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation | pdf |
| general | PICD: Versatile Perceptual Image Compression with Diffusion Rendering | pdf |
| motion | VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and Editing | pdf |
| general | Six-CD: Benchmarking Concept Removals for Text-to-image Diffusion Models | pdf |
| general | Black-Box Forgery Attacks on Semantic Watermarks for Diffusion Models | pdf |
| general | VidSeg: Training-free Video Semantic Segmentation based on Diffusion Models | pdf |
| motion | PersonaBooth: Personalized Text-to-Motion Generation | pdf |
| general | Star with Bilinear Mapping | pdf |
| general | DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models | pdf |
| gsplat | Time of the Flight of the Gaussians: Optimizing Depth Indirectly in Dynamic Radiance Fields | pdf |
| general | Align3R: Aligned Monocular Depth Estimation for Dynamic Videos | pdf |
| general | Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment Recognition | pdf |
| general | Anomize: Better Open Vocabulary Video Anomaly Detection | pdf |
| general | Efficient Diffusion as Low Light Enhancer | pdf |
| general | HyperNVD: Accelerating Neural Video Decomposition via Hypernetworks | pdf |
| general | Instant Adversarial Purification with Adversarial Consistency Distillation | pdf |
| general | Feature Selection for Latent Factor Models | pdf |
| general | Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing | pdf |
| general | Decoupling Training-Free Guided Diffusion by ADMM | pdf |
| general | SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusion | pdf |
| general | Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes | pdf |
| general | CLIP-driven Coarse-to-fine Semantic Guidance for Fine-grained Open-set Semi-supervised Learning | pdf |
| general | A Simple Data Augmentation for Feature Distribution Skewed Federated Learning | pdf |
| general | GLane3D: Detecting Lanes with Graph of 3D Keypoints | pdf |
| general | Minimal Interaction Seperated Tuning: A New Paradigm for Visual Adaptation | pdf |
| general | Attraction Diminishing and Distributing for Few-Shot Class-Incremental Learning | pdf |
| gsplat | 4DTAM: Non-Rigid Tracking and Mapping via Dynamic Surface Gaussians | pdf |
| general | Unseen Visual Anomaly Generation | pdf |
| general | T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting | pdf |
| general | ReNeg: Learning Negative Embedding with Reward Guidance | pdf |
| motion | MotionPro: A Precise Motion Controller for Image-to-Video Generation | pdf |
| general | Goku: Flow Based Video Generative Foundation Models | pdf |
| general | WISH: Weakly Supervised Instance Segmentation using Heterogeneous Labels | pdf |
| general | Good, Cheap, and Fast: Overfitted Image Compression with Wasserstein Distortion | pdf |
| general | Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model | pdf |
| general | V2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object Detection | pdf |
| general | TAROT: Towards Essentially Domain-Invariant Robustness with Theoretical Justification | pdf |
| general | Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis | pdf |
| general | APT: Adaptive Personalized Training for Diffusion Models with Limited Data | pdf |
| general | SCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute Prompting | pdf |
| general | Tracktention: Leveraging Point Tracking to Attend Videos Faster and Better | pdf |
| general | DocVLM: Make Your VLM an Efficient Reader | pdf |
| general | Revisiting Source-Free Domain Adaptation: Insights into Representativeness, Generalization, and Variety | pdf |
| general | Adaptive Unimodal Regulation for Balanced Multimodal Information Acquisition | pdf |
| general | FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views | pdf |
| gsplat | Improving Gaussian Splatting with Localized Points Management | pdf |
| general | One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models | pdf |
| general | Domain Adaptive Diabetic Retinopathy Grading with Model Absence and Flowing Data | pdf |
| general | LoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Models | pdf |
| general | SEAL: Semantic Attention Learning for Long Video Representation | pdf |
| general | SCFlow2: Plug-and-Play Object Pose Refiner with Shape-Constraint Scene Flow | pdf |
| motion | FlipSketch: Flipping Static Drawings to Text-Guided Sketch Animations | pdf |
| general | SketchAgent: Language-Driven Sequential Sketch Generation | pdf |
| general | DRAWER: Digital Reconstruction and Articulation With Environment Realism | pdf |
| general | GoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View Synthesis | pdf |
| general | Deep Change Monitoring: A Hyperbolic Representative Learning Framework and a Dataset for Long-term Fine-grained Tree Change Detection | pdf |
| general | ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On | pdf |
| general | MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval | pdf |
| general | VolFormer: Explore More Comprehensive Cube Interaction for Hyperspectral Image Restoration and Beyond | pdf |
| general | BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation | pdf |
| general | SmartCLIP: Modular Vision-language Alignment with Identification Guarantees | pdf |
| general | Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers | pdf |
| general | RoboGround: Robotic Manipulation with Grounded Vision-Language Priors | pdf |
| general | Improving Transferable Targeted Attacks with Feature Tuning Mixup | pdf |
| worlds | DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation | pdf |
| beings | HuPerFlow: A Comprehensive Benchmark for Human vs. Machine Motion Estimation Comparison | pdf |
| general | MetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt Tuning | pdf |
| general | Subnet-Aware Dynamic Supernet Training for Neural Architecture Search | pdf |
| worlds | EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance | pdf |
| beings | Controllable Human Image Generation with Personalized Multi-Garments | pdf |
| general | UHD-processer: Unified UHD Image Restoration with Progressive Frequency Learning and Degradation-aware Prompts | pdf |
| gsplat | GBC-Splat: Generalizable Gaussian-Based Clothed Human Digitalization under Sparse RGB Cameras | pdf |
| general | AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers | pdf |
| general | A Unified Model for Compressed Sensing MRI Across Undersampling Patterns | pdf |
| worlds | TSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-Resolution | pdf |
| general | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass | pdf |
| general | StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements | pdf |
| general | CTRL-O: Language-Controllable Object-Centric Visual Representation Learning | pdf |
| general | Text Augmented Correlation Transformer For Few-shot Classification & Segmentation | pdf |
| general | Unified Dense Prediction of Video Diffusion | pdf |
| general | Towards Million-Scale Adversarial Robustness Evaluation With Stronger Individual Attacks | pdf |
| general | Temporal Action Detection Model Compression by Progressive Block Drop | pdf |
| general | Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction | pdf |
| general | DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment | pdf |
| general | Learning Affine Correspondences by Integrating Geometric Constraints | pdf |
| general | UCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label Learning | pdf |
| general | Geometry in Style: 3D Stylization via Surface Normal Deformation | pdf |
| general | PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models | pdf |
| general | Multiple Object Tracking as ID Prediction | pdf |
| general | PIDLoc: Cross-View Pose Optimization Network Inspired by PID Controllers | pdf |
| general | DreamOmni: Unified Image Generation and Editing | pdf |
| general | Hash3D: Training-free Acceleration for 3D Generation | pdf |
| worlds | Learning Hazing to Dehazing: Towards Realistic Haze Generation for Real-World Image Dehazing | pdf |
| general | RUBIK: A Structured Benchmark for Image Matching across Geometric Challenges | pdf |
| general | Fast and Accurate Gigapixel Pathological Image Classification with Hierarchical Distillation Multi-Instance Learning | pdf |
| gsplat | IncEventGS: Pose-Free Gaussian Splatting from a Single Event Camera | pdf |
| general | OpenMIBOOD: Open Medical Imaging Benchmarks for Out-Of-Distribution Detection | pdf |
| general | FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models | pdf |
| general | When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning | pdf |
| beings | UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing | pdf |
| motion | POMP: Physics-consistent Motion Generative Model through Phase Manifolds | pdf |
| general | Reasoning to Attend: Try to Understand How <SEG> Token Works | pdf |
| general | ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams | pdf |
| general | DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution | pdf |
| general | Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy | pdf |
| worlds | VideoWorld: Exploring Knowledge Learning from Unlabeled Videos | pdf |
| general | 3D-SLNR: A Super Lightweight Neural Representation for Large-scale 3D Mapping | pdf |
| general | STINR: Deciphering Spatial Transcriptomics via Implicit Neural Representation | pdf |
| general | RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models | pdf |
| general | Unsupervised Discovery of Facial Landmarks and Head Pose | pdf |
| general | Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learning | pdf |
| general | Stabilizing and Accelerating Autofocus with Expert Trajectory Regularized Deep Reinforcement Learning | pdf |
| general | Repurposing Stable Diffusion Attention for Training-Free Unsupervised Interactive Segmentation | pdf |
| gsplat | GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector | pdf |
| general | DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation | pdf |
| general | Simulator HC: Regression-based Online Simulation of Starting Problem-Solution Pairs for Homotopy Continuation in Geometric Vision | pdf |
| general | Dynamic Integration of Task-Specific Adapters for Class Incremental Learning | pdf |
| general | EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision | pdf |
| general | DiverseFlow: Sample-Efficient Diverse Mode Coverage in Flows | pdf |
| worlds | MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction | pdf |
| general | 3D-MVP: 3D Multiview Pretraining for Manipulation | pdf |
| general | Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations | pdf |
| general | Mimir: Improving Video Diffusion Models for Precise Text Understanding | pdf |
| general | UCM-VeID V2: A Richer Dataset and A Pre-training Method for UAV Cross-Modality Vehicle Re-Identification | pdf |
| beings | GeoAvatar: Geometrically-Consistent Multi-Person Avatar Reconstruction from Sparse Multi-View Videos | pdf |
| gsplat | DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting | pdf |
| gsplat | Speedy-Splat: Fast 3D Gaussian Splatting with Sparse Pixels and Sparse Primitives | pdf |
| general | Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos | pdf |
| beings | ODHSR: Online Dense 3D Reconstruction of Humans and Scenes from Monocular Videos | pdf |
| general | SpiritSight Agent: Advanced GUI Agent with One Look | pdf |
| general | Zero-Shot Monocular Scene Flow Estimation in the Wild | pdf |
| motion | MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities | pdf |
| general | Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher Distillation | pdf |
| general | MMRL: Multi-Modal Representation Learning for Vision-Language Models | pdf |
| general | Anchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2D | pdf |
| general | Breaking the Low-Rank Dilemma of Linear Attention | pdf |
| general | Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning | pdf |
| general | Unity in Diversity: Video Editing via Gradient-Latent Purification | pdf |
| general | Revealing Key Details to See Differences: A Novel Prototypical Perspective for Skeleton-based Action Recognition | pdf |
| general | Finsler Multi-Dimensional Scaling: Manifold Learning for Asymmetric Dimensionality Reduction and Embedding | pdf |
| general | VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection | pdf |
| motion | Cross-Modal Distillation for 2D/3D Multi-Object Discovery from 2D Motion | pdf |
| general | Bridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical Flow | pdf |
| general | Inversion Circle Interpolation: Diffusion-based Image Augmentation for Data-scarce Classification | pdf |
| general | TSP-Mamba: The Travelling Salesman Problem Meets Mamba for Image Super-resolution and Beyond | pdf |
| general | RENO: Real-Time Neural Compression for 3D LiDAR Point Clouds | pdf |
| general | FADE: Frequency-Aware Diffusion Model Factorization for Video Editing | pdf |
| beings | Data Synthesis with Diverse Styles for Face Recognition via 3DMM-Guided Diffusion | pdf |
| general | Geometric Knowledge-Guided Localized Global Distribution Alignment for Federated Learning | pdf |
| general | GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing | pdf |
| general | Birth and Death of a Rose | pdf |
| general | MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural Representation | pdf |
| general | MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation | pdf |
| general | Be More Specific: Evaluating Object-centric Realism in Synthetic Images | pdf |
| general | Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuning | pdf |
| beings | SFDM: Robust Decomposition of Geometry and Reflectance for Realistic Face Rendering from Sparse-view Images | pdf |
| general | Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion | pdf |
| general | QMambaBSR: Burst Image Super-Resolution with Query State Space Model | pdf |
| general | Multi-Group Proportional Representations for Text-to-Image Models | pdf |
| general | Towards Generalizable Trajectory Prediction using Dual-Level Representation Learning and Adaptive Prompting | pdf |
| general | CoMatcher: Multi-View Collaborative Feature Matching | pdf |
| beings | Towards a Universal Synthetic Video Detector: From Face or Background Manipulations to Fully AI-Generated Content | pdf |
| beings | A Focused Human Body Model for Accurate Anthropometric Measurements Extraction | pdf |
| general | ACE: Anti-Editing Concept Erasure in Text-to-Image Models | pdf |
| motion | Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors | pdf |
| general | Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time Adaptation | pdf |
| general | LaTexBlend: Scaling Multi-concept Customized Generation with Latent Textual Blending | pdf |
| general | DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification | pdf |
| general | Let's Verify and Reinforce Image Generation Step by Step | pdf |
| general | All-Optical Nonlinear Diffractive Deep Network for Ultrafast Image Denoising | pdf |
| general | UNOPose: Unseen Object Pose Estimation with an Unposed RGB-D Reference Image | pdf |
| meshing | HybridMQA: Exploring Geometry-Texture Interactions for Colored Mesh Quality Assessment | pdf |
| general | SIR-DIFF: Sparse Image Sets Restoration with Multi-View Diffusion Model | pdf |
| general | Reversible Decoupling Network for Single Image Reflection Removal | pdf |
| general | Hierarchical Features Matter: A Deep Exploration of Progressive Parameterization Method for Dataset Distillation | pdf |
| general | GLASS: Guided Latent Slot Diffusion for Object-Centric Learning | pdf |
| general | SASep: Saliency-Aware Structured Separation of Geometry and Feature for Open Set Learning on Point Clouds | pdf |
| general | Low-Biased General Annotated Dataset Generation | pdf |
| general | Generative Hard Example Augmentation for Semantic Point Cloud Segmentation | pdf |
| general | ETAP: Event-based Tracking of Any Point | pdf |
| general | Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge | pdf |
| meshing | Volumetric Surfaces: Representing Fuzzy Geometries with Layered Meshes | pdf |
| general | STEPS: Sequential Probability Tensor Estimation for Text-to-Image Hard Prompt Search | pdf |
| general | VIRES: Video Instance Repainting via Sketch and Text Guided Generation | pdf |
| general | MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling | pdf |
| gsplat | From Sparse to Dense: Camera Relocalization with Scene-Specific Detector from Feature Gaussian Splatting | pdf |
| beings | StableAnimator: High-Quality Identity-Preserving Human Image Animation | pdf |
| general | OODD: Test-time Out-of-Distribution Detection with Dynamic Dictionary | pdf |
| general | BIMBA: Selective-Scan Compression for Long-Range Video Question Answering | pdf |
| general | Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment | pdf |
| general | Prof. Robot: Differentiable Robot Rendering Without Static and Self-Collisions | pdf |
| gsplat | DropGaussian: Structural Regularization for Sparse-view Gaussian Splatting | pdf |
| general | Blurred LiDAR for Sharper 3D: Robust Handheld 3D Scanning with Diffuse LiDAR and RGB | pdf |
| general | Novel View Synthesis with Pixel-Space Diffusion Models | pdf |
| general | Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark Dataset | pdf |
| general | Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents | pdf |
| general | Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages | pdf |
| gsplat | TAGA: Self-supervised Learning for Template-free Animatable Gaussian Articulated Model | pdf |
| gsplat | Horizon-GS: Unified 3D Gaussian Splatting for Large-Scale Aerial-to-Ground Scenes | pdf |
| general | LotusFilter: Fast Diverse Nearest Neighbor Search via a Learned Cutoff Table | pdf |
| general | Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation | pdf |
| general | SoMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation Learning | pdf |
| gsplat | Ref-GS: Directional Factorization for 2D Gaussian Splatting | pdf |
| general | VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents | pdf |
| general | Concept Lancet: Image Editing with Compositional Representation Transplant | pdf |
| gsplat | Generative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D Reconstruction | pdf |
| general | Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis | pdf |
| beings | Are Images Indistinguishable to Humans Also Indistinguishable to Classifiers? | pdf |
| general | Understanding Multi-layered Transmission Matrices | pdf |
| gsplat | GS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Tracking | pdf |
| motion | AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models | pdf |
| worlds | Dual Energy-Based Model with Open-World Uncertainty Estimation for Out-of-distribution Detection | pdf |
| general | DTGBrepGen: A Novel B-rep Generative Model through Decoupling Topology and Geometry | pdf |
| general | Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generation | pdf |
| general | Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval | pdf |
| general | TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation | pdf |
| general | NoT: Federated Unlearning via Weight Negation | pdf |
| general | RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings | pdf |
| beings | SimMotionEdit: Text-Based Human Motion Editing with Motion Similarity Prediction | pdf |
| general | From Head to Tail: Efficient Black-box Model Inversion Attack via Long-tailed Learning | pdf |
| general | SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model | pdf |
| general | Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation | pdf |
| general | Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera | pdf |
| general | ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models | pdf |
| general | Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues | pdf |
| general | OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations | pdf |
| worlds | LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models | pdf |
| general | Point Clouds Meets Physics: Dynamic Acoustic Field Fitting Network for Point Cloud Understanding | pdf |
| general | Faster Parameter-Efficient Tuning with Token Redundancy Reduction | pdf |
| general | Panorama Generation From NFoV Image Done Right | pdf |
| gsplat | Sparse Point Cloud Patches Rendering via Splitting 2D Gaussians | pdf |
| general | Distilling Monocular Foundation Model for Fine-grained Depth Completion | pdf |
| general | AniGrad: Anisotropic Gradient-Adaptive Sampling for 3D Reconstruction From Monocular Video | pdf |
| general | Less Attention is More: Prompt Transformer for Generalized Category Discovery | pdf |
| motion | AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward | pdf |
| gsplat | DoF-Gaussian: Controllable Depth-of-Field for 3D Gaussian Splatting | pdf |
| general | Reconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in Dual | pdf |
| general | Hierarchical Flow Diffusion for Efficient Frame Interpolation | pdf |
| general | BASKET: A Large-Scale Video Dataset for Fine-Grained Skill Estimation | pdf |
| general | Arbitrary-steps Image Super-resolution via Diffusion Inversion | pdf |
| general | Dynamic Neural Surfaces for Elastic 4D Shape Representation and Analysis | pdf |
| general | ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems | pdf |
| general | Incomplete Multi-View Multi-label Learning via Disentangled Representation and Label Semantic Embedding | pdf |
| general | AutoURDF: Unsupervised Robot Modeling from Point Cloud Frames Using Cluster Registration | pdf |
| general | Golden Cudgel Network for Real-Time Semantic Segmentation | pdf |
| general | Multi-modal Contrastive Learning with Negative Sampling Calibration for Phenotypic Drug Discovery | pdf |
| general | R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning | pdf |
| gsplat | SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis | pdf |
| general | Boltzmann Attention Sampling for Image Analysis with Small Objects | pdf |
| gsplat | Generalized Recorrupted-to-Recorrupted: Self-Supervised Learning Beyond Gaussian Noise | pdf |
| motion | Dynamic Motion Blending for Versatile Motion Editing | pdf |
| general | StdGEN: Semantic-Decomposed 3D Character Generation from Single Images | pdf |
| general | Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting | pdf |
| general | FFR: Frequency Feature Rectification for Weakly Supervised Semantic Segmentation | pdf |
| general | Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding | pdf |
| general | Sonata: Self-Supervised Learning of Reliable Point Representations | pdf |
| general | DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation | pdf |
| general | DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters | pdf |
| general | A Unified Approach to Interpreting Self-supervised Pre-training Methods for 3D Point Clouds via Interactions | pdf |
| general | Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image Segmentation | pdf |
| general | STEREO: A Two-Stage Framework for Adversarially Robust Concept Erasing from Text-to-Image Diffusion Models | pdf |
| general | GenVDM: Generating Vector Displacement Maps From a Single Image | pdf |
| general | Effective SAM Combination for Open-Vocabulary Semantic Segmentation | pdf |
| worlds | Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection | pdf |
| general | UniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic Programming | pdf |
| general | Turbo3D: Ultra-fast Text-to-3D Generation | pdf |
| meshing | SUM Parts: Benchmarking Part-Level Semantic Segmentation of Urban Meshes | pdf |
| beings | MODA: Motion-Drift Augmentation for Inertial Human Motion Analysis | pdf |
| general | Higher-Order Ratio Cycles for Fast and Globally Optimal Shape Matching | pdf |
| general | Hyperdimensional Uncertainty Quantification for Multimodal Uncertainty Fusion in Autonomous Vehicles Perception | pdf |
| gsplat | GIFStream: 4D Gaussian-based Immersive Video with Feature Stream | pdf |
| general | Multi-Scale Neighborhood Occupancy Masked Autoencoder for Self-Supervised Learning in LiDAR Point Clouds | pdf |
| general | PLeaS - Merging Models with Permutations and Least Squares | pdf |
| general | Incremental Object Keypoint Learning | pdf |
| general | InteractVLM: 3D Interaction Reasoning from 2D Foundational Models | pdf |
| general | Attribute-Missing Multi-view Graph Clustering | pdf |
| general | Pose-Guided Temporal Enhancement for Robust Low-Resolution Hand Reconstruction | pdf |
| general | ReDiffDet: Rotation-equivariant Diffusion Model for Oriented Object Detection | pdf |
| general | Unlocking Generalization Power in LiDAR Point Cloud Registration | pdf |
| general | LoRA Recycle: Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs | pdf |
| general | Event Fields: Capturing Light Fields at High Speed, Resolution, and Dynamic Range | pdf |
| general | HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing Imagery | pdf |
| gsplat | GBlobs: Explicit Local Structure via Gaussian Blobs for Improved Cross-Domain LiDAR-based 3D Object Detection | pdf |
| gsplat | 3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial Representations | pdf |
| general | MambaIRv2: Attentive State Space Restoration | pdf |
| general | Floating No More: Object-Ground Reconstruction from a Single Image | pdf |
| general | Pattern Analogies: Learning to Perform Programmatic Image Edits by Analogy | pdf |
| general | STAR-Edge: Structure-aware Local Spherical Curve Representation for Thin-walled Edge Extraction from Unstructured Point Clouds | pdf |
| general | Boosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty Estimation | pdf |
| general | pFedMxF: Personalized Federated Class-Incremental Learning with Mixture of Frequency Aggregation | pdf |
| general | Efficient Transfer Learning for Video-language Foundation Models | pdf |
| general | Radio Frequency Ray Tracing with Neural Object Representation for Enhanced RF Modeling | pdf |
| general | Neuro-3D: Towards 3D Visual Decoding from EEG Signals | pdf |
| general | Probing the Mid-level Vision Capabilities of Self-Supervised Learning | pdf |
| general | Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction | pdf |
| general | Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAI | pdf |
| general | ZoomLDM: Latent Diffusion Model for Multi-scale Image Generation | pdf |
| gsplat | GaussianUDF: Inferring Unsigned Distance Functions through 3D Gaussian Splatting | pdf |
| general | CrossSDF: 3D Reconstruction of Thin Structures From Cross-Sections | pdf |
| general | DV-Matcher: Deformation-based Non-rigid Point Cloud Matching Guided by Pre-trained Visual Features | pdf |
| general | Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance Grounding | pdf |
| general | Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement | pdf |
| general | FLAIR: VLM with Fine-grained Language-informed Image Representations | pdf |
| general | GG-SSMs: Graph-Generating State Space Models | pdf |
| general | Continuous Adverse Weather Removal via Degradation-Aware Distillation | pdf |
| general | Exploiting Temporal State Space Sharing for Video Semantic Segmentation | pdf |
| gsplat | High-fidelity 3D Object Generation from Single Image with RGBN-Volume Gaussian Reconstruction Model | pdf |
| gsplat | Steepest Descent Density Control for Compact 3D Gaussian Splatting | pdf |
| beings | Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing | pdf |
| general | Robust 3D Shape Reconstruction in Zero-Shot from a Single Image in the Wild | pdf |
| general | BOE-ViT: Boosting Orientation Estimation with Equivariance in Self-Supervised 3D Subtomogram Alignment | pdf |
| general | Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency | pdf |
| general | Beyond Local Sharpness: Communication-Efficient Global Sharpness-aware Minimization for Federated Learning | pdf |
| motion | Parameterized Blur Kernel Prior Learning for Local Motion Deblurring | pdf |
| general | Scene4U: Hierarchical Layered 3D Scene Reconstruction from Single Panoramic Image for Your Immerse Exploration | pdf |
| general | ACAttack: Adaptive Cross Attacking RGB-T Tracker via Multi-Modal Response Decoupling | pdf |
| general | DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos | pdf |
| gsplat | HoGS: Unified Near and Far Object Reconstruction via Homogeneous Gaussian Splatting | pdf |
| general | SmartEraser: Remove Anything from Images using Masked-Region Guidance | pdf |
| general | Sample- and Parameter-Efficient Auto-Regressive Image Models | pdf |
| general | Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment | pdf |
| general | BiLoRA: Almost-Orthogonal Parameter Spaces for Continual Learning | pdf |
| meshing | Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation | pdf |
| worlds | SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments | pdf |
| general | Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient | pdf |
| general | AerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesis | pdf |
| general | Visual Representation Learning through Causal Intervention for Controllable Image Editing | pdf |
| general | Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis | pdf |
| general | A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation | pdf |
| gsplat | Deformable Radial Kernel Splatting | pdf |
| general | Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection | pdf |
| general | HalLoc: Token-level Localization of Hallucinations for Vision Language Models | pdf |
| general | DiffPortrait360: Consistent Portrait Diffusion for 360 View Synthesis | pdf |
| general | SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation Sparsity | pdf |
| general | From Slow Bidirectional to Fast Autoregressive Video Diffusion Models | pdf |
| general | Noise Diffusion for Enhancing Semantic Faithfulness in Text-to-Image Synthesis | pdf |
| general | MonoInstance: Enhancing Monocular Priors via Multi-view Instance Alignment for Neural Rendering and Reconstruction | pdf |
| general | CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models | pdf |
| general | Exploring Semantic Feature Discrimination for Perceptual Image Super-Resolution and Opinion-Unaware No-Reference Image Quality Assessment | pdf |
| general | Distilling Long-tailed Datasets | pdf |
| general | Gaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoders | pdf |
| general | Incorporating Dense Knowledge Alignment into Unified Multimodal Representation Models | pdf |
| beings | Boost Your Human Image Generation Model via Direct Preference Optimization | pdf |
| general | Learning to Highlight Audio by Watching Movies | pdf |
| general | Unified Uncertainty-Aware Diffusion for Multi-Agent Trajectory Modeling | pdf |
| general | WeGen: A Unified Model for Interactive Multimodal Generation as We Chat | pdf |
| gsplat | HRAvatar: High-Quality and Relightable Gaussian Head Avatar | pdf |
| general | A Distractor-Aware Memory for Visual Object Tracking with SAM2 | pdf |
| general | Activating Sparse Part Concepts for 3D Class Incremental Learning | pdf |
| general | ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding | pdf |
| general | BFANet: Revisiting 3D Semantic Segmentation with Boundary Feature Analysis | pdf |
| general | Beyond Words: Augmenting Discriminative Richness via Diffusions in Unsupervised Prompt Learning | pdf |
| general | Unlocking the Potential of Unlabeled Data in Semi-Supervised Domain Generalization | pdf |
| general | Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks | pdf |
| general | Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusion | pdf |
| general | Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark | pdf |
| gsplat | Mitigating Ambiguities in 3D Classification with Gaussian Splatting | pdf |
| general | DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings Learning | pdf |
| general | Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing | pdf |
| general | DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation | pdf |
| general | DUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teachers | pdf |
| general | Black Hole-Driven Identity Absorbing in Diffusion Models | pdf |
| general | HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models | pdf |
| motion | Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer | pdf |
| general | SeqMvRL: A Sequential Fusion Framework for Multi-view Representation Learning | pdf |
| general | BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models | pdf |
| general | VLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement Learning | pdf |
| general | NeISF++: Neural Incident Stokes Field for Polarized Inverse Rendering of Conductors and Dielectrics | pdf |
| general | Non-Natural Image Understanding with Advancing Frequency-based Vision Encoders | pdf |
| general | Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens | pdf |
| gsplat | SplatFlow: Self-Supervised Dynamic Gaussian Splatting in Neural Motion Flow Field for Autonomous Driving | pdf |
| worlds | AesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Models | pdf |
| general | FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity | pdf |
| general | Chebyshev Attention Depth Permutation Texture Network with Latent Texture Attribute Loss | pdf |
| general | Decentralized Diffusion Models | pdf |
| general | AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea | pdf |
| general | DNF: Unconditional 4D Generation with Dictionary-based Neural Fields | pdf |
| general | ARM: Appearance Reconstruction Model for Relightable 3D Generation | pdf |
| general | Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels | pdf |
| meshing | TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing | pdf |
| general | Generating 3D-Consistent Videos from Unposed Internet Photos | pdf |
| general | Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation | pdf |
| general | ViUniT: Visual Unit Tests for More Robust Visual Programming | pdf |
| general | DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations | pdf |
| general | beta-FFT: Nonlinear Interpolation and Differentiated Training Strategies for Semi-Supervised Medical Image Segmentation | pdf |
| general | Dynamic Group Normalization: Spatio-Temporal Adaptation to Evolving Data Statistics | pdf |
| general | SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding | pdf |
| general | Uncertain Multimodal Intention and Emotion Understanding in the Wild | pdf |
| general | VidTwin: Video VAE with Decoupled Structure and Dynamics | pdf |
| general | CL-LoRA: Continual Low-Rank Adaptation for Rehearsal-Free Class-Incremental Learning | pdf |
| beings | Design2GarmentCode: Turning Design Concepts to Tangible Garments Through Program Synthesis | pdf |
| gsplat | Efficient Dynamic Scene Editing via 4D Gaussian-based Static-Dynamic Separation | pdf |
| general | Unlearning through Knowledge Overwriting: Reversible Federated Unlearning via Selective Sparse Adapter | pdf |
| general | SocialMOIF: Multi-Order Intention Fusion for Pedestrian Trajectory Prediction | pdf |
| general | HistoFS: Non-IID Histopathologic Whole Slide Image Classification via Federated Style Transfer with RoI-Preserving | pdf |
| gsplat | SGSST: Scaling Gaussian Splatting Style Transfer | pdf |
| general | Learning Bijective Surface Parameterization for Inferring Signed Distance Functions from Sparse Point Clouds with Grid Deformation | pdf |
| general | Balancing Two Classifiers via A Simplex ETF Structure for Model Calibration | pdf |
| general | DAMM-Diffusion: Learning Divergence-Aware Multi-Modal Diffusion Model for Nanoparticles Distribution Prediction | pdf |
| general | U-Know-DiffPAN: An Uncertainty-aware Knowledge Distillation Diffusion Framework with Details Enhancement for PAN-Sharpening | pdf |
| gsplat | RelationField: Relate Anything in Radiance Fields | pdf |
| beings | Let Humanoids Hike! Integrative Skill Development on Complex Trails | pdf |
| general | BF-STVSR: B-Splines and Fourier---Best Friends for High Fidelity Spatial-Temporal Video Super-Resolution | pdf |
| worlds | DIO: Decomposable Implicit 4D Occupancy-Flow World Model | pdf |
| general | SLADE: Shielding against Dual Exploits in Large Vision-Language Models | pdf |
| beings | Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input | pdf |
| general | FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component Analysis | pdf |
| general | Mind the Time: Temporally-Controlled Multi-Event Video Generation | pdf |
| general | Audio-Visual Semantic Graph Network for Audio-Visual Event Localization | pdf |
| motion | Video Motion Transfer with Diffusion Transformers | pdf |
| general | Unified Reconstruction of Static and Dynamic Scenes from Events | pdf |
| general | Automatic Spectral Calibration of Hyperspectral Images: Method, Dataset and Benchmark | pdf |
| general | Point-to-Region Loss for Semi-Supervised Point-Based Crowd Counting | pdf |
| beings | Move-in-2D: 2D-Conditioned Human Motion Generation | pdf |
| general | MATCHA: Towards Matching Anything | pdf |
| general | CTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D Diffusion | pdf |
| general | Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question Answering | pdf |
| general | SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding | pdf |
| general | Fitted Neural Lossless Image Compression | pdf |
| general | JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration | pdf |
| general | F-LMM: Grounding Frozen Large Multimodal Models | pdf |
| general | EntityErasure: Erasing Entity Cleanly via Amodal Entity Segmentation and Completion | pdf |
| general | Joint Out-of-Distribution Filtering and Data Discovery Active Learning | pdf |
| general | Finding Local Diffusion Schrodinger Bridge using Kolmogorov-Arnold Network | pdf |
| general | CorrBEV: Multi-View 3D Object Detection by Correlation Learning with Multi-modal Prototypes | pdf |
| general | Completion as Enhancement: A Degradation-Aware Selective Image Guided Network for Depth Completion | pdf |
| worlds | Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation | pdf |
| gsplat | Real-time High-fidelity Gaussian Human Avatars with Position-based Interpolation of Spatially Distributed MLPs | pdf |
| general | RoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured Environments | pdf |
| general | DEFOM-Stereo: Depth Foundation Model Based Stereo Matching | pdf |
| general | DiskVPS: Vanishing Point Detector via Hough Transform in a Disk Region | pdf |
| general | Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding | pdf |
| general | Towards Autonomous Micromobility through Scalable Urban Simulation | pdf |
| general | Language-Assisted Debiasing and Smoothing for Foundation Model-Based Semi-Supervised Learning | pdf |
| general | EdgeMovingNet: Edge-preserving Point Cloud Reconstruction via Joint Geometry Features | pdf |
| general | Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment | pdf |
| gsplat | Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object Detection | pdf |
| general | Enhancing Diversity for Data-free Quantization | pdf |
| general | From Alexnet to Transformers: Measuring the Non-linearity of Deep Neural Networks with Affine Optimal Transport | pdf |
| general | Prompt2Perturb (P2P): Text-Guided Diffusion-Based Adversarial Attack on Breast Ultrasound Images | pdf |
| general | COAP: Memory-Efficient Training with Correlation-Aware Gradient Projection | pdf |
| general | Gyro-based Neural Single Image Deblurring | pdf |
| general | Improved Monocular Depth Prediction Using Distance Transform Over Pre-semantic Contours with Self-supervised Neural Networks | pdf |
| beings | Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body | pdf |
| general | Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation | pdf |
| general | ROLL: Robust Noisy Pseudo-label Learning for Multi-View Clustering with Noisy Correspondence | pdf |
| general | Towards In-the-wild 3D Plane Reconstruction from a Single Image | pdf |
| general | PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction | pdf |
| general | CheXwhatsApp: A Dataset for Exploring Challenges in the Diagnosis of Chest X-rays through Mobile Devices | pdf |
| general | Degradation-Aware Feature Perturbation for All-in-One Image Restoration | pdf |
| general | GenDeg: Diffusion-based Degradation Synthesis for Generalizable All-In-One Image Restoration | pdf |
| general | The Power of Context: How Multimodality Improves Image Super-Resolution | pdf |
| general | Detect Any Mirrors: Boosting Learning Reliability on Large-Scale Unlabeled Data with an Iterative Data Engine | pdf |
| gsplat | 4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models | pdf |
| beings | MotionMap: Representing Multimodality in Human Pose Forecasting | pdf |
| general | Factored-NeuS: Reconstructing Surfaces, Illumination, and Materials of Possibly Glossy Objects | pdf |
| gsplat | GaussianSpa: An "Optimizing-Sparsifying" Simplification Framework for Compact and High-Quality 3D Gaussian Splatting | pdf |
| general | Navigating the Unseen: Zero-shot Scene Graph Generation via Capsule-Based Equivariant Features | pdf |
| general | VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models | pdf |
| general | ASHiTA: Automatic Scene-grounded HIerarchical Task Analysis | pdf |
| general | Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Models | pdf |
| general | RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation | pdf |
| general | Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis | pdf |
| general | A Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image Segmentation | pdf |
| general | FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models | pdf |
| general | GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation | pdf |
| general | Learning from Neighbors: Category Extrapolation for Long-Tail Learning | pdf |
| general | Material Anything: Generating Materials for Any 3D Object via Diffusion | pdf |
| general | ImagineFSL: Self-Supervised Pretraining Matters on Imagined Base Set for VLM-based Few-shot Learning | pdf |
| general | Continuous Locomotive Crowd Behavior Generation | pdf |
| general | Project-Probe-Aggregate: Efficient Fine-Tuning for Group Robustness | pdf |
| general | Implicit Bias Injection Attacks against Text-to-Image Diffusion Models | pdf |
| general | ROICtrl: Boosting Instance Control for Visual Generation | pdf |
| general | Cropper: Vision-Language Model for Image Cropping through In-Context Learning | pdf |
| motion | ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model | pdf |
| general | ICE: Intrinsic Concept Extraction from a Single Image via Diffusion Models | pdf |
| general | ASIGN: An Anatomy-aware Spatial Imputation Graphic Network for 3D Spatial Transcriptomics | pdf |
| general | MultiMorph: On-demand Atlas Construction | pdf |
| general | Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding | pdf |
| general | Spiking Transformer: Introducing Accurate Addition-Only Spiking Self-Attention for Transformer | pdf |
| worlds | MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes | pdf |
| general | Symbolic Representation for Any-to-Any Generative Tasks | pdf |
| general | Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations | pdf |
| general | MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations | pdf |
| gsplat | ArticulatedGS: Self-supervised Digital Twin Modeling of Articulated Objects using 3D Gaussian Splatting | pdf |
| general | Leveraging 3D Geometric Priors in 2D Rotation Symmetry Detection | pdf |
| general | Noise Calibration and Spatial-Frequency Interactive Network for STEM Image Enhancement | pdf |
| beings | Homogeneous Dynamics Space for Heterogeneous Humans | pdf |
| general | TailedCore: Few-Shot Sampling for Unsupervised Long-Tail Noisy Anomaly Detection | pdf |
| general | Satellite Observations Guided Diffusion Model for Accurate Meteorological States at Arbitrary Resolution | pdf |
| general | Reconstructing People, Places, and Cameras | pdf |
| general | InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model Alignment | pdf |
| general | Identifying and Mitigating Spurious Correlation in Multi-Task Learning | pdf |
| general | Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment | pdf |
| general | CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation | pdf |
| meshing | PMNI: Pose-free Multi-view Normal Integration for Reflective and Textureless Surface Reconstruction | pdf |
| gsplat | LeanGaussian: Breaking Pixel or Point Cloud Correspondence in Modeling 3D Gaussians | pdf |
| worlds | Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks | pdf |
| general | A Unified Latent Schrodinger Bridge Diffusion Model for Unsupervised Anomaly Detection and Localization | pdf |
| general | MambaVision: A Hybrid Mamba-Transformer Vision Backbone | pdf |
| general | Multi-Label Prototype Visual Spatial Search for Weakly Supervised Semantic Segmentation | pdf |
| general | Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features | pdf |
| gsplat | Learnable Infinite Taylor Gaussian for Dynamic View Rendering | pdf |
| general | SaMam: Style-aware State Space Model for Arbitrary Image Style Transfer | pdf |
| general | Making Old Film Great Again: Degradation-aware State Space Model for Old Film Restoration | pdf |
| motion | MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion | pdf |
| beings | Deterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion Synthesis | pdf |
| general | CacheQuant: Comprehensively Accelerated Diffusion Models | pdf |
| worlds | Open-World Objectness Modeling Unifies Novel Object Detection | pdf |
| beings | MotionPRO: Exploring the Role of Pressure in Human MoCap and Beyond | pdf |
| general | DiffVsgg: Diffusion-Driven Online Video Scene Graph Generation | pdf |
| general | Towards Smart Point-and-Shoot Photography | pdf |
| general | Prototype-Based Image Prompting for Weakly Supervised Histopathological Image Segmentation | pdf |
| beings | Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation | pdf |
| general | SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language | pdf |
| general | Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion | pdf |
| general | SoftShadow: Leveraging Soft Masks for Penumbra-Aware Shadow Removal | pdf |
| general | VTON-HandFit: Virtual Try-on for Arbitrary Hand Pose Guided by Hand Priors Embedding | pdf |
| general | Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion | pdf |
| general | POSTA: A Go-to Framework for Customized Artistic Poster Generation | pdf |
| general | NSD-Imagery: A Benchmark Dataset for Extending fMRI Vision Decoding Methods to Mental Imagery | pdf |
| general | VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models | pdf |
| motion | Just Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection | pdf |
| motion | Efficient Motion-Aware Video MLLM | pdf |
| general | Zero-Shot 4D Lidar Panoptic Segmentation | pdf |
| general | ADU: Adaptive Detection of Unknown Categories in Black-Box Domain Adaptation | pdf |
| general | EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion | pdf |
| general | Unsupervised Foundation Model-Agnostic Slide-Level Representation Learning | pdf |
| general | UNIALIGN: Scaling Multimodal Alignment within One Unified Model | pdf |
| general | ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions | pdf |
| general | Exploration-Driven Generative Interactive Environments | pdf |
| general | DreamText: High Fidelity Scene Text Synthesis | pdf |
| general | ProKeR: A Kernel Perspective on Few-Shot Adaptation of Large Vision-Language Models | pdf |
| general | MonoTAKD: Teaching Assistant Knowledge Distillation for Monocular 3D Object Detection | pdf |
| general | Easy-editable Image Vectorization with Multi-layer Multi-scale Distributed Visual Feature Embedding | pdf |
| general | Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration | pdf |
| general | Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens | pdf |
| motion | SpectroMotion: Dynamic 3D Reconstruction of Specular Scenes | pdf |
| general | VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction | pdf |
| general | MVBoost: Boost 3D Reconstruction with Multi-View Refinement | pdf |
| general | Category-Agnostic Neural Object Rigging | pdf |
| general | POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation | pdf |
| general | MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis | pdf |
| general | Mimic In-Context Learning for Multimodal Tasks | pdf |
| general | Vision-Language Models Do Not Understand Negation | pdf |
| gsplat | NexusGS: Sparse View Synthesis with Epipolar Depth Priors in 3D Gaussian Splatting | pdf |
| general | HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight Trajectories | pdf |
| general | RICCARDO: Radar Hit Prediction and Convolution for Camera-Radar 3D Object Detection | pdf |
| beings | BLADE: Single-view Body Mesh Estimation through Accurate Depth Estimation | pdf |
| motion | MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation | pdf |
| gsplat | ReCap: Better Gaussian Relighting with Cross-Environment Captures | pdf |
| general | Vision-Language Embodiment for Monocular Depth Estimation | pdf |
| general | Frequency Dynamic Convolution for Dense Image Prediction | pdf |
| general | IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification | pdf |
| general | Consistency Posterior Sampling for Diverse Image Synthesis | pdf |
| general | IMFine: 3D Inpainting via Geometry-guided Multi-view Refinement | pdf |
| general | DeepCompress-ViT: Rethinking Model Compression to Enhance Efficiency of Vision Transformers at the Edge | pdf |
| general | EvOcc: Accurate Semantic Occupancy for Automated Driving Using Evidence Theory | pdf |
| general | Towards Continual Universal Segmentation | pdf |
| gsplat | PGC: Physics-Based Gaussian Cloth from a Single Pose | pdf |
| beings | OFER: Occluded Face Expression Reconstruction | pdf |
| worlds | Cubify Anything: Scaling Indoor 3D Object Detection | pdf |
| general | DEAL: Data-Efficient Adversarial Learning for High-Quality Infrared Imaging | pdf |
| general | BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects | pdf |
| general | CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models | pdf |
| gsplat | FreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene Reconstruction | pdf |
| general | Show and Tell: Visually Explainable Deep Neural Nets via Spatially-Aware Concept Bottleneck Models | pdf |
| general | Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene | pdf |
| general | Knowledge Bridger: Towards Training-Free Missing Modality Completion | pdf |
| beings | TexGarment: Consistent Garment UV Texture Generation via Efficient 3D Structure-Guided Diffusion Transformer | pdf |
| worlds | Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video Deraining | pdf |
| general | TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time Correction | pdf |
| general | VSNet: Focusing on the Linguistic Characteristics of Sign Language | pdf |
| general | Learning to Sample Effective and Diverse Prompts for Text-to-Image Generation | pdf |
| general | Multi-modal Medical Diagnosis via Large-small Model Collaboration | pdf |
| motion | Image Referenced Sketch Colorization Based on Animation Creation Workflow | pdf |
| beings | GaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D Reconstruction | pdf |
| beings | ProbPose: A Probabilistic Approach to 2D Human Pose Estimation | pdf |
| worlds | MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation | pdf |
| general | ABC-Former: Auxiliary Bimodal Cross-domain Transformer with Interactive Channel Attention for White Balance | pdf |
| general | Fingerprinting Denoising Diffusion Probabilistic Models | pdf |
| general | NightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene Segmentation | pdf |
| beings | UMFN: Unified Multi-Domain Face Normalization for Joint Cross-domain Prototype Learning and Heterogeneous Face Recognition | pdf |
| beings | LUCAS: Layered Universal Codec Avatars | pdf |
| general | D^3: Scaling Up Deepfake Detection by Learning from Discrepancy | pdf |
| general | Jailbreaking the Non-Transferable Barrier via Test-Time Data Disguising | pdf |
| general | 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination | pdf |
| general | Generative Zero-Shot Composed Image Retrieval | pdf |
| general | Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards | pdf |
| general | Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Mutimodal Models | pdf |
| general | Omnidirectional Multi-Object Tracking | pdf |
| general | Potential Field Based Deep Metric Learning | pdf |
| general | Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data | pdf |
| general | Directional Label Diffusion Model for Learning from Noisy Labels | pdf |
| general | Learning Endogenous Attention for Incremental Object Detection | pdf |
| worlds | StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation | pdf |
| general | HomoGen: Enhanced Video Inpainting via Homography Propagation and Diffusion | pdf |
| general | Do ImageNet-trained Models Learn Shortcuts? The Impact of Frequency Shortcuts on Generalization | pdf |
| beings | HORP: Human-Object Relation Priors Guided HOI Detection | pdf |
| general | Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs | pdf |