Nishi Library › Research Papers

Research Papers — CVF Open Access

Compiled by nx_cvf_index from fetched CVF open-access listings. Each paper links to its open-access abstract page and PDF at the Computer Vision Foundation (copyright the authors/CVF). Families are title-keyword tags serving the procgen program: worlds/procgen · gaussian-splatting · beings/avatars · motion · meshing. COVERAGE IS STATED, NEVER SILENT: only the proceedings listed below are ingested so far; the remaining CVF proceedings (2013-2026, ~30k papers) land via the standing fetch + crawl lanes and appear here as their listings are banked.

99worlds185gsplat175beings76motion23meshing2313general2871total papers

CVPR 2025 - Day 1 (2025-06-13)

familypaperlinks
generalTowards Source-Free Machine Unlearningpdf
generalUni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Videopdf
generalHyperbolic Category Discoverypdf
beingsThe Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motionpdf
generalCALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Modelspdf
generalWords or Vision: Do Vision-Language Models Have Blind Faith in Text?pdf
generalLearning to Detect Objects from Multi-Agent LiDAR Scans without Manual Labelspdf
generalDeepLA-Net: Very Deep Local Aggregation Networks for Point Cloud Analysispdf
generalMulti-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practicespdf
generalAPHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformerspdf
generalAdaptCMVC: Robust Adaption to Incremental Views in Continual Multi-view Clusteringpdf
generalUA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial Referencespdf
generalBinarized Mamba-Transformer for Lightweight Quad Bayer HybridEVS Demosaicingpdf
generalInterpretable Image Classification via Non-parametric Part Prototype Learningpdf
beingsDAGSM: Disentangled Avatar Generation with GS-enhanced Meshpdf
worldsEstimating Body and Hand Motion in an Ego-sensed Worldpdf
generalEvaluating Vision-Language Models as Evaluators in Path Planningpdf
generalFree on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EMpdf
generalSGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI Detectionpdf
generalGalaxy Walker: Geometry-aware VLMs For Galaxy-scale Understandingpdf
generalSnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback Optimizationpdf
motionExploring Timeline Control for Facial Motion Generationpdf
gsplatGAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusionpdf
beingsAI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness Benchmarkpdf
generalEnhancing Video-LLM Reasoning via Agent-of-Thoughts Distillationpdf
generalDe^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze Estimationpdf
generalReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuningpdf
generalSelf-Expansion of Pre-trained Models with Mixture of Adapters for Continual Learningpdf
generalBrain-Inspired Spiking Neural Networks for Energy-Efficient Object Detectionpdf
generalMedusa: A Multi-Scale High-order Contrastive Dual-Diffusion Approach for Multi-View Clusteringpdf
generalMambaOut: Do We Really Need Mamba for Vision?pdf
generalSeurat: From Moving Points to Depthpdf
generalImg-Diff: Contrastive Data Synthesis for Multimodal Large Language Modelspdf
generalThe Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generationpdf
generalDnLUT: Ultra-Efficient Color Image Denoising via Channel-Aware Lookup Tablespdf
motionBiM-VFI: Bidirectional Motion Field-Guided Frame Interpolation for Video with Non-uniform Motionspdf
generalSATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision Transformerspdf
generalNested Diffusion Models Using Hierarchical Latent Priorspdf
generalA Theory of Learning Unified Model via Knowledge Integration from Label Space Varying Domainspdf
generalHiLoTs: High-Low Temporal Sensitive Representation Learning for Semi-Supervised LiDAR Segmentation in Autonomous Drivingpdf
generalDKDM: Data-Free Knowledge Distillation for Diffusion Models with Any Architecturepdf
generalSymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimizationpdf
generalDebiasing Multimodal Large Language Models via Noise-Aware Preference Optimizationpdf
gsplatFeat2GS: Probing Visual Foundation Models with Gaussian Splattingpdf
generalLSNet: See Large, Focus Smallpdf
generalDynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenespdf
generalDocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understandingpdf
generalEDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolationpdf
generalHandling Spatial-Temporal Data Heterogeneity for Federated Continual Learning via Tail Anchorpdf
gsplatDeSiRe-GS: 4D Street Gaussians for Static-Dynamic Decomposition and Surface Reconstruction for Urban Driving Scenespdf
beingsREWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioningpdf
generalDoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cyclespdf
gsplatGaussian Splashing: Unified Particles for Versatile Motion Synthesis and Renderingpdf
generalImprove Representation for Imbalanced Regression through Geometric Constraintspdf
generalPartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Modelpdf
generalDiffFNO: Diffusion Fourier Neural Operatorpdf
generalZero-Shot Styled Text Image Generation, but Make It Autoregressivepdf
generalLeveraging Perturbation Robustness to Enhance Out-of-Distribution Detectionpdf
motionSALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editingpdf
generalLookingGlass: Generative Anamorphoses via Laplacian Pyramid Warpingpdf
generalShowMak3r: Compositional TV Show Reconstructionpdf
generalCADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature Leveragingpdf
generalVideoDirector: Precise Video Editing via Text-to-Video Modelspdf
generalVISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentationpdf
generalGA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context Encodingpdf
gsplatRigGS: Rigging of 3D Gaussians for Modeling Articulated Objects in Videospdf
generalNoise Modeling in One Hour: Minimizing Preparation Efforts for Self-supervised Low-Light RAW Image Denoisingpdf
generalHigh Dynamic Range Video Compression: A Large-Scale Benchmark Dataset and A Learned Bit-depth Scalable Compression Algorithmpdf
generalDivPrune: Diversity-based Visual Token Pruning for Large Multimodal Modelspdf
general3D-AVS: LiDAR-based 3D Auto-Vocabulary Segmentationpdf
beingsMEGA: Masked Generative Autoencoder for Human Mesh Recoverypdf
generalDisentangling Safe and Unsafe Image Corruptions via Anisotropy and Localitypdf
worldsPrometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generationpdf
generalSphereUFormer: A U-Shaped Transformer for Spherical 360 Perceptionpdf
generalBeyond Clean Training Data: A Versatile and Model-Agnostic Framework for Out-of-Distribution Detection with Contaminated Training Datapdf
generalFreeUV: Ground-Truth-Free Realistic Facial UV Texture Recovery via Cross-Assembly Inference Strategypdf
generalHarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronizationpdf
generalStyleMaster: Stylize Your Video with Artistic Generation and Translationpdf
generalUnsupervised Continual Domain Shift Learning with Multi-Prototype Modelingpdf
generalOmniGuard: Hybrid Manipulation Localization via Augmented Versatile Deep Image Watermarkingpdf
generalOpen-Canopy: Towards Very High Resolution Forest Monitoringpdf
generalVision-Language Model IP Protection via Prompt-based Learningpdf
generalKiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generationpdf
generalKoala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Contentpdf
generalVASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsificationpdf
generalSPARC: Score Prompting and Adaptive Fusion for Zero-Shot Multi-Label Recognition in Vision-Language Modelspdf
generalErase Diffusion: Empowering Object Removal Through Calibrating Diffusion Pathwayspdf
generalPrompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysispdf
generalInstruction-based Image Manipulation by Watching How Things Movepdf
generalFerret: An Efficient Online Continual Learning Framework under Varying Memory Constraintspdf
generalVidComposition: Can MLLMs Analyze Compositions in Compiled Videos?pdf
generalSelf-Supervised Learning for Color Spike Camera Reconstructionpdf
generalFrom Elements to Design: A Layered Approach for Automatic Graphic Design Compositionpdf
generalSALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysispdf
generalDA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformerspdf
generalTowards Lossless Implicit Neural Representation via Bit Plane Decompositionpdf
gsplatiSegMan: Interactive Segment-and-Manipulate 3D Gaussianspdf
generalBlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devicespdf
generalUnraveling Normal Anatomy via Fluid-Driven Anomaly Randomizationpdf
generalTaming Teacher Forcing for Masked Autoregressive Video Generationpdf
generalRevisiting Backdoor Attacks against Large Vision-Language Models from Domain Shiftpdf
generalTCFG: Tangential Damping Classifier-free Guidancepdf
generalMatAnyone: Stable Video Matting with Consistent Memory Propagationpdf
generalMolmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Modelspdf
generalMMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perceptionpdf
generalT2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generationpdf
generalMultimodal Autoregressive Pre-training of Large Vision Encoderspdf
generalAKiRa: Augmentation Kit on Rays for Optical Video Generationpdf
generalTFCustom: Customized Image Generation with Time-Aware Frequency Feature Guidancepdf
generalSketchFusion: Learning Universal Sketch Features through Fusing Foundation Modelspdf
generalBridging the Vision-Brain Gap with an Uncertainty-Aware Blur Priorpdf
generalAffordDP: Generalizable Diffusion Policy with Transferable Affordancepdf
generalHMAR: Efficient Hierarchical Masked Auto-Regressive Image Generationpdf
generalDKC: Differentiated Knowledge Consolidation for Cloth-Hybrid Lifelong Person Re-identificationpdf
generalEnhancing Facial Privacy Protection via Weakening Diffusion Purificationpdf
generalORIDa: Object-centric Real-world Image Composition Datasetpdf
generalImage Generation Diversity Issues and How to Tame Thempdf
generalAnnotation Ambiguity Aware Semi-Supervised Medical Image Segmentationpdf
beingsCAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Modelspdf
beingsCORE4D: A 4D Human-Object-Human Interaction Dataset for Collaborative Object REarrangementpdf
gsplatPOp-GS: Next Best View in 3D-Gaussian Splatting with P-Optimalitypdf
generalCritic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoningpdf
generalMaRI: Material Retrieval Integration across Domainspdf
generalQ-Bench-Video: Benchmark the Video Quality Understanding of LMMspdf
generalGlossy Object Reconstruction with Cost-effective Polarized Acquisitionpdf
generalL-SWAG: Layer-Sample Wise Activation with Gradients Information for Zero-Shot NAS on Vision Transformerspdf
generalCommonsense Video Question Answering through Video-Grounded Entailment Tree Reasoningpdf
generalLifelong Knowledge Editing for Vision Language Models with Low-Rank Mixture-of-Expertspdf
generalPartGen: Part-level 3D Generation and Reconstruction with Multi-view Diffusion Modelspdf
generalSINR: Sparsity Driven Compressed Implicit Neural Representationspdf
generalManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learningpdf
generalShining Yourself: High-Fidelity Ornaments Virtual Try-on with Diffusion Modelpdf
generalUniversal Domain Adaptation for Semantic Segmentationpdf
gsplatHyperGS: Hyperspectral 3D Gaussian Splattingpdf
generalLMO: Linear Mamba Operator for MRI Reconstructionpdf
generalAnomalyNCD: Towards Novel Anomaly Class Discovery in Industrial Scenariospdf
generalYour Large Vision-Language Model Only Needs A Few Attention Heads For Visual Groundingpdf
generalGes3ViG : Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understandingpdf
generalProgressive Focused Transformer for Single Image Super-Resolutionpdf
generalVladVA: Discriminative Fine-tuning of LVLMspdf
beingsHumanMM: Global Human Motion Recovery from Multi-shot Videospdf
generalRemoving Reflections from RAW Photospdf
generalAMR-Transformer: Enabling Efficient Long-range Interaction for Complex Neural Fluid Simulationpdf
generalBlurry-Edges: Photon-Limited Depth Estimation from Defocused Boundariespdf
generalMICAS: Multi-grained In-Context Adaptive Sampling for 3D Point Cloud Processingpdf
generalGoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Drivingpdf
motionColabSfM: Collaborative Structure-from-Motion by Point Cloud Registrationpdf
generalMangaNinja: Line Art Colorization with Precise Reference Followingpdf
gsplatNonisotropic Gaussian Diffusion for Realistic 3D Human Motion Predictionpdf
generalPICO: Reconstructing 3D People In Contact with Objectspdf
generalLinguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognitionpdf
generalScaling up Image Segmentation across Data and Taskspdf
generalBridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planningpdf
generalBlood Flow Speed Estimation with Optical Coherence Tomography Angiography Imagespdf
generalDreamTrack: Dreaming the Future for Multimodal Visual Object Trackingpdf
generalOmniStyle: Filtering High Quality Style Transfer Data at Scalepdf
generalCross-View Completion Models are Zero-shot Correspondence Estimatorspdf
generalMulti-party Collaborative Attention Control for Image Customizationpdf
generalHOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videospdf
generalDELT: A Simple Diversity-driven EarlyLate Training for Dataset Distillationpdf
generalRoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concretepdf
generalBeyond Image Classification: A Video Benchmark and Dual-Branch Hybrid Discrimination Framework for Compositional Zero-Shot Learningpdf
generalABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image Objectspdf
generalDo Your Best and Get Enough Rest for Continual Learningpdf
generalEnhancing Few-Shot Class-Incremental Learning via Training-Free Bi-Level Modality Calibrationpdf
generalMUSt3R: Multi-view Network for Stereo 3D Reconstructionpdf
generalHybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language Modelspdf
generalA New Statistical Model of Star Speckles for Learning to Detect and Characterize Exoplanets in Direct Imaging Observationspdf
generalCoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Modelspdf
generalWAVE: Weight Templates for Adaptive Initialization of Variable-sized Modelspdf
generalCXPMRG-Bench: Pre-training and Benchmarking for X-ray Medical Report Generation on CheXpert Plus Datasetpdf
generalEvent-Equalized Dense Video Captioningpdf
generalEDCFlow: Exploring Temporally Dense Difference Maps for Event-based Optical Flow Estimationpdf
generalLibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributionspdf
generalLost in Translation, Found in Context: Sign Language Translation with Contextual Cuespdf
generalSynchronized Video-to-Audio Generation via Mel Quantization-Continuum Decompositionpdf
gsplatFATE: Full-head Gaussian Avatar with Textural Editing from Monocular Videopdf
generalTouch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstructionpdf
generalVITED: Video Temporal Evidence Distillationpdf
generalTemporal Score Analysis for Understanding and Correcting Diffusion Artifactspdf
motionGo-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noisepdf
generalVision-Language Gradient Descent-driven All-in-One Deep Unfolding Networkspdf
general3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformerpdf
generalBoost the Inference with Co-training: A Depth-guided Mutual Learning Framework for Semi-supervised Medical Polyp Segmentationpdf
worldsFrom Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-Identificationpdf
general4Deform: Neural Surface Deformation for Robust Shape Interpolationpdf
generalDense Match Summarization for Faster Two-view Estimationpdf
generalAlign-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video Editingpdf
generalLION-FS: Fast & Slow Video-Language Thinker as Online Video Assistantpdf
motionMotion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Levelpdf
generalToward Robust Neural Reconstruction from Sparse Point Setspdf
gsplatGPAvatar: High-fidelity Head Avatars by Learning Efficient Gaussian Projectionspdf
generalPIAD: Pose and Illumination agnostic Anomaly Detectionpdf
generalTwo is Better than One: Efficient Ensemble Defense for Robust and Compact Modelspdf
generalTiled Diffusionpdf
generalDescriptor-In-Pixel : Point-Feature Tracking For Pixel Processor Arrayspdf
gsplatUVGS: Reimagining Unstructured 3D Gaussian Splatting using UV Mappingpdf
beingsInterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generationpdf
generalTAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognitionpdf
generalBioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational Pathologypdf
gsplatGauCho: Gaussian Distributions with Cholesky Decomposition for Oriented Object Detectionpdf
generalNo Thing, Nothing: Highlighting Safety-Critical Classes for Robust LiDAR Semantic Segmentation in Adverse Weatherpdf
generalMind the Gap: Detecting Black-box Adversarial Attacks in the Making through Query Update Analysispdf
worldsGaussianWorld: Gaussian World Model for Streaming 3D Occupancy Predictionpdf
generalICP: Immediate Compensation Pruning for Mid-to-high Sparsitypdf
generalVinaBench: Benchmark for Faithful and Consistent Visual Narrativespdf
generalDual Diffusion for Unified Image Generation and Understandingpdf
generalWeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentationpdf
gsplat4DGC: Rate-Aware 4D Gaussian Compression for Efficient Streamable Free-Viewpoint Videopdf
gsplatGASP: Gaussian Avatars with Synthetic Priorspdf
generalCOSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptationpdf
generalHigh-quality Point Cloud Oriented Normal Estimation via Hybrid Angular and Euclidean Distance Encodingpdf
generalPrior-free 3D Object Trackingpdf
generalProgressive Correspondence Regenerator for Robust 3D Registrationpdf
generalCross-Modal 3D Representation with Multi-View Images and Point Cloudspdf
generalDecompositional Neural Scene Reconstruction with Generative Diffusion Priorpdf
generalLearning Visual Generative Priors without Textpdf
generalLifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inferencepdf
generalPixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approachpdf
generalMr. DETR: Instructive Multi-Route Training for Detection Transformerspdf
generalHearing Hands: Generating Sounds from Physical Interactions in 3D Scenespdf
generalAirRoom: Objects Matter in Room Reidentificationpdf
generalDefMamba: Deformable Visual State Space Modelpdf
generalHOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generationpdf
gsplatVoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Predictionpdf
beingsControlFace: Harnessing Facial Parametric Control for Face Riggingpdf
generalImputation-free and Alignment-free: Incomplete Multi-view Clustering Driven by Consensus Semantic Learningpdf
generalSensitivity-Aware Efficient Fine-Tuning via Compact Dynamic-Rank Adaptationpdf
generalA Selective Re-learning Mechanism for Hyperspectral Fusion Imagingpdf
generalAutoregressive Sequential Pretraining for Visual Trackingpdf
beingsPromptHMR: Promptable Human Mesh Recoverypdf
generalVISTREAM: Improving Computation Efficiency of Visual Streaming Perception via Law-of-Charge-Conservation Inspired Spiking Neural Networkpdf
generalSTPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Groundingpdf
generalRashomon Sets for Prototypical-Part Networks: Editing Interpretable Models in Real-Timepdf
beingsEnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty Modelingpdf
generalTuning the Frequencies: Robust Training for Sinusoidal Neural Networkspdf
beingsReal-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Texturespdf
generalLarge Self-Supervised Models Bridge the Gap in Domain Adaptive Object Detectionpdf
generalEvaluating Model Perception of Color Illusions in Photorealistic Scenespdf
generalDo Visual Imaginations Improve Vision-and-Language Navigation Agents?pdf
generalHotSpot: Signed Distance Function Optimization with an Asymptotically Sufficient Conditionpdf
generalLibra-Merging: Importance-redundancy and Pruning-merging Trade-off for Acceleration Plug-in in Large Vision-Language Modelpdf
generalLEDiff: Latent Exposure Diffusion for HDR Generationpdf
generalVideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulationpdf
generalZero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusionpdf
generalUnveiling Visual Perception in Language Models: An Attention Head Analysis Approachpdf
generalSemiDAViL: Semi-supervised Domain Adaptation with Vision-Language Guidance for Semantic Segmentationpdf
generaldFLMoE: Decentralized Federated Learning via Mixture of Experts for Medical Data Analysispdf
beingsReconstructing Humans with a Biomechanically Accurate Skeletonpdf
generalAdaCM^2: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reductionpdf
generalVGGT: Visual Geometry Grounded Transformerpdf
generalSilent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Modelspdf
generalVisual Consensus Prompting for Co-Salient Object Detectionpdf
generalQuantization without Tearspdf
worldsPHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric Videospdf
generalTowards Consistent Multi-Task Learning: Unlocking the Potential of Task-Specific Parameterspdf
worldsSceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generationpdf
generalHiFi-Portrait: Zero-shot Identity-preserved Portrait Generation with High-fidelity Multi-face Fusionpdf
generalFloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesispdf
generalRAD: Region-Aware Diffusion Models for Image Inpaintingpdf
generalUnderstanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Spacepdf
gsplatTexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian Splattingpdf
generalConcept Replacer: Replacing Sensitive Concepts in Diffusion Models via Precision Localizationpdf
generalA Regularization-Guided Equivariant Approach for Image Restorationpdf
generalDeep Fair Multi-View Clustering with Attention KANpdf
generalLineArt: A Knowledge-guided Training-free High-quality Appearance Transfer for Design Drawing with Diffusion Modelpdf
generalVideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understandingpdf
generalZero-Shot Image Restoration Using Few-Step Guidance of Consistency Models (and Beyond)pdf
generalSimilarity-Guided Layer-Adaptive Vision Transformer for UAV Trackingpdf
generalLidarGait++: Learning Local Features and Size Awareness from LiDAR Point Clouds for 3D Gait Recognitionpdf
generalCoarse Correspondences Boost Spatial-Temporal Reasoning in Multimodal Language Modelpdf
generalFoundationStereo: Zero-Shot Stereo Matchingpdf
generalUniNet: A Contrastive Learning-guided Unified Framework with Feature Selection for Anomaly Detectionpdf
generalMoEdit: On Learning Quantity Perception for Multi-object Image Editingpdf
beingsSeeing More with Less: Human-like Representations in Vision Modelspdf
beingsModeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identificationpdf
generalAeroGen: Enhancing Remote Sensing Object Detection with Diffusion-Driven Data Generationpdf
generalTra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioningpdf
generalStyle Quantization for Data-Efficient GAN Trainingpdf
generalLocalizing Events in Videos with Multimodal Queriespdf
generalPhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachabilitypdf
generalCleanDIFT: Diffusion Features without Noisepdf
generalMAD: Memory-Augmented Detection of 3D Objectspdf
generalDoppelgangers and Adversarial Vulnerabilitypdf
generalPrecise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalancepdf
generalSteady Progress Beats Stagnation: Mutual Aid of Foundation and Conventional Models in Mixed Domain Semi-Supervised Medical Image Segmentationpdf
worldsARKit LabelMaker: A New Scale for Indoor 3D Scene Understandingpdf
generalStreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Textpdf
generalAFL: A Single-Round Analytic Approach for Federated Learning with Pre-trained Modelspdf
generalBOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understandingpdf
generalReference-Based 3D-Aware Image Editing with Triplanespdf
generalOne is Plenty: A Polymorphic Feature Interpreter for Immutable Heterogeneous Collaborative Perceptionpdf
beingsSegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectoriespdf
generalSceneCrafter: Controllable Multi-View Driving Scene Editingpdf
beingsHeatFormer: A Neural Optimizer for Multiview Human Mesh Recoverypdf
generalGPS as a Control Signal for Image Generationpdf
generalCPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathologypdf
gsplatMAGiC-SLAM: Multi-Agent Gaussian Globally Consistent SLAMpdf
generalNTClick: Achieving Precise Interactive Segmentation With Noise-tolerant Clickspdf
generalMVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Modelpdf
gsplatHiMoR: Monocular Deformable Gaussian Reconstruction with Hierarchical Motion Representationpdf
generalEnhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute Recognitionpdf
beingsHSI-GPT: A General-Purpose Large Scene-Motion-Language Model for Human Scene Interactionpdf
beingsVid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Priorpdf
worldsRoomPainter: View-Integrated Diffusion for Consistent Indoor Scene Texturingpdf
worldsIRIS: Inverse Rendering of Indoor Scenes from Low Dynamic Range Imagespdf
gsplatRoGSplat: Learning Robust Generalizable Human Gaussian Splatting from Sparse Multi-View Imagespdf
generalEnliveningGS: Active Locomotion of 3DGSpdf
generalKnowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognitionpdf
generalLPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic Segmentationpdf
worldsPhysGen3D: Crafting a Miniature Interactive World from a Single Imagepdf
generalDocopilot: Improving Multimodal Models for Document-Level Understandingpdf
generalSelf-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolutionpdf
generalLATTE-MV: Learning to Anticipate Table Tennis Hits from Monocular Videospdf
generalDistilling Spatially-Heterogeneous Distortion Perception for Blind Image Quality Assessmentpdf
general2DMamba: Efficient State Space Model for Image Representation with Applications on Giga-Pixel Whole Slide Image Classificationpdf
generalUnboxed: Geometrically and Temporally Consistent Video Outpaintingpdf
beingsK-Sort Arena: Efficient and Reliable Benchmarking for Generative Models via K-wise Human Preferencespdf
motionDense-SfM: Structure from Motion with Dense Consistent Matchingpdf
generalSketchy Bounding-box Supervision for 3D Instance Segmentationpdf
generalStreetCrafter: Street View Synthesis with Controllable Video Diffusion Modelspdf
beingsLearning Person-Specific Animatable Face Models from In-the-Wild Images via a Shared Base Modelpdf
beingsTIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion Generationpdf
generalHybrid Reciprocal Transformer with Triplet Feature Alignment for Scene Graph Generationpdf
generalGradient Inversion Attacks on Parameter-Efficient Fine-Tuningpdf
generalUPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluationpdf
generalMBQ: Modality-Balanced Quantization for Large Vision-Language Modelspdf
generalVideoDPO: Omni-Preference Alignment for Video Diffusion Generationpdf
generalAssociative Transformerpdf
beingsChatGarment: Garment Estimation, Generation and Editing via Large Language Modelspdf
generalRDD: Robust Feature Detector and Descriptor using Deformable Transformerpdf
generalBuilding Vision Models upon Heat Conductionpdf
generalLT3SD: Latent Trees for 3D Scene Diffusionpdf
meshingCraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry Refinerpdf
beingsGIF: Generative Inspiration for Face Recognition at Scalepdf
generalSKDream: Controllable Multi-view and 3D Generation with Arbitrary Skeletonspdf
generalOptimus-2: Multimodal Minecraft Agent with Goal-Observation-Action Conditioned Policypdf
worldsClassic Video Denoising in a Machine Learning World: Robust, Fast, and Controllablepdf
generalPopulation Normalization for Federated Learningpdf
generalRipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safetypdf
generalESCAPE: Equivariant Shape Completion via Anchor Point Encodingpdf
generalSatellite to GroundScape - Large-scale Consistent Ground View Generation from Satellite Viewspdf
generalVariance-Based Membership Inference Attacks Against Large-Scale Image Captioning Modelspdf
generalLearning Dynamic Collaborative Network for Semi-supervised 3D Vessel Segmentationpdf
generalTemporal Alignment-Free Video Matching for Few-shot Action Recognitionpdf
generalOSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIPpdf
generalVLog: Video-Language Models by Generative Retrieval of Narration Vocabularypdf
generalForensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language Modelspdf
generalForensic Self-Descriptions Are All You Need for Zero-Shot Detection, Open-Set Source Attribution, and Clustering of AI-generated Imagespdf
gsplatFlexDrive: Toward Trajectory Flexibility in Driving Scene Gaussian Splatting Reconstruction and Renderingpdf
gsplatTaming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputspdf
generalAugmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answeringpdf
generalDifIISR: A Diffusion Model with Gradient Guidance for Infrared Image Super-Resolutionpdf
gsplat3D-GSW: 3D Gaussian Splatting for Robust Watermarkingpdf
generalOpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generationpdf
generalDual Exposure Stereo for Extended Dynamic Range 3D Imagingpdf
generalPAVE: Patching and Adapting Video Large Language Modelspdf
generalGenerative Image Layer Decomposition with Visual Effectspdf
generalAR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusionpdf
generalRevisiting Audio-Visual Segmentation with Vision-Centric Transformerpdf
generalHOIGPT: Learning Long-Sequence Hand-Object Interaction with Language Modelspdf
generalTaxonomy-Aware Evaluation of Vision-Language Modelspdf
generalActive Event-based Stereo Visionpdf
motionSMILE: Infusing Spatial and Motion Semantics in Masked Video Learningpdf
generalVideo Language Model Pretraining with Spatio-temporal Maskingpdf
generalSynthetic Data is an Elegant GIFT for Continual Vision-Language Modelspdf
generalTimestep Embedding Tells: It's Time to Cache for Video Diffusion Modelpdf
generalHypergraph Vision Transformers: Images are More than Nodes, More than Edgespdf
generalBinarized Neural Network for Multi-spectral Image Fusionpdf
gsplatGaussianIP: Identity-Preserving Realistic 3D Human Generation via Human-Centric Diffusion Priorpdf
generalFineVQ: Fine-Grained User Generated Content Video Quality Assessmentpdf
generalUnveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectlypdf
generalMANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipationpdf
generalMETASCENES: Towards Automated Replica Creation for Real-world 3D Scanspdf
generalRobust Multimodal Survival Prediction with Conditional Latent Differentiation Variational AutoEncoderpdf
generalZero-Shot Blind-spot Image Denoising via Implicit Neural Samplingpdf
generalTrack4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generationpdf
beingsInteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsingpdf
generalWonderland: Navigating 3D Scenes from a Single Imagepdf
generalTowards Satellite Image Road Graph Extraction: A Global-Scale Dataset and A Novel Methodpdf
generalSuperLightNet: Lightweight Parameter Aggregation Network for Multimodal Brain Tumor Segmentationpdf
generalSelf-Supervised Spatial Correspondence Across Modalitiespdf
generalMOS-Attack: A Scalable Multi-objective Adversarial Attack Frameworkpdf
motionMotion Modes: What Could Happen Next?pdf
generalFiner-CAM: Spotting the Difference Reveals Finer Details for Visual Explanationpdf
generalThe Change You Want To Detect: Semantic Change Detection In Earth Observation With Hybrid Data Generationfpdf
generalWeakly Supervised Semantic Segmentation via Progressive Confidence Region Expansionpdf
generalRELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representationspdf
generalHOTFormerLoc: Hierarchical Octree Transformer for Versatile Lidar Place Recognition Across Ground and Aerial Viewspdf
generalUniK3D: Universal Camera Monocular 3D Estimationpdf
motionConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transferpdf
generalAG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identificationpdf
generalEBS-EKF: Accurate and High Frequency Event-based Star Trackingpdf
worldsBenchmarking Object Detectors under Real-World Distribution Shifts in Satellite Imagerypdf
generalSAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image Pretrainingpdf
generalClosed-Loop Supervised Fine-Tuning of Tokenized Traffic Modelspdf
generalCoLLM: A Large Language Model for Composed Image Retrievalpdf
generalGraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learningpdf
generalMMVU: Measuring Expert-Level Multi-Discipline Video Understandingpdf
motionEgoLM: Multi-Modal Language Model of Egocentric Motionspdf
generalLong Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curationpdf
generalLearning Heterogeneous Tissues with Mixture of Experts for Gigapixel Whole Slide Imagespdf
generalDisentangled Pose and Appearance Guidance for Multi-Pose Generationpdf
generalMind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-Mismatchpdf
generalElectromyography-Informed Facial Expression Reconstruction for Physiological-Based Synthesis and Analysispdf
beingsImproving the Transferability of Adversarial Attacks on Face Recognition with Diverse Parameters Augmentationpdf
generalAdapting to Observation Length of Trajectory Prediction via Contrastive Learningpdf
generalFine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentationpdf
generalNitroFusion: High-Fidelity Single-Step Diffusion through Dynamic Adversarial Trainingpdf
generalCMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based Frameworkpdf
generalRC-AutoCalib: An End-to-End Radar-Camera Automatic Calibration Networkpdf
generalArgus: A Compact and Versatile Foundation Model for Visionpdf
generalSampling Innovation-Based Adaptive Compressive Sensingpdf
motionMotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Modelspdf
generalART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generationpdf
generalArcPro: Architectural Programs for Structured 3D Abstraction of Sparse Pointspdf
gsplatHardware-Rasterized Ray-Based Gaussian Splattingpdf
generalFreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency Debiasingpdf
generalMulti-subject Open-set Personalization in Video Generationpdf
motionWav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animationpdf
generalFG^2: Fine-Grained Cross-View Localization by Fine-Grained Feature Matchingpdf
generalDistilled Prompt Learning for Incomplete Multimodal Survival Predictionpdf
generalLearning Conditional Space-Time Prompt Distributions for Video Class-Incremental Learningpdf
generalHyperbolic Safety-Aware Vision-Language Modelspdf
gsplatSinGS: Animatable Single-Image Human Gaussian Splats with Kinematic Priorspdf
generalMulti-Modal Synergistic Implicit Image Enhancement for Efficient Optical Flow Estimationpdf
generalGenerating Multimodal Driving Scenes via Next-Scene Predictionpdf
generalSymmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generationpdf
generalPosterMaker: Towards High-Quality Product Poster Generation with Accurate Text Renderingpdf
generalRethinking Personalized Aesthetics Assessment: Employing Physique Aesthetics Assessment as An Exemplificationpdf
generalYou See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scalepdf
generalPEACE: Empowering Geologic Map Holistic Understanding with MLLMspdf
generalConceptGuard: Continual Personalized Text-to-Image Generation with Forgetting and Confusion Mitigationpdf
generalMARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creationpdf
generalERUPT: Efficient Rendering with Unposed Patch Transformerpdf
gsplatRethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splattingpdf
generalQuad-Pixel Image Defocus Deblurring: A New Benchmark and Modelpdf
generalGEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistencypdf
generalDynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semanticspdf
generalESC: Erasing Space Concept for Knowledge Deletionpdf
generalTemporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networkspdf
generalTheoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference Systemspdf
generalFlowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolutionpdf
generalInterpretable Generative Models through Post-hoc Concept Bottleneckspdf
generalWatermarking One for All: A Robust Watermarking Scheme Against Partial Image Theftpdf
generalMinding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion Segmentationpdf
motionIM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Mannerpdf
generalLink-based Contrastive Learning for One-Shot Unsupervised Domain Adaptationpdf
generalUniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object Detectionpdf
generalRethinking the Adversarial Robustness of Multi-Exit Neural Networks in an Attack-Defense Gamepdf
generalVideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guidepdf
generalOmniStereo: Real-time Omnidireactional Depth Estimation with Multiview Fisheye Cameraspdf
gsplatDroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone Imagerypdf
generalSDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Predictionpdf
beingsEfficient Video Face Enhancement with Enhanced Spatial-Temporal Consistencypdf
generalIDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generationpdf
generalLoTUS: Large-Scale Machine Unlearning with a Taste of Uncertaintypdf
generalSleeperMark: Towards Robust Watermark against Fine-Tuning Text-to-image Diffusion Modelspdf
generalIllumination Spectrum Estimation for Multispectral Images via Surface Reflectance Modeling and Spatial-Spectral Feature Generationpdf
generalTowards Explainable and Unprecedented Accuracy in Matching Challenging Finger Crease Patternspdf
generalNeural Hierarchical Decomposition for Single Image Plant Modelingpdf
generalDual-Agent Optimization framework for Cross-Domain Few-Shot Segmentationpdf
generalSACB-Net: Spatial-awareness Convolutions for Medical Image Registrationpdf
generalText Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Mapspdf
generalDCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusionpdf
beingsAIpparel: A Multimodal Foundation Model for Digital Garmentspdf
generalPO3AD: Predicting Point Offsets toward Better 3D Point Cloud Anomaly Detectionpdf
generalICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Modelspdf
generalPreciseCam: Precise Camera Control for Text-to-Image Generationpdf
generalSET: Spectral Enhancement for Tiny Object Detectionpdf
generalDifferentiable Inverse Rendering with Interpretable Basis BRDFspdf
generalEquiPose: Exploiting Permutation Equivariance for Relative Camera Pose Estimationpdf
beingsFace Forgery Video Detection via Temporal Forgery Cue Unravelingpdf
generalTemporally Consistent Object-Centric Learning by Contrasting Slotspdf
generalMC^2: Multi-concept Guidance for Customized Multi-concept Generationpdf
generalMulti-modal Vision Pre-training for Medical Image Analysispdf
generalSTEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Trainingpdf
generalLIM: Large Interpolator Model for Dynamic Reconstructionpdf
generalAutoPresent: Designing Structured Visuals from Scratchpdf
worldsVisionArena: 230k Real World User-VLM Conversations with Preference Labelspdf
generalFAM Diffusion: Frequency and Attention Modulation for High-Resolution Image Generation with Stable Diffusionpdf
beingsMultiGO: Towards Multi-level Geometry Learning for Monocular 3D Textured Human Reconstructionpdf
generalGenerative Photomontagepdf
generalMulti-view Reconstruction via SfM-guided Monocular Depth Estimationpdf
beingsHuMoCon: Concept Discovery for Human Motion Understandingpdf
worldsFreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free Promptspdf
generalRethinking Correspondence-based Category-Level Object Pose Estimationpdf
generalCurriculum Direct Preference Optimization for Diffusion and Consistency Modelspdf
generalPersonalized Preference Fine-tuning of Diffusion Modelspdf
generalNN-Former: Rethinking Graph Structure in Neural Architecture Representationpdf
generalA Unified Image-Dense Annotation Generation Model for Underwater Scenespdf
gsplatNTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on Thermodynamicspdf
generalFSHNet: Fully Sparse Hybrid Network for 3D Object Detectionpdf
generalJTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV Systemspdf
worldsHaWoR: World-Space Hand Motion Reconstruction from Egocentric Videospdf
generalHigh Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous Flightpdf
worldsGenerative Gaussian Splatting for Unbounded 3D City Generationpdf
generalGeoMM: On Geodesic Perspective for Multi-modal Learningpdf
generalVISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoningpdf
generalAdaptive Dropout: Unleashing Dropout across Layers for Generalizable Image Super-Resolutionpdf
generalBreaking the Memory Barrier of Contrastive Loss via Tile-Based Strategypdf
generalLearning Phase Distortion with Selective State Space Models for Video Turbulence Mitigationpdf
generalRoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Trainingpdf
generalDistraction is All You Need for Multimodal Large Language Model Jailbreakingpdf
generalLearning to Normalize on the SPD Manifold under Bures-Wasserstein Geometrypdf
generalSAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentationpdf
generalBEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidancepdf
beingsFinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal Guidancepdf
generalSeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restorationpdf
generalBeyond Generation: A Diffusion-based Low-level Feature Extractor for Detecting AI-generated Imagespdf
generalOptical-Flow Guided Prompt Optimization for Coherent Video Generationpdf
generalCode-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detectionpdf
generalRethinking Reconstruction and Denoising in the Dark: New Perspective, General Architecture and Beyondpdf
generalFederated Learning with Domain Shift Eraserpdf
generalDiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generationpdf
beingsLink to the Past: Temporal Propagation for Fast 3D Human Reconstruction from Monocular Videopdf
generalDeterministic Certification of Graph Neural Networks against Graph Poisoning Attacks with Arbitrary Perturbationspdf
generalA3: Few-shot Prompt Learning of Unlearnable Examples with Cross-Modal Adversarial Feature Alignmentpdf
generalMVPaint: Synchronized Multi-View Diffusion for Painting Anything 3Dpdf
generalASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Groundingpdf
motionMotiF: Making Text Count in Image Animation with Motion Focal Losspdf
generalTowards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse Weatherpdf
gsplatMaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Maskspdf
generalSoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Bindingpdf
generalCLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representationpdf
generalNavigating Image Restoration with VAR's Distribution Alignment Priorpdf
generalDissecting and Mitigating Diffusion Bias via Mechanistic Interpretabilitypdf
generalGraph Neural Network Combining Event Stream and Periodic Aggregation for Low-Latency Event-based Visionpdf
generalArtFormer: Controllable Generation of Diverse 3D Articulated Objectspdf
generalBridging Gait Recognition and Large Language Models Sequence Modelingpdf
beingsDiSRT-In-Bed: Diffusion-Based Sim-to-Real Transfer Framework for In-Bed Human Mesh Recoverypdf
generalCOUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shiftspdf
generalHOT: Hadamard-based Optimized Trainingpdf
generalTokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generationpdf
generalSnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Devicepdf
generalAdapting Dense Matching for Homography Estimation with Grid-based Accelerationpdf
generalCoSDH: Communication-Efficient Collaborative Perception via Supply-Demand Awareness and Intermediate-Late Hybridizationpdf
generalStereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Failpdf
generalOrder-Robust Class Incremental Learning: Graph-Driven Dynamic Similarity Groupingpdf
generalWhere the Devil Hides: Deepfake Detectors Can No Longer Be Trustedpdf
generalCaMuViD: Calibration-Free Multi-View Detectionpdf
generalProsody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbingpdf
generalHVI: A New Color Space for Low-light Image Enhancementpdf
generalDualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstructionpdf
generalOne Diffusion to Generate Them Allpdf
generalCoSER: Towards Consistent Dense Multiview Text-to-Image Generator for 3D Creationpdf
generalUNEM: UNrolled Generalized EM for Transductive Few-Shot Learningpdf
generalG3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulationpdf
generalExplaining Domain Shifts in Language: Concept Erasing for Interpretable Image Classificationpdf
gsplatTextured Gaussians for Enhanced 3D Scene Appearance Modelingpdf
generalNeighborRetr: Balancing Hub Centrality in Cross-Modal Retrievalpdf
worldsGlobal-Local Tree Search in VLMs for 3D Indoor Scene Generationpdf
generalGFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow Networkspdf
generalMAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretrainingpdf
generalSegment Any-Quality Images with Generative Latent Space Enhancementpdf
worldsCityWalker: Learning Embodied Urban Navigation from Web-Scale Videospdf
generalLearning Visual Composition through Improved Semantic Guidancepdf
generalJanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generationpdf
generalVisual Prompting for One-shot Controllable Video Editing without Inversionpdf
generalAVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual Learningpdf
generalFlash-Split: 2D Reflection Removal with Flash Cues and Latent Diffusion Separationpdf
generalAttention IoU: Examining Biases in CelebA using Attention Mapspdf
generalHEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluatorpdf
motionSegment Any Motion in Videospdf
generalTask-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Groundingpdf
generalPatchDEMUX: A Certifiably Robust Framework for Multi-label Classifiers Against Adversarial Patchespdf
generalEgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answeringpdf
generalToken Cropr: Faster ViTs for Quite a Few Taskspdf
generalSTCOcc: Sparse Spatial-Temporal Cascade Renovation for 3D Occupancy and Scene Flow Predictionpdf
generalResilient Sensor Fusion Under Adverse Sensor Failures via Multi-Modal Expert Fusionpdf
generalMambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothingpdf
worldsIndoorGS: Geometric Cues Guided Gaussian Splatting for Indoor Scene Reconstructionpdf
generalPoint-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysispdf
generalStop Walking in Circles! Bailing Out Early in Projected Gradient Descentpdf
generalMoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulationpdf
generalStable-SCore: A Stable Registration-based Framework for 3D Shape Correspondencepdf
generalBeyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and Harmonizationpdf
generalAlign-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Large Model Enhancementpdf
generalPose Priors from Language Modelspdf
generalLogoSP: Local-global Grouping of Superpoints for Unsupervised Semantic Segmentation of 3D Point Cloudspdf
generalExploring Intrinsic Normal Prototypes within a Single Image for Universal Anomaly Detectionpdf
generalAugmenting Perceptual Super-Resolution via Image Quality Predictorspdf
generalTurboFill: Adapting Few-step Text-to-image Model for Fast Image Inpaintingpdf
beingsStochastic Human Motion Prediction with Memory of Action Transition and Action Characteristicpdf
generalPerception Tokens Enhance Visual Reasoning in Multimodal Language Modelspdf
beingsX-Dyna: Expressive Dynamic Human Image Animationpdf
generalTowards Effective and Sparse Adversarial Attack on Spiking Neural Networks via Breaking Invisible Surrogate Gradientspdf
generalTowards Understanding and Quantifying Uncertainty for Text-to-Image Generationpdf
generalLLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understandingpdf
generalInsight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Modelspdf
generalMaIR: A Locality- and Continuity-Preserving Mamba for Image Restorationpdf
generalRSAR: Restricted State Angle Resolver and Rotated SAR Benchmarkpdf
motionContinuous Space-Time Video Resampling with Invertible Motion Steganographypdf
generalProtoDepth: Unsupervised Continual Depth Completion with Prototypespdf
beingsParaHome: Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactionspdf
generalAdapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholdspdf
beingsOpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generationpdf
generalTrack Any Anomalous Object:A Granular Video Anomaly Detection Pipelinepdf
generalObject-aware Sound Source Localization via Audio-Visual Scene Understandingpdf
generalSerialGen: Personalized Image Generation by First Standardization Then Personalizationpdf
generalAugmented Deep Contexts for Spatially Embedded Video Codingpdf
generalProximal Algorithm Unrolling: Flexible and Efficient Reconstruction Networks for Single-Pixel Imagingpdf
beingsImage Quality Assessment: From Human to Machine Preferencepdf
generalContext-Aware Multimodal Pretrainingpdf
generalTask-driven Image Fusion with Learnable Fusion Losspdf
generalLamRA: Large Multimodal Model as Your Advanced Retrieval Assistantpdf
generalCoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generationpdf
generalMoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervisionpdf
generalBeyond Background Shift: Rethinking Instance Replay in Continual Semantic Segmentationpdf
generalScaleLSD: Scalable Deep Line Segment Detection Streamlinedpdf
generalRevisiting MAE Pre-training for 3D Medical Image Segmentationpdf
beingsChatHuman: Chatting about 3D Humans with Toolspdf
generalScalable Autoregressive Monocular Depth Estimationpdf
generalRecurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrievalpdf
generalCamouflage Anything: Learning to Hide using Controlled Out-painting and Representation Engineeringpdf
generalTest-Time Fine-Tuning of Image Compression Models for Multi-Task Adaptabilitypdf
beingsDynPose: Largely Improving the Efficiency of Human Pose Estimation by a Simple Dynamic Frameworkpdf
generalVideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videospdf
generalDistinguish Then Exploit: Source-free Open Set Domain Adaptation via Weight Barcode Estimation and Sparse Label Assignmentpdf
generalMulti-Sensor Object Anomaly Detection: Unifying Appearance, Geometry, and Internal Propertiespdf
meshingFancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformationpdf
gsplatPlanarSplatting: Accurate Planar Surface Reconstruction in 3 Minutespdf
generalOmni-ID: Holistic Identity Representation Designed for Generative Taskspdf
generalMIRE: Matched Implicit Neural Representationspdf
generalAeSPa : Attention-guided Self-supervised Parallel Imaging for MRI Reconstructionpdf
generalRobSense: A Robust Multi-modal Foundation Model for Remote Sensing with Static, Temporal, and Incomplete Data Adaptabilitypdf
gsplatMAC-Ego3D: Multi-Agent Gaussian Consensus for Real-Time Collaborative Ego-Motion and Photorealistic 3D Reconstructionpdf
generalOnline Video Understanding: OVBench and VideoChat-Onlinepdf
generalLightLoc: Learning Outdoor LiDAR Localization at Light Speedpdf
generalAccurate Differential Operators for Hybrid Neural Fieldspdf
generalFeedEdit: Text-Based Image Editing with Dynamic Feedback Regulationpdf
generalClassifier-guided CLIP Distillation for Unsupervised Multi-label Classificationpdf
beingsUMotion: Uncertainty-driven Human Motion Estimation from Inertial and Ultra-wideband Unitspdf
generalScene Map-based Prompt Tuning for Navigation Instruction Generationpdf
gsplatDropoutGS: Dropping Out Gaussians for Better Sparse-view Renderingpdf
generalBridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language Modelspdf
generalEnhancing Dataset Distillation via Non-Critical Region Refinementpdf
gsplatPUP 3D-GS: Principled Uncertainty Pruning for 3D Gaussian Splattingpdf
worldsScribbleLight: Single Image Indoor Relighting with Scribblespdf
generalInsightEdit: Towards Better Instruction Following for Image Editingpdf
generalOne-for-More: Continual Diffusion Model for Anomaly Detectionpdf
generalOmni-RGPT: Unifying Image and Video Region-level Understanding via Token Markspdf
generalEDM: Equirectangular Projection-Oriented Dense Kernelized Feature Matchingpdf
generalEZSR: Event-based Zero-Shot Recognitionpdf
beingsSVFR: A Unified Framework for Generalized Video Face Restorationpdf
generalDecoupling Fine Detail and Global Geometry for Compressed Depth Map Super-Resolutionpdf
generalDigital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Datasetpdf
meshingMeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentationpdf
beingsDeClotH: Decomposable 3D Cloth and Human Body Reconstruction from a Single Imagepdf
generalNeuron: Learning Context-Aware Evolving Representations for Zero-Shot Skeleton Action Recognitionpdf
motionHigh-Fidelity Relightable Monocular Portrait Animation with Lighting-Controllable Video Diffusion Modelpdf
generalPlug-and-Play PPO: An Adaptive Point Prompt Optimizer Making SAM Greaterpdf
generalEchoONE: Segmenting Multiple Echocardiography Planes in One Modelpdf
generalEasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wildpdf
generalPatch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perceptionpdf
worldsPosterO: Structuring Layout Trees to Enable Language Models in Generalized Content-Aware Layout Generationpdf
generalOne2Any: One-Reference 6D Pose Estimation for Any Objectpdf
generalContextual AD Narration with Interleaved Multimodal Sequencepdf
generalMNE-SLAM: Multi-Agent Neural SLAM for Mobile Robotspdf
generalTensoFlow: Tensorial Flow-based Sampler for Inverse Renderingpdf
generalFRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answeringpdf
generalLSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferencespdf
generalExploring Temporally-Aware Features for Point Trackingpdf
generalV^2Dial: Unification of Video and Visual Dialog via Multimodal Expertspdf
generalDetail-Preserving Latent Diffusion for Stable Shadow Removalpdf
generalCrossOver: 3D Scene Cross-Modal Alignmentpdf
generalRethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Predictionpdf
generalScalable Video-to-Dataset Generation for Cross-Platform Mobile Agentspdf
generalTokenize Image Patches: Global Context Fusion for Effective Haze Removal in Large Imagespdf
generalANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interactionpdf
generalMET3R: Measuring Multi-View Consistency in Generated Imagespdf
generalSegmenting Maxillofacial Structures in CBCT Volumespdf
general3D Dental Model Segmentation with Geometrical Boundary Preservingpdf
generalVideoGigaGAN: Towards Detail-rich Video Super-Resolutionpdf
generalGLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentationpdf
generalTowards RAW Object Detection in Diverse Conditionspdf
generalFLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-trainingpdf
generalAdapter Merging with Centroid Prototype Mapping for Scalable Class-Incremental Learningpdf
worldsOpenSDI: Spotting Diffusion-Generated Images in the Open Worldpdf
generalLux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Datasetpdf
generalDiG: Scalable and Efficient Diffusion Models with Gated Linear Attentionpdf
gsplatMonocular and Generalizable Gaussian Talking Head Animationpdf
generalLocally Orderless Images for Optimization in Differentiable Renderingpdf
generalPlug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Controlpdf
generalFine-Grained Erasure in Text-to-Image Diffusion-based Foundation Modelspdf
generalDeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Modelspdf
beingsALIEN: Implicit Neural Representations for Human Motion Prediction under Arbitrary Latencypdf
generalSufficient Invariant Learning for Distribution Shiftpdf
generalDomain Generalization in CLIP via Learning with Diverse Text Promptspdf
generalIterIS: Iterative Inference-Solving Alignment for LoRA Mergingpdf
generalEfficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise Replacementpdf
generalPrEditor3D: Fast and Precise 3D Shape Editingpdf
generalComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matricespdf
generalLOCORE: Image Re-ranking with Long-Context Sequence Modelingpdf
generalLiVOS: Light Video Object Segmentation with Gated Linear Matchingpdf
generalPolarized Color Screen Mattingpdf
generalGOAL: Global-local Object Alignment Learningpdf
generalPost-pre-training for Modality Alignment in Vision-Language Foundation Modelspdf
beingsSynthLight: Portrait Relighting with Diffusion Model by Learning to Re-render Synthetic Facespdf
generalPseudo Visible Feature Fine-Grained Fusion for Thermal Object Detectionpdf
generalNVILA: Efficient Frontier Visual Language Modelspdf
generalSemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spottingpdf
generalNoPain: No-box Point Cloud Attack via Optimal Transport Singular Boundarypdf
beingsFADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillationpdf
gsplatGeometry Field Splatting with Gaussian Surfelspdf
generalPS-EIP: Robust Photometric Stereo Based on Event Interval Profilepdf
generalGenPC: Zero-shot Point Cloud Completion via 3D Generative Priorspdf
generalLatent Drifting in Diffusion Models for Counterfactual Medical Image Synthesispdf
generalRethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking Transformerspdf
generalHeMoRa: Unsupervised Heuristic Consensus Sampling for Robust Point Cloud Registrationpdf
generalReducing Class-wise Confusion for Incremental Learning with Disentangled Manifoldspdf
motionAniMo: Species-Aware Model for Text-Driven Animal Motion Generationpdf
generalEditAR: Unified Conditional Generation with Autoregressive Modelspdf
generalInstance-wise Supervision-level Optimization in Active Learningpdf
generalBHViT: Binarized Hybrid Vision Transformerpdf
generalPathways on the Image Manifold: Image Editing via Video Generationpdf
gsplatDeSplat: Decomposed Gaussian Splatting for Distractor-Free Renderingpdf
generalStable Flow: Vital Layers for Training-Free Image Editingpdf
beingsTokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generationpdf
generalCholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Toolspdf
generalConditional Balance: Improving Multi-Conditioning Trade-Offs in Image Generationpdf
beingsKeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolationpdf
generalContext-Enhanced Memory-Refined Transformer for Online Action Detectionpdf
generalData-Free Group-Wise Fully Quantized Winograd Convolution via Learnable Scalespdf
worldsGEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Controlpdf
generalA Dataset for Semantic Segmentation in the Presence of Unknownspdf
generalHierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understandingpdf
generalCASP: Compression of Large Multimodal Models Based on Attention Sparsitypdf
generalUNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generationpdf
generalTowards Cost-Effective Learning: A Synergy of Semi-Supervised and Active Learningpdf
generalGenerative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesispdf
generalAdvancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 Datasetpdf
gsplatEnvGS: Modeling View-Dependent Appearance with Environment Gaussianpdf
generalProvoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learningpdf
generalMonoDGP: Monocular 3D Object Detection with Decoupled-Query and Geometry-Error Priorspdf
generalFlexible Group Count Enables Hassle-Free Structured Pruningpdf
beingsEasyCraft: A Robust and Efficient Framework for Automatic Avatar Craftingpdf
meshingMeshArt: Generating Articulated Meshes with Structure-Guided Transformerspdf
generalAdaptive Non-Uniform Timestep Sampling for Accelerating Diffusion Model Trainingpdf
generalExplainable Saliency: Articulating Reasoning with Contextual Prioritizationpdf
generalCompass Control: Multi Object Orientation Control for Text-to-Image Generationpdf
generalUnleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulationpdf
generalVideoGEM: Training-free Action Grounding in Videospdf
motionStructure-from-Motion with a Non-Parametric Camera Modelpdf
beingsLAL: Enhancing 3D Human Motion Prediction with Latency-aware Auxiliary Learningpdf
generalRefPose: Leveraging Reference Geometric Correspondences for Accurate 6D Pose Estimation of Unseen Objectspdf
generalRelation3D : Enhancing Relation Modeling for Point Cloud Instance Segmentationpdf
beingsShape My Moves: Text-Driven Shape-Aware Synthesis of Human Motionspdf
worldsBeyond Human Perception: Understanding Multi-Object World from Monocular Viewpdf
gsplatLIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fieldspdf
generalLatent Space Super-Resolution for Higher-Resolution Image Generation with Diffusion Modelspdf
generalRethinking Noisy Video-Text Retrieval via Relation-aware Alignmentpdf
generalScaling Vision Pre-Training to 4K Resolutionpdf
beingsGarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulationpdf
generalImproving Editability in Image Generation with Layer-wise Memorypdf
generalSimplification Is All You Need against Out-of-Distribution Overconfidencepdf
generalSpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Inputpdf
gsplatLOD-GS: Achieving Levels of Detail using Scalable Gaussian Souppdf
generalThe Devil is in Low-Level Features for Cross-Domain Few-Shot Segmentationpdf
generalDiscrete to Continuous: Generating Smooth Transition Poses from Sign Language Observationspdf
generalUnified Medical Lesion Segmentation via Self-referring Indicatorpdf
generalUnveiling Differences in Generative Models: A Scalable Differential Clustering Approachpdf
generalSemantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentationpdf
generalPhyS-EdiT: Physics-aware Semantic Image Editing with Text Descriptionpdf
worldsSceneDiffuser++: City-Scale Traffic Simulation via a Generative World Modelpdf
gsplatGaussian Splatting Feature Fields for (Privacy-Preserving) Visual Localizationpdf
beingsGazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activitiespdf
generalRecover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal Transportpdf
generalDyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understandingpdf
generalFrom Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibrationpdf
beingsHERA: Hybrid Explicit Representation for Ultra-Realistic Head Avatarspdf
generalCoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantificationpdf
generalHierarchical Adaptive Filtering Network for Text Image Specular Highlight Removalpdf
generalLearning Extremely High Density Crowds as Active Matterspdf
beingsEchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animationpdf
gsplatGaussian Splatting for Efficient Satellite Image Photogrammetrypdf
generalTowards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long Textpdf
generalParallel Sequence Modeling via Generalized Spatial Propagation Networkpdf
generalNADER: Neural Architecture Design via Multi-Agent Collaborationpdf
generalFortifying Federated Learning Towards Trustworthiness via Auditable Data Valuation and Verifiable Client Contributionpdf
gsplatUniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splattingpdf
generalCharm: The Missing Piece in ViT Fine-Tuning for Image Aesthetic Assessmentpdf
generalSeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Groundingpdf
generalLinear Attention Modeling for Learned Image Compressionpdf
generalAsynchronous Collaborative Graph Representation for Frames and Eventspdf
worldsReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restorationpdf
generalGenFusion: Closing the Loop between Reconstruction and Generation via Videospdf
generalThe PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour Recognitionpdf
motionSonic: Shifting Focus to Global Audio Perception in Portrait Animationpdf
worldsMultitwine: Multi-Object Compositing with Text and Layout Controlpdf
generalVideo Depth without Video Modelspdf
generalPointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learningpdf
beingsHumanRig: Learning Automatic Rigging for Humanoid Character in a Large Scale Datasetpdf
gsplatGaussHDR: High Dynamic Range Gaussian Splatting via Learning Unified 3D and 2D Local Tone Mappingpdf
worldsChannel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor Scenespdf
generalTargeted Forgetting of Image Subgroups in CLIP Modelspdf
generalSeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Modelpdf
generalDSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot Segmentationpdf
generalSemAlign3D: Semantic Correspondence between RGB-Images through Aligning 3D Object-Class Representationspdf
generalDPU: Dynamic Prototype Updating for Multimodal Out-of-Distribution Detectionpdf
generalTowards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipationpdf
generalBoosting Domain Incremental Learning: Selecting the Optimal Parameters is All You Needpdf
generalPerceptual Inductive Bias Is What You Need Before Contrastive Learningpdf
beingsFaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMspdf
generalEffiDec3D: An Optimized Decoder for High-Performance and Efficient 3D Medical Image Segmentationpdf
generalExploring Historical Information for RGBE Visual Tracking with Mambapdf
worldsArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediarypdf
generalImproving Sound Source Localization with Joint Slot Attention on Image and Audiopdf
meshingFeature-Preserving Mesh Decimation for Normal Integrationpdf
generalDexHandDiff: Interaction-aware Diffusion Planning for Adaptive Dexterous Manipulationpdf
generalVerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awarenesspdf
generalMemories of Forgotten Conceptspdf
generalPSBD: Prediction Shift Uncertainty Unlocks Backdoor Detectionpdf
motionPhoenix: A Motion-based Self-Reflection Framework for Fine-grained Robotic Action Correctionpdf
generalMultirate Neural Image Compression with Adaptive Lattice Vector Quantizationpdf
generalEventFly: Event Camera Perception from Ground to the Skypdf
generalCH3Depth: Efficient and Flexible Depth Foundation Model with Flow Matchingpdf
generalPow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priorspdf
generalMambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Trackingpdf
generalAdaptive Parameter Selection for Tuning Vision-Language Modelspdf
generalLearning-enabled Polynomial Lyapunov Function Synthesis via High-Accuracy Counterexample-Guided Frameworkpdf
generalPatient-Level Anatomy Meets Scanning-Level Physics: Personalized Federated Low-Dose CT Denoising Empowered by Large Language Modelpdf
gsplatExploiting Deblurring Networks for Radiance Fieldspdf
generalRethinking Lanes and Points in Complex Scenarios for Monocular 3D Lane Detectionpdf
generalPSA-SSL: Pose and Size-aware Self-Supervised Learning on LiDAR Point Cloudspdf
generalSnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Trainingpdf
generalCommunity Forensics: Using Thousands of Generators to Train Fake Image Detectorspdf
motionModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modelingpdf
generalQuaffure: Real-Time Quasi-Static Neural Hair Simulationpdf
worldsLumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relightingpdf
generalDiC: Rethinking Conv3x3 Designs in Diffusion Modelspdf
gsplatMoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffoldspdf
generalNot Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Modelspdf
beingsTokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenizationpdf
generalShapeShifter: 3D Variations Using Multiscale and Sparse Point-Voxel Diffusionpdf
beingsFRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few Imagespdf
gsplatReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoningpdf
worldsWonderWorld: Interactive 3D Scene Generation from a Single Imagepdf
generalA Lightweight UDF Learning Framework for 3D Reconstruction Based on Local Shape Functionspdf
generalDiffCAM: Data-Driven Saliency Maps by Capturing Feature Differencespdf
motionFrom Sparse Signal to Smooth Motion: Real-Time Motion Generation with Rolling Prediction Modelspdf
generalImagine and Seek: Improving Composed Image Retrieval with an Imagined Proxypdf
generalEMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotionspdf
generalReversing Flow for Image Restorationpdf
generalShadow Generation Using Diffusion Model with Geometry Priorpdf
generalRethinking Epistemic and Aleatoric Uncertainty for Active Open-Set Annotation: An Energy-Based Approachpdf
generalAny3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask Trackingpdf
generalFDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image Editingpdf
generalMMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modelingpdf
generalROS-SAM: High-Quality Interactive Segmentation for Remote Sensing Moving Objectpdf
generalMI-DETR: An Object Detection Model with Multi-time Inquiries Mechanismpdf
generalSynthetic Visual Genomepdf
generalStop Learning it all to Mitigate Visual Hallucination, Focus on the Hallucination Target.pdf
generalSeeing the Abstract: Translating the Abstract Language for Vision Language Modelspdf
generalOne-Step Event-Driven High-Speed Autofocuspdf
generalPanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial Augmentationpdf
beingsTowards High-fidelity 3D Talking Avatar with Personalized Dynamic Texturepdf
worldsScene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Modelpdf
worldsJiSAM: Alleviate Labeling Burden and Corner Case Problems in Autonomous Driving via Minimal Real-World Datapdf
generalOSMamba: Omnidirectional Spectral Mamba with Dual-Domain Prior Generator for Exposure Correctionpdf
generalRandAR: Decoder-only Autoregressive Visual Generation in Random Orderspdf
generalVid2Sim: Realistic and Interactive Simulation from Video for Urban Navigationpdf
generalType-R: Automatically Retouching Typos for Text-to-Image Generationpdf
generalVideo-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understandingpdf
generalSingle Domain Generalization for Few-Shot Counting via Universal Representation Matchingpdf
generalDiscovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learningpdf
generalDo We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wildpdf
generalA General Adaptive Dual-level Weighting Mechanism for Remote Sensing Pansharpeningpdf
generalRASP: Revisiting 3D Anamorphic Art for Shadow-Guided Packing of Irregular Objectspdf
generalDiffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image Captioningpdf
generalMODfinity: Unsupervised Domain Adaptation with Multimodal Information Flow Intertwiningpdf
generalTowards Universal Soccer Video Understandingpdf
generalEnhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Modelspdf
generalNVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Imagespdf
generalEfficient Personalization of Quantized Diffusion Model without Backpropagationpdf
generalAlias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Spacepdf
generalKVQ: Boosting Video Quality Assessment via Saliency-guided Local Perceptionpdf
generalLearning Flow Fields in Attention for Controllable Person Image Generationpdf
generalEarly-Bird Diffusion: Investigating and Leveraging Timestep-Aware Early-Bird Tickets in Diffusion Models for Efficient Trainingpdf
generalDIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videospdf
generalEfficientLLaVA: Generalizable Auto-Pruning for Large Vision-language Modelspdf
generalOnlineAnySeg: Online Zero-Shot 3D Segmentation by Visual Foundation Model Guided 2D Mask Mergingpdf
generalTora: Trajectory-oriented Diffusion Transformer for Video Generationpdf
gsplatMorpheus: Text-Driven 3D Gaussian Splat Shape and Color Stylizationpdf
generalGeneralized Diffusion Detector: Mining Robust Features from Diffusion Models for Domain-Generalized Detectionpdf
motionEnhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Modelpdf
generalNeuro-Symbolic Evaluation of Text-to-Video Models using Formal Verificationpdf
generalSpherical Manifold Guided Diffusion Model for Panoramic Image Generationpdf
generalRethinking Query-based Transformer for Continual Image Segmentationpdf
generalSlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understandingpdf
generalText-guided Sparse Voxel Pruning for Efficient 3D Visual Groundingpdf
generalCommon3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature Spacepdf
generalECVC: Exploiting Non-Local Correlations in Multiple Frames for Contextual Video Compressionpdf
generalLinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexitypdf
generalTowards Open-Vocabulary Audio-Visual Event Localizationpdf
gsplatS2Gaussian: Sparse-View Super-Resolution 3D Gaussian Splattingpdf
generalHIIF: Hierarchical Encoding based Implicit Image Function for Continuous Super-resolutionpdf
motionMotion Prompting: Controlling Video Generation with Motion Trajectoriespdf
generalVERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Modelspdf
generalCCIN: Compositional Conflict Identification and Neutralization for Composed Image Retrievalpdf
generalOverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernelspdf
generalFlash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Localitypdf
beingsComprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and Harmonizationpdf
gsplatMAtCha Gaussians: Atlas of Charts for High-Quality Geometry and Photorealism From Sparse Viewspdf
generalExtreme Rotation Estimation in the Wildpdf
generalTraversing Distortion-Perception Tradeoff using a Single Score-Based Generative Modelpdf
generalTask-Agnostic Guided Feature Expansion for Class-Incremental Learningpdf
motionLet's Chorus: Partner-aware Hybrid Song-Driven 3D Head Animationpdf
generalTwinner: Shining Light on Digital Twins in a Few Snapspdf
generalMonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular Detectionpdf
gsplatSGCR: Spherical Gaussians for Efficient 3D Curve Reconstructionpdf
generalAdvancing Semantic Future Prediction through Multimodal Visual Sequence Transformerspdf
generalOn the Out-Of-Distribution Generalization of Large Multimodal Modelspdf
generalEnhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report Generationpdf
generalScaling Inference Time Compute for Diffusion Modelspdf
generalChat-based Person Retrieval via Dialogue-Refined Cross-Modal Alignmentpdf
generalAVF-MAE++: Scaling Affective Video Facial Masked Autoencoders via Efficient Audio-Visual Self-Supervised Learningpdf
motionDiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusionpdf
beingsGroundingFace: Fine-grained Face Understanding via Pixel Grounding Multimodal Large Language Modelpdf
generalDynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understandingpdf
generalV-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agentspdf
generalID-Patch: Robust ID Association for Group Photo Personalizationpdf
gsplatiG-6DoF: Model-free 6DoF Pose Estimation for Unseen Object via Iterative 3D Gaussian Splattingpdf
generalForestLPR: LiDAR Place Recognition in Forests Attentioning Multiple BEV Density Imagespdf
generalAdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimizationpdf
generalSLVR: Super-Light Visual Reconstruction via Blueprint Controllable Convolutions and Exploring Feature Diversity Representationpdf
generalLayered Image Vectorization via Semantic Simplificationpdf
generalHearing Anywhere in Any Environmentpdf
generalAutomated Proof of Polynomial Inequalities via Reinforcement Learningpdf
generalSAM-I2V: Upgrading SAM to Support Promptable Video Segmentation with Less than 0.2% Training Costpdf
motionHow Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactionspdf
generalJoint Vision-Language Social Bias Removal for CLIPpdf
generalMV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Secondspdf
generalExplicit Depth-Aware Blurry Video Frame Interpolation Guided by Differential Curvespdf
generalMonSter: Marry Monodepth to Stereo Unleashes Powerpdf
generalA Tale of Two Classes: Adapting Supervised Contrastive Learning to Binary Imbalanced Datasetspdf
generalLearning Class Prototypes for Unified Sparse-Supervised 3D Object Detectionpdf
worldsOpen-World Amodal Appearance Completionpdf
generalRivuletMLP: An MLP-based Architecture for Efficient Compressed Video Quality Enhancementpdf
generalReanimating Images using Neural Representations of Dynamic Stimulipdf
generalDepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videospdf
beingsRethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detectorpdf
generalActive Hyperspectral Imaging Using an Event Camerapdf
gsplatBridging the Gap between Gaussian Diffusion Models and Universal Quantization for Image Compressionpdf
generalSAMBLE: Shape-Specific Point Cloud Sampling for an Optimal Trade-Off Between Local Detail and Global Uniformitypdf
generalDriving by the Rules: A Benchmark for Integrating Traffic Sign Regulations into Vectorized HD Mappdf
generalS4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual Representationpdf
generalScience-T2I: Addressing Scientific Illusions in Image Synthesispdf
generalMoST: Efficient Monarch Sparse Tuning for 3D Representation Learningpdf
generalRe-thinking Temporal Search for Long-Form Video Understandingpdf
generalWhen Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven Approachpdf
generalBIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligencepdf
generalQuery Efficient Black-Box Visual Prompting with Subspace Learningpdf
generalImproving Autoregressive Visual Generation with Cluster-Oriented Token Predictionpdf
beingsDual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans?pdf
worldsSolving Instance Detection from an Open-World Perspectivepdf
worldsPercept, Memory, and Imagine: World Feature Simulating for Open-Domain Unknown Object Detectionpdf
generalEfficient Depth Estimation for Unstable Stereo Camera Systems on AR Glassespdf
gsplatLiDAR-RT: Gaussian-based Ray Tracing for Dynamic LiDAR Re-simulationpdf
generalLarge-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generatorpdf
gsplatFlow-NeRF: Joint Learning of Geometry, Poses, and Dense Flow within Unified Neural Representationspdf
motionConsistent and Controllable Image Animation with Motion Diffusion Modelspdf
generalAA-CLIP: Enhancing Zero-Shot Anomaly Detection via Anomaly-Aware CLIPpdf
gsplatHybridGS: Decoupling Transients and Statics with 2D and 3D Gaussian Splattingpdf
generalChannel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image Derainingpdf
generalMobileMamba: Lightweight Multi-Receptive Visual Mamba Networkpdf
generalSimLTD: Simple Supervised and Semi-Supervised Long-Tailed Object Detectionpdf
generalHyperSeg: Hybrid Segmentation Assistant with Fine-grained Visual Perceiverpdf
generalDiffusion-based Event Generation for High-Quality Image Deblurringpdf
generalBalanced Rate-Distortion Optimization in Learned Image Compressionpdf
generalBridge the Gap: From Weak to Full Supervision for Temporal Action Localization with PseudoFormerpdf

CVPR 2025 - Day 2 (2025-06-14)

familypaperlinks
generalDynScene: Scalable Generation of Dynamic Robotic Manipulation Scenes for Embodied AIpdf
generalDiffLocks: Generating 3D Hair from a Single Image using Diffusion Modelspdf
generalHarnessing Frequency Spectrum Insights for Image Copyright Protection Against Diffusion Modelspdf
generalIDEA-Bench: How Far are Generative Models from Professional Designing?pdf
generalPhD: A ChatGPT-Prompted Visual Hallucination Evaluation Datasetpdf
worldsClimbingCap: Multi-Modal Dataset and Method for Rock Climbing in World Coordinatepdf
generalA Bias-Free Training Paradigm for More General AI-generated Image Detectionpdf
generalFALCON: Fairness Learning via Contrastive Attention Approach to Continual Semantic Scene Understandingpdf
beingsCertified Human Trajectory Predictionpdf
generalTransformers without Normalizationpdf
beingsHiPART: Hierarchical Pose AutoRegressive Transformer for Occluded 3D Human Pose Estimationpdf
beingsFrom Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speechpdf
generalDFM: Differentiable Feature Matching for Anomaly Detectionpdf
generalPointSR: Self-Regularized Point Supervision for Drone-View Object Detectionpdf
worldsv-CLR: View-Consistent Learning for Open-World Instance Segmentationpdf
generalReloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localizationpdf
generalJanus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generationpdf
generalMagicArticulate: Make Your 3D Models Articulation-Readypdf
generalDual Prompting Image Restoration with Diffusion Transformerspdf
generalDepthCues: Evaluating Monocular Depth Perception in Large Vision Modelspdf
gsplatSpecTRe-GS: Modeling Highly Specular Surfaces with Reflected Nearby Objects by Tracing Rays in 3D Gaussian Splattingpdf
generalAuraFusion360: Augmented Unseen Region Alignment for Reference-based 360deg Unbounded Scene Inpaintingpdf
generalLanguage-Guided Image Tokenization for Generationpdf
generalHyperbolic Uncertainty-Aware Few-Shot Incremental Point Cloud Segmentationpdf
beingsD^3-Human: Dynamic Disentangled Digital Human from Monocular Videopdf
generalCurriculum Coarse-to-Fine Selection for High-IPC Dataset Distillationpdf
generalBADGR: Bundle Adjustment Diffusion Conditioned by Gradients for Wide-Baseline Floor Plan Reconstructionpdf
generalThree Cars Approaching within 100m! Enhancing Distant Geometry by Tri-Axis Voxel Scanning for Camera-based Semantic Scene Completionpdf
generalDViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehensionpdf
generalSpiking Transformer with Spatial-Temporal Attentionpdf
generalPerceptual Video Compression with Neural Wrappingpdf
generalViKIENet: Towards Efficient 3D Object Detection with Virtual Key Instance Enhanced Networkpdf
generalData-free Universal Adversarial Perturbation with Pseudo-semantic Priorpdf
beingsFRAME: Floor-aligned Representation for Avatar Motion from Egocentric Videopdf
generalGeneralized Zero-Shot Classification via Semantics-Free Inter-Class Feature Generationpdf
generalMulti-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEspdf
generalMaDCoW: Marginal Distortion Correction for Wide-Angle Photography with Arbitrary Objectspdf
generalAny6D: Model-free 6D Pose Estimation of Novel Objectspdf
generalDrVideo: Document Retrieval Based Long Video Understandingpdf
generalBuffer Anytime: Zero-Shot Video Depth and Normal from Image Priorspdf
beingsPSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshingpdf
generalHiding Images in Diffusion Models by Editing Learned Score Functionspdf
generalWeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba Diffusionpdf
generalMUST: The First Dataset and Unified Framework for Multispectral UAV Single Object Trackingpdf
generalTightening Robustness Verification of MaxPool-based Neural Networks via Minimizing the Over-Approximation Zonepdf
generalPhysicsGen: Can Generative Models Learn from Images to Predict Complex Physical Relations?pdf
generalSpectral Informed Mamba for Robust Point Cloud Processingpdf
generalBlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representationspdf
generalD2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective Recognition.pdf
generalLaVin-DiT: Large Vision Diffusion Transformerpdf
generalCAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Imagepdf
generalAesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimizationpdf
gsplatBARD-GS: Blur-Aware Reconstruction of Dynamic Scenes via Gaussian Splattingpdf
generalDiN: Diffusion Model for Robust Medical VQA with Semantic Noisy Labelspdf
beingsS^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion Priorspdf
generalFSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understandingpdf
generalKeep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic Segmentationpdf
generalLLM-driven Multimodal and Multi-Identity Listening Head Generationpdf
meshingOffsetOPT: Explicit Surface Reconstruction without Normalspdf
generalAny-Resolution AI-Generated Image Detection by Spectral Learningpdf
generalSTOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understandingpdf
motionTimeTracker: Event-based Continuous Point Tracking for Video Frame Interpolation with Non-linear Motionpdf
worldsShading Meets Motion: Self-supervised Indoor 3D Reconstruction Via Simultaneous Shape-from-Shading and Structure-from-Motionpdf
generalBelieving is Seeing: Unobserved Object Detection using Generative Modelspdf
generalNLPrompt: Noise-Label Prompt Learning for Vision-Language Modelspdf
gsplatPBR-NeRF: Inverse Rendering with Physics-Based Neural Fieldspdf
generalNo Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image Recognitionpdf
generalClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large Language Modelspdf
generalTacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusionpdf
generalPhysical Plausibility-aware Trajectory Prediction via Locomotion Embodimentpdf
beingsAvatarArtist: Open-Domain 4D Avatarizationpdf
generalUsing Powerful Prior Knowledge of Diffusion Model in Deep Unfolding Networks for Image Compressive Sensingpdf
generalUniGoal: Towards Universal Zero-shot Goal-oriented Navigationpdf
generalNoise-Consistent Siamese-Diffusion for Medical Image Synthesis and Segmentationpdf
generalDefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspectionpdf
generalLess is More: Efficient Image Vectorization with Adaptive Parameterizationpdf
generalFedMIA: An Effective Membership Inference Attack Exploiting "All for One" Principle in Federated Learningpdf
generalDPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid Frameworkpdf
generalDocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learningpdf
generalSpatiotemporal Skip Guidance for Enhanced Video Diffusion Samplingpdf
generalODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language Modelspdf
generalMasking meets Supervision: A Strong Learning Alliancepdf
worldsDI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creationpdf
generalNotes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question Answeringpdf
generalUniRestore: Unified Perceptual and Task-Oriented Image Restoration Model Using Diffusion Priorpdf
generalCondensing Action Segmentation Datasets via Generative Network Inversionpdf
generalCan Generative Video Models Help Pose Estimation?pdf
generalDriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Drivingpdf
meshingHigh-Fidelity Lightweight Mesh Reconstruction from Point Cloudspdf
generalMDP: Multidimensional Vision Model Pruning with Latency Constraintpdf
beingsOSDFace: One-Step Diffusion Model for Face Restorationpdf
generalTask Singular Vectors: Reducing Task Interference in Model Mergingpdf
generalSelf-Evolving Visual Concept Library using Vision-Language Criticspdf
generalBoosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal Transportpdf
generalEffective Cloud Removal for Remote Sensing Images by an Improved Mean-Reverting Denoising Model with Elucidated Design Spacepdf
generalOpticalNet: An Optical Imaging Dataset and Benchmark Beyond the Diffraction Limitpdf
generalEmpowering Vector Graphics with Consistently Arbitrary Viewing and View-dependent Visibilitypdf
generalDIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-IDpdf
generalHyperPose: Hypernetwork-Infused Camera Pose Localization and an Extended Cambridge Landmarks Datasetpdf
generalMono3DVLT: Monocular-Video-Based 3D Visual Language Trackingpdf
generalTowards Universal Dataset Distillation via Task-Driven Diffusionpdf
meshingParametric Point Cloud Completion for Polygonal Surface Reconstructionpdf
generalSyncSDE: A Probabilistic Framework for Diffusion Synchronizationpdf
generalMCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generationpdf
generalDual Semantic Guidance for Open Vocabulary Semantic Segmentationpdf
generalGeneralizable Object Keypoint Localization from Generative Priorspdf
generalFedCALM: Conflict-aware Layer-wise Mitigation for Selective Aggregation in Deeper Personalized Federated Learningpdf
generalCaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Boothpdf
gsplatFlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splattingpdf
generalGeneralizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuningpdf
generalT2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generationpdf
generalMake It Count: Text-to-Image Generation with an Accurate Number of Objectspdf
generalTraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent Perceptionpdf
generalDreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Cachingpdf
generalFlexUOD: The Answer to Real-world Unsupervised Image Outlier Detectionpdf
generalFocusing on Tracks for Online Multi-Object Trackingpdf
generalIdentity-preserving Distillation Sampling by Fixed-Point Iteratorpdf
generalWiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wildpdf
generalBiomedCoOp: Learning to Prompt for Biomedical Vision-Language Modelspdf
beingsMV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimationpdf
generalAnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalitiespdf
worldsOVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?pdf
gsplatGuardSplat: Efficient and Robust Watermarking for 3D Gaussian Splattingpdf
generalRoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narrativespdf
worldsIs Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generationpdf
gsplatLookCloser: Frequency-aware Radiance Field for Tiny-Detail Scenepdf
worldsConvex Relaxation for Robust Vanishing Point Estimation in Manhattan Worldpdf
gsplatFruitNinja: 3D Object Interior Texture Generation with Gaussian Splattingpdf
generalTake the Bull by the Horns: Learning to Segment Hard Samplespdf
generalEIDT-V: Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video Generationpdf
generalReproducible Vision-Language Models Meet Concepts Out of Pre-Trainingpdf
motionThrough-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generationpdf
generalMAGE : Single Image to Material-Aware 3D via the Multi-View G-Buffer Estimation Modelpdf
generalMESC-3D:Mining Effective Semantic Cues for 3D Reconstruction from a Single Imagepdf
generalAdvancing Multiple Instance Learning with Continual Learning for Whole Slide Imagingpdf
generalDecoupled Distillation to Erase: A General Unlearning Method for Any Class-centric Taskspdf
generalTAET: Two-Stage Adversarial Equalization Training on Long-Tailed Distributionspdf
generalFew-shot Personalized Scanpath Predictionpdf
generalMamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Modelspdf
generalMamba as a Bridge: Where Vision Foundation Models Meet Vision Language Models for Domain-Generalized Semantic Segmentationpdf
beingsVision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D Scenespdf
beingsChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generationpdf
generalCLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair Losspdf
generalObjectMover: Generative Object Movement with Video Priorpdf
beingsMLLM-as-a-Judge for Image Safety without Human Labelingpdf
generalLearning to Filter Outlier Edges in Global SfMpdf
beingsForensics Adapter: Adapting CLIP for Generalizable Face Forgery Detectionpdf
generalKAC: Kolmogorov-Arnold Classifier for Continual Learningpdf
generalBOOTPLACE: Bootstrapped Object Placement with Detection Transformerspdf
generalFASTer: Focal token Acquiring-and-Scaling Transformer for Long-term 3D Objection Detectionpdf
generalGeometry-guided Online 3D Video Synthesis with Multi-View Temporal Consistencypdf
worldsPoint2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instancespdf
gsplatCoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Imagespdf
generalSemantic and Sequential Alignment for Referring Video Object Segmentationpdf
generalContinual SFT Matches Multimodal RLHF with Negative Supervisionpdf
generalSemantic-guided Cross-Modal Prompt Learning for Skeleton-based Zero-shot Action Recognitionpdf
generalChatGen: Automatic Text-to-Image Generation From FreeStyle Chattingpdf
generalVEU-Bench: Towards Comprehensive Understanding of Video Editingpdf
generalDecouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image Compressionpdf
generalYo'Chameleon: Personalized Vision and Language Generationpdf
generalPatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolutionpdf
generalFluxSpace: Disentangled Semantic Editing in Rectified Flow Modelspdf
generalAdversarial Domain Prompt Tuning and Generation for Single Domain Generalizationpdf
generalShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware Promptspdf
generalAuto-Encoded Supervision for Perceptual Image Super-Resolutionpdf
generalSilence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generationpdf
worldsIterative Predictor-Critic Code Decoding for Real-World Image Dehazingpdf
generalFilter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuningpdf
generalGradient-Guided Annealing for Domain Generalizationpdf
generalMicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Researchpdf
generalEvent-based Video Super-Resolution via State Space Modelspdf
generalMasked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene Understandingpdf
generalVidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understandingpdf
generalCARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interactionpdf
generalPaint by Inpaint: Learning to Add Image Objects by Removing Them Firstpdf
generalPMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba Adapterpdf
generalLC-Mamba: Local and Continuous Mamba with Shifted Windows for Frame Interpolationpdf
worldsZero-Shot Head Swapping in Real-World Scenariospdf
generalCAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignmentpdf
generalCOBRA: COmBinatorial Retrieval Augmentation for Few-Shot Adaptationpdf
meshingProbeSDF: Light Field Probes For Neural Surface Reconstructionpdf
generalHybrid Concept Bottleneck Modelspdf
generalDual Consolidation for Pre-Trained Model-Based Domain-Incremental Learningpdf
beingsRORem: Training a Robust Object Remover with Human-in-the-Looppdf
generalAll Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languagespdf
beingsVideo-Bench: Human-Aligned Video Generation Benchmarkpdf
generalMergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantizationpdf
generalAnyattack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Modelspdf
gsplatJoint Optimization of Neural Radiance Fields and Continuous Camera Motion from a Monocular Videopdf
gsplatIRGS: Inter-Reflective Gaussian Splatting with 2D Gaussian Ray Tracingpdf
beingsInterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactionspdf
generalEfficient Test-time Adaptive Object Detection via Sensitivity-Guided Pruningpdf
generalA Data-Centric Revisit of Pre-Trained Vision Models for Robot Learningpdf
generalVisual Agentic AI for Spatial Reasoning with a Dynamic APIpdf
generalFeature Spectrum Learning for Remote Sensing Change Detectionpdf
worldsDriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representationpdf
generalLoKi: Low-dimensional KAN for Efficient Fine-tuning Image Modelspdf
gsplatDr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding Registrationpdf
generalConsistent Normal Orientation for 3D Point Clouds via Least Squares on Delaunay Graphpdf
generalATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpaintingpdf
generalOptimizing for the Shortest Path in Denoising Diffusion Modelpdf
generalAntidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perceptionpdf
generalDynamic Pseudo Labeling via Gradient Cutting for High-Low Entropy Explorationpdf
generalVODiff: Controlling Object Visibility Order in Text-to-Image Generationpdf
generalCAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model Generationpdf
generalLeveraging SD Map to Augment HD Map-based Trajectory Predictionpdf
generalONDA-Pose: Occlusion-Aware Neural Domain Adaptation for Self-Supervised 6D Object Pose Estimationpdf
generalQ-PART: Quasi-Periodic Adaptive Regression with Test-time Training for Pediatric Left Ventricular Ejection Fraction Regressionpdf
generalComposing Parts for Expressive Object Generationpdf
generalCarPlanner: Consistent Auto-regressive Trajectory Planning for Large-Scale Reinforcement Learning in Autonomous Drivingpdf
generalApply Hierarchical-Chain-of-Generation to Complex Attributes Text-to-3D Generationpdf
generalSOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Drivingpdf
generalShift the Lens: Environment-Aware Unsupervised Camouflaged Object Detectionpdf
generalDriveScape: High-Resolution Driving Video Generation by Multi-View Feature Fusionpdf
generalTraining-free Neural Architecture Search through Variance of Knowledge of Deep Network Weightspdf
generalEvery SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyondpdf
generalRAEncoder: A Label-Free Reversible Adversarial Examples Encoder for Dataset Intellectual Property Protectionpdf
motionLatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand Diffusion.pdf
generalChapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMspdf
generalDistribution Prototype Diffusion Learning for Open-set Supervised Anomaly Detectionpdf
generalFull-DoF Egomotion Estimation for Event Cameras Using Geometric Solverspdf
generalTeaching Large Language Models to Regress Accurate Image Quality Scores Using Score Distributionpdf
generalYour Scale Factors are My Weapon: Targeted Bit-Flip Attacks on Vision Transformers via Scale Factor Manipulationpdf
generalMarten: Visual Question Answering with Mask Generation for Multi-modal Document Understandingpdf
generalMamba-Reg: Vision Mamba Also Needs Registerspdf
beingsVisual Persona: Foundation Model for Full-Body Human Customizationpdf
gsplatSOGS: Second-Order Anchor for Advanced 3D Gaussian Splattingpdf
generalMExD: An Expert-Infused Diffusion Model for Whole-Slide Image Classificationpdf
generalLet Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of Samplespdf
gsplatMoDec-GS: Global-to-Local Motion Decomposition and Temporal Interval Adjustment for Compact Dynamic 3D Gaussian Splattingpdf
generalDaCapo: Score Distillation as Stacked Bridge for Fast and High-quality 3D Editingpdf
generalNumber it: Temporal Grounding Videos like Flipping Mangapdf
generalSyncVP: Joint Diffusion for Synchronous Multi-Modal Video Predictionpdf
generalHUSH: Holistic Panoramic 3D Scene Understanding using Spherical Harmonicspdf
generalSkillMimic: Learning Basketball Interaction Skills from Demonstrationspdf
gsplatRGBAvatar: Reduced Gaussian Blendshapes for Online Modeling of Head Avatarspdf
generalEEE-Bench: A Comprehensive Multimodal Electrical And Electronics Engineering Benchmarkpdf
generalA Unified Framework for Heterogeneous Semi-supervised Learningpdf
gsplatFree360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed Viewspdf
generalOpen Ad-hoc Categorization with Contextualized Feature Learningpdf
generalDynamic Updates for Language Adaptation in Visual-Language Trackingpdf
generalMulti-focal Conditioned Latent Diffusion for Person Image Synthesispdf
worldsUncertainty Meets Diversity: A Comprehensive Active Learning Framework for Indoor 3D Object Detectionpdf
generalIdentity-Clothing Similarity Modeling for Unsupervised Clothing Change Person Re-Identificationpdf
generalOccMamba: Semantic Occupancy Prediction with State Space Modelspdf
generalCheb-GR: Rethinking K-nearest Neighbor Search in Re-ranking for Person Re-identificationpdf
generalSpotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Drivingpdf
generalIs `Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuningpdf
generalGCC: Generative Color Constancy via Diffusing a Color Checkerpdf
generalOn Denoising Walking Videos for Gait Recognitionpdf
generalConformal Prediction for Zero-Shot Modelspdf
motionPhysAnimator: Physics-Guided Generative Cartoon Animationpdf
generalFIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximationpdf
generalBACON: Improving Clarity of Image Captions via Bag-of-Concept Graphspdf
generalVasTSD: Learning 3D Vascular Tree-state Space Diffusion Model for Angiography Synthesispdf
gsplatPanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splattingpdf
generalWISNet: Pseudo Label Generation on Unbalanced and Patch Annotated Waste Imagespdf
beingsMixerMDM: Learnable Composition of Human Motion Diffusion Modelspdf
generalHand-held Object Reconstruction from RGB Video with Dynamic Interactionpdf
beingsAudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformerspdf
generalThinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spacespdf
generalA Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMspdf
generalSemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Modelspdf
beingsArc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidancepdf
generalSeeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenespdf
generalStructure from Collisionpdf
generalCrab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperationpdf
generalNullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projectionpdf
generalOralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware Miningpdf
gsplatSplatAD: Real-Time Lidar and Camera Rendering with 3D Gaussian Splatting for Autonomous Drivingpdf
generalAudio-Visual Instance Segmentationpdf
generalUniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimationpdf
generalRL-RC-DoT: A Block-level RL agent for Task-Aware Video Compressionpdf
generalRecognition-Synergistic Scene Text Editingpdf
beingsWildAvatar: Learning In-the-wild 3D Avatars from the Webpdf
generalRectified Diffusion Guidance for Conditional Generationpdf
generalIAAO: Interactive Affordance Learning for Articulated Objects in 3D Environmentspdf
generalRaSS: Improving Denoising Diffusion Samplers with Reinforced Active Sampling Schedulerpdf
generalOSV: One Step is Enough for High-Quality Image to Video Generationpdf
generalFuzzy Multimodal Learning for Trusted Cross-modal Retrievalpdf
generalGUI-Xplore: Empowering Generalizable GUI Agents with One Explorationpdf
generalFew-Shot Recognition via Stage-Wise Retrieval-Augmented Finetuningpdf
gsplatRestorGS: Depth-aware Gaussian Splatting for Efficient 3D Scene Restorationpdf
general4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusionpdf
generalZ-Magic: Zero-shot Multiple Attributes Guided Image Creatorpdf
generalOn the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free Approachpdf
beingsTowards General Visual-Linguistic Face Forgery Detectionpdf
generalMovie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Promptspdf
generalLongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videospdf
generalMitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Keypdf
generalSimpler Diffusion: 1.5 FID on ImageNet512 with Pixel-space Diffusionpdf
worldsSTING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspectionpdf
generalNot All Parameters Matter: Masking Diffusion Models for Enhancing Generation Abilitypdf
generalComplexity Experts are Task-Discriminative Learners for Any Image Restorationpdf
generalGenerative Omnimatte: Learning to Decompose Video into Layerspdf
general5%>100%: Breaking Performance Shackles of Full Fine-Tuning on Visual Recognition Taskspdf
worldsReal-IAD D3: A Real-World 2D/Pseudo-3D/3D Dataset for Industrial Anomaly Detectionpdf
generalATP: Adaptive Threshold Pruning for Efficient Data Encoding in Quantum Neural Networkspdf
motionDecoupled Motion Expression Video Segmentationpdf
generalK-LoRA: Unlocking Training-Free Fusion of Any Subject and Style LoRAspdf
generalWF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Modelpdf
generalXLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?pdf
generalEfficient Data Driven Mixture-of-Expert Extraction from Trained Networkspdf
generalStyleSSP: Sampling StartPoint Enhancement for Training-free Diffusion-based Method for Style Transferpdf
beingsMotions as Queries: One-Stage Multi-Person Holistic Human Motion Capturepdf
generalAMO Sampler: Enhancing Text Rendering with Overshootingpdf
generalImViD: Immersive Volumetric Videos for Enhanced VR Engagementpdf
generalI2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video Modelspdf
generalSaliuitl: Ensemble Salience Guided Recovery of Adversarial Patches against CNNspdf
generalOPTICAL: Leveraging Optimal Transport for Contribution Allocation in Dataset Distillationpdf
generalShow and Segment: Universal Medical Image Segmentation via In-Context Learningpdf
generalCADCrafter: Generating Computer-Aided Design Models from Unconstrained Imagespdf
generalGenerative Multiview Relighting for 3D Reconstruction under Extreme Illumination Variationpdf
generalDyMO: Training-Free Diffusion Model Alignment with Dynamic Multi-Objective Schedulingpdf
generalGENIUS: A Generative Framework for Universal Multimodal Searchpdf
meshingSF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglementpdf
generalTowards Precise Embodied Dialogue Localization via Causality Guided Diffusionpdf
generalEfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Dualitypdf
generalA4A: Adapter for Adapter Transfer via All-for-All Mapping for Cross-Architecture Modelspdf
generalViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentationpdf
generalA Universal Scale-Adaptive Deformable Transformer for Image Restoration across Diverse Artifactspdf
generalTowards Precise Scaling Laws for Video Diffusion Transformerspdf
generalSPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Trackingpdf
generalAnyCam: Learning to Recover Camera Poses and Intrinsics from Casual Videospdf
generalPixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Visionpdf
generalCan Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and Understandingpdf
generalTowards Efficient Foundation Model for Zero-shot Amodal Segmentationpdf
generalScaling Properties of Diffusion Models For Perceptual Taskspdf
generalExact: Exploring Space-Time Perceptive Clues for Weakly Supervised Satellite Image Time Series Semantic Segmentationpdf
generalPolarFree: Polarization-based Reflection-Free Imagingpdf
generalSeeking Consistent Flat Minima for Better Domain Generalization via Refining Loss Landscapespdf
generalMultimodalStudio: A Heterogeneous Sensor Dataset and Framework for Neural Rendering across Multiple Imaging Modalitiespdf
generalMuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image Translationpdf
generalImage Quality Assessment: Investigating Causal Perceptual Effects with Abductive Counterfactual Inferencepdf
generalPos3R: 6D Pose Estimation for Unseen Objects Made Easypdf
generalRaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera Fusionpdf
generalUnderstanding Multi-Task Activities from Single-Task Videospdf
motionCo-Speech Gesture Video Generation with Implicit Motion-Audio Entanglementpdf
generalTransPixeler: Advancing Text-to-Video Generation with Transparencypdf
generalWhat's in the Image? A Deep-Dive into the Vision of Vision Language Modelspdf
generalFreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenespdf
generalSeq2Time: Sequential Knowledge Transfer for Video LLM Temporal Groundingpdf
generalGPVK-VL: Geometry-Preserving Virtual Keyframes for Visual Localization under Large Viewpoint Changespdf
generalEnhancing Online Continual Learning with Plug-and-Play State Space Model and Class-Conditional Mixture of Discretizationpdf
generalGRAPHGPT-O: Synergistic Multimodal Comprehension and Generation on Graphspdf
generalModel Poisoning Attacks to Federated Learning via Multi-Round Consistencypdf
gsplatTaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian Splattingpdf
beingsStacking Brick by Brick: Aligned Feature Isolation for Incremental Face Forgery Detectionpdf
generalCO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AIpdf
generalGENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulationpdf
generalLocalized Concept Erasure for Text-to-Image Diffusion Models Using Training-Free Gated Low-Rank Adaptationpdf
generalCamera Resection from Known Line Pencils and a Radially Distorted Scanlinepdf
worldsSPC-GS: Gaussian Splatting with Semantic-Prompt Consistency for Indoor Open-World Free-view Synthesis from Sparse Inputspdf
generalM3amba: Memory Mamba is All You Need for Whole Slide Image Classificationpdf
generalRedefining <Creative> in Dictionary: Towards an Enhanced Semantic Understanding of Creative Generationpdf
generalDinomaly: The Less Is More Philosophy in Multi-Class Unsupervised Anomaly Detectionpdf
generalDetect-and-Guide: Self-regulation of Diffusion Models for Safe Text-to-Image Generation via Guideline Token Optimizationpdf
worldsMirrorVerse: Pushing Diffusion Models to Realistically Reflect the Worldpdf
gsplatEAP-GS: Efficient Augmentation of Pointcloud for 3D Gaussian Splatting in Few-shot Scene Reconstructionpdf
generalEmpowering Large Language Models with 3D Situation Awarenesspdf
generalEchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insightspdf
generalInteractive Medical Image Segmentation: A Benchmark Dataset and Baselinepdf
generalGigaHands: A Massive Annotated Dataset of Bimanual Hand Activitiespdf
generalAutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashingpdf
gsplatFrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned Priorspdf
generalPioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuningpdf
gsplatCompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D Gaussianspdf
generalFIRE: Robust Detection of Diffusion-Generated Images via Frequency-Guided Reconstruction Errorpdf
generalAssessing and Learning Alignment of Unimodal Vision and Language Modelspdf
generalAction Detail Matters: Refining Video Recognition with Local Action Queriespdf
generalGenerative Map Priors for Collaborative BEV Semantic Segmentationpdf
generalCoherent 3D Portrait Video Reconstruction via Triplane Fusionpdf
generalManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Graspingpdf
generalFedCS: Coreset Selection for Federated Learningpdf
generalDual-Granularity Semantic Guided Sparse Routing Diffusion Model for General Pansharpeningpdf
generalOmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contextspdf
generalSOAP: Vision-Centric 3D Semantic Scene Completion with Scene-Adaptive Decoder and Occluded Region-Aware View Projectionpdf
worldsSimVS: Simulating World Inconsistencies for Robust View Synthesispdf
generalFrom Zero to Detail: Deconstructing Ultra-High-Definition Image Restoration from Progressive Spectral Perspectivepdf
generalCOSMOS: Cross-Modality Self-Distillation for Vision Language Pre-trainingpdf
worldsLifting Motion to the 3D World via 2D Diffusionpdf
generalTAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language Modelspdf
generalActive Data Curation Effectively Distills Large-Scale Multimodal Modelspdf
generalSCSA: A Plug-and-Play Semantic Continuous-Sparse Attention for Arbitrary Semantic Style Transferpdf
generalCan't Slow Me Down: Learning Robust and Hardware-Adaptive Object Detectors against Latency Attacks for Edge Devicespdf
generalSAM2Object: Consolidating View Consistency via SAM2 for Zero-Shot 3D Instance Segmentationpdf
generalCDI: Copyrighted Data Identification in Diffusion Modelspdf
generalCRISP: Object Pose and Shape Estimation with Test-Time Adaptationpdf
gsplatCreating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian Splattingpdf
generalSim-to-Real Causal Transfer: A Metric Learning Approach to Causally-Aware Interaction Representationspdf
generalTripartite Weight-Space Ensemble for Few-Shot Class-Incremental Learningpdf
generalPerLA: Perceptive 3D Language Assistantpdf
generalPhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generationpdf
generalMask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generationpdf
generalJamMa: Ultra-lightweight Local Feature Matching with Joint Mambapdf
generalDyCoke: Dynamic Compression of Tokens for Fast Video Large Language Modelspdf
generalMammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alpspdf
motionDiffusion-based Realistic Listening Head Generation via Hybrid Motion Modelingpdf
meshingSAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokenspdf
worldsUniScene: Unified Occupancy-centric Driving Scene Generationpdf
generalLearning from Streaming Video with Orthogonal Gradientspdf
generalClassifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifierspdf
beingsAn Image-like Diffusion Method for Human-Object Interaction Detectionpdf
gsplatCOB-GS: Clear Object Boundaries in 3DGS Segmentation Based on Boundary-Adaptive Gaussian Splittingpdf
generalPEER Pressure: Model-to-Model Regularization for Single Source Domain Generalizationpdf
generalRevisiting Fairness in Multitask Learning: A Performance-Driven Approach for Variance Reductionpdf
motionVideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Modelspdf
generalCompositional Caching for Training-free Open-vocabulary Attribute Detectionpdf
generalVI^3NR: Variance Informed Initialization for Implicit Neural Representationspdf
generalM-LLM Based Video Frame Selection for Efficient Video Understandingpdf
generalSearch and Detect: Training-Free Long Tail Object Detection via Web-Image Retrievalpdf
generalUnleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulationpdf
generalDiffusion Model is Effectively Its Own Teacherpdf
generalUnCommon Objects in 3Dpdf
worldsLearning Textual Prompts for Open-World Semi-Supervised Learningpdf
generalLongDiff: Training-Free Long Video Generation in One Gopdf
generalMask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentationpdf
generalMPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Drivingpdf
generalByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Waypdf
generalMasked Point-Entity Contrast for Open-Vocabulary 3D Scene Understandingpdf
generalOn the Generalization of Handwritten Text Recognition Modelspdf
generalInsTaG: Learning Personalized 3D Talking Head from Few-Second Videopdf
generalBenchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioningpdf
generalRotation-Equivariant Self-Supervised Method in Image Denoisingpdf
generalFlashSloth : Lightning Multimodal Large Language Models via Embedded Visual Compressionpdf
generalT2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Drivingpdf
generalRealEdit: Reddit Edits As a Large-scale Empirical Dataset for Image Transformationspdf
generalVideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Steppdf
gsplat3D-HGS: 3D Half-Gaussian Splattingpdf
generalScale Efficient Training for Large Datasetspdf
generalDecoder Gradient Shield: Provable and High-Fidelity Prevention of Gradient-Based Box-Free Watermark Removalpdf
generalConvex Combination Star Shape Prior for Data-driven Image Semantic Segmentationpdf
generalParameter-efficient Fine-tuning in Hyperspherical Space for Open-vocabulary Semantic Segmentationpdf
generalRelative Pose Estimation through Affine Corrections of Monocular Depth Priorspdf
beingsZero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusionpdf
worldsOcclusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object Recognitionpdf
generalConical Visual Concentration for Efficient Large Vision-Language Modelspdf
generalFoundations of the Theory of Performance-Based Rankingpdf
gsplatBIGS: Bimanual Category-agnostic Interaction Reconstruction from Monocular Videos via 3D Gaussian Splattingpdf
generalFrequency-Biased Synergistic Design for Image Compression and Compensationpdf
gsplatSparse Voxels Rasterization: Real-time High-fidelity Radiance Field Renderingpdf
generalMambaIC: State Space Models for High-Performance Learned Image Compressionpdf
gsplatInstant Gaussian Stream: Fast and Generalizable Streaming of Dynamic Scene Reconstruction via Gaussian Splattingpdf
beingsLocality-Aware Zero-Shot Human-Object Interaction Detectionpdf
generalTwo by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulationpdf
generalSGFormer: Satellite-Ground Fusion for 3D Semantic Scene Completionpdf
generalRandom Conditioning for Diffusion Model Compression with Distillationpdf
gsplatHierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generationpdf
generalHeterogeneous Skeleton-Based Action Representation Learningpdf
motionAnyMap: Learning a General Camera Model for Structure-from-Motion with Unknown Distortion in Dynamic Scenespdf
generalLanguage Guided Concept Bottleneck Models for Interpretable Continual Learningpdf
worldsRe-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Modelpdf
generalOdd-One-Out: Anomaly Detection by Comparing with Neighborspdf
generalD^3CTTA: Domain-Dependent Decorrelation for Continual Test-Time Adaption of 3D LiDAR Segmentationpdf
generalA Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Trainingpdf
generalEmpowering LLMs to Understand and Generate Complex Vector Graphicspdf
gsplatPanoGS: Gaussian-based Panoptic Segmentation for 3D Open Vocabulary Scene Understandingpdf
generalMLVU: Benchmarking Multi-task Long Video Understandingpdf
generalRecovering Dynamic 3D Sketches from Videospdf
gsplatEigenGS Representation: From Eigenspace to Gaussian Image Spacepdf
generalMaSS13K: A Matting-level Semantic Segmentation Benchmarkpdf
generalEnhancing Testing-Time Robustness for Trusted Multi-View Classification in the Wildpdf
generalROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language Modelspdf
generalnnWNet: Rethinking the Use of Transformers in Biomedical Image Segmentation and Calling for a Unified Evaluation Benchmarkpdf
generalVELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailmentpdf
generalSeeing is Not Believing: Adversarial Natural Object Optimization for Hard-Label 3D Scene Attackspdf
generalLessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual Recognitionpdf
beingsPippo: High-Resolution Multi-View Humans from a Single Imagepdf
generalH2ST: Hierarchical Two-Sample Tests for Continual Out-of-Distribution Detectionpdf
generalMoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoderspdf
generalCamFreeDiff: Camera-free Image to Panorama Generation with Diffusion Modelpdf
generalImproving Visual and Downstream Performance of Low-Light Enhancer with Vision Foundation Models Collaborationpdf
generalFineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputspdf
generalDivot: Diffusion Powers Video Tokenizer for Comprehension and Generationpdf
generalTowards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Modelspdf
generalVideo-Guided Foley Sound Generation with Multimodal Controlspdf
generalF^3OCUS - Federated Finetuning of Vision-Language Foundation Models with Optimal Client Layer Updating Strategy via Multi-objective Meta-Heuristicspdf
general3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformationpdf
generalCan Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?pdf
generalg3D-LF: Generalizable 3D-Language Feature Fields for Embodied Taskspdf
generalUniReal: Universal Image Generation and Editing via Learning Real-world Dynamicspdf
generalExploring Contextual Attribute Density in Referring Expression Countingpdf
generalSegMAN: Omni-scale Context Modeling with State Space Models and Local Attention for Semantic Segmentationpdf
generalOmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flowspdf
generalSLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videospdf
beingsSemGeoMo: Dynamic Contextual Human Motion Generation with Semantic and Geometric Guidancepdf
worldsDetecting Open World Objects via Partial Attribute Assignmentpdf
generalNeural Inverse Rendering from Propagating Lightpdf
gsplatDecoupledGaussian: Object-Scene Decoupling for Physics-Based Interactionpdf
gsplatDashGaussian: Optimizing 3D Gaussian Splatting in 200 Secondspdf
generalPACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Modelspdf
generalR-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localizationpdf
generalStyle Evolving along Chain-of-Thought for Unknown-Domain Object Detectionpdf
gsplatOmniSplat: Taming Feed-Forward 3D Gaussian Splatting for Omnidirectional Images with Editable Capabilitiespdf
generalUniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detectionpdf
worldsRemote Photoplethysmography in Real-World and Extreme Lighting Scenariospdf
generalMulti-Modal Contrastive Masked Autoencoders: A Two-Stage Progressive Pre-training Approach for RGBD Datasetspdf
generalFont-Agent: Enhancing Font Understanding with Large Language Modelspdf
generalSecret Lies in Color: Enhancing AI-Generated Images Detection with Color Distribution Analysispdf
generalCross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understandingpdf
generalSVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situationpdf
generalMixture of Submodules for Domain Adaptive Person Searchpdf
generalSharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillationpdf
generalEvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with Eventspdf
worldsSeeing A 3D World in A Grain of Sandpdf
beingsMoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillationpdf
generalReason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrievalpdf
generalUniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Graspingpdf
generalApollo: An Exploration of Video Understanding in Large Multimodal Modelspdf
generalSkip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselvespdf
generalPatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generationpdf
motionMegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videospdf
generalRobust-MVTON: Learning Cross-Pose Feature Alignment and Fusion for Robust Multi-View Virtual Try-Onpdf
generalIdentity-Preserving Text-to-Video Generation by Frequency Decompositionpdf
gsplatFreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocitypdf
generalMOS: Modeling Object-Scene Associations in Generalized Category Discoverypdf
generalTest-time Augmentation Improves Efficiency in Conformal Predictionpdf
generalStoryGPT-V: Large Language Models as Consistent Story Visualizerspdf
generalEdge-SD-SR: Low Latency and Parameter Efficient On-device Super-Resolution with Stable Diffusion via Bidirectional Conditioningpdf
generalINFP: Audio-Driven Interactive Head Generation in Dyadic Conversationspdf
gsplatEVPGS: Enhanced View Prior Guidance for Splatting-based Extrapolated View Synthesispdf
generalGREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Groundingpdf
generalAdapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal Perceptionpdf
generalMASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representationspdf
generalUWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsingpdf
generalMosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph Learningpdf
generalFirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placementpdf
generalEnd-to-End Implicit Neural Representations for Classificationpdf
generalUNICL-SAM: Uncertainty-Driven In-Context Segmentation with Part Prototype Discoverypdf
motionLayered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videospdf
generalDiffusion Self-Distillation for Zero-Shot Customized Image Generationpdf
generalUncertainty-guided Perturbation for Image Super-Resolution Diffusion Modelpdf
beingsTowards Human-Understandable Multi-Dimensional Concept Discoverypdf
generalConText-CIR: Learning from Concepts in Text for Composed Image Retrievalpdf
generalPerturb-and-Revise: Flexible 3D Editing with Generative Trajectoriespdf
generalLearning Compatible Multi-Prize Subnetworks for Asymmetric Retrievalpdf
generalLoRACLR: Contrastive Adaptation for Customization of Diffusion Modelspdf
generalOpportunistic Single-Photon Time of Flightpdf
generalArgus: Vision-Centric Reasoning with Grounded Chain-of-Thoughtpdf
generalBootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representationspdf
generalEncapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesispdf
generalRetrieving Semantics from the Deep: an RAG Solution for Gesture Synthesispdf
generalImproving Personalized Search with Regularized Low-Rank Parameter Updatespdf
generalHyperLoRA: Parameter-Efficient Adaptive Generation for Portrait Synthesispdf
generalEchoMatch: Partial-to-Partial Shape Matching via Correspondence Reflectionpdf
generalCross-Modal Interactive Perception Network with Mamba for Lung Tumor Segmentation in PET-CT Imagespdf
generalSuperPC: A Single Diffusion Model for Point Cloud Completion, Upsampling, Denoising, and Colorizationpdf
generalMaintaining Consistent Inter-Class Topology in Continual Test-Time Adaptationpdf
gsplatGeneralized Gaussian Entropy Model for Point Cloud Attribute Compression with Dynamic Likelihood Intervalspdf
generalSelf-Learning Hyperspectral and Multispectral Image Fusion via Adaptive Residual Guided Subspace Diffusion Modelpdf
beingsStickMotion: Generating 3D Human Motions by Drawing a Stickmanpdf
generalEnduring, Efficient and Robust Trajectory Prediction Attack in Autonomous Driving via Optimization-Driven Multi-Frame Perturbation Frameworkpdf
generalToward Generalized Image Quality Assessment: Relaxing the Perfect Reference Quality Assumptionpdf
generalHazy Low-Quality Satellite Video Restoration Via Learning Optimal Joint Degradation Patterns and Continuous-Scale Super-Resolution Reconstructionpdf
generalOvercoming Shortcut Problem in VLM for Robust Out-of-Distribution Detectionpdf
generalRoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Roboticspdf
gsplatBG-Triangle: Bezier Gaussian Triangle for 3D Vectorization and Renderingpdf
generalTKG-DM: Training-free Chroma Key Content Generation Diffusion Modelpdf
generalLift3D Policy: Lifting 2D Foundation Models for Robust 3D Robotic Manipulationpdf
generalMulti-View Pose-Agnostic Change Localization with Zero Labelspdf
generalAccelerating Diffusion Transformer via Increment-Calibrated Caching with Channel-Aware Singular Value Decompositionpdf
worldsA Simple yet Effective Layout Token in Large Language Models for Document Understandingpdf
generalReconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Modelspdf
beingsMEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attentionpdf
generalFree Lunch Enhancements for Multi-modal Crowd Countingpdf
gsplatEVolSplat: Efficient Volume-based Gaussian Splatting for Urban View Synthesispdf
generalPIDSR: Complementary Polarized Image Demosaicing and Super-Resolutionpdf
generalMegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Datapdf
generalHandOS: 3D Hand Reconstruction in One Stagepdf
generalAll-Day Multi-Camera Multi-Target Trackingpdf
beingsEnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Spacepdf
generalStarVector: Generating Scalable Vector Graphics Code from Images and Textpdf
generalExplaining in Diffusion: Explaining a Classifier with Diffusion Semanticspdf
generalAttention Distillation: A Unified Approach to Visual Characteristics Transferpdf
generalFrom Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editingpdf
generalDreamRelation: Bridging Customization and Relation Generationpdf
gsplatDepth-Guided Bundle Sampling for Efficient Generalizable Neural Radiance Field Reconstructionpdf
generalTinyFusion: Diffusion Transformers Learned Shallowpdf
gsplatSVG-IR: Spatially-Varying Gaussian Splatting for Inverse Renderingpdf
meshingScaling Mesh Generation via Compressive Tokenizationpdf
generalTowards Optimizing Large-Scale Multi-Graph Matching in Bioimagingpdf
generalPS-Diffusion: Photorealistic Subject-Driven Image Editing with Disentangled Control and Attentionpdf
generalExploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Modelspdf
generalFew-shot Implicit Function Generation via Equivariancepdf
generalvesselFM: A Foundation Model for Universal 3D Blood Vessel Segmentationpdf
generalClassifier-Free Guidance Inside the Attraction Basin May Cause Memorizationpdf
generalMagma: A Foundation Model for Multimodal AI Agentspdf
generalVolume Tells: Dual Cycle-Consistent Diffusion for 3D Fluorescence Microscopy De-noising and Super-Resolutionpdf
generalMatrix3D: Large Photogrammetry Model All-in-Onepdf
general3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancementpdf
generalInvestigating the Role of Weight Decay in Enhancing Nonconvex SGDpdf
generalMarkushGrapher: Joint Visual and Textual Recognition of Markush Structurespdf
generalDetecting Backdoor Attacks in Federated Learning via Direction Alignment Inspectionpdf
motionBlockDance: Reuse Structurally Similar Spatio-Temporal Features to Accelerate Diffusion Transformerspdf
generalMamba-Adaptor: State Space Model Adaptor for Visual Recognitionpdf
generalRobust Message Embedding via Attention Flow-Based Steganographypdf
generalCompositional Targeted Multi-Label Universal Perturbationspdf
generalPatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo Anomaliespdf
generalNeural Video Compression with Context Modulationpdf
generalOn-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only Eventspdf
generalLearning with Noisy Triplet Correspondence for Composed Image Retrievalpdf
generalParallelized Autoregressive Visual Generationpdf
generalCGMatch: A Different Perspective of Semi-supervised Learningpdf
generalFIction: 4D Future Interaction Prediction from Videopdf
generalD^2iT: Dynamic Diffusion Transformer for Accurate Image Generationpdf
motionAniDoc: Animation Creation Made Easierpdf
generalLiSu: A Dataset and Method for LiDAR Surface Normal Estimationpdf
motionSpk2SRImgNet: Super-Resolve Dynamic Scene from Spike Stream via Motion Aligned Collaborative Filteringpdf
generalVideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in Videospdf
generalZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Graspingpdf
generalPDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic Manipulationpdf
generalDense Dispersed Structured Light for Hyperspectral 3D Imaging of Dynamic Scenespdf
generalMM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environmentspdf
generalDora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoderspdf
generalOnce-Tuning-Multiple-Variants: Tuning Once and Expanded as Multiple Vision-Language Model Variantspdf
generalReconstructing Animals and the Wildpdf
generalDiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Drivingpdf
generalDVHGNN: Multi-Scale Dilated Vision HGNN for Efficient Vision Recognitionpdf
beingsReconstructing In-the-Wild Open-Vocabulary Human-Object Interactionspdf
generalGROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skillpdf
gsplatGauSTAR: Gaussian Surface Tracking and Reconstructionpdf
generalTraining-free Dense-Aligned Diffusion Guidance for Modular Conditional Image Synthesispdf
generalTADFormer: Task-Adaptive Dynamic TransFormer for Efficient Multi-Task Learningpdf
generalCamPoint: Boosting Point Cloud Segmentation with Virtual Camerapdf
generalMERGE: Multi-faceted Hierarchical Graph-based GNN for Gene Expression Prediction from Whole Slide Histopathology Imagespdf
generalSPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Modelspdf
generalPTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion Modelpdf
generalPreserving Clusters in Prompt Learning for Unsupervised Domain Adaptationpdf
generalAttend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Accelerationpdf
generalFlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulationpdf
generalVisual Lexicon: Rich Image Features in Language Spacepdf
generalTest-Time Visual In-Context Tuningpdf
generalPrior Does Matter: Visual Navigation via Denoising Diffusion Bridge Modelspdf
generalSegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Imagespdf
generalDo We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?pdf
generalHarnessing Global-Local Collaborative Adversarial Perturbation for Anti-Customizationpdf
generalAcc3D: Accelerating Single Image to 3D Diffusion Models via Edge Consistency Guided Score Distillationpdf
generalSoft Self-labeling and Potts Relaxations for Weakly-supervised Segmentationpdf
generalMVSAnywhere: Zero-Shot Multi-View Stereopdf
generalGenerating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Visionpdf
generalBIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literaturepdf
generalStructure-Aware Correspondence Learning for Relative Pose Estimationpdf
generalPyTorchGeoNodes: Enabling Differentiable Shape Programs for 3D Shape Reconstructionpdf
generalFIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze Estimationpdf
generalShape Abstraction via Marching Differentiable Support Functionspdf
generalScaling Down Text Encoders of Text-to-Image Diffusion Modelspdf
generalPOT: Prototypical Optimal Transport for Weakly Supervised Semantic Segmentationpdf
worldsSKE-Layout: Spatial Knowledge Enhanced Layout Generation with LLMspdf
gsplatGaussian Eigen Models for Human Headspdf
general4D-Fly: Fast 4D Reconstruction from a Single Monocular Videopdf
generalComplementary Advantages: Exploiting Cross-Field Frequency Correlation for NIR-Assisted Image Denoisingpdf
generalEval3D: Interpretable and Fine-grained Evaluation for 3D Generationpdf
generalDiffLO: Semantic-Aware LiDAR Odometry with Diffusion-Based Refinementpdf
generalStyle-Editor: Text-driven Object-centric Style Editingpdf
generalTransfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scenepdf
generalFastVLM: Efficient Vision Encoding for Vision Language Modelspdf
generalVISTA3D: A Unified Segmentation Foundation Model For 3D Medical Imagingpdf
generalS2D-LFE: Sparse-to-Dense Light Field Event Generationpdf
generalThe Art of Deception: Color Visual Illusions and Diffusion Modelspdf
meshingProgressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Datapdf
beingsDo Computer Vision Foundation Models Learn the Low-level Characteristics of the Human Visual System?pdf
generalOnline Task-Free Continual Learning via Dynamic Expansionable Memory Distributionpdf
generalRethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level Taskspdf
generalSVDC: Consistent Direct Time-of-Flight Video Depth Completion with Frequency Selective Fusionpdf
generalRethinking Training for De-biasing Text-to-Image Generation: Unlocking the Potential of Stable Diffusionpdf
generalInstant3dit: Multiview Inpainting for Fast Editing of 3D Objectspdf
generalSTDD: Spatio-Temporal Dual Diffusion for Video Generationpdf
generalImplicit Correspondence Learning for Image-to-Point Cloud Registrationpdf
generalILIAS: Instance-Level Image retrieval At Scalepdf
generalGeoDepth: From Point-to-Depth to Plane-to-Depth Modeling for Self-Supervised Monocular Depth Estimationpdf
generalSSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split Optimizationpdf
gsplatUSP-Gaussian: Unifying Spike-based Image Reconstruction, Pose Correction and Gaussian Splattingpdf
generalHolmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularitypdf
generalQuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edgepdf
generalReWind: Understanding Long Videos with Instructed Learnable Memorypdf
gsplatDirectTriGS: Triplane-based Gaussian Splatting Field Representation for 3D Generationpdf
motionMake-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Characterspdf
generalSubspace Constraint and Contribution Estimation for Heterogeneous Federated Learningpdf
worldsNeRFPrior: Learning Neural Radiance Field as a Prior for Indoor Scene Reconstructionpdf
generalTowards Training-free Anomaly Detection with Vision and Language Foundation Modelspdf
beingsDynamic Content Prediction with Motion-aware Priors for Blind Face Video Restorationpdf
generalExploring Sparse MoE in GANs for Text-conditioned Image Synthesispdf
generalEfficient Event-Based Object Detection: A Hybrid Neural Network with Spatial and Temporal Attentionpdf
generalHUNet: Homotopy Unfolding Network for Image Compressive Sensingpdf
generalSee Further When Clear: Curriculum Consistency Modelpdf
generalPassionSR: Post-Training Quantization with Adaptive Scale in One-Step Diffusion based Image Super-Resolutionpdf
gsplatRainyGS: Efficient Rain Synthesis with Physically-Based Gaussian Splattingpdf
generalThree-view Focal Length Recovery From Homographiespdf
generalRAP: Retrieval-Augmented Personalization for Multimodal Large Language Modelspdf
generalDistilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentationpdf
generalStereo4D: Learning How Things Move in 3D from Internet Stereo Videospdf
generalFoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generationpdf
generalInterDyn: Controllable Interactive Dynamics with Video Diffusion Modelspdf
generalLLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Modelspdf
generalMagicQuill: An Intelligent Interactive Image Editing Systempdf
worldsOpen-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spacespdf
generalBoosting Adversarial Transferability through Augmentation in Hypothesis Spacepdf
meshingViiNeuS: Volumetric Initialization for Implicit Neural Surface Reconstruction of Urban Scenes with Limited Image Overlappdf
generalModel Diagnosis and Correction via Linguistic and Implicit Attribute Editingpdf
generalUniPhy: Learning a Unified Constitutive Model for Inverse Physics Simulationpdf
generalSTAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural Networkspdf
generalKnowledge Memorization and Rumination for Pre-trained Model-based Class-Incremental Learningpdf
generalVideo-ColBERT: Contextualized Late Interaction for Text-to-Video Retrievalpdf
generalVisual and Semantic Prompt Collaboration for Generalized Zero-Shot Learningpdf
generalVidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modelingpdf
beingsHuman-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identificationpdf
generalExposure-slot: Exposure-centric Representations Learning with Slot-in-Slot Attention for Region-aware Exposure Correctionpdf
generalEdgeDiff: Edge-aware Diffusion Network for Building Reconstruction from Point Cloudspdf
generalDeNVeR: Deformable Neural Vessel Representations for Unsupervised Video Vessel Segmentationpdf
generalTask-Aware Clustering for Prompting Vision-Language Modelspdf
generalFSboard: Over 3 Million Characters of ASL Fingerspelling Collected via Smartphonespdf
generalLight Transport-aware Diffusion Posterior Sampling for Single-View Reconstruction of 3D Volumespdf
generalSTiL: Semi-supervised Tabular-Image Learning for Comprehensive Task-Relevant Information Exploration in Multimodal Classificationpdf
generalAuto Cherry-Picker: Learning from High-quality Generative Data Driven by Languagepdf
generalReRAW: RGB-to-RAW Image Reconstruction via Stratified Sampling for Efficient Object Detection on the Edgepdf
motionHunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animationpdf
generalZero-shot 3D Question Answering via Voxel-based Dynamic Token Compressionpdf
generalEnhanced then Progressive Fusion with View Graph for Multi-View Clusteringpdf
generalCorrecting Deviations from Normality: A Reformulated Diffusion Model for Multi-Class Unsupervised Anomaly Detectionpdf
generalContinuous 3D Perception Model with Persistent Statepdf
worldsLP-Diff: Towards Improved Restoration of Real-World Degraded License Platepdf
generalFilmComposer: LLM-Driven Music Production for Silent Film Clipspdf
generalEventPSR: Surface Normal and Reflectance Estimation from Photometric Stereo Using an Event Camerapdf
generalCASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional Videopdf
beingsGazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual Foragingpdf
generalFOCUS: Knowledge-enhanced Adaptive Visual Compression for Few-shot Whole Slide Image Classificationpdf
generalGRAE-3DMOT: Geometry Relation-Aware Encoder for Online 3D Multi-Object Trackingpdf
generalAutomatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compressionpdf
generalMAR-3D: Progressive Masked Auto-regressor for High-Resolution 3D Generationpdf
beingsSynthetic Prior for Few-Shot Drivable Head Avatar Inversionpdf
generalReasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems Approachpdf
generalDFormerv2: Geometry Self-Attention for RGBD Semantic Segmentationpdf
beingsGroomLight: Hybrid Inverse Rendering for Relightable Human Hair Appearance Modelingpdf
generalSea-ing in Low-lightpdf
generalGenerative Modeling of Class Probability for Multi-Modal Representation Learningpdf
generalVisionZip: Longer is Better but Not Necessary in Vision Language Modelspdf
generalBlenderGym: Benchmarking Foundational Model Systems for Graphics Editingpdf
generalVoteFlow: Enforcing Local Rigidity in Self-Supervised Scene Flowpdf
generalUncertainty Weighted Gradients for Model Calibrationpdf
gsplatFFaceNeRF: Few-shot Face Editing in Neural Radiance Fieldspdf
generalMinimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance Segmentationpdf
generalLayer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformerspdf
generalZero-shot RGB-D Point Cloud Registration with Pre-trained Large Vision Modelpdf
generalDistinctAD: Distinctive Audio Description Generation in Contextspdf
generalCL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answeringpdf
generalPoint Cloud Upsampling Using Conditional Diffusion Module with Adaptive Noise Suppressionpdf
generalTrajectory Mamba: Efficient Attention-Mamba Forecasting Model Based on Selective SSMpdf
beingsVLOGGER: Multimodal Diffusion for Embodied Avatar Synthesispdf
generalDEIM: DETR with Improved Matching for Fast Convergencepdf
beingsHuman Motion Instruction Tuningpdf
generalA Flag Decomposition for Hierarchical Datasetspdf
generalRCP-Bench: Benchmarking Robustness for Collaborative Perception Under Diverse Corruptionspdf
generalOlympus: A Universal Task Router for Computer Vision Taskspdf
generalCircumventing Shortcuts in Audio-visual Deepfake Detection Datasets with Unsupervised Learningpdf
generalImage Over Text: Transforming Formula Recognition Evaluation with Character Detection Matchingpdf
generalImproving Semi-Supervised Semantic Segmentation with Sliced-Wasserstein Feature Alignment and Uniformitypdf
general3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoningpdf
worldsNavigation World Modelspdf
generalConformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generationpdf
beingsReconstructing Close Human Interaction with Appearance and Proxemics Reasoningpdf
generalScenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environmentspdf
generalPoly-Autoregressive Prediction for Modeling Interactionspdf
beingsPoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimationpdf
generalDecision SpikeFormer: Spike-Driven Transformer for Decision Makingpdf
generalTheory-Inspired Deep Multi-View Multi-Label Learning with Incomplete Views and Noisy Labelspdf
generalEMOE: Modality-Specific Enhanced Dynamic Emotion Expertspdf
generalGenerative Video Propagationpdf
generalFrom Multimodal LLMs to Generalist Embodied Agents: Methods and Lessonspdf
generalMosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentationpdf
generalT-CIL: Temperature Scaling using Adversarial Perturbation for Calibration in Class-Incremental Learningpdf
generalLoRA Subtraction for Drift-Resistant Space in Exemplar-Free Continual Learningpdf
generalAniMer: Animal Pose and Shape Estimation Using Family Aware Transformerpdf
generalCo-op: Correspondence-based Novel Object Pose Estimationpdf
generalCATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-Resolutionpdf
generalRayFlow: Instance-Aware Diffusion Acceleration via Adaptive Flow Trajectoriespdf
generalSelf-Supervised Large Scale Point Cloud Completion for Archaeological Site Restorationpdf
generalChain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attackspdf
generalRate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty Estimationpdf
generalThin-Shell-SfT: Fine-Grained Monocular Non-rigid 3D Surface Tracking with Neural Deformation Fieldspdf
generalDeCLIP: Decoupled Learning for Open-Vocabulary Dense Perceptionpdf
generalSocialGesture: Delving into Multi-person Gesture Understandingpdf
generalMulti-modal Topology-embedded Graph Learning for Spatially Resolved Genes Prediction from Pathology Images with Prior Gene Similarity Informationpdf
gsplatQuestion-Aware Gaussian Experts for Audio-Visual Question Answeringpdf
generalAdaptive Rectangular Convolution for Remote Sensing Pansharpeningpdf
generalUIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion Modelspdf
generalFisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentationpdf
generalAdMiT: Adaptive Multi-Source Tuning in Dynamic Environmentspdf
generalUnbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasingpdf
generalUnleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulationpdf
generalRevisiting Generative Replay for Class Incremental Object Detectionpdf
generalBridging Viewpoint Gaps: Geometric Reasoning Boosts Semantic Correspondencepdf
generalSpatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesispdf
beingsMobileH2R: Learning Generalizable Human to Mobile Robot Handover Exclusively from Scalable and Diverse Synthetic Datapdf
beingsPERSE: Personalized 3D Generative Avatars from A Single Portraitpdf
generalDynamic Stereotype Theory Induced Micro-expression Recognition with Oriented Deformationpdf
generalACL: Activating Capability of Linear Attention for Image Restorationpdf
generalMARBLE: Material Recomposition and Blending in CLIP-Spacepdf
generalEfficient Visual State Space Model for Image Deblurringpdf
generalEnhancing 3D Gaze Estimation in the Wild using Weak Supervision with Gaze Following Labelspdf
generalReward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Rewardpdf
generalDetecting Out-of-Distribution Through the Lens of Neural Collapsepdf
gsplatSparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse Viewspdf
generalVinTAGe: Joint Video and Text Conditioning for Holistic Audio Generationpdf
gsplatEfficient Decoupled Feature 3D Gaussian Splatting via Hierarchical Compressionpdf
generalCountLLM: Towards Generalizable Repetitive Action Counting via Large Language Modelpdf
generalSPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Imagespdf
generalFocus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generationpdf
generalLabel Shift Meets Online Learning: Ensuring Consistent Adaptation with Universal Dynamic Regretpdf
generalA Physics-Informed Blur Learning Framework for Imaging Systemspdf
generalTowards Practical Real-Time Neural Video Compressionpdf
gsplatDepthSplat: Connecting Gaussian Splatting and Depthpdf
generalDynamic Camera Poses and Where to Find Thempdf
generalOmniGen: Unified Image Generationpdf
generalQuCOOP: A Versatile Framework for Solving Composite and Binary-Parametrised Problems on Quantum Annealerspdf
meshingMesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured Meshespdf
generalSILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generationpdf
generalCalibrated Multi-Preference Optimization for Aligning Diffusion Modelspdf
gsplatAdvancing Adversarial Robustness in GNeRFs: The IL2-NeRF Attackpdf
generalPolarNeXt: Rethink Instance Segmentation with Polar Representationpdf
generalSAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything Modelpdf
generalDarkIR: Robust Low-Light Image Restorationpdf
generalR2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action Plannerpdf
generalFrom Prototypes to General Distributions: An Efficient Curriculum for Masked Image Modelingpdf
generalDifference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generationpdf
generalMTADiffusion: Mask Text Alignment Diffusion Model for Object Inpaintingpdf
generalGrounding 3D Object Affordance with Language Instructions, Visual Observations and Interactionspdf
generalImage is All You Need to Empower Large-scale Diffusion Models for In-Domain Generationpdf
generalEvolving High-Quality Rendering and Reconstruction in a Unified Framework with Contribution-Adaptive Regularizationpdf
generalNoiseCtrl: A Sampling-Algorithm-Agnostic Conditional Generation Method for Diffusion Modelspdf
generalKMD: Koopman Multi-modality Decomposition for Generalized Brain Tumor Segmentation under Incomplete Modalitiespdf
generalDORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-Resolutionpdf
generalFractal Calibration for Long-tailed Object Detectionpdf
generalM3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world Settingspdf
generalFedSPA: Generalizable Federated Graph Learning under Homophily Heterogeneitypdf
generalGazeGene: Large-scale Synthetic Gaze Dataset with 3D Eyeball Annotationspdf
generalVideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priorspdf
gsplatGaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understandingpdf
generalContinuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directionspdf
generalSimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignmentpdf
generalImproved Video VAE for Latent Video Diffusion Modelpdf
generalEfficient Video Super-Resolution for Real-time Rendering with Decoupled G-buffer Guidancepdf
generalLearned Image Compression with Dictionary-based Entropy Modelpdf
generalFireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Modelpdf
generalDL2G: Degradation-guided Local-to-Global Restoration for Eyeglass Reflection Removalpdf
generalMFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and Forecastingpdf
generalThe Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion Modelspdf
generalLeveraging Global Stereo Consistency for Category-Level Shape and 6D Pose Estimation from Stereo Imagespdf
generalAlphaPre: Amplitude-Phase Disentanglement Model for Precipitation Nowcastingpdf
generalDetection-Friendly Nonuniformity Correction: A Union Framework for Infrared UAV Target Detectionpdf
generalArticulated Kinematics Distillation from Video Diffusion Modelspdf
generalExpertAF: Expert Actionable Feedback from Videopdf
gsplatVolumetrically Consistent 3D Gaussian Rasterizationpdf
generalThe Impact Label Noise and Choice of Threshold has on Cross-Entropy and Soft-Dice in Image Segmentationpdf
generalLLaVA-Critic: Learning to Evaluate Multimodal Modelspdf
generalVILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledgepdf
generalRepurposing Pre-trained Video Diffusion Models for Event-based Video Interpolationpdf
generalLarge-scale Multi-view Tensor Clustering with Implicit Linear Kernelspdf
generalQ-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Contentpdf
generalDual Focus-Attention Transformer for Robust Point Cloud Registrationpdf
generalForming Auxiliary High-confident Instance-level Loss to Promote Learning from Label Proportionspdf
generalProgress-Aware Video Frame Captioningpdf
generalSMTPD: A New Benchmark for Temporal Prediction of Social Media Popularitypdf
generalLearning on Model Weights using Tree Expertspdf
generalImage Reconstruction from Readout-Multiplexed Single-Photon Detector Arrayspdf
generalTowards Transformer-Based Aligned Generation with Self-Coherence Guidancepdf
generalAccurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-Distillationpdf
generalDART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report Generationpdf
generalOn the Consistency of Video Large Language Models in Temporal Comprehensionpdf
generalLess is More: Efficient Model Merging with Binary Task Switchpdf
generalOne-Minute Video Generation with Test-Time Trainingpdf
generalInteractionMap: Improving Online Vectorized HDMap Construction with Interactionpdf
worldsROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Promptingpdf
generalRLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthinesspdf
gsplatEditSplat: Multi-View Fusion and Attention-Guided Optimization for View-Consistent 3D Scene Editing with 3D Gaussian Splattingpdf
generalOne-shot 3D Object Canonicalization based on Geometric and Semantic Consistencypdf
generalInst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuningpdf
generalProHOC: Probabilistic Hierarchical Out-of-Distribution Classification via Multi-Depth Networkspdf
generalCLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIPpdf
generalGraph-Embedded Structure-Aware Perceptual Hashing for Neural Network Protection and Piracy Detectionpdf
generalInterleaved-Modal Chain-of-Thoughtpdf
generalEnhancing Adversarial Transferability with Checkpoints of a Single Model's Trainingpdf
generalO-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Modelspdf
generalAnalyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose Estimationpdf
gsplatFeature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fieldspdf
generalHyperspectral Pansharpening via Diffusion Models with Iteratively Zero-Shot Guidancepdf
generalEASEMVC:Efficient Dual Selection Mechanism for Deep Multi-View Clusteringpdf
generalDSPNet: Dual-vision Scene Perception for Robust 3D Question Answeringpdf
generalIceDiff: High Resolution and High-Quality Arctic Sea Ice Forecasting with Generative Diffusion Priorpdf
generalDTOS: Dynamic Time Object Sensing with Large Multimodal Modelpdf
generalHow to Merge Your Multimodal Models Over Time?pdf
generalIdentifying and Mitigating Position Bias of Multi-image Vision-Language Modelspdf
generalExploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentationpdf
generalShowUI: One Vision-Language-Action Model for GUI Visual Agentpdf
generalInfinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesispdf
beingsHumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generationpdf
generalReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videospdf
generalArtiFade: Learning to Generate High-quality Subject from Blemished Imagespdf
generalPrompting Depth Anything for 4K Resolution Accurate Metric Depth Estimationpdf
generalGET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discoverypdf
generalTest-Time Domain Generalization via Universe Learning: A Multi-Graph Matching Approach for Medical Image Segmentationpdf
generalDeDe: Detecting Backdoor Samples for SSL Encoders via Decoderspdf
beingsTowards Scalable Human-aligned Benchmark for Text-guided Image Editingpdf
generalCoeff-Tuning: A Graph Filter Subspace View for Tuning Attention-Based Large Modelspdf
generalSelf-Supervised Cross-View Correspondence with Predictive Cycle Consistencypdf
beingsCryptoFace: End-to-End Encrypted Face Recognitionpdf
generalRelation-Rich Visual Document Generator for Visual Information Extractionpdf
generalPromptHash:Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrievalpdf
generalUniversal Scene Graph Generationpdf
generalSplit Adaptation for Pre-trained Vision Transformerspdf
generalSpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Modelspdf
generalLearning Occlusion-Robust Vision Transformers for Real-Time UAV Trackingpdf
generalPlug-and-Play Versatile Compressed Video Enhancementpdf
generalUltraFusion: Ultra High Dynamic Imaging using Exposure Fusionpdf
generalNoise-Resistant Video Anomaly Detection via RGB Error-Guided Multiscale Predictive Coding and Dynamic Memorypdf
generalGroupMamba: Efficient Group-Based Visual State Space Modelpdf
generalEscaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spacespdf
generalActiveGAMER: Active GAussian Mapping through Efficient Renderingpdf
generalPositive2Negative: Breaking the Information-Lossy Barrier in Self-Supervised Single Image Denoisingpdf
generalSeaLion: Semantic Part-Aware Latent Point Diffusion Models for 3D Generationpdf
gsplatToward Real-world BEV Perception: Depth Uncertainty Estimation via Gaussian Splattingpdf
generalEfficient Fine-Tuning and Concept Suppression for Pruned Diffusion Modelspdf
gsplatWildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environmentspdf
beingsRePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformancepdf
worldsCheXWorld: Exploring Image World Modeling for Radiograph Representation Learningpdf
generalTowards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Methodpdf
generalAIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMMpdf
generalAutoregressive Distillation of Diffusion Transformerspdf
generalOmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraintspdf
generalDiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrievalpdf
generalVisual-Instructed Degradation Diffusion for All-in-One Image Restorationpdf
generalInsightful Instance Features for 3D Instance Segmentationpdf
generalEmoDubber: Towards High Quality and Emotion Controllable Movie Dubbingpdf
generalA Hubness Perspective on Representation Learning for Graph-Based Multi-View Clusteringpdf
generalSpatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulationpdf
generalZeroVO: Visual Odometry with Minimal Assumptionspdf
generalVideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLMpdf
generalHoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embeddingpdf
generalGen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objectspdf
worldsSkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modelingpdf
generalAdaDARE-gamma: Balancing Stability and Plasticity in Multi-modal LLMs through Efficient Adaptationpdf
generalLeviTor: 3D Trajectory Oriented Image-to-Video Synthesispdf
beingsSapiensID: Foundation for Human Recognitionpdf
motionExtrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Thinkpdf
beingsFreeCloth: Free-form Generation Enhances Challenging Clothed Human Modelingpdf
gsplatInstanceGaussian: Appearance-Semantic Joint Gaussian Representation for 3D Instance-Level Perceptionpdf
generalCSC-PA: Cross-image Semantic Correlation via Prototype Attentions for Single-network Semi-supervised Breast Tumor Segmentationpdf
generalVisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Drivingpdf
generalDetecting Adversarial Data Using Perturbation Forgerypdf
generalCoA: Towards Real Image Dehazing via Compression-and-Adaptationpdf
generalTopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Modelpdf
generalLearned Binocular-Encoding Optics for RGBD Imaging Using Joint Stereo and Focus Cuespdf
beingsMobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devicespdf
motionLight3R-SfM: Towards Feed-forward Structure-from-Motionpdf
generalRobotic Visual Instructionpdf
generalMASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priorspdf
generalViewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learningpdf
generalCross-modal Information Flow in Multimodal Large Language Modelspdf
generalKeyframe-Guided Creative Video Inpaintingpdf
generalEdgeTAM: On-Device Track Anything Modelpdf
generalEarthDial: Turning Multi-sensory Earth Observations to Interactive Dialoguespdf
generalVideo Summarization with Large Language Modelspdf
generalSketchtopia: A Dataset and Foundational Agents for Benchmarking Asynchronous Multimodal Communication with Iconic Feedbackpdf
generalConsistency-aware Self-Training for Iterative-based Stereo Matchingpdf
generalMV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contextspdf
generalGeneralized Few-shot 3D Point Cloud Segmentation with Vision-Language Modelpdf

CVPR 2025 - Day 3 (2025-06-15)

familypaperlinks
generalDeterministic Image-to-Image Translation via Denoising Brownian Bridge Models with Dual Approximatorspdf
generalTask Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignmentpdf
generalCross-modal Causal Relation Alignment for Video Question Groundingpdf
generalDiffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion Modelspdf
gsplatOmni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstructionpdf
general3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusionpdf
worldsMissing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrievalpdf
generalDiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generationpdf
generalNarrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captionspdf
generalCARL: A Framework for Equivariant Image Registrationpdf
gsplatFlashGS: Efficient 3D Gaussian Splatting for Large-scale and High-resolution Renderingpdf
generalChat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Modelspdf
generalInference-Scale Complexity in ANN-SNN Conversion for High-Performance and Low-Power Applicationspdf
beingsMVDoppler-Pose: Multi-Modal Multi-View mmWave Sensing for Long-Distance Self-Occluded Human Walking Pose Estimationpdf
generalTopNet: Transformer-Efficient Occupancy Prediction Network for Octree-Structured Point Cloud Geometry Compressionpdf
generalGain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring Classespdf
generalM^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentationpdf
generalEverything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignmentpdf
generalMulti-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain Adaptationpdf
motionA Polarization-Aided Transformer for Image Deblurring via Motion Vector Decompositionpdf
generalCocoER: Aligning Multi-Level Feature by Competition and Coordination for Emotion Recognitionpdf
generalEnhancing Creative Generation on Stable Diffusion-based Modelspdf
generalDenoising Functional Maps: Diffusion Models for Shape Correspondencepdf
generalProReflow: Progressive Reflow with Decomposed Velocitypdf
generalDevil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attentionpdf
generalMetaShadow: Object-Centered Shadow Detection, Removal, and Synthesispdf
generalTANGO: Training-free Embodied AI Agents for Open-world Taskspdf
generalStealthy Backdoor Attack in Self-Supervised Learning Vision Encoders for Large Vision Language Modelspdf
generalSAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenespdf
generalGIVEPose: Gradual Intra-class Variation Elimination for RGB-based Category-Level Object Pose Estimationpdf
beingsSketch Down the FLOPs: Towards Efficient Networks for Human Sketchpdf
generalRethinking Decoder Design: Improving Biomarker Segmentation Using Depth-to-Space Restoration and Residual Linear Attentionpdf
generalSynTab-LLaVA: Enhancing Multimodal Table Understanding with Decoupled Synthesispdf
beingsEdit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative Editingpdf
generalImproving Accuracy and Calibration via Differentiated Deep Mutual Learningpdf
generalInfighting in the Dark: Multi-Label Backdoor Attack in Federated Learningpdf
generalTartan IMU: A Light Foundation Model for Inertial Positioning in Roboticspdf
generalEvent Ellipsometer: Event-based Mueller-Matrix Video Imagingpdf
generalEnd-to-End HOI Reconstruction Transformer with Graph-based Encodingpdf
beingsDisco4D: Disentangled 4D Human Generation and Animation from a Single Imagepdf
beingsIDOL: Instant Photorealistic 3D Human Creation from a Single Imagepdf
generalSketchVideo: Sketch-based Video Generation and Editingpdf
generalTaste More, Taste Better: Diverse Data and Strong Model Boost Semi-Supervised Crowd Countingpdf
beingsAnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Modelspdf
generalLatent Space Imagingpdf
generalBalanced Direction from Multifarious Choices: Arithmetic Meta-Learning for Domain Generalizationpdf
generalAnatomical Consistency and Adaptive Prior-informed Transformation for Multi-contrast MR Image Synthesis via Diffusion Modelpdf
generalSeCap: Self-Calibrating and Adaptive Prompts for Cross-view Person Re-Identification in Aerial-Ground Networkspdf
generalDon't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Drivingpdf
worldsNeural Motion Simulator Pushing the Limit of World Models in Reinforcement Learningpdf
worldsAdversarial Diffusion Compression for Real-World Image Super-Resolutionpdf
generalDiSciPLE: Learning Interpretable Programs for Scientific Visual Discoverypdf
generalSOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characterspdf
generalEntropyMark: Towards More Harmless Backdoor Watermark via Entropy-based Constraint for Open-source Dataset Copyright Protectionpdf
generalAdaptive Markup Language Generation for Contextually-Grounded Visual Document Understandingpdf
generalTowards Universal AI-Generated Image Detection by Variational Information Bottleneck Networkpdf
generalHSI: A Holistic Style Injector for Arbitrary Style Transferpdf
generalV2V3D: View-to-View Denoised 3D Reconstruction for Light Field Microscopypdf
gsplatSplatter-360: Generalizable 360 Gaussian Splatting for Wide-baseline Panoramic Imagespdf
generalTowards Understanding How Knowledge Evolves in Large Vision-Language Modelspdf
generalA Unified, Resilient, and Explainable Adversarial Patch Detectorpdf
generalStructured 3D Latents for Scalable and Versatile 3D Generationpdf
generalSelf-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjectspdf
generalAdv-CPG: A Customized Portrait Generation Framework with Facial Adversarial Attackspdf
generalFish-Vista: A Multi-Purpose Dataset for Understanding & Identification of Traits from Imagespdf
generalPCM : Picard Consistency Model for Fast Parallel Sampling of Diffusion Modelspdf
gsplatCoMapGS: Covisibility Map-based Gaussian Splatting for Sparse Novel View Synthesispdf
generalTraining Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training?pdf
generalImproving the Training of Data-Efficient GANs via Quality Aware Dynamic Discriminator Rejection Samplingpdf
motionMotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generationpdf
generalAdvancing Generalizable Tumor Segmentation with Anomaly-Aware Open-Vocabulary Attention Maps and Frozen Foundation Diffusion Modelspdf
generalTowards Generalizable Scene Change Detectionpdf
generalIncomplete Multi-modal Brain Tumor Segmentation via Learnable Sorting State Space Modelpdf
generalFedAWA: Adaptive Optimization of Aggregation Weights in Federated Learning Using Client Vectorspdf
beingsRethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregressionpdf
generalStretching Each Dollar: Diffusion Training from Scratch on a Micro-Budgetpdf
beingsGuiding Human-Object Interactions with Rich Geometry and Relationspdf
generalCADDreamer: CAD Object Generation from Single-view Imagespdf
generalWhere's the Liability in the Generative Era? Recovery-based Black-Box Detection of AI-Generated Contentpdf
generalDiTASK: Multi-Task Fine-Tuning with Diffeomorphic Transformationspdf
worldsOW-OVD: Unified Open World and Open Vocabulary Object Detectionpdf
generalImproving Diffusion Inverse Problem Solving with Decoupled Noise Annealingpdf
generalDesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Modelspdf
generalSAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAEpdf
generalDual-Interrelated Diffusion Model for Few-Shot Anomaly Image Generationpdf
generalInteractive Medical Image Analysis with Concept-based Similarity Reasoningpdf
generalh-Edit: Effective and Flexible Diffusion-Based Editing via Doob's h-Transformpdf
beingsAre Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?pdf
generalSpectral State Space Model for Rotation-Invariant Visual Representation Learningpdf
generalSharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and Manipulationpdf
generalURWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image Restorationpdf
generalFunctionality Understanding and Segmentation in 3D Scenespdf
generalDragin3D: Image Editing by Dragging in 3D Spacepdf
generalTowards Stable and Storage-efficient Dataset Distillation: Matching Convexified Trajectorypdf
generalTSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual Segmentationpdf
generalInvisible Backdoor Attack against Self-supervised Learningpdf
meshingPerceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation Metricspdf
generalBWFormer: Building Wireframe Reconstruction from Airborne LiDAR Point Cloud with Transformerpdf
generalDiffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Modelspdf
generalOmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoningpdf
gsplatMeGA: Hybrid Mesh-Gaussian Head Avatar for High-Fidelity Rendering and Head Editingpdf
generalComprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformerspdf
generalDataset Distillation with Neural Characteristic Function: A Minmax Perspectivepdf
beingsFree-viewpoint Human Animation with Pose-correlated Reference Selectionpdf
generalPillarHist: A Quantization-aware Pillar Feature Encoder based on Height-aware Histogrampdf
generalSemantic and Expressive Variations in Image Captions Across Languagespdf
generalATP-LLaVA: Adaptive Token Pruning for Large Vision Language Modelspdf
generalADD: Attribution-Driven Data Augmentation Framework for Boosting Image Super-Resolutionpdf
generalCroCoDL: Cross-device Collaborative Dataset for Localizationpdf
generalCLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCRpdf
generalWhat Makes a Good Dataset for Knowledge Distillation?pdf
generalRectification-specific Supervision and Constrained Estimator for Online Stereo Rectificationpdf
generalShape and Texture: What Influences Reliable Optical Flow Estimation?pdf
generalPrecise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matterspdf
beingsHOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generationpdf
generalOrder-One Rolling Shutter Cameraspdf
generalAnimate and Sound an Imagepdf
generalFoveated Instance Segmentationpdf
generalEmphasizing Discriminative Features for Dataset Distillation in Complex Scenariospdf
generalSegment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentationpdf
generalTask-Specific Gradient Adaptation for Few-Shot One-Class Classificationpdf
gsplat3D Gaussian Inpainting with Depth-Guided Cross-View Consistencypdf
generalFloxels: Fast Unsupervised Voxel Based Scene Flow Estimationpdf
generalLiveCC: Learning Video LLM with Streaming Speech Transcription at Scalepdf
generalFlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Computepdf
generalHyperGLM: HyperGraph for Video Scene Graph Generation and Anticipationpdf
beingsFSFM: A Generalizable Face Security Foundation Model via Self-Supervised Facial Representation Learningpdf
generalAlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignmentpdf
generalVideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Modelspdf
generalOne Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusionpdf
generalCan Text-to-Video Generation help Video-Language Alignment?pdf
generalWeakly Supervised Contrastive Adversarial Training for Learning Robust Features from Semi-supervised Datapdf
generalFrom Poses to Identity: Training-Free Person Re-Identification via Feature Centralizationpdf
generalMIMO: A Medical Vision Language Model with Visual Referring Multimodal Input and Pixel Grounding Multimodal Outputpdf
generalBias for Action: Video Implicit Neural Representations with Bias Modulationpdf
generalSegment Anything, Even Occludedpdf
generalLOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot Learningpdf
generalUniversal Actions for Enhanced Embodied Foundation Modelspdf
generalFaithDiff: Unleashing Diffusion Priors for Faithful Image Super-resolutionpdf
generalScene-agnostic Pose Regression for Visual Localizationpdf
generalDivide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial Purificationpdf
generalSEC-Prompt:SEmantic Complementary Prompting for Few-Shot Class-Incremental Learningpdf
generalLiMoE: Mixture of LiDAR Representation Learners from Automotive Scenespdf
beingsPI-HMR: Towards Robust In-bed Temporal Human Shape Reconstruction with Contact Pressure Sensingpdf
generalCheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulationpdf
generalSEEN-DA: SEmantic ENtropy guided Domain-aware Attention for Domain Adaptive Object Detectionpdf
generalBlind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion Modelpdf
generalMind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreakingpdf
worldsGEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Controlpdf
generalScene-Centric Unsupervised Panoptic Segmentationpdf
generalLearning Physics From Video: Unsupervised Physical Parameter Estimation for Continuous Dynamical Systemspdf
generalProAPO: Progressively Automatic Prompt Optimization for Visual Classificationpdf
generalBlack Swan: Abductive and Defeasible Video Reasoning in Unpredictable Eventspdf
gsplatRNG: Relightable Neural Gaussianspdf
gsplatTowards Realistic Example-based Modeling via 3D Gaussian Stitchingpdf
gsplatGenerative Sparse-View Gaussian Splattingpdf
generalGenerative Inbetweening through Frame-wise Conditions-Driven Video Generationpdf
generalDexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awarenesspdf
generalCustAny: Customizing Anything from A Single Examplepdf
generalPoseTraj: Pose-Aware Trajectory Control in Video Diffusionpdf
generalVL2Lite: Task-Specific Knowledge Distillation from Large Vision-Language Models to Lightweight Networkspdf
generalStageDesigner: Artistic Stage Generation for Scenography via Theater Scriptspdf
generalInterpreting Object-level Foundation Models via Visual Precision Searchpdf
generalFoley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flowspdf
generalAll-directional Disparity Estimation for Real-world QPD Imagespdf
generalUsing Diffusion Priors for Video Amodal Segmentationpdf
motionDyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camerapdf
generalThe Scene Language: Representing Scenes with Programs, Words, and Embeddingspdf
beingsLearning Physics-Based Full-Body Human Reaching and Grasping from Brief Walking Referencespdf
generalEmoEdit: Evoking Emotions through Image Manipulationpdf
generalSparseAlign: a Fully Sparse Framework for Cooperative Object Detectionpdf
generalData Distributional Properties As Inductive Bias for Systematic Generalizationpdf
generalTopoCellGen: Generating Histopathology Cell Topology with a Diffusion Modelpdf
generalMeta-Learning Hyperparameters for Parameter Efficient Fine-Tuningpdf
meshingTriTex: Learning Texture from a Single Mesh via Triplane Semantic Featurespdf
generalWavelet and Prototype Augmented Query-based Transformer for Pixel-level Surface Defect Detectionpdf
generalAlignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answeringpdf
generalLanguage-Guided Audio-Visual Learning for Long-Term Sports Assessmentpdf
motionTeller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generationpdf
beingsPersonaHOI: Effortlessly Improving Face Personalization in Human-Object Interaction Generationpdf
generalVideo-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Modelspdf
gsplatMonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Modelspdf
generalHybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentationpdf
generalProbability Density Geodesics in Image Diffusion Latent Spacepdf
generalEgoLife: Towards Egocentric Life Assistantpdf
generalBrepGiff: Lightweight Generation of Complex B-rep with 3D GAT Diffusionpdf
generalTowards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency Partitionpdf
generalJoint Scheduling of Causal Prompts and Tasks for Multi-Task Learningpdf
gsplat3DGUT: Enabling Distorted Cameras and Secondary Rays in Gaussian Splattingpdf
generalIt's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Datapdf
generalOpen Set Label Shift with Test Time Out-of-Distribution Referencepdf
gsplatGaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Predictionpdf
generalFlexible Frame Selection for Efficient Video Reasoningpdf
generalEventGPT: Event Stream Understanding with Multimodal Large Language Modelspdf
generalMITracker: Multi-View Integration for Visual Object Trackingpdf
generalNot Only Text: Exploring Compositionality of Visual Representations in Vision-Language Modelspdf
generalExploring Scene Affinity for Semi-Supervised LiDAR Semantic Segmentationpdf
generalMinority-Focused Text-to-Image Generation via Prompt Optimizationpdf
generalMANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny Objectspdf
generalSCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in Structurespdf
generalCollaborative Tree Search for Enhancing Embodied Multi-Agent Collaborationpdf
generalText-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual Abductionpdf
generalAdapting Text-to-Image Generation with Feature Difference Instruction for Generic Image Restorationpdf
generalReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence Learningpdf
generalPreconditioners for the Stochastic Training of Neural Fieldspdf
worldsECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmarkpdf
gsplatSfM-Free 3D Gaussian Splatting via Hierarchical Trainingpdf
generalCASAGPT: Cuboid Arrangement and Scene Assembly for Interior Designpdf
generalMINIMA: Modality Invariant Image Matchingpdf
gsplat3D Convex Splatting: Radiance Field Rendering with 3D Smooth Convexespdf
general3D Prior Is All You Need: Cross-Task Few-shot 2D Gaze Estimationpdf
generalSeriesBench: A Benchmark for Narrative-Driven Drama Series Understandingpdf
generalWeakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Modelspdf
generalGliaNet: Adaptive Neural Network Structure Learning with Glia-Drivenpdf
generalEntitySAM: Segment Everything in Videopdf
generalGS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstructionpdf
generalVideo Depth Anything: Consistent Depth Estimation for Super-Long Videospdf
generalInstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Captionpdf
gsplatLuminance-GS: Adapting 3D Gaussian Splatting to Challenging Lighting Conditions with View-Adaptive Curve Adjustmentpdf
gsplatEventSplat: 3D Gaussian Splatting from Moving Event Cameras for Real-time Renderingpdf
gsplat3D Student Splatting and Scoopingpdf
worldsWorld-consistent Video Diffusion with Explicit 3D Modelingpdf
generalLearning Partonomic 3D Reconstruction from Image Collectionspdf
generalODA-GAN: Orthogonal Decoupling Alignment GAN Assisted by Weakly-supervised Learning for Virtual Immunohistochemistry Stainingpdf
generalEVOS: Efficient Implicit Neural Training via EVOlutionary Selectorpdf
generalMEET: Towards Memory-Efficient Temporal Sparse Deep Neural Networkspdf
generalProbabilistic Prompt Distribution Learning for Animal Pose Estimationpdf
generalMitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attentionpdf
generalUniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplinespdf
gsplatMani-GS: Gaussian Splatting Manipulation with Triangular Meshpdf
generalBooW-VTON: Boosting In-the-Wild Virtual Try-On via Mask-Free Pseudo Data Trainingpdf
generalSupervising Sound Localization by In-the-wild Egomotionpdf
generalAutoLUT: LUT-Based Image Super-Resolution with Automatic Sampling and Adaptive Residual Learningpdf
gsplatAniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian Reconstructionpdf
generalIM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular VideosCpdf
gsplatDynaMoDe-NeRF: Motion-aware Deblurring Neural Radiance Field for Dynamic Scenespdf
generalUrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene Simulationpdf
generalDiff-Palm: Realistic Palmprint Generation with Polynomial Creases and Intra-Class Variation Controllable Diffusion Modelspdf
gsplatChain of Semantics Programming in 3D Gaussian Splatting Representation for 3D Vision Groundingpdf
motionMVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait Animationpdf
generalAccelerating Multimodal Large Language Models by Searching Optimal Vision Token Reductionpdf
generalMatrix-Free Shared Intrinsics Bundle Adjustmentpdf
generalUncertainty-Instructed Structure Injection for Generalizable HD Map Constructionpdf
generalColor Alignment in Diffusionpdf
generalLLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Livingpdf
generalLanguage-Guided Salient Object Rankingpdf
generalTowards More General Video-based Deepfake Detection through Facial Component Guided Adaptation for Foundation Modelpdf
generalSP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic Promptspdf
generalVoCo-LLaMA: Towards Vision Compression with Large Language Modelspdf
generalFocal Split: Untethered Snapshot Depth from Differential Defocuspdf
generalPURA: Parameter Update-Recovery Test-Time Adaption for RGB-T Trackingpdf
generalTowards All-in-One Medical Image Re-Identificationpdf
generalIntegral Fast Fourier Color Constancypdf
generalResCLIP: Residual Attention for Training-free Dense Vision-language Inferencepdf
generalDispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reactionpdf
generalBayesian Test-Time Adaptation for Vision-Language Modelspdf
generalCausal Composition Diffusion Model for Closed-loop Traffic Generationpdf
generalChange3D: Revisiting Change Detection and Captioning from A Video Modeling Perspectivepdf
generalAttribute-formed Class-specific Concept Space: Endowing Language Bottleneck Model with Better Interpretability and Scalabilitypdf
generalCustomized Condition Controllable Generation for Video Soundtrackpdf
beingsProjAttacker: A Configurable Physical Adversarial Attack for Face Recognition via Projectorpdf
generalWISE: A Framework for Gigapixel Whole-Slide-Image Lossless Compressionpdf
generalGromov-Wasserstein Problem with Cyclic Symmetrypdf
beingsSimAvatar: Simulation-Ready Avatars with Layered Hair and Clothingpdf
generalTest-Time Backdoor Detection for Object Detection Modelspdf
generalSDBF: Steep-Decision-Boundary Fingerprinting for Hard-Label Tampering Detection of DNN Modelspdf
generalDistilling Multi-modal Large Language Models for Autonomous Drivingpdf
generalHD-EPIC: A Highly-Detailed Egocentric Video Datasetpdf
generalAdvancing Myopia To Holism: Fully Contrastive Language-Image Pre-trainingpdf
beingsH-MoRe: Learning Human-centric Motion Representation for Action Analysispdf
generalHierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric Learningpdf
generalEffortless Active Labeling for Long-Term Test-Time Adaptationpdf
generalLeveraging Temporal Cues for Semi-Supervised Multi-View 3D Object Detectionpdf
generalLogits DeConfusion with CLIP for Few-Shot Learningpdf
generalPay Attention to the Foreground in Object-Centric Learningpdf
generalFluidNexus: 3D Fluid Reconstruction and Prediction from a Single Videopdf
generalDeformCL: Learning Deformable Centerline Representation for Vessel Extraction in 3D Medical Imagepdf
worldsOCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation Triadpdf
generalSPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstructionpdf
beingsVidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulationpdf
generalWhich Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videospdf
generalAdaptive Keyframe Sampling for Long Video Understandingpdf
generalPerson De-reidentification: A Variation-guided Identity Shift Modelingpdf
generalDiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformerpdf
generalFlorence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusionpdf
generalRealistic Test-Time Adaptation of Vision-Language Modelspdf
gsplatSelfSplat: Pose-Free and 3D Prior-Free Generalizable 3D Gaussian Splattingpdf
generalEnhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise Schedulingpdf
generalExploring Simple Open-Vocabulary Semantic Segmentationpdf
generalMP-GUI: Modality Perception with MLLMs for GUI Understandingpdf
generalImproving Adversarial Transferability on Vision Transformers via Forward Propagation Refinementpdf
generalSeeing What Matters: Empowering CLIP with Patch Generation-to-Selectionpdf
generalErasing Undesirable Influence in Diffusion Modelspdf
generalClosest Neighbors are Harmful for Lightweight Masked Auto-encoderspdf
generalDecouple-Then-Merge: Finetune Diffusion Models as Multi-Task Learningpdf
worldsHELVIPAD: A Real-World Dataset for Omnidirectional Stereo Depth Estimationpdf
generalTowards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistencypdf
generalPractical Solutions to the Relative Pose of Three Calibrated Cameraspdf
generalPARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Modelspdf
generalRoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twinspdf
motionAnimateAnything: Consistent and Controllable Animation for Video Generationpdf
generalPRaDA: Projective Radial Distortion Averagingpdf
generalGenAssets: Generating in-the-wild 3D Assets in Latent Spacepdf
generalLow-Rank Adaptation in Multilinear Operator Networks for Security-Preserving Incremental Learningpdf
generalFiRe: Fixed-points of Restoration Priors for Solving Inverse Problemspdf
generalCoMBO: Conflict Mitigation via Branched Optimization for Class Incremental Segmentationpdf
generalRecurrent Feature Mining and Keypoint Mixup Padding for Category-Agnostic Pose Estimationpdf
generalUnbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarkspdf
generalEmbodied Scene Understanding for Vision Language Models via MetaVQApdf
generalLearning Temporally Consistent Video Depth from Video Diffusion Priorspdf
generalSamba: A Unified Mamba-based Framework for General Salient Object Detectionpdf
beingsLesionLocator: Zero-Shot Universal Tumor Segmentation and Tracking in 3D Whole-Body Imagingpdf
gsplatDOF-GS: Adjustable Depth-of-Field 3D Gaussian Splatting for Post-Capture Refocusing, Defocus Rendering and Blur Removalpdf
generalThe Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographerspdf
motionSynergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generationpdf
generalGraphI2P: Image-to-Point Cloud Registration with Exploring Pattern of Correspondence via Graph Learningpdf
generalSoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizerpdf
generalDPC: Dual-Prompt Collaboration for Tuning Vision-Language Modelspdf
generalAIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic Datapdf
generalRobust Multi-Object 4D Generation for In-the-wild Videospdf
generalMono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-trainingpdf
generalFLAVC: Learned Video Compression with Feature Level Attentionpdf
generalAn End-to-End Robust Point Cloud Semantic Segmentation Network with Single-Step Conditional Diffusion Modelspdf
generalPCDreamer: Point Cloud Completion Through Multi-view Diffusion Priorspdf
generalYour ViT is Secretly an Image Segmentation Modelpdf
generalCross-Rejective Open-Set SAR Image Registrationpdf
gsplatSplineGS: Robust Motion-Adaptive Spline for Real-Time Dynamic 3D Gaussians from Monocular Videopdf
beingsMulti-modal Knowledge Distillation-based Human Trajectory Forecastingpdf
generalShiftwiseConv: Small Convolutional Kernel with Large Kernel Effectpdf
generalObject-Shot Enhanced Grounding Network for Egocentric Videopdf
generalEv-3DOD: Pushing the Temporal Boundaries of 3D Object Detection with Event Cameraspdf
generalNearly Zero-Cost Protection Against Mimicry by Personalized Diffusion Modelspdf
generalThe Devil is in Temporal Token: High Quality Video Reasoning Segmentationpdf
gsplatLITA-GS: Illumination-Agnostic Novel View Synthesis via Reference-Free 3D Gaussian Splatting and Physical Priorspdf
generalT-FAKE: Synthesizing Thermal Images for Facial Landmarkingpdf
generalMulti-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representationpdf
generalPICD: Versatile Perceptual Image Compression with Diffusion Renderingpdf
motionVideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and Editingpdf
generalSix-CD: Benchmarking Concept Removals for Text-to-image Diffusion Modelspdf
generalBlack-Box Forgery Attacks on Semantic Watermarks for Diffusion Modelspdf
generalVidSeg: Training-free Video Semantic Segmentation based on Diffusion Modelspdf
motionPersonaBooth: Personalized Text-to-Motion Generationpdf
generalStar with Bilinear Mappingpdf
generalDIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Modelspdf
gsplatTime of the Flight of the Gaussians: Optimizing Depth Indirectly in Dynamic Radiance Fieldspdf
generalAlign3R: Aligned Monocular Depth Estimation for Dynamic Videospdf
generalSeek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment Recognitionpdf
generalAnomize: Better Open Vocabulary Video Anomaly Detectionpdf
generalEfficient Diffusion as Low Light Enhancerpdf
generalHyperNVD: Accelerating Neural Video Decomposition via Hypernetworkspdf
generalInstant Adversarial Purification with Adversarial Consistency Distillationpdf
generalFeature Selection for Latent Factor Modelspdf
generalPreserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editingpdf
generalDecoupling Training-Free Guided Diffusion by ADMMpdf
generalSwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step Diffusionpdf
generalLearning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenespdf
generalCLIP-driven Coarse-to-fine Semantic Guidance for Fine-grained Open-set Semi-supervised Learningpdf
generalA Simple Data Augmentation for Feature Distribution Skewed Federated Learningpdf
generalGLane3D: Detecting Lanes with Graph of 3D Keypointspdf
generalMinimal Interaction Seperated Tuning: A New Paradigm for Visual Adaptationpdf
generalAttraction Diminishing and Distributing for Few-Shot Class-Incremental Learningpdf
gsplat4DTAM: Non-Rigid Tracking and Mapping via Dynamic Surface Gaussianspdf
generalUnseen Visual Anomaly Generationpdf
generalT2ICount: Enhancing Cross-modal Understanding for Zero-Shot Countingpdf
generalReNeg: Learning Negative Embedding with Reward Guidancepdf
motionMotionPro: A Precise Motion Controller for Image-to-Video Generationpdf
generalGoku: Flow Based Video Generative Foundation Modelspdf
generalWISH: Weakly Supervised Instance Segmentation using Heterogeneous Labelspdf
generalGood, Cheap, and Fast: Overfitted Image Compression with Wasserstein Distortionpdf
generalPeriod-LLM: Extending the Periodic Capability of Multimodal Large Language Modelpdf
generalV2X-R: Cooperative LiDAR-4D Radar Fusion with Denoising Diffusion for 3D Object Detectionpdf
generalTAROT: Towards Essentially Domain-Invariant Robustness with Theoretical Justificationpdf
generalUnveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysispdf
generalAPT: Adaptive Personalized Training for Diffusion Models with Limited Datapdf
generalSCAP: Transductive Test-Time Adaptation via Supportive Clique-based Attribute Promptingpdf
generalTracktention: Leveraging Point Tracking to Attend Videos Faster and Betterpdf
generalDocVLM: Make Your VLM an Efficient Readerpdf
generalRevisiting Source-Free Domain Adaptation: Insights into Representativeness, Generalization, and Varietypdf
generalAdaptive Unimodal Regulation for Balanced Multimodal Information Acquisitionpdf
generalFLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Viewspdf
gsplatImproving Gaussian Splatting with Localized Points Managementpdf
generalOne-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Modelspdf
generalDomain Adaptive Diabetic Retinopathy Grading with Model Absence and Flowing Datapdf
generalLoRASculpt: Sculpting LoRA for Harmonizing General and Specialized Knowledge in Multimodal Large Language Modelspdf
generalSEAL: Semantic Attention Learning for Long Video Representationpdf
generalSCFlow2: Plug-and-Play Object Pose Refiner with Shape-Constraint Scene Flowpdf
motionFlipSketch: Flipping Static Drawings to Text-Guided Sketch Animationspdf
generalSketchAgent: Language-Driven Sequential Sketch Generationpdf
generalDRAWER: Digital Reconstruction and Articulation With Environment Realismpdf
generalGoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View Synthesispdf
generalDeep Change Monitoring: A Hyperbolic Representative Learning Framework and a Dataset for Long-term Fine-grained Tree Change Detectionpdf
generalITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-Onpdf
generalMultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrievalpdf
generalVolFormer: Explore More Comprehensive Cube Interaction for Hyperspectral Image Restoration and Beyondpdf
generalBizGen: Advancing Article-level Visual Text Rendering for Infographics Generationpdf
generalSmartCLIP: Modular Vision-language Alignment with Identification Guaranteespdf
generalQ-DiT: Accurate Post-Training Quantization for Diffusion Transformerspdf
generalRoboGround: Robotic Manipulation with Grounded Vision-Language Priorspdf
generalImproving Transferable Targeted Attacks with Feature Tuning Mixuppdf
worldsDrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulationpdf
beingsHuPerFlow: A Comprehensive Benchmark for Human vs. Machine Motion Estimation Comparisonpdf
generalMetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt Tuningpdf
generalSubnet-Aware Dynamic Supernet Training for Neural Architecture Searchpdf
worldsEchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidancepdf
beingsControllable Human Image Generation with Personalized Multi-Garmentspdf
generalUHD-processer: Unified UHD Image Restoration with Progressive Frequency Learning and Degradation-aware Promptspdf
gsplatGBC-Splat: Generalizable Gaussian-Based Clothed Human Digitalization under Sparse RGB Cameraspdf
generalAC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformerspdf
generalA Unified Model for Compressed Sensing MRI Across Undersampling Patternspdf
worldsTSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-Resolutionpdf
generalFast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Passpdf
generalStyleStudio: Text-Driven Style Transfer with Selective Control of Style Elementspdf
generalCTRL-O: Language-Controllable Object-Centric Visual Representation Learningpdf
generalText Augmented Correlation Transformer For Few-shot Classification & Segmentationpdf
generalUnified Dense Prediction of Video Diffusionpdf
generalTowards Million-Scale Adversarial Robustness Evaluation With Stronger Individual Attackspdf
generalTemporal Action Detection Model Compression by Progressive Block Droppdf
generalPursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interactionpdf
generalDINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignmentpdf
generalLearning Affine Correspondences by Integrating Geometric Constraintspdf
generalUCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label Learningpdf
generalGeometry in Style: 3D Stylization via Surface Normal Deformationpdf
generalPVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Modelspdf
generalMultiple Object Tracking as ID Predictionpdf
generalPIDLoc: Cross-View Pose Optimization Network Inspired by PID Controllerspdf
generalDreamOmni: Unified Image Generation and Editingpdf
generalHash3D: Training-free Acceleration for 3D Generationpdf
worldsLearning Hazing to Dehazing: Towards Realistic Haze Generation for Real-World Image Dehazingpdf
generalRUBIK: A Structured Benchmark for Image Matching across Geometric Challengespdf
generalFast and Accurate Gigapixel Pathological Image Classification with Hierarchical Distillation Multi-Instance Learningpdf
gsplatIncEventGS: Pose-Free Gaussian Splatting from a Single Event Camerapdf
generalOpenMIBOOD: Open Medical Imaging Benchmarks for Out-Of-Distribution Detectionpdf
generalFactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Modelspdf
generalWhen the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learningpdf
beingsUniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editingpdf
motionPOMP: Physics-consistent Motion Generative Model through Phase Manifoldspdf
generalReasoning to Attend: Try to Understand How <SEG> Token Workspdf
generalReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streamspdf
generalDynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolutionpdf
generalPlaying the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategypdf
worldsVideoWorld: Exploring Knowledge Learning from Unlabeled Videospdf
general3D-SLNR: A Super Lightweight Neural Representation for Large-scale 3D Mappingpdf
generalSTINR: Deciphering Spatial Transcriptomics via Implicit Neural Representationpdf
generalRADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Modelspdf
generalUnsupervised Discovery of Facial Landmarks and Head Posepdf
generalInstruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learningpdf
generalStabilizing and Accelerating Autofocus with Expert Trajectory Regularized Deep Reinforcement Learningpdf
generalRepurposing Stable Diffusion Attention for Training-Free Unsupervised Interactive Segmentationpdf
gsplatGO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detectorpdf
generalDPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentationpdf
generalSimulator HC: Regression-based Online Simulation of Starting Problem-Solution Pairs for Homotopy Continuation in Geometric Visionpdf
generalDynamic Integration of Task-Specific Adapters for Class Incremental Learningpdf
generalEgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Visionpdf
generalDiverseFlow: Sample-Efficient Diverse Mode Coverage in Flowspdf
worldsMaskGWM: A Generalizable Driving World Model with Video Mask Reconstructionpdf
general3D-MVP: 3D Multiview Pretraining for Manipulationpdf
generalEnhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representationspdf
generalMimir: Improving Video Diffusion Models for Precise Text Understandingpdf
generalUCM-VeID V2: A Richer Dataset and A Pre-training Method for UAV Cross-Modality Vehicle Re-Identificationpdf
beingsGeoAvatar: Geometrically-Consistent Multi-Person Avatar Reconstruction from Sparse Multi-View Videospdf
gsplatDiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splattingpdf
gsplatSpeedy-Splat: Fast 3D Gaussian Splatting with Sparse Pixels and Sparse Primitivespdf
generalOmnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videospdf
beingsODHSR: Online Dense 3D Reconstruction of Humans and Scenes from Monocular Videospdf
generalSpiritSight Agent: Advanced GUI Agent with One Lookpdf
generalZero-Shot Monocular Scene Flow Estimation in the Wildpdf
motionMG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularitiespdf
generalRetaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher Distillationpdf
generalMMRL: Multi-Modal Representation Learning for Vision-Language Modelspdf
generalAnchor-Aware Similarity Cohesion in Target Frames Enables Predicting Temporal Moment Boundaries in 2Dpdf
generalBreaking the Low-Rank Dilemma of Linear Attentionpdf
generalEmbracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learningpdf
generalUnity in Diversity: Video Editing via Gradient-Latent Purificationpdf
generalRevealing Key Details to See Differences: A Novel Prototypical Perspective for Skeleton-based Action Recognitionpdf
generalFinsler Multi-Dimensional Scaling: Manifold Learning for Asymmetric Dimensionality Reduction and Embeddingpdf
generalVideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selectionpdf
motionCross-Modal Distillation for 2D/3D Multi-Object Discovery from 2D Motionpdf
generalBridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical Flowpdf
generalInversion Circle Interpolation: Diffusion-based Image Augmentation for Data-scarce Classificationpdf
generalTSP-Mamba: The Travelling Salesman Problem Meets Mamba for Image Super-resolution and Beyondpdf
generalRENO: Real-Time Neural Compression for 3D LiDAR Point Cloudspdf
generalFADE: Frequency-Aware Diffusion Model Factorization for Video Editingpdf
beingsData Synthesis with Diverse Styles for Face Recognition via 3DMM-Guided Diffusionpdf
generalGeometric Knowledge-Guided Localized Global Distribution Alignment for Federated Learningpdf
generalGlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editingpdf
generalBirth and Death of a Rosepdf
generalMetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural Representationpdf
generalMovieBench: A Hierarchical Movie Level Dataset for Long Video Generationpdf
generalBe More Specific: Evaluating Object-centric Realism in Synthetic Imagespdf
generalCorrelative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuningpdf
beingsSFDM: Robust Decomposition of Geometry and Reflectance for Realistic Face Rendering from Sparse-view Imagespdf
generalOuroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusionpdf
generalQMambaBSR: Burst Image Super-Resolution with Query State Space Modelpdf
generalMulti-Group Proportional Representations for Text-to-Image Modelspdf
generalTowards Generalizable Trajectory Prediction using Dual-Level Representation Learning and Adaptive Promptingpdf
generalCoMatcher: Multi-View Collaborative Feature Matchingpdf
beingsTowards a Universal Synthetic Video Detector: From Face or Background Manipulations to Fully AI-Generated Contentpdf
beingsA Focused Human Body Model for Accurate Anthropometric Measurements Extractionpdf
generalACE: Anti-Editing Concept Erasure in Text-to-Image Modelspdf
motionSynthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priorspdf
generalHierarchical Knowledge Prompt Tuning for Multi-task Test-Time Adaptationpdf
generalLaTexBlend: Scaling Multi-concept Customized Generation with Latent Textual Blendingpdf
generalDejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classificationpdf
generalLet's Verify and Reinforce Image Generation Step by Steppdf
generalAll-Optical Nonlinear Diffractive Deep Network for Ultrafast Image Denoisingpdf
generalUNOPose: Unseen Object Pose Estimation with an Unposed RGB-D Reference Imagepdf
meshingHybridMQA: Exploring Geometry-Texture Interactions for Colored Mesh Quality Assessmentpdf
generalSIR-DIFF: Sparse Image Sets Restoration with Multi-View Diffusion Modelpdf
generalReversible Decoupling Network for Single Image Reflection Removalpdf
generalHierarchical Features Matter: A Deep Exploration of Progressive Parameterization Method for Dataset Distillationpdf
generalGLASS: Guided Latent Slot Diffusion for Object-Centric Learningpdf
generalSASep: Saliency-Aware Structured Separation of Geometry and Feature for Open Set Learning on Point Cloudspdf
generalLow-Biased General Annotated Dataset Generationpdf
generalGenerative Hard Example Augmentation for Semantic Point Cloud Segmentationpdf
generalETAP: Event-based Tracking of Any Pointpdf
generalBeyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledgepdf
meshingVolumetric Surfaces: Representing Fuzzy Geometries with Layered Meshespdf
generalSTEPS: Sequential Probability Tensor Estimation for Text-to-Image Hard Prompt Searchpdf
generalVIRES: Video Instance Repainting via Sketch and Text Guided Generationpdf
generalMIMO: Controllable Character Video Synthesis with Spatial Decomposed Modelingpdf
gsplatFrom Sparse to Dense: Camera Relocalization with Scene-Specific Detector from Feature Gaussian Splattingpdf
beingsStableAnimator: High-Quality Identity-Preserving Human Image Animationpdf
generalOODD: Test-time Out-of-Distribution Detection with Dynamic Dictionarypdf
generalBIMBA: Selective-Scan Compression for Long-Range Video Question Answeringpdf
generalDiff2Flow: Training Flow Matching Models via Diffusion Model Alignmentpdf
generalProf. Robot: Differentiable Robot Rendering Without Static and Self-Collisionspdf
gsplatDropGaussian: Structural Regularization for Sparse-view Gaussian Splattingpdf
generalBlurred LiDAR for Sharper 3D: Robust Handheld 3D Scanning with Diffuse LiDAR and RGBpdf
generalNovel View Synthesis with Pixel-Space Diffusion Modelspdf
generalObject Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark Datasetpdf
generalDocument Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documentspdf
generalRethinking Few-Shot Adaptation of Vision-Language Models in Two Stagespdf
gsplatTAGA: Self-supervised Learning for Template-free Animatable Gaussian Articulated Modelpdf
gsplatHorizon-GS: Unified 3D Gaussian Splatting for Large-Scale Aerial-to-Ground Scenespdf
generalLotusFilter: Fast Diverse Nearest Neighbor Search via a Learned Cutoff Tablepdf
generalThink Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulationpdf
generalSoMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation Learningpdf
gsplatRef-GS: Directional Factorization for 2D Gaussian Splattingpdf
generalVDocRAG: Retrieval-Augmented Generation over Visually-Rich Documentspdf
generalConcept Lancet: Image Editing with Compositional Representation Transplantpdf
gsplatGenerative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D Reconstructionpdf
generalVideo-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysispdf
beingsAre Images Indistinguishable to Humans Also Indistinguishable to Classifiers?pdf
generalUnderstanding Multi-layered Transmission Matricespdf
gsplatGS-DiT: Advancing Video Generation with Dynamic 3D Gaussian Fields through Efficient Dense 3D Point Trackingpdf
motionAnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Modelspdf
worldsDual Energy-Based Model with Open-World Uncertainty Estimation for Out-of-distribution Detectionpdf
generalDTGBrepGen: A Novel B-rep Generative Model through Decoupling Topology and Geometrypdf
generalSchedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generationpdf
generalLearning Audio-guided Video Representation with Gated Attention for Video-Text Retrievalpdf
generalTASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulationpdf
generalNoT: Federated Unlearning via Weight Negationpdf
generalRANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddingspdf
beingsSimMotionEdit: Text-Based Human Motion Editing with Motion Similarity Predictionpdf
generalFrom Head to Tail: Efficient Black-box Model Inversion Attack via Long-tailed Learningpdf
generalSIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Modelpdf
generalObject-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulationpdf
generalDepth Any Camera: Zero-Shot Metric Depth Estimation from Any Camerapdf
generalShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Modelspdf
generalSound Bridge: Associating Egocentric and Exocentric Videos via Audio Cuespdf
generalOmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotationspdf
worldsLayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Modelspdf
generalPoint Clouds Meets Physics: Dynamic Acoustic Field Fitting Network for Point Cloud Understandingpdf
generalFaster Parameter-Efficient Tuning with Token Redundancy Reductionpdf
generalPanorama Generation From NFoV Image Done Rightpdf
gsplatSparse Point Cloud Patches Rendering via Splitting 2D Gaussianspdf
generalDistilling Monocular Foundation Model for Fine-grained Depth Completionpdf
generalAniGrad: Anisotropic Gradient-Adaptive Sampling for 3D Reconstruction From Monocular Videopdf
generalLess Attention is More: Prompt Transformer for Generalized Category Discoverypdf
motionAToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Rewardpdf
gsplatDoF-Gaussian: Controllable Depth-of-Field for 3D Gaussian Splattingpdf
generalReconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in Dualpdf
generalHierarchical Flow Diffusion for Efficient Frame Interpolationpdf
generalBASKET: A Large-Scale Video Dataset for Fine-Grained Skill Estimationpdf
generalArbitrary-steps Image Super-resolution via Diffusion Inversionpdf
generalDynamic Neural Surfaces for Elastic 4D Shape Representation and Analysispdf
generalComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systemspdf
generalIncomplete Multi-View Multi-label Learning via Disentangled Representation and Label Semantic Embeddingpdf
generalAutoURDF: Unsupervised Robot Modeling from Point Cloud Frames Using Cluster Registrationpdf
generalGolden Cudgel Network for Real-Time Semantic Segmentationpdf
generalMulti-modal Contrastive Learning with Negative Sampling Calibration for Phenotypic Drug Discoverypdf
generalR-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuningpdf
gsplatSplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesispdf
generalBoltzmann Attention Sampling for Image Analysis with Small Objectspdf
gsplatGeneralized Recorrupted-to-Recorrupted: Self-Supervised Learning Beyond Gaussian Noisepdf
motionDynamic Motion Blending for Versatile Motion Editingpdf
generalStdGEN: Semantic-Decomposed 3D Character Generation from Single Imagespdf
generalSpatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecastingpdf
generalFFR: Frequency Feature Rectification for Weakly Supervised Semantic Segmentationpdf
generalVideo-XL: Extra-Long Vision Language Model for Hour-Scale Video Understandingpdf
generalSonata: Self-Supervised Learning of Reliable Point Representationspdf
generalDriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generationpdf
generalDRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characterspdf
generalA Unified Approach to Interpreting Self-supervised Pre-training Methods for 3D Point Clouds via Interactionspdf
generalEnhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image Segmentationpdf
generalSTEREO: A Two-Stage Framework for Adversarially Robust Concept Erasing from Text-to-Image Diffusion Modelspdf
generalGenVDM: Generating Vector Displacement Maps From a Single Imagepdf
generalEffective SAM Combination for Open-Vocabulary Semantic Segmentationpdf
worldsTowards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detectionpdf
generalUniAP: Unifying Inter- and Intra-Layer Automatic Parallelism by Mixed Integer Quadratic Programmingpdf
generalTurbo3D: Ultra-fast Text-to-3D Generationpdf
meshingSUM Parts: Benchmarking Part-Level Semantic Segmentation of Urban Meshespdf
beingsMODA: Motion-Drift Augmentation for Inertial Human Motion Analysispdf
generalHigher-Order Ratio Cycles for Fast and Globally Optimal Shape Matchingpdf
generalHyperdimensional Uncertainty Quantification for Multimodal Uncertainty Fusion in Autonomous Vehicles Perceptionpdf
gsplatGIFStream: 4D Gaussian-based Immersive Video with Feature Streampdf
generalMulti-Scale Neighborhood Occupancy Masked Autoencoder for Self-Supervised Learning in LiDAR Point Cloudspdf
generalPLeaS - Merging Models with Permutations and Least Squarespdf
generalIncremental Object Keypoint Learningpdf
generalInteractVLM: 3D Interaction Reasoning from 2D Foundational Modelspdf
generalAttribute-Missing Multi-view Graph Clusteringpdf
generalPose-Guided Temporal Enhancement for Robust Low-Resolution Hand Reconstructionpdf
generalReDiffDet: Rotation-equivariant Diffusion Model for Oriented Object Detectionpdf
generalUnlocking Generalization Power in LiDAR Point Cloud Registrationpdf
generalLoRA Recycle: Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAspdf
generalEvent Fields: Capturing Light Fields at High Speed, Resolution, and Dynamic Rangepdf
generalHyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing Imagerypdf
gsplatGBlobs: Explicit Local Structure via Gaussian Blobs for Improved Cross-Domain LiDAR-based 3D Object Detectionpdf
gsplat3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial Representationspdf
generalMambaIRv2: Attentive State Space Restorationpdf
generalFloating No More: Object-Ground Reconstruction from a Single Imagepdf
generalPattern Analogies: Learning to Perform Programmatic Image Edits by Analogypdf
generalSTAR-Edge: Structure-aware Local Spherical Curve Representation for Thin-walled Edge Extraction from Unstructured Point Cloudspdf
generalBoosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty Estimationpdf
generalpFedMxF: Personalized Federated Class-Incremental Learning with Mixture of Frequency Aggregationpdf
generalEfficient Transfer Learning for Video-language Foundation Modelspdf
generalRadio Frequency Ray Tracing with Neural Object Representation for Enhanced RF Modelingpdf
generalNeuro-3D: Towards 3D Visual Decoding from EEG Signalspdf
generalProbing the Mid-level Vision Capabilities of Self-Supervised Learningpdf
generalEfficient Long Video Tokenization via Coordinate-based Patch Reconstructionpdf
generalDerivative-Free Diffusion Manifold-Constrained Gradient for Unified XAIpdf
generalZoomLDM: Latent Diffusion Model for Multi-scale Image Generationpdf
gsplatGaussianUDF: Inferring Unsigned Distance Functions through 3D Gaussian Splattingpdf
generalCrossSDF: 3D Reconstruction of Thin Structures From Cross-Sectionspdf
generalDV-Matcher: Deformation-based Non-rigid Point Cloud Matching Guided by Pre-trained Visual Featurespdf
generalReasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance Groundingpdf
generalAdaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancementpdf
generalFLAIR: VLM with Fine-grained Language-informed Image Representationspdf
generalGG-SSMs: Graph-Generating State Space Modelspdf
generalContinuous Adverse Weather Removal via Degradation-Aware Distillationpdf
generalExploiting Temporal State Space Sharing for Video Semantic Segmentationpdf
gsplatHigh-fidelity 3D Object Generation from Single Image with RGBN-Volume Gaussian Reconstruction Modelpdf
gsplatSteepest Descent Density Control for Compact 3D Gaussian Splattingpdf
beingsOptimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofingpdf
generalRobust 3D Shape Reconstruction in Zero-Shot from a Single Image in the Wildpdf
generalBOE-ViT: Boosting Orientation Estimation with Equivariance in Self-Supervised 3D Subtomogram Alignmentpdf
generalAdventurer: Optimizing Vision Mamba Architecture Designs for Efficiencypdf
generalBeyond Local Sharpness: Communication-Efficient Global Sharpness-aware Minimization for Federated Learningpdf
motionParameterized Blur Kernel Prior Learning for Local Motion Deblurringpdf
generalScene4U: Hierarchical Layered 3D Scene Reconstruction from Single Panoramic Image for Your Immerse Explorationpdf
generalACAttack: Adaptive Cross Attacking RGB-T Tracker via Multi-Modal Response Decouplingpdf
generalDeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videospdf
gsplatHoGS: Unified Near and Far Object Reconstruction via Homogeneous Gaussian Splattingpdf
generalSmartEraser: Remove Anything from Images using Masked-Region Guidancepdf
generalSample- and Parameter-Efficient Auto-Regressive Image Modelspdf
generalRobust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignmentpdf
generalBiLoRA: Almost-Orthogonal Parameter Spaces for Continual Learningpdf
meshingVid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulationpdf
worldsSceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environmentspdf
generalCollaborative Decoding Makes Visual Auto-Regressive Modeling Efficientpdf
generalAerialMegaDepth: Learning Aerial-Ground Reconstruction and View Synthesispdf
generalVisual Representation Learning through Causal Intervention for Controllable Image Editingpdf
generalExploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesispdf
generalA Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generationpdf
gsplatDeformable Radial Kernel Splattingpdf
generalBayesian Prompt Flow Learning for Zero-Shot Anomaly Detectionpdf
generalHalLoc: Token-level Localization of Hallucinations for Vision Language Modelspdf
generalDiffPortrait360: Consistent Portrait Diffusion for 360 View Synthesispdf
generalSURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation Sparsitypdf
generalFrom Slow Bidirectional to Fast Autoregressive Video Diffusion Modelspdf
generalNoise Diffusion for Enhancing Semantic Faithfulness in Text-to-Image Synthesispdf
generalMonoInstance: Enhancing Monocular Priors via Multi-view Instance Alignment for Neural Rendering and Reconstructionpdf
generalCAT4D: Create Anything in 4D with Multi-View Video Diffusion Modelspdf
generalExploring Semantic Feature Discrimination for Perceptual Image Super-Resolution and Opinion-Unaware No-Reference Image Quality Assessmentpdf
generalDistilling Long-tailed Datasetspdf
generalGaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoderspdf
generalIncorporating Dense Knowledge Alignment into Unified Multimodal Representation Modelspdf
beingsBoost Your Human Image Generation Model via Direct Preference Optimizationpdf
generalLearning to Highlight Audio by Watching Moviespdf
generalUnified Uncertainty-Aware Diffusion for Multi-Agent Trajectory Modelingpdf
generalWeGen: A Unified Model for Interactive Multimodal Generation as We Chatpdf
gsplatHRAvatar: High-Quality and Relightable Gaussian Head Avatarpdf
generalA Distractor-Aware Memory for Visual Object Tracking with SAM2pdf
generalActivating Sparse Part Concepts for 3D Class Incremental Learningpdf
generalProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Groundingpdf
generalBFANet: Revisiting 3D Semantic Segmentation with Boundary Feature Analysispdf
generalBeyond Words: Augmenting Discriminative Richness via Diffusions in Unsupervised Prompt Learningpdf
generalUnlocking the Potential of Unlabeled Data in Semi-Supervised Domain Generalizationpdf
generalSteering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreakspdf
generalNeural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusionpdf
generalTowards Natural Language-Based Document Image Retrieval: New Dataset and Benchmarkpdf
gsplatMitigating Ambiguities in 3D Classification with Gaussian Splattingpdf
generalDH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings Learningpdf
generalUnveil Inversion and Invariance in Flow Transformer for Versatile Image Editingpdf
generalDyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentationpdf
generalDUNE: Distilling a Universal Encoder from Heterogeneous 2D and 3D Teacherspdf
generalBlack Hole-Driven Identity Absorbing in Diffusion Modelspdf
generalHiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Modelspdf
motionHallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformerpdf
generalSeqMvRL: A Sequential Fusion Framework for Multi-view Representation Learningpdf
generalBadToken: Token-level Backdoor Attacks to Multi-modal Large Language Modelspdf
generalVLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement Learningpdf
generalNeISF++: Neural Incident Stokes Field for Polarized Inverse Rendering of Conductors and Dielectricspdf
generalNon-Natural Image Understanding with Advancing Frequency-based Vision Encoderspdf
generalGenerative Multimodal Pretraining with Discrete Diffusion Timestep Tokenspdf
gsplatSplatFlow: Self-Supervised Dynamic Gaussian Splatting in Neural Motion Flow Field for Autonomous Drivingpdf
worldsAesthetiQ: Enhancing Graphic Layout Design via Aesthetic-Aware Preference Alignment of Multi-modal Large Language Modelspdf
generalFINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularitypdf
generalChebyshev Attention Depth Permutation Texture Network with Latent Texture Attribute Losspdf
generalDecentralized Diffusion Modelspdf
generalAnyEdit: Mastering Unified High-Quality Image Editing for Any Ideapdf
generalDNF: Unconditional 4D Generation with Dictionary-based Neural Fieldspdf
generalARM: Appearance Reconstruction Model for Relightable 3D Generationpdf
generalGround-V: Teaching VLMs to Ground Complex Instructions in Pixelspdf
meshingTreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencingpdf
generalGenerating 3D-Consistent Videos from Unposed Internet Photospdf
generalParameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformationpdf
generalViUniT: Visual Unit Tests for More Robust Visual Programmingpdf
generalDualTalk: Dual-Speaker Interaction for 3D Talking Head Conversationspdf
generalbeta-FFT: Nonlinear Interpolation and Differentiated Training Strategies for Semi-Supervised Medical Image Segmentationpdf
generalDynamic Group Normalization: Spatio-Temporal Adaptation to Evolving Data Statisticspdf
generalSynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Foldingpdf
generalUncertain Multimodal Intention and Emotion Understanding in the Wildpdf
generalVidTwin: Video VAE with Decoupled Structure and Dynamicspdf
generalCL-LoRA: Continual Low-Rank Adaptation for Rehearsal-Free Class-Incremental Learningpdf
beingsDesign2GarmentCode: Turning Design Concepts to Tangible Garments Through Program Synthesispdf
gsplatEfficient Dynamic Scene Editing via 4D Gaussian-based Static-Dynamic Separationpdf
generalUnlearning through Knowledge Overwriting: Reversible Federated Unlearning via Selective Sparse Adapterpdf
generalSocialMOIF: Multi-Order Intention Fusion for Pedestrian Trajectory Predictionpdf
generalHistoFS: Non-IID Histopathologic Whole Slide Image Classification via Federated Style Transfer with RoI-Preservingpdf
gsplatSGSST: Scaling Gaussian Splatting Style Transferpdf
generalLearning Bijective Surface Parameterization for Inferring Signed Distance Functions from Sparse Point Clouds with Grid Deformationpdf
generalBalancing Two Classifiers via A Simplex ETF Structure for Model Calibrationpdf
generalDAMM-Diffusion: Learning Divergence-Aware Multi-Modal Diffusion Model for Nanoparticles Distribution Predictionpdf
generalU-Know-DiffPAN: An Uncertainty-aware Knowledge Distillation Diffusion Framework with Details Enhancement for PAN-Sharpeningpdf
gsplatRelationField: Relate Anything in Radiance Fieldspdf
beingsLet Humanoids Hike! Integrative Skill Development on Complex Trailspdf
generalBF-STVSR: B-Splines and Fourier---Best Friends for High Fidelity Spatial-Temporal Video Super-Resolutionpdf
worldsDIO: Decomposable Implicit 4D Occupancy-Flow World Modelpdf
generalSLADE: Shielding against Dual Exploits in Large Vision-Language Modelspdf
beingsEgo4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Inputpdf
generalFreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component Analysispdf
generalMind the Time: Temporally-Controlled Multi-Event Video Generationpdf
generalAudio-Visual Semantic Graph Network for Audio-Visual Event Localizationpdf
motionVideo Motion Transfer with Diffusion Transformerspdf
generalUnified Reconstruction of Static and Dynamic Scenes from Eventspdf
generalAutomatic Spectral Calibration of Hyperspectral Images: Method, Dataset and Benchmarkpdf
generalPoint-to-Region Loss for Semi-Supervised Point-Based Crowd Countingpdf
beingsMove-in-2D: 2D-Conditioned Human Motion Generationpdf
generalMATCHA: Towards Matching Anythingpdf
generalCTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D Diffusionpdf
generalSeparation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question Answeringpdf
generalSF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understandingpdf
generalFitted Neural Lossless Image Compressionpdf
generalJarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restorationpdf
generalF-LMM: Grounding Frozen Large Multimodal Modelspdf
generalEntityErasure: Erasing Entity Cleanly via Amodal Entity Segmentation and Completionpdf
generalJoint Out-of-Distribution Filtering and Data Discovery Active Learningpdf
generalFinding Local Diffusion Schrodinger Bridge using Kolmogorov-Arnold Networkpdf
generalCorrBEV: Multi-View 3D Object Detection by Correlation Learning with Multi-modal Prototypespdf
generalCompletion as Enhancement: A Degradation-Aware Selective Image Guided Network for Depth Completionpdf
worldsAround the World in 80 Timesteps: A Generative Approach to Global Visual Geolocationpdf
gsplatReal-time High-fidelity Gaussian Human Avatars with Position-based Interpolation of Spatially Distributed MLPspdf
generalRoboSense: Large-scale Dataset and Benchmark for Egocentric Robot Perception and Navigation in Crowded and Unstructured Environmentspdf
generalDEFOM-Stereo: Depth Foundation Model Based Stereo Matchingpdf
generalDiskVPS: Vanishing Point Detector via Hough Transform in a Disk Regionpdf
generalSeeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decodingpdf
generalTowards Autonomous Micromobility through Scalable Urban Simulationpdf
generalLanguage-Assisted Debiasing and Smoothing for Foundation Model-Based Semi-Supervised Learningpdf
generalEdgeMovingNet: Edge-preserving Point Cloud Reconstruction via Joint Geometry Featurespdf
generalHarnessing Frozen Unimodal Encoders for Flexible Multimodal Alignmentpdf
gsplatFeature Information Driven Position Gaussian Distribution Estimation for Tiny Object Detectionpdf
generalEnhancing Diversity for Data-free Quantizationpdf
generalFrom Alexnet to Transformers: Measuring the Non-linearity of Deep Neural Networks with Affine Optimal Transportpdf
generalPrompt2Perturb (P2P): Text-Guided Diffusion-Based Adversarial Attack on Breast Ultrasound Imagespdf
generalCOAP: Memory-Efficient Training with Correlation-Aware Gradient Projectionpdf
generalGyro-based Neural Single Image Deblurringpdf
generalImproved Monocular Depth Prediction Using Distance Transform Over Pre-semantic Contours with Self-supervised Neural Networkspdf
beingsIs this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-bodypdf
generalAutomated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluationpdf
generalROLL: Robust Noisy Pseudo-label Learning for Multi-View Clustering with Noisy Correspondencepdf
generalTowards In-the-wild 3D Plane Reconstruction from a Single Imagepdf
generalPQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Predictionpdf
generalCheXwhatsApp: A Dataset for Exploring Challenges in the Diagnosis of Chest X-rays through Mobile Devicespdf
generalDegradation-Aware Feature Perturbation for All-in-One Image Restorationpdf
generalGenDeg: Diffusion-based Degradation Synthesis for Generalizable All-In-One Image Restorationpdf
generalThe Power of Context: How Multimodality Improves Image Super-Resolutionpdf
generalDetect Any Mirrors: Boosting Learning Reliability on Large-Scale Unlabeled Data with an Iterative Data Enginepdf
gsplat4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Modelspdf
beingsMotionMap: Representing Multimodality in Human Pose Forecastingpdf
generalFactored-NeuS: Reconstructing Surfaces, Illumination, and Materials of Possibly Glossy Objectspdf
gsplatGaussianSpa: An "Optimizing-Sparsifying" Simplification Framework for Compact and High-Quality 3D Gaussian Splattingpdf
generalNavigating the Unseen: Zero-shot Scene Graph Generation via Capsule-Based Equivariant Featurespdf
generalVL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Modelspdf
generalASHiTA: Automatic Scene-grounded HIerarchical Task Analysispdf
generalDiscovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Modelspdf
generalRoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigationpdf
generalBringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysispdf
generalA Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image Segmentationpdf
generalFedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Modelspdf
generalGCE-Pose: Global Context Enhancement for Category-level Object Pose Estimationpdf
generalLearning from Neighbors: Category Extrapolation for Long-Tail Learningpdf
generalMaterial Anything: Generating Materials for Any 3D Object via Diffusionpdf
generalImagineFSL: Self-Supervised Pretraining Matters on Imagined Base Set for VLM-based Few-shot Learningpdf
generalContinuous Locomotive Crowd Behavior Generationpdf
generalProject-Probe-Aggregate: Efficient Fine-Tuning for Group Robustnesspdf
generalImplicit Bias Injection Attacks against Text-to-Image Diffusion Modelspdf
generalROICtrl: Boosting Instance Control for Visual Generationpdf
generalCropper: Vision-Language Model for Image Cropping through In-Context Learningpdf
motionScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Modelpdf
generalICE: Intrinsic Concept Extraction from a Single Image via Diffusion Modelspdf
generalASIGN: An Anatomy-aware Spatial Imputation Graphic Network for 3D Spatial Transcriptomicspdf
generalMultiMorph: On-demand Atlas Constructionpdf
generalOctopus: Alleviating Hallucination via Dynamic Contrastive Decodingpdf
generalSpiking Transformer: Introducing Accurate Addition-Only Spiking Self-Attention for Transformerpdf
worldsMOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenespdf
generalSymbolic Representation for Any-to-Any Generative Taskspdf
generalProtecting Your Video Content: Disrupting Automated Video-based LLM Annotationspdf
generalMedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representationspdf
gsplatArticulatedGS: Self-supervised Digital Twin Modeling of Articulated Objects using 3D Gaussian Splattingpdf
generalLeveraging 3D Geometric Priors in 2D Rotation Symmetry Detectionpdf
generalNoise Calibration and Spatial-Frequency Interactive Network for STEM Image Enhancementpdf
beingsHomogeneous Dynamics Space for Heterogeneous Humanspdf
generalTailedCore: Few-Shot Sampling for Unsupervised Long-Tail Noisy Anomaly Detectionpdf
generalSatellite Observations Guided Diffusion Model for Accurate Meteorological States at Arbitrary Resolutionpdf
generalReconstructing People, Places, and Cameraspdf
generalInPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model Alignmentpdf
generalIdentifying and Mitigating Spurious Correlation in Multi-Task Learningpdf
generalImmune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignmentpdf
generalCustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillationpdf
meshingPMNI: Pose-free Multi-view Normal Integration for Reflective and Textureless Surface Reconstructionpdf
gsplatLeanGaussian: Breaking Pixel or Point Cloud Correspondence in Modeling 3D Gaussianspdf
worldsModeling Multiple Normal Action Representations for Error Detection in Procedural Taskspdf
generalA Unified Latent Schrodinger Bridge Diffusion Model for Unsupervised Anomaly Detection and Localizationpdf
generalMambaVision: A Hybrid Mamba-Transformer Vision Backbonepdf
generalMulti-Label Prototype Visual Spatial Search for Weakly Supervised Semantic Segmentationpdf
generalDoppelgangers++: Improved Visual Disambiguation with Geometric 3D Featurespdf
gsplatLearnable Infinite Taylor Gaussian for Dynamic View Renderingpdf
generalSaMam: Style-aware State Space Model for Arbitrary Image Style Transferpdf
generalMaking Old Film Great Again: Degradation-aware State Space Model for Old Film Restorationpdf
motionMP-SfM: Monocular Surface Priors for Robust Structure-from-Motionpdf
beingsDeterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion Synthesispdf
generalCacheQuant: Comprehensively Accelerated Diffusion Modelspdf
worldsOpen-World Objectness Modeling Unifies Novel Object Detectionpdf
beingsMotionPRO: Exploring the Role of Pressure in Human MoCap and Beyondpdf
generalDiffVsgg: Diffusion-Driven Online Video Scene Graph Generationpdf
generalTowards Smart Point-and-Shoot Photographypdf
generalPrototype-Based Image Prompting for Weakly Supervised Histopathological Image Segmentationpdf
beingsMitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulationpdf
generalSpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Languagepdf
generalMono2Stereo: A Benchmark and Empirical Study for Stereo Conversionpdf
generalSoftShadow: Leveraging Soft Masks for Penumbra-Aware Shadow Removalpdf
generalVTON-HandFit: Virtual Try-on for Arbitrary Hand Pose Guided by Hand Priors Embeddingpdf
generalUni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusionpdf
generalPOSTA: A Go-to Framework for Customized Artistic Poster Generationpdf
generalNSD-Imagery: A Benchmark Dataset for Extending fMRI Vision Decoding Methods to Mental Imagerypdf
generalVLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Modelspdf
motionJust Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detectionpdf
motionEfficient Motion-Aware Video MLLMpdf
generalZero-Shot 4D Lidar Panoptic Segmentationpdf
generalADU: Adaptive Detection of Unknown Categories in Black-Box Domain Adaptationpdf
generalEmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusionpdf
generalUnsupervised Foundation Model-Agnostic Slide-Level Representation Learningpdf
generalUNIALIGN: Scaling Multimodal Alignment within One Unified Modelpdf
generalShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructionspdf
generalExploration-Driven Generative Interactive Environmentspdf
generalDreamText: High Fidelity Scene Text Synthesispdf
generalProKeR: A Kernel Perspective on Few-Shot Adaptation of Large Vision-Language Modelspdf
generalMonoTAKD: Teaching Assistant Knowledge Distillation for Monocular 3D Object Detectionpdf
generalEasy-editable Image Vectorization with Multi-layer Multi-scale Distributed Visual Feature Embeddingpdf
generalAcquire and then Adapt: Squeezing out Text-to-Image Model for Image Restorationpdf
generalDevils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lenspdf
motionSpectroMotion: Dynamic 3D Reconstruction of Specular Scenespdf
generalVTON 360: High-Fidelity Virtual Try-On from Any Viewing Directionpdf
generalMVBoost: Boost 3D Reconstruction with Multi-View Refinementpdf
generalCategory-Agnostic Neural Object Riggingpdf
generalPOPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentationpdf
generalMMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesispdf
generalMimic In-Context Learning for Multimodal Taskspdf
generalVision-Language Models Do Not Understand Negationpdf
gsplatNexusGS: Sparse View Synthesis with Epipolar Depth Priors in 3D Gaussian Splattingpdf
generalHyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight Trajectoriespdf
generalRICCARDO: Radar Hit Prediction and Convolution for Camera-Radar 3D Object Detectionpdf
beingsBLADE: Single-view Body Mesh Estimation through Accurate Depth Estimationpdf
motionMoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animationpdf
gsplatReCap: Better Gaussian Relighting with Cross-Environment Capturespdf
generalVision-Language Embodiment for Monocular Depth Estimationpdf
generalFrequency Dynamic Convolution for Dense Image Predictionpdf
generalIDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identificationpdf
generalConsistency Posterior Sampling for Diverse Image Synthesispdf
generalIMFine: 3D Inpainting via Geometry-guided Multi-view Refinementpdf
generalDeepCompress-ViT: Rethinking Model Compression to Enhance Efficiency of Vision Transformers at the Edgepdf
generalEvOcc: Accurate Semantic Occupancy for Automated Driving Using Evidence Theorypdf
generalTowards Continual Universal Segmentationpdf
gsplatPGC: Physics-Based Gaussian Cloth from a Single Posepdf
beingsOFER: Occluded Face Expression Reconstructionpdf
worldsCubify Anything: Scaling Indoor 3D Object Detectionpdf
generalDEAL: Data-Efficient Adversarial Learning for High-Quality Infrared Imagingpdf
generalBimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objectspdf
generalCoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Modelspdf
gsplatFreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene Reconstructionpdf
generalShow and Tell: Visually Explainable Deep Neural Nets via Spatially-Aware Concept Bottleneck Modelspdf
generalLearning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scenepdf
generalKnowledge Bridger: Towards Training-Free Missing Modality Completionpdf
beingsTexGarment: Consistent Garment UV Texture Generation via Efficient 3D Structure-Guided Diffusion Transformerpdf
worldsSemi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video Derainingpdf
generalTIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time Correctionpdf
generalVSNet: Focusing on the Linguistic Characteristics of Sign Languagepdf
generalLearning to Sample Effective and Diverse Prompts for Text-to-Image Generationpdf
generalMulti-modal Medical Diagnosis via Large-small Model Collaborationpdf
motionImage Referenced Sketch Colorization Based on Animation Creation Workflowpdf
beingsGaPT-DAR: Category-level Garments Pose Tracking via Integrated 2D Deformation and 3D Reconstructionpdf
beingsProbPose: A Probabilistic Approach to 2D Human Pose Estimationpdf
worldsMIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generationpdf
generalABC-Former: Auxiliary Bimodal Cross-domain Transformer with Interactive Channel Attention for White Balancepdf
generalFingerprinting Denoising Diffusion Probabilistic Modelspdf
generalNightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene Segmentationpdf
beingsUMFN: Unified Multi-Domain Face Normalization for Joint Cross-domain Prototype Learning and Heterogeneous Face Recognitionpdf
beingsLUCAS: Layered Universal Codec Avatarspdf
generalD^3: Scaling Up Deepfake Detection by Learning from Discrepancypdf
generalJailbreaking the Non-Transferable Barrier via Test-Time Data Disguisingpdf
general3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucinationpdf
generalGenerative Zero-Shot Composed Image Retrievalpdf
generalTowards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewardspdf
generalSpatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Mutimodal Modelspdf
generalOmnidirectional Multi-Object Trackingpdf
generalPotential Field Based Deep Metric Learningpdf
generalEnhancing Vision-Language Compositional Understanding with Multimodal Synthetic Datapdf
generalDirectional Label Diffusion Model for Learning from Noisy Labelspdf
generalLearning Endogenous Attention for Incremental Object Detectionpdf
worldsStarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generationpdf
generalHomoGen: Enhanced Video Inpainting via Homography Propagation and Diffusionpdf
generalDo ImageNet-trained Models Learn Shortcuts? The Impact of Frequency Shortcuts on Generalizationpdf
beingsHORP: Human-Object Relation Priors Guided HOI Detectionpdf
generalBuilding a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMspdf