Code and implementation guidelines for the paper ✨Counting Anything. Project Page: https://mengqi-lei.github.io/count-anything-projectpage/
-
Updated
Aug 12, 2026 - Python
Code and implementation guidelines for the paper ✨Counting Anything. Project Page: https://mengqi-lei.github.io/count-anything-projectpage/
OpenVision (ICCV 2025), OpenVision 2 (CVPR 2026), and OpenVision 3
A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explore how LMMs represent, transform, and align multimodal information internally.
[--branch main] Face Security Foundation Model via Self-Supervised Facial Representation Learning (CVPR 2025). [--branch FSVFM-extension-R1] FS-VFM extension with scalable face-security visual backbones, Linear Probing, FS-Adapter, and DF40 evaluation support.
Recognize Any Regions
[ICLR 2026] FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph Diffusion
A Vision Foundation Model for Cine Cardiac Magnetic Resonance Imaging
One-Shot Open Affordance Learning with Foundation Models (CVPR 2024)
[WAICA-26 Best Student Paper] Official repository of "Enhancing Vision Foundation Models via Multimodal Continual Pre-Training"
MonoDINO-DETR: Depth-Enhanced Monocular 3D Object Detection Using a Vision Foundation Model
This repo collects some latest research work of Generative AI. It provides simple implementations to understand the ideas and some follow-up discussions to inspire future work.
"Boosting Gaze Object Prediction via Pixel-level Supervision from Vision Foundation Model"
Implementation of CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation.
A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
Codebase for probing VFMs and Feature Upsamplers using Intractive Segmentation.
Simple Gradio application integrated with Hugging Face Multimodals to support visual question answering chatbot and more features
A Training Objective for Interpretable Monosemantic Representations
[TGRS 2026] EntSeg: Entropy-Guided Pseudolabel Denoising and Masked Image Consistency for Cross-Domain Remote Sensing Segmentation
Florence-2 quick test
Add a description, image, and links to the vision-foundation-model topic page so that developers can more easily learn about it.
To associate your repository with the vision-foundation-model topic, visit your repo's landing page and select "manage topics."