HuggingFaceTB/SmolVLM-256M-Instruct Image-Text-to-Text β’ 0.3B β’ Updated Apr 8, 2025 β’ 987k β’ 396
Qwen/Qwen2.5-VL-7B-Instruct Image-Text-to-Text β’ 8B β’ Updated Apr 6, 2025 β’ 8.63M β’ β’ 1.67k
Running on Zero Agents Featured 2.02k Chat With Janus-Pro-7B π 2.02k A unified multimodal understanding and generation model.
Running on CPU Upgrade Agents 1.02k Open VLM Leaderboard π 1.02k VLMEvalKit Evaluation Results Collection