Instructions to use omron-sinicx/CLAP-Qwen3.5-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use omron-sinicx/CLAP-Qwen3.5-2B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("omron-sinicx/CLAP-Qwen3.5-2B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
CLAP-Qwen3.5-2B
CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding checkpoint: Qwen/Qwen3.5-2B fully fine-tuned with the CLAP recipe on LIBERO benchmark tasks.
CLAP (Causal Language-Action Prediction) converts a pretrained VLM into a VLA with no architectural changes: it prepends each numeric action-token sequence with a natural-language action description (a "language-action plan"), causally conditioning precise action-token prediction on that plan while staying close to the VLM's pretrained language distribution.
Quick Start
1. Download
pip install huggingface_hub
hf download omron-sinicx/CLAP-Qwen3.5-2B --local-dir ./CLAP-Qwen3.5-2B
2. Load with transformers
import pickle
import torch
from transformers import Qwen2_5_VLForConditionalGeneration, Qwen2_5_VLProcessor
ckpt_dir = "./CLAP-Qwen3.5-2B"
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
f"{ckpt_dir}/model_final", torch_dtype=torch.bfloat16, device_map="auto",
)
processor = Qwen2_5_VLProcessor.from_pretrained(f"{ckpt_dir}/model_final")
# Load dataset stats (required for action denormalization)
with open(f"{ckpt_dir}/dataset_stats.pkl", "rb") as f:
dataset_stats = pickle.load(f)
3. Load with the CLAP training framework
from rv_train.train import get_pretrained_model
model, cfg = get_pretrained_model("./CLAP-Qwen3.5-2B", device=0)
model.eval()
dataset_stats.pkl
Action normalization statistics computed from the training dataset. Required at inference time to denormalize model outputs back to the original action space.
import pickle
with open("dataset_stats.pkl", "rb") as f:
stats = pickle.load(f)
# stats contains mean/std for action dimensions
Checkpoint
This is the main branch weights, corresponding to training step 18000.
Training Details
- Base Model:
Qwen/Qwen3.5-2B - Method: CLAP (Causal Language-Action Prediction), Full Fine-Tuning
- Prefix masking augmentation: disabled (
action_mask_aug_per: 0.0) — "without prefix masking" variant - Dataset: LIBERO (via RoboVerse)
- Framework: CLAP LIBERO Training
Files
| File | Description |
|---|---|
model_final/model.safetensors |
Full model weights |
model_final/config.json |
Model configuration |
model_final/tokenizer.json |
Tokenizer |
dataset_stats.pkl |
Action normalization statistics (required for inference) |
config.yaml |
Training configuration |
Citation
@inproceedings{ishitoya2026clap,
title={CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding},
author={Yuri Ishitoya and Jeremy Siburian and Masashi Hamaya and Haruto Suzuki and Kuniaki Saito and Toshihiko Fukushima and Cristian C. Beltran-Hernandez and Mai Nishimura},
booktitle={10th Annual Conference on Robot Learning},
year={2026}
}
License
CC-BY-NC-4.0 (following the upstream CLAP license). Subject to Qwen License for the base model.
- Downloads last month
- 6