Instructions to use MCG-NJU/TimeLens2-2B-SFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MCG-NJU/TimeLens2-2B-SFT with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("MCG-NJU/TimeLens2-2B-SFT") model = AutoModelForMultimodalLM.from_pretrained("MCG-NJU/TimeLens2-2B-SFT", device_map="auto") - Notebooks
- Google Colab
- Kaggle
TimeLens2-2B-SFT
This repository contains the supervised fine-tuning (SFT) checkpoint for TimeLens2-2B, a video multimodal language model for temporal grounding. Given a video and a text query, the model is trained to return all time intervals that contain the requested visual evidence.
It is initialized from Qwen3-VL-2B-Instruct and fine-tuned with long-context temporal-grounding supervision. This checkpoint is the starting point for TimeLens2 rollout generation and GRPO training.
For the released TimeLens2 model, see TimeLens2-2B.
Training data and recipe
SFT uses examples from TimeLens2-93K, TimeLens-100K, and the official Ego4D-NLQ-v2 training split. Training retains full-video context and supervises variable-cardinality temporal interval sets. The model sees diverse grounding instructions and timestamp response formats.
The 2B SFT recipe trains for one epoch with packed sequences up to 160K tokens, a global batch size of 128, and a learning rate of 1e-5. The visual encoder is frozen during training.
Intended use
- Generate off-policy rollouts for the TimeLens2 GRPO stage.
- Reproduce or extend the SFT-to-GRPO training pipeline.
- Study long-context video temporal grounding before reinforcement learning.
This is an SFT checkpoint. It is not the final GRPO-optimized release reported in the TimeLens2 paper.
Loading the checkpoint
This checkpoint is hosted in a Hugging Face dataset repository. Download it locally, then load the directory with Transformers:
from huggingface_hub import snapshot_download
from transformers import AutoModelForImageTextToText, AutoProcessor
model_dir = snapshot_download(
repo_id="MCG-NJU/TimeLens2-2B-SFT",
repo_type="dataset",
)
model = AutoModelForImageTextToText.from_pretrained(
model_dir,
torch_dtype="auto",
device_map="auto",
)
processor = AutoProcessor.from_pretrained(model_dir)
See the official repository for the SFT and GRPO training recipes, data annotations, rollout configuration, and evaluation code.
Citation
@article{zhu2026timelens2,
title={TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs},
author={Zhu, Yuhan and Ma, Changlian and Zeng, Xiangyu and Li, Xinhao and Zhang, Zhiqiu and Li, Songze and Zhang, Jun and Jiang, Tianxiang and Yang, Yuandong and Yan, Ziang and others},
journal={arXiv preprint arXiv:2607.17423},
year={2026}
}
- Downloads last month
- 34