Qwen2.5-VL
阿里云 released the 多模态 model · Launched on 2025-01
Open Source
多模态
开源
API 可用
模型简介
Qwen2.5-VL is a vision-language model from Alibaba's Tongyi Qianwen team, supporting deep understanding of images, videos, and documents. It can process images of any size with dynamic resolution, support video analysis of up to several hours, and accurately locate object coordinates in images. The 72B version reaches industry-leading levels in multimodal benchmarks, and is one of the benchmarks for open-source vision-language models.
核心特性
- 动态分辨率视觉理解
- 长视频分析
- 物体定位与 OCR
- 多规格可选
核心优势
- 视觉理解能力领先
- 视频分析出色
- 开源灵活
技术规格
参数量
3B / 7B / 72B
上下文窗口
128K tokens
API 支持
Yes
开源状态
Open Source
发布日期
2025-01
提供商
阿里云
价格信息
免费开源;API 调用按量付费
基准测试
MMBench 87.5%, DocVQA 96.2%, 多模态基准领先