Skills MCP Model 博客 提交 Skills

Qwen2.5-VL

阿里云 released the 多模态 model · Launched on 2025-01

Open Source 多模态 开源 API 可用

Qwen2.5-VL is a vision-language model from Alibaba's Tongyi Qianwen team, supporting deep understanding of images, videos, and documents. It can process images of any size with dynamic resolution, support video analysis of up to several hours, and accurately locate object coordinates in images. The 72B version reaches industry-leading levels in multimodal benchmarks, and is one of the benchmarks for open-source vision-language models.

核心特性

  • 动态分辨率视觉理解
  • 长视频分析
  • 物体定位与 OCR
  • 多规格可选

核心优势

  • 视觉理解能力领先
  • 视频分析出色
  • 开源灵活
参数量
3B / 7B / 72B
上下文窗口
128K tokens
API 支持
Yes
开源状态
Open Source
发布日期
2025-01
提供商
阿里云

免费开源;API 调用按量付费

MMBench 87.5%, DocVQA 96.2%, 多模态基准领先

← 返回大模型全景图

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

完全免费,取消任意时间。我们不会发送垃圾邮件。