Skills MCP Model 博客 提交 Skills

Phi-4-multimodal

Microsoft released the 多模态 model · Launched on 2025-05

Open Source 多模态 开源 API 可用

Phi-4-multimodal is an ultra-lightweight multimodal model launched by Microsoft, with only 5.6B parameters yet supporting full modality processing of speech, vision, and text. It can run in real-time on mobile devices, supporting speech recognition, image understanding, and text generation. Phi-4-multimodal sets a new benchmark in small-model multimodal field, making it the preferred multimodal solution for edge AI devices.

核心特性

  • 5.6B 超轻量
  • 语音+视觉+文本全模态
  • 128K 上下文
  • 移动端运行

核心优势

  • 全模态融合
  • 移动端部署
  • 参数量极小
参数量
5.6B
上下文窗口
128K tokens
API 支持
Yes
开源状态
Open Source
发布日期
2025-05
提供商
Microsoft

免费开源

在 5.6B 参数规模下实现全模态能力,业界首创

← 返回大模型全景图

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

完全免费,取消任意时间。我们不会发送垃圾邮件。