Phi-4-multimodal
Microsoft released the 多模态 model · Launched on 2025-05
Open Source
多模态
开源
API 可用
模型简介
Phi-4-multimodal is an ultra-lightweight multimodal model launched by Microsoft, with only 5.6B parameters yet supporting full modality processing of speech, vision, and text. It can run in real-time on mobile devices, supporting speech recognition, image understanding, and text generation. Phi-4-multimodal sets a new benchmark in small-model multimodal field, making it the preferred multimodal solution for edge AI devices.
核心特性
- 5.6B 超轻量
- 语音+视觉+文本全模态
- 128K 上下文
- 移动端运行
核心优势
- 全模态融合
- 移动端部署
- 参数量极小
技术规格
参数量
5.6B
上下文窗口
128K tokens
API 支持
Yes
开源状态
Open Source
发布日期
2025-05
提供商
Microsoft
价格信息
免费开源
基准测试
在 5.6B 参数规模下实现全模态能力,业界首创