Grounded multimodal language model that links generated text to regions in images.
- Jun 20, 2023
- —
- textimage
- —
- —
- 1
Releases
Jun 20, 2023Kosmos-2 grounded multimodal model introducedResearch paper
Microsoft Research introduces Kosmos-2, linking generated language to object regions in images through grounding.

