# Multimodal Machine Learning Applications

**Type:** Topics  
**Canonical URL:** https://scholariq.org/topics/multimodal-machine-learning-applications-2/

## Facts

| Field | Value |
| --- | --- |
| Citations | 715,808 |
| Description | This cluster of papers focuses on the development and improvement of visual question answering systems, image captioning techniques, and neural networks for understanding and generating descriptions of images and videos. The research involves semantic reasoning, multimodal fusion, scene graph generation, attention mechanisms, and deep learning approaches to bridge the gap between vision and language. |
| Domain | Physical Sciences |
| Field | Computer Science |
| OpenAlex ID | https://openalex.org/T11714 |
| Works | 70,709 |

## Topic researchers

Showing 12 of 20.

- [Kaiming He](https://scholariq.org/researchers/kaiming-he/)
- [Yoshua Bengio](https://scholariq.org/researchers/yoshua-bengio/)
- [Geoffrey E. Hinton](https://scholariq.org/researchers/geoffrey-e-hinton/)
- [Shaoqing Ren](https://scholariq.org/researchers/shaoqing-ren/)
- [Ross Girshick](https://scholariq.org/researchers/ross-girshick/)
- [Xiangyu Zhang](https://scholariq.org/researchers/xiangyu-zhang/)
- [Jian Sun](https://scholariq.org/researchers/jian-sun-2/)
- [Andrew Zisserman](https://scholariq.org/researchers/andrew-zisserman/)
- [Xiaogang Wang](https://scholariq.org/researchers/xiaogang-wang/)
- [Li Fei-Fei](https://scholariq.org/researchers/li-fei-fei/)
- [Ilya Sutskever](https://scholariq.org/researchers/ilya-sutskever/)
- [Trevor Darrell](https://scholariq.org/researchers/trevor-darrell/)

---
Source: ScholarIQ — public research metadata, principally OpenAlex. See https://scholariq.org/sources/ for provenance and https://scholariq.org/methodology/ for what these figures mean.
